VLDB 2026 Research / reviewers in the wild / expert
Guang Yang 0006
dblp:25/5712-6
· DBLP profile ↗
137ranked-venue papers
7as first author
116since 2021 · last 2026
0000-0001-7344-7733ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 82 · 2 first-author · 67 since 2021Artificial intelligence and machine learning · 47 · 2 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 2 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report EvaluationabstractAutomatic medical report generation has the potential to support clinical diagnosis, reduce the workload of radiologists, and demonstrate potential for enhancing diagnostic consistency. However, current evaluation metrics often fail to reflect the clinical reliability of generated reports. Overlap-based methods overlook fine-grained details (e.g., location, severity), diagnostic metrics are constrained by fixed vocabularies. Some diagnostic metrics are limited by fixed vocabularies or templates, reducing their ability to capture diverse clinical expressions. LLM-based metrics lack interpretable reasoning, limiting trust in clinical settings. Therefore, we propose a Granular Explainable Multi-Agent Score (GEMA-Score) in this paper, which conducts both objective quantification and subjective evaluation through a large language model-based multi-agent workflow. Our GEMA-Score parses structured reports and employs stable calculations through interactive exchanges of information among agents to assess disease diagnosis, location, severity, and uncertainty. Additionally, an LLM-based scoring agent evaluates completeness, readability, and clinical terminology while providing explanatory feedback. Extensive experiments show that GEMA-Score achieves the highest correlation with human experts on public datasets (Kendall = 0.69 on ReXVal; 0.45 on RadEvalX), demonstrating improved clinical scoring reliability. Zhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng, Huichi Zhou, Zihao Jin, Zhifan Gao, Dominic C. Marshall, Yingying Fang, Guang Yang 0006 |
AAAI | 11 |
| 2026 | 3D Wavelet-Based Structural Priors for Controlled Diffusion in Whole-Body Low-Dose PET Denoising
Peiyuan Jing, Chun-Wun Cheng, Zhenxuan Zhang, Liutao Yang, Thiago Lima 0001, Klaus Strobel, Antoine Leimgruber, Angelica I. Avilés-Rivero, Guang Yang 0006, Javier A. Montoya-Zegarra |
ICPR (9) | 10 |
| 2026 | A global-to-local state space model with context-mixing dynamic kernels for medical image classification
Yuehui Liao, Pengxia Yue, Panfei Li, Guang Yang 0006, Qun Jin, Shiqing Zhang, Xiaobo Lai, Qi Tian 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Adaptive personalized federated learning for left atrium segmentation from multi-center LGE CMR images
Zhe Liu 0004, Yuyang Xin, Guang Yang 0006, Qiaoying Teng, Xiongfeng Cao, Guozhong Du, Jun Chen 0030, Lingyun Zu |
Expert Syst. Appl. | 3 |
| 2026 | SeLoRA: Self-expanding LoRA for high-quality and efficient medical image synthesis
Hongwei Li 0004, Wei Pang 0001, Giorgos Papanastasiou, Guang Yang 0006, Ehsan Mohammadi Pasand, Theodore Harrison-Drummond, Chengjia Wang |
Expert Syst. Appl. | 5 |
| 2026 | A hybrid CNN-Mamba state space model with pyramid-pooled skip connections for prostate tumor segmentation
Xueting Wei, Yuehui Liao, Shiqing Zhang, Guang Yang 0006, Qun Jin, Xiaobo Lai, Qi Tian 0001 |
Expert Syst. Appl. | 5 |
| 2026 | RectMamba: Exploring state space models with entropy-divergence framework for noisy label rectification
Ningwei Wang, Weiqiang Jin, Haixia Bi, Guang Yang 0006 |
Neurocomputing | 4 |
| 2026 | An efficient, scalable, and adaptable plug-and-play temporal attention module for motion-guided cardiac segmentation with sparse temporal labelsabstractUNet and DT-VNet. Integrating TAM into SAM yields a temporal SAM that reduces Hausdorff distance (HD) from 3.99 mm to 3.51 mm on the CAMUS dataset, while integrating TAM into a pre-trained MedSAM reduces HD from 3.04 to 2.06 pixels after fine-tuning on the EchoNet-Dynamic dataset. On the ACDC 3D dataset, our TAM-UNet and TAM-DT-VNet achieve substantial reductions in HD, from 7.97 mm to 4.23 mm and 6.87 mm to 4.74 mm, respectively. Additionally, TAM's training does not require segmentation of ground truths from all time frames and can be achieved with sparse temporal annotation. TAM is thus a robust, generalizable, and adaptable solution for motion-awareness enhancement that is easily scaled from 2D to 3D. The code is available at https://github.com/kamruleee51/TAM. Md. Kamrul Hasan 0002, Guang Yang 0006, Choon Hwai Yap |
Medical Image Anal. | 2 |
| 2026 | Reason like a radiologist: Chain-of-thought and reinforcement learning for verifiable report generationabstractRadiology report generation is critical for efficiency, but current models often lack the structured reasoning of experts and the ability to explicitly ground findings in anatomical evidence, which limits clinical trust and explainability. This paper introduces BoxMed-RL, a unified training framework to generate spatially verifiable and explainable chest X-ray reports. BoxMed-RL advances chest X-ray report generation through two integrated phases: (1) Pretraining Phase. BoxMed-RL learns radiologist-like reasoning through medical concept learning and enforces spatial grounding with reinforcement learning. (2) Downstream Adapter Phase. Pretrained weights are frozen while a lightweight adapter ensures fluency and clinical credibility. Experiments on two widely used public benchmarks (MIMIC-CXR and IU X-Ray) demonstrate that BoxMed-RL achieves an average 7 % improvement in both METEOR and ROUGE-L metrics compared to state-of-the-art methods. An average 5 % improvement in large language model-based metrics further underscores BoxMed-RL's robustness in generating high-quality reports. Related code and training templates are publicly available at https://github.com/ayanglab/BoxMed-RL. Peiyuan Jing, Kinhei Lee, Zhenxuan Zhang, Huichi Zhou, Zhengqing Yuan, Zhifan Gao, Lei Zhu 0003, Giorgos Papanastasiou, Yingying Fang, Guang Yang 0006 |
Medical Image Anal. | 10 |
| 2026 | Explicit differentiable slicing and global deformation for cardiac mesh reconstructionabstractThree-dimensional (3D) mesh reconstruction of the cardiac anatomy from medical images is useful for shape and motion measurements and biophysics simulations. However, 3D medical images are often acquired as 2D slices that are sparsely sampled (e.g., large slice spacing) and noisy, and 3D mesh reconstruction on such data is a challenging task. Traditional voxel-based approaches utilize non-differentiable pre- and post-processing that compromises fidelity to images, while mesh-level deep learning approaches require large 3D mesh annotations that are difficult to obtain. Differentiable cross-domain supervision from 2D images to 3D meshes is therefore crucial for enabling end-to-end optimization in medical imaging. While there have been attempts to approximate the voxelization and slicing of meshes that are being optimized, there has not yet been a method for directly using 2D slices to supervise 3D mesh reconstruction in a differentiable manner. Here, we propose a novel explicit differentiable voxelization and slicing (DVS) algorithm allowing gradient backpropagation to a 3D mesh from its slices, which facilitates refined mesh optimization directly supervised by the losses defined on 2D images. Further, we propose an innovative framework for extracting patient-specific left ventricle (LV) meshes from medical images by coupling DVS with a graph harmonic deformation (GHD) mesh morphing descriptor of cardiac shape that naturally preserves mesh quality and smoothness during optimization. The proposed framework achieves state-of-the-art performance in cardiac mesh reconstruction tasks from densely sampled (CT) as well as sparsely sampled (MRI stack with few slices) images, outperforming alternatives, including Marching Cubes, statistical shape models, algorithms with vertex-based mesh morphing algorithms and alternative methods for image-supervision of mesh reconstruction. Experimental results demonstrate that our method achieves an overall Dice score of 90% during a sparse fitting on multi-datasets. The proposed method can further quantify clinically useful parameters such as ejection fraction and global myocardial strains, closely matching the ground truth and outperforming the traditional voxel-based approach in sparse images. Yihao Luo, Dario Sesia, Fanwen Wang, Yinzhe Wu 0001, Wenhao Ding, Md. Kamrul Hasan 0002, Fadong Shi, Anoop Shah, Amit Kaura, Jamil Mayet, Guang Yang 0006, Choon Hwai Yap |
Medical Image Anal. | 12 |
| 2026 | Test-time generative augmentation for medical image segmentation
Xiao Ma 0011, Yuhui Tao, Zetian Zhang, Yuhan Zhang 0001, Xi Wang 0013, Sheng Zhang 0024, Zexuan Ji, Yizhe Zhang 0001, Qiang Chen 0004, Guang Yang 0006 |
Medical Image Anal. | 10 |
| 2026 | Simultaneous multi-slice Cardiac Diffusion Tensor Imaging with variable CAIPIRINHA shifts and artefact-aware AIabstractCardiac Diffusion Tensor Imaging (cDTI) provides unique insights into myocardial microstructure in-vivo but requires averaging multiple repetitions for adequate signal quality, leading to prohibitively long acquisition times. Standard acceleration strategies, such as reducing repetitions and employing simultaneous multi-slice (SMS) imaging, are limited by low signal-to-noise ratio (SNR) and inter-slice leakage artefacts, respectively. We introduce ORCAS, a unified framework that synergistically combines a novel variable CAIPIRINHA acquisition with an artefact-aware AI reconstruction to overcome these challenges. The variable CAIPIRINHA scheme decoheres SMS artefacts across repetitions, while our dual-domain deep learning model simultaneously suppresses these artefacts and combats the low SNR from fewer repetitions. The model is guided by patient-specific single-band auxiliary data to preserve anatomical fidelity. Validated on ex-vivo hearts with and without anomalies, ORCAS achieves an over 18-fold acceleration by combining these strategies, reducing a whole-heart scan from over two hours to under 7 min. This is accomplished while reducing errors in key biomarkers, such as Fractional Anisotropy, by up to 64%. The framework preserves essential microstructural properties and the delineation of abnormalities, representing a significant step towards the clinical translation of whole-heart cDTI. • Novel variable CAIPIRINHA reduces SMS artefacts in cardiac DTI. • AI framework achieves 18×acceleration while preserving biomarkers. • Reduces DTI errors by 64% compared to conventional reconstruction. • Enables whole-heart cDTI in under 7 min vs over 2 h. • Preserves abnormalities even at extreme acceleration factors. Michael Tänzer, Eun Ji Lim, Huaqi Qiu, Camila Munoz, Andrew D. Scott, Dudley Pennell, Pedro F. Ferreira, Daniel Rueckert, Guang Yang 0006, Sonia Nielles-Vallespin |
Medical Image Anal. | 9 |
| 2026 | From noisy labels to intrinsic structure: A geometric-structural dual-guided framework for noise-robust medical image segmentation
Tao Wang 0085, Zhenxuan Zhang, Yuanbo Zhou, Xinlin Zhang, Yuanbin Chen, Tao Tan 0002, Guang Yang 0006, Tong Tong 0001 |
Medical Image Anal. | 7 |
| 2026 | Dynamical multi-order responses and global semantic-infused adversarial learning: A robust airway segmentation methodabstractAutomated airway segmentation in computerized tomography (CT) images is crucial for the accurate diagnosis of lung diseases. However, the scarcity of manual annotations hinders the efficacy of supervised learning, while unconstrained intensities and sample imbalance lead to discontinuity and false-negative issues. To address these challenges, we propose a novel airway segmentation model named Dynamical Multi-order responses and Global Semantic-infused Adversarial network (DMGSA), integrating the unsupervised and supervised learning in parallel to alleviate the label scarcity of airway. In the unsupervised branch, (1) we propose several novel strategies of Dynamic Mask-Ratio (DMR) to empower the model to perceive context information of varying sizes, mimicking the laws of human learning vividly; (2) we present a novel target of Multi-Order Normalized Responses (MONR), exploiting the distinct order exponential operation of raw images and oriented gradients to enhance the textural representations of bronchioles; (3) we introduce the Adversarial Learning (AL) on the top of MONR module to discern nuances between real and fake images, focusing on capturing the textural features of terminal bronchioles. For the supervised branch, we propose an innovative Generalized Mean pooling based Global Semantic-infused (GMGS) module to ulteriorly improve the robustness. Ultimately, we have verified the method performance and robustness by training on normal lung disease datasets, while testing on lung cancer, COVID-19 and Lung fibrosis datasets. All experimental results have proved that our method exceeds state-of-the-art methods significantly. Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Yongkai Liu, Giorgos Papanastasiou, Zhifan Gao, Shuo Li 0001, Simon Walsh, Guang Yang 0006 |
Medical Image Anal. | 10 |
| 2026 | Hybrid aggregation strategy with double inverted residual blocks for lightweight salient object detection
Mingfeng Jiang, Xian Fang, Jiatong Chen, Yaming Wang, Guang Yang 0006 |
Neural Networks | 6 |
| 2026 | From Coarse to Continuous: Progressive Refinement Implicit Neural Representation for Motion-Robust Anisotropic MRI ReconstructionabstractIn motion-robust magnetic resonance imaging (MRI), slice-to-volume reconstruction is critical for recovering anatomically consistent 3D brain volumes from 2D slices, especially under accelerated acquisitions or patient motion. However, this task remains challenging due to hierarchical structural disruptions. It includes local detail loss from k-space undersampling, global structural aliasing caused by motion, and volumetric anisotropy. Therefore, we propose a progressive refinement implicit neural representation (PR-INR) framework. Our PR-INR unifies motion correction, structural refinement, and volumetric synthesis within a geometry-aware coordinate space. Specifically, a motion-aware diffusion module is first employed to generate coarse volumetric reconstructions that suppress motion artifacts and preserve global anatomical structures. Then, we introduce an implicit detail restoration module that performs residual refinement by aligning spatial coordinates with visual features. It corrects local structures and enhances boundary precision. Further, a voxel continuous-aware representation module represents the image as a continuous function over 3D coordinates. It enables accurate inter-slice completion and high-frequency detail recovery. We evaluate PR-INR on five public MRI datasets under various motion conditions (3% and 5% displacement), undersampling rates (4x and 8x) and slice resolutions (scale = 5). Experimental results demonstrate that PR-INR outperforms state-of-the-art methods in both quantitative reconstruction metrics and visual quality. It further shows generalization and robustness across diverse unseen domains. Zhenxuan Zhang, Lipei Zhang, Yanqi Cheng, Zi Wang 0005, Fanwen Wang, Haosen Zhang, Yinzhe Wu 0001, Angelica I. Avilés-Rivero, Zhifan Gao, Guang Yang 0006, Peter J. Lally |
IEEE Trans. Image Process. | 12 |
| 2026 | MSAFed: Generalized Multi-Stage and Adaptive Federated Learning for Test-Time Medical SegmentationabstractFederated learning (FL) enables collaborative model training across multiple medical centers without sharing data, offering significant promise for privacy-preserving AI in healthcare. However, FL models often lack generalization across all participating clients (inside FL) and perform poorly when deployed to unseen clients (outside FL), particularly in heterogeneous domains. Current test-time adaptation methods for outside FL fail to address biases in personalized models toward source distributions, limiting their clinical applications. To tackle these challenges, we propose MSAFed, a generalized multi-stage adaptive FL framework that enhances both inside generalization and outside test-time adaptation. During pretraining, intra-client and inter-client contrastive learning with prototype-aware aggregation produces a generalized global model. An adaptive learning rate strategy further improves inside FL generalization. For unseen clients, source knowledge, including adaptive learning rates and prototypes, is leveraged to dynamically adapt the network architecture during test time. Experiments on three real-world multi-center medical datasets demonstrate the effectiveness of MSAFed, achieving superior performance on both inside and outside FL tasks. Jiajie Jin, Xuanmin Chen, Liyan Ma, Shihui Ying, Guang Yang 0006, Tieyong Zeng |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | 4-D Reconstruction of Fetal Left Ventricle From Echocardiography via 2.5-D Radial Segmentation and Graph-Fourier Reconstruction
Md. Kamrul Hasan 0002, Haziq Shahard, Lucas Iijima, Nida Ruseckaite, Yihao Luo, Iris Scharnreitner, Andreas Tulzer, Bin Liu 0040, Guang Yang 0006, Choon Hwai Yap |
IEEE Trans. Medical Imaging | 10 |
| 2026 | EviVLM: When Evidential Learning Meets Vision Language Model for Medical Image SegmentationabstractThe disparity between image and text representations, often referred to as the modality gap, remains a significant obstacle for Vision Language Models (VLMs) in medical image segmentation. This gap complicates multi-modal fusion, thereby restricting segmentation performance. To address this challenge, we propose Evidence-driven Vision Language Model (EviVLM)-a novel paradigm that integrates Evidential Learning (EL) into VLMs to systematically measure and mitigate the modality gap for enhanced multi-modal fusion. To drive this paradigm, an Evidence Affinity Map Generator (EAMG) is proposed to collect complementary cross-modal evidences by learning a global cross-modal affinity map, thus refining modality-specific evidence embedding. An Evidence Differential Similarity Learning (EDSL) is further proposed to collect consistent cross-modal evidences by performing Bias-Variance Decomposition on differential matrix derived from bidirectional similarity matrices between image and text evidence embeddings. Finally, the subjective logic is used for mapping the collected evidences to opinions, and the Dempster-Shafer's theory based combination rule is introduced for opinion aggregation, thereby quantifying the modality gap and facilitating effective multi-modal integration. Experimental results on three public medical image segmentation datasets validate that the proposed EviVLM can achieve state-of-the-art performance. Code is available at: https://github.com/QingtaoPan/EviVLM. Qingtao Pan, Zhengrong Li, Guang Yang 0006, Bing Ji 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2026 | Toward Modality- and Sampling-Universal Learning Strategies for Accelerating Cardiovascular Imaging: Summary of the CMRxRecon2024 ChallengeabstractCardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the clinical reference standard for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging sequences, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen modalities and robustness to diverse undersampling patterns. We introduced the largest public multi-modality CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging. Fanwen Wang, Zi Wang 0005, Yan Li 0064, Chen Qin, Shuo Wang 0011, Kunyuan Guo, Mengting Sun, Mingkai Huang, Michael Tänzer, Qirong Li, Yinzhe Wu 0001, Haosen Zhang, Kian Anvari Hamedani, Yuntong Lyu, Longyu Sun, Tianxing He, Lizhen Lan, Qiong Yao, Bingyu Xin, Dimitris N. Metaxas, Narges Razizadeh, Shahabedin Nabavi, George Yiasemis, Jonas Teuwen, Daniel B. Ennis, Zhihao Xue, Ruru Xu, Ilkay Öksüz, Donghang Lyu, Yanxin Huang, Xinrui Guo, Ruqian Hao, Jaykumar H. Patel, Guanke Cai, Binghua Chen, Sha Hua, Zhensen Chen, Qi Dou 0001, Xiahai Zhuang, Wenjia Bai, Harry Qin, He Wang 0016, Claudia Prieto, Michael Markl 0001, Alistair A. Young, Hao Li 0082, Xihong Hu, Lianming Wu, Xiaobo Qu 0001, Guang Yang 0006, Chengyan Wang |
IEEE Trans. Medical Imaging | 62 |
| 2026 | LCM-Net: LLM-Driven Cross-Modality MoE Feature Fusion Network for Cancer Survival AnalysisabstractCancer survival analysis aims to predict survival outcomes to evaluate the efficacy and prognosis of treatment. Although current approaches have designed diverse cross-modal learning methods to integrate genetic data and pathology images, they are frequently hindered by data redundancy. Pattern representation in high-dimensional genetic data remains a significant hurdle. Pathology data analysis is computationally intensive because of the giga-pixel resolution. Moreover, the heterogeneity of data types poses a barrier to extending multimodal fusion methods. To address the aforementioned issues, we propose a novel LLM-driven Cross-Modality MoE-feature Fusion Network (LCM-Net) with three innovative modules for boosting cancer survival prediction. Specifically, the Genomic Language Alignment (GLA) module integrates genomic features with learnable prompts. Utilizing large language models, it encodes genomic information into concise and semantically relevant representations. Then, we devise the Pathological Feature Refinement (PFR) module to serve as a plug-and-play component that filters out irrelevant regions in pathology images. Finally, we propose a Multimodal Expert Integration (MEI) module to effectively leverage the capabilities of different experts, integrating the processed features from both the genomic and pathological domains. Extensive experiments on five public datasets demonstrate that our approach outperforms state-of-the-art methods, and the ablation study confirms the effectiveness of the proposed modules. Our code is publicly available at https://github.com/script-Yang/LCM-Net. Sicheng Yang 0001, Haipeng Zhou, Weiming Wang 0002, Shifu Chen, Guang Yang 0006, Huazhu Fu, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 6 |
| 2026 | Cyclic Self-Supervised Diffusion for Ultra Low-Field to High-Field MRI SynthesisabstractSynthesizing high-quality images from low-field MRI holds significant potential. Low-field MRI is cheaper, more accessible, and safer, but suffers from low resolution and poor signal-to-noise ratio. This synthesis process can reduce reliance on costly acquisitions and expand data availability. However, synthesizing high-field MRI still suffers from a clinical fidelity gap. There is a need to preserve anatomical fidelity, enhance fine-grained structural details, and bridge domain gaps in image contrast. To address these issues, we propose a cyclic self-supervised diffusion (CSS-Diff) framework for high-field MRI synthesis from real low-field MRI data. Our core idea is to reformulate diffusion-based synthesis under a cycle-consistent constraint. It enforces anatomical preservation throughout the generative process rather than just relying on paired pixel-level supervision. The CSS-Diff framework further incorporates two novel processes. The slice-wise gap perception network aligns inter-slice inconsistencies via contrastive learning. The local structure correction network enhances local feature restoration through self-reconstruction of masked and perturbed patches. Extensive experiments on cross-field synthesis tasks demonstrate the effectiveness of our method, achieving state-of-the-art performance (e.g., $31.80~\pm ~2.70$ dB in PSNR, $0.943~\pm ~0.102$ in SSIM, and $0.0864~\pm ~0.0689$ in LPIPS). Beyond pixel-wise fidelity, our method also preserves fine-grained anatomical structures compared with the original low-field MRI (e.g., left cerebral white matter error drops from 12.1% to 2.1%, cortex from 4.2% to 3.7%). To conclude, our CSS-Diff can synthesize images that are both quantitatively reliable and anatomically consistent. The code is available at: https://github.com/ayanglab/CSS-Diff. Zhenxuan Zhang, Peiyuan Jing, Zi Wang 0005, Ula Briski, Coraline Beitone, Yinzhe Wu 0001, Fanwen Wang, Liutao Yang, Zhifan Gao, Zhaolin Chen, Kh Tohidul Islam, Guang Yang 0006, Peter J. Lally |
IEEE Trans. Medical Imaging | 14 |
| 2025 | Agnostic Biomolecular Binding Affinity Prediction via Frame Averaging Graph TransformerabstractPredicting binding affinity between biomolecules is a critical task in drug discovery, where deep learning methods have achieved significant progress. However, many existing approaches employ modality-specific network architectures, limiting their direct applicability across diverse biomolecular interaction types. In this work, we propose F3Affinity, a novel structure-based graph transformer that is agnostic to biomolecular interaction type in its architectural design. Leveraging a frame averaging technique, our model flexibly learns SE(3)-invariant representations of input structures. We further demonstrate that F3Affinity can be independently applied to various binding affinity prediction benchmarks without requiring any pre-trained embeddings, including protein-ligand, protein-protein, and protein-nucleic acid interactions, and achieving competitive performance. These results validate its broad applicability and effectiveness across multiple biomolecular interaction types. Krinos Li, Lucas He, Xianglu Xiao, Shenglong Deng, Zijun Zhong, Guang Yang 0006 |
BIBM | 7 |
| 2025 | A Simple Data Augmentation for Feature Distribution Skewed Federated LearningabstractFederated Learning (FL) facilitates collaborative learning among multiple clients in a distributed manner and ensures the security of privacy. However, its performance inevitably degrades with non-Independent and Identically Distributed (non-IID) data. In this paper, we focus on the feature distribution skewed FL scenario, a common non-IID situation in real-world applications where data from different clients exhibit varying underlying distributions. This variation leads to feature shift, which is a key issue of this scenario. While previous works have made notable progress, few pay attention to the data itself, i.e., the root of this issue. The primary goal of this paper is to mitigate feature shift from the perspective of data. To this end, we propose a simple yet remarkably effective input-level data augmentation method, namely FedRDN, which randomly injects the statistical information of the local distribution from the entire federation into the client’s data. This is beneficial to improve the generalization of local feature representations, thereby mitigating feature shift. Moreover, our FedRDN is a plug-and-play component, which can be seamlessly integrated into the data augmentation flow with only a few lines of code. Extensive experiments on several datasets show that the performance of various representative FL methods can be further improved by integrating our FedRDN, demonstrating its effectiveness, strong compatibility and generalizability. Code is available at https://github.com/IAMJackYan/FedRDN. Yunlu Yan, Huazhu Fu, Yuexiang Li, Jinheng Xie, Jun Ma 0008, Guang Yang 0006, Lei Zhu 0003 |
CVPR | 6 |
| 2025 | Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report GenerationabstractDespite significant advancements in automated report generation, the opaqueness of text interpretability continues to cast doubt on the reliability of the content produced. This paper introduces a novel approach to identify specific image features in X-ray images that influence the outputs of report generation models. Specifically, we propose Cyclic Vision-Language Manipulator (CVLM), a module to generate a manipulated X-ray from an original X-ray and its report from a designated report generator. The essence of CVLM is that cycling manipulated X-rays to the report generator produces altered reports aligned with the alterations pre-injected into the reports for X-ray generation, achieving the term ``cyclic manipulation''. This process allows direct comparison between original and manipulated X-rays, clarifying the critical image features driving changes in reports and enabling model users to assess the reliability of the generated texts. Empirical evaluations demonstrate that CVLM can identify more precise and reliable features compared to existing explanation methods, significantly enhancing the transparency and applicability of AI-generated reports. Yingying Fang, Zihao Jin, Shaojie Guo, Jinda Liu, Zhiling Yue, Yijian Gao, Junzhi Ning, Simon Walsh, Guang Yang 0006 |
IJCAI | 10 |
| 2025 | Surgical-MambaLLM: Mamba2-Enhanced Multimodal Large Language Model for VQLA in Robotic Surgery
Pengfei Hao, Hongqiu Wang, Shuaibo Li, Zhaohu Xing, Guang Yang 0006, Kaishun Wu, Lei Zhu 0003 |
MICCAI (9) | 5 |
| 2025 | A Parallel Network for LRCT Segmentation and Uncertainty Mitigation with Fuzzy SetsabstractAccurate segmentation of airways in Low-Resolution CT (LRCT) scans is vital for diagnostics in scenarios such as reduced radiation exposure, emergency response, or limited resources. Yet manual annotation is labor-intensive and prone to variability, while existing automated methods often fail to capture small airway branches in lower-resolution 3D data. To address this, we introduce \textbf{FuzzySR}, a parallel framework that merges super-resolution (SR) and segmentation. By concurrently producing high-resolution reconstructions and precise airway masks, it enhances anatomic fidelity and captures delicate bronchi. FuzzySR employs a deep fuzzy set mechanism, leveraging learnable $t$-distribution and triangular membership functions via cross-attention. Through parameters $\mu$, $\sigma$, and $d_f$, it preserves uncertain features and mitigates boundary noise. Extensive evaluations on lung cancer, COVID-19, and pulmonary fibrosis datasets confirm FuzzySR’s superior segmentation accuracy on LRCT, surpassing even high-resolution baselines. By uniting fuzzy-logic-driven uncertainty handling with SR-based resolution enhancement, FuzzySR effectively bridges the gap for robust airway delineation from LRCT data. Yang Nan 0002, Xiaodan Xing, Yingying Fang, Simon Walsh, Guang Yang 0006 |
UAI | 6 |
| 2025 | DMRN: A Dynamical Multi-Order Response Network for the Robust Lung Airway SegmentationabstractAutomated airway segmentation in CT images is crucial for lung diseases' diagnosis. However, manual annotation scarcity hinders supervised learning efficacy, while unlimited intensities and sample imbalance lead to discontinuity and false-negative issues. To address these challenges, we propose a novel airway segmentation model named Dynamical Multi-order Response Network (DMRN), integrating the unsupervised and supervised learning in parallel to alleviate the label scarcity of airway. In the unsupervised branch, (1) we propose several novel strategies of Dynamic Mask-Ratio (DMR) to enable the model to perceive context information of varying sizes, mimicking the laws of human learning vividly; (2) we present a novel target of Multi-Order Normalized Responses (MONR), exploiting the distinct order exponential operation of raw images and oriented gradients to enhance the textural representations of bronchioles. For the supervised branch, we directly predict the final full segmentation map by the large-ratio cube-masked input instead of full input. Ultimately, we verify the method performance and robustness by training on normal lung disease datasets, while testing on lung cancer, COVID-19 and Lung fibrosis datasets. All experimental results have proved that our method exceeds state-of-the-art methods significantly. Code will be released in the future. Sheng Zhang 0024, Jinge Wu, Junzhi Ning, Guang Yang 0006 |
WACV | 4 |
| 2025 | McCaD: Multi-Contrast MRI Conditioned, Adaptive Adversarial Diffusion Model for High-Fidelity MRI Synthesis
Sanuwani Dayarathna, Kh Tohidul Islam, Bohan Zhuang, Guang Yang 0006, Jianfei Cai 0001, Meng Law, Zhaolin Chen |
WACV | 4 |
| 2025 | Self-Adaptive LLM Instructions Optimization for Aspect-Based Sentiment Analysis by Incorporating Emotion-Oriented In-ContextsabstractABSTRACT Aspect‐based Sentiment Analysis (ABSA) is a vital NLP task that identifies sentiment towards specific entities or aspect terms within a text. Recently, large language models (LLMs) have shown impressive capabilities in semantic comprehension and logical inference. However, LLM hallucinations pose challenges in accurately determining sentiment polarity for aspect terms, leading to performance issues. Moreover, current ABSA methods often fail to fully leverage the vast prior knowledge embedded within LLMs, resulting in suboptimal classification outcomes for specific aspects. Inspired by these challenges, we propose the BYD‐OBS‐ABSA framework—‘Beyond Simple Observations, Embracing Comprehensive Contextual Insights’ for ABSA tasks. This framework leverages unique in‐context constraints, backgrounds, and analogical reasoning to address LLM hallucinations and uses self‐adaptive bootstrap instructions optimization to enhance LLM predictions. BYD‐OBS‐ABSA integrates various in‐context augmentation strategies, including emotion‐oriented backgrounds, constraints, and analogical reasoning. BYD‐OBS‐ABSA further improves initial LLM instructions through adaptive iterative optimization using a random search bootstrap algorithm, maximizing the benefits of LLM prompting. Extensive zero/few‐shot experiments with GPT‐3.5‐turbo across six public datasets validate the effectiveness and robustness of our framework, even surpassing human judgment in certain scenarios. Weiqiang Jin, Bohang Shi, Ningwei Wang, Biao Zhao 0003, Guang Yang 0006 |
Comput. Intell. | 7 |
| 2025 | A prompting multi-task learning-based veracity dissemination consistency reasoning augmentation for few-shot fake news detection
Weiqiang Jin, Ningwei Wang, Tao Tao 0005, Mengying Jiang, Yebei Xing, Biao Zhao 0003, Haibin Duan, Guang Yang 0006 |
Eng. Appl. Artif. Intell. | 9 |
| 2025 | Dynamic mask stitching-guided region consistency for semi-supervised 3D medical image segmentation
Dongsheng Ruan, Yang Li 0097, Tao Tan 0002, Lianming Wu, Guang Yang 0006, Mingfeng Jiang |
Expert Syst. Appl. | 6 |
| 2025 | Artificial immunofluorescence in a flash: Rapid synthetic imaging from brightfield through residual diffusionabstractImmunofluorescent (IF) imaging is crucial for visualising biomarker expressions, cell morphology and assessing the effects of drug treatments on sub-cellular components. IF imaging needs extra staining process and often requiring cell fixation, therefore it may also introduce artefacts and alter endogenous cell morphology. Some IF stains are expensive or not readily available hence hindering experiments. Recent diffusion models, which synthesise high-fidelity IF images from easy-to-acquire brightfield (BF) images, offer a promising solution but are hindered by training instability and slow inference times due to the noise diffusion process. This paper presents a novel method for the conditional synthesis of IF images directly from BF images along with cell segmentation masks. Our approach employs a Residual Diffusion process that enhances stability and significantly reduces inference time. We performed a critical evaluation against other image-to-image synthesis models, including UNets, GANs, and advanced diffusion models. Our model demonstrates significant improvements in image quality ( p < 0 . 05 in MSE, PSNR, and SSIM), inference speed (26 times faster than competing diffusion models), and accurate segmentation results for both nuclei and cell bodies (0.77 and 0.63 mean IOU for nuclei and cell true positives, respectively). This paper is a substantial advancement in the field, providing robust and efficient tools for cell image analysis. • We introduce a novel diffusion model to synthesise fluorescence images from brightfield images. • CellResDM improves quality, speed, and segmentation accuracy, surpassing existing models. • CellResDM model can simultaneously generates IF images and cell/nuclei segmentation. Xiaodan Xing, Chunling Tang, Siofra Murdoch, Giorgos Papanastasiou, Yunzhe Guo, Xianglu Xiao, Jan Oscar Cross-Zamirski, Carola-Bibiane Schönlieb, Kristina Xiao Liang, Zhangming Niu, Evandro Fei Fang, Yinhai Wang, Guang Yang 0006 |
Neurocomputing | 13 |
| 2025 | Enhancing global sensitivity and uncertainty quantification in medical image reconstruction with Monte Carlo arbitrary-masked mambaabstractDeep learning has been extensively applied in medical image reconstruction, where Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) represent the predominant paradigms, each possessing distinct advantages and inherent limitations: CNNs exhibit linear complexity with local sensitivity, whereas ViTs demonstrate quadratic complexity with global sensitivity. The emerging Mamba has shown superiority in learning visual representation, which combines the advantages of linear scalability and global sensitivity. In this study, we introduce MambaMIR, an Arbitrary-Masked Mamba-based model with wavelet decomposition for joint medical image reconstruction and uncertainty estimation. A novel Arbitrary Scan Masking (ASM) mechanism "masks out" redundant information to introduce randomness for further uncertainty estimation. Compared to the commonly used Monte Carlo (MC) dropout, our proposed MC-ASM provides an uncertainty map without the need for hyperparameter tuning and mitigates the performance drop typically observed when applying dropout to low-level tasks. For further texture preservation and better perceptual quality, we employ the wavelet transformation into MambaMIR and explore its variant based on the Generative Adversarial Network, namely MambaMIR-GAN. Comprehensive experiments have been conducted for multiple representative medical image reconstruction tasks, demonstrating that the proposed MambaMIR and MambaMIR-GAN outperform other baseline and state-of-the-art methods in different reconstruction tasks, where MambaMIR achieves the best reconstruction fidelity and MambaMIR-GAN has the best perceptual quality. In addition, our MC-ASM provides uncertainty maps as an additional tool for clinicians, while mitigating the typical performance drop caused by the commonly used dropout. Liutao Yang, Fanwen Wang, Yinzhe Wu 0001, Yang Nan 0002, Weiwen Wu, Chengyan Wang, Kuangyu Shi, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Daoqiang Zhang, Guang Yang 0006 |
Medical Image Anal. | 12 |
| 2025 | Shadow defense against gradient inversion attack in federated learningabstractFederated learning (FL) has emerged as a transformative framework for privacy-preserving distributed training, allowing clients to collaboratively train a global model without sharing their local data. This is especially crucial in sensitive fields like healthcare, where protecting patient data is paramount. However, privacy leakage remains a critical challenge, as the communication of model updates can be exploited by potential adversaries. Gradient inversion attacks (GIAs), for instance, allow adversaries to approximate the gradients used for training and reconstruct training images, thus stealing patient privacy. Existing defense mechanisms obscure gradients, yet lack a nuanced understanding of which gradients or types of image information are most vulnerable to such attacks. These indiscriminate calibrated perturbations result in either excessive privacy protection degrading model accuracy, or insufficient one failing to safeguard sensitive information. Therefore, we introduce a framework that addresses these challenges by leveraging a shadow model with interpretability for identifying sensitive areas. This enables a more targeted and sample-specific noise injection. Specially, our defensive strategy achieves discrepancies of 3.73 in PSNR and 0.2 in SSIM compared to the circumstance without defense on the ChestXRay dataset, and 2.78 in PSNR and 0.166 in the EyePACS dataset. Moreover, it minimizes adverse effects on model performance, with less than 1% F1 reduction compared to SOTA methods. Our extensive experiments, conducted across diverse types of medical images, validate the generalization of the proposed framework. The stable defense improvements for FedAvg are consistently over 1.5% times in LPIPS and SSIM. It also offers a universal defense against various GIA types, especially for these sensitive areas in images. Liyan Ma, Guang Yang 0006 |
Medical Image Anal. | 3 |
| 2025 | A lung structure and function information-guided residual diffusion model for predicting idiopathic pulmonary fibrosis progression
Caiwen Jiang, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Simon Walsh, Guang Yang 0006, Dinggang Shen |
Medical Image Anal. | 7 |
| 2025 | From challenges and pitfalls to recommendations and opportunities: Implementing federated learning in healthcareabstractFederated learning holds great potential for enabling large-scale healthcare research and collaboration across multiple centers while ensuring data privacy and security are not compromised. Although numerous recent studies suggest or utilize federated learning based methods in healthcare, it remains unclear which ones have potential clinical utility. This review paper considers and analyzes the most recent studies up to May 2024 that describe federated learning based methods in healthcare. After a thorough review, we find that the vast majority are not appropriate for clinical use due to their methodological flaws and/or underlying biases which include but are not limited to privacy concerns, generalization issues, and communication costs. As a result, the effectiveness of federated learning in healthcare is significantly compromised. To overcome these challenges, we provide recommendations and promising opportunities that might be implemented to resolve these problems and improve the quality of model development in federated learning with healthcare. • Evaluate recent FL technologies in healthcare, focusing on challenges and pitfalls. • Offer a taxonomic analysis of FL in healthcare across critical aspects. • Recommend strategies for improving FL and ensuring reproducibility. • Highlight trends and opportunities to enhance FL workflow. Ming Li 0005, Zeyu Tang 0001, Guang Yang 0006 |
Medical Image Anal. | 5 |
| 2025 | The state-of-the-art in cardiac MRI reconstruction: Results of the CMRxRecon challenge in MICCAI 2023
Chen Qin, Shuo Wang 0011, Fanwen Wang, Yan Li 0064, Zi Wang 0005, Kunyuan Guo, Ouyang Cheng, Michael Tänzer, Longyu Sun, Mengting Sun, Zhang Shi, Sha Hua, Hao Li 0082, Zhensen Chen, Bingyu Xin, Dimitris N. Metaxas, George Yiasemis, Jonas Teuwen, Weitian Chen, Yidong Zhao, Yanwei Pang, Artem Razumov, Dmitry V. Dylov, Quan Dou, Yuyang Xue, Yuning Du, Julia Dietlmeier, Carles García-Cabrera, Ziad Al-Haj Hemidi, Nora Vogt, Ying-Hua Chu, Weibo Chen, Wenjia Bai, Xiahai Zhuang, Harry Qin, Lianming Wu, Guang Yang 0006, Xiaobo Qu 0001, He Wang 0016, Chengyan Wang |
Medical Image Anal. | 47 |
| 2025 | Revisiting medical image retrieval via knowledge consolidationabstractAs artificial intelligence and digital medicine increasingly permeate healthcare systems, robust governance frameworks are essential to ensure ethical, secure, and effective implementation. In this context, medical image retrieval becomes a critical component of clinical data management, playing a vital role in decision-making and safeguarding patient information. Existing methods usually learn hash functions using bottleneck features, which fail to produce representative hash codes from blended embeddings. Although contrastive hashing has shown superior performance, current approaches often treat image retrieval as a classification task, using category labels to create positive/negative pairs. Moreover, many methods fail to address the out-of-distribution (OOD) issue when models encounter external OOD queries or adversarial attacks. In this work, we propose a novel method to consolidate knowledge of hierarchical features and optimization functions. We formulate the knowledge consolidation by introducing Depth-aware Representation Fusion (DaRF) and Structure-aware Contrastive Hashing (SCH). DaRF adaptively integrates shallow and deep representations into blended features, and SCH incorporates image fingerprints to enhance the adaptability of positive/negative pairings. These blended features further facilitate OOD detection and content-based recommendation, contributing to a secure AI-driven healthcare environment. Moreover, we present a content-guided ranking to improve the robustness and reproducibility of retrieval results. Our comprehensive assessments demonstrate that the proposed method could effectively recognize OOD samples and significantly outperform existing approaches in medical image retrieval (p < 0 . 05 ). In particular, our method achieves a 5.6–38.9% improvement in mean Average Precision on the anatomical radiology dataset. • Structure-aware pairing using image fingerprints to address over-centralized issues. • A novel model to consolidate hierarchical embeddings for representation learning. • Addressing ill-posed gradient issues introduced by relaxed Hamming distance. • A self-supervised OOD detection module by evaluating image reconstruction disparity. • Content-guided ranking mechanism for robust and precise retrieval. Yang Nan 0002, Huichi Zhou, Xiaodan Xing, Giorgos Papanastasiou, Lei Zhu 0003, Zhifan Gao, Alejandro F. Frangi, Guang Yang 0006 |
Medical Image Anal. | 8 |
| 2025 | One for multiple: Physics-informed synthetic data boosts generalizable deep learning for fast MRI reconstruction
Zi Wang 0005, Xiaotong Yu, Chengyan Wang, Weibo Chen, Ying-Hua Chu, Rushuai Li, Peiyong Li, Haiwei Han, Taishan Kang, Jianzhong Lin, Shufu Chang, Zhang Shi, Sha Hua, Yan Li 0064, Liuhong Zhu, Jianjun Zhou 0004, Meijing Lin, Jiefeng Guo, Congbo Cai, Zhong Chen 0005, Di Guo 0003, Guang Yang 0006, Xiaobo Qu 0001 |
Medical Image Anal. | 27 |
| 2025 | Diff-UNet: A diffusion embedded network for robust 3D medical image segmentation
Zhaohu Xing, Huazhu Fu, Guang Yang 0006, Lequan Yu, Bai Ying Lei, Lei Zhu 0003 |
Medical Image Anal. | 4 |
| 2025 | Representation-driven sampling and adaptive policy resetting for improving multi-Agent reinforcement learning
Weiqiang Jin, Xingwu Tian, Ningwei Wang, Baohai Wu, Bohang Shi, Biao Zhao 0003, Guang Yang 0006 |
Neural Networks | 7 |
| 2025 | Data augmentation strategies for semi-supervised medical image segmentation
Dongsheng Ruan, Yang Li 0097, Yongquan Wu, Tao Tan 0002, Guang Yang 0006, Mingfeng Jiang |
Pattern Recognit. | 7 |
| 2025 | Unpaired translation of chest X-ray images for lung opacity diagnosis via adaptive activation masks and cross-domain alignmentabstractChest X-ray radiographs (CXRs) play a pivotal role in diagnosing and monitoring cardiopulmonary diseases. However, lung opacities in CXRs frequently obscure anatomical structures, impeding clear identification of lung borders and complicating localisation of pathology. This challenge significantly hampers segmentation accuracy and precise lesion identification, crucial for diagnosis. To tackle these issues, our study proposes an unpaired CXR translation framework that converts CXRs with lung opacities into counterparts without lung opacities while preserving semantic features. Central to our approach is the use of adaptive activation masks to selectively modify opacity regions in lung CXRs. Cross-domain alignment ensures translated CXRs without opacity issues align with feature maps and prediction labels from a pre-trained CXR lesion classifier, facilitating the interpretability of the translation process. We validate our method using RSNA, MIMIC-CXR-JPG and JSRT datasets, demonstrating superior translation quality through lower Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) scores compared to existing methods (FID: 67.18 vs. 210.4, KID: 0.01604 vs. 0.225). Evaluation on RSNA opacity, MIMIC acute respiratory distress syndrome (ARDS) patient CXRs and JSRT CXRs shows our method enhances segmentation accuracy of lung borders and improves lesion classification, further underscoring its potential in clinical settings (RSNA: mIoU: 76.58% vs. 62.58%, Sensitivity: 85.58% vs. 77.03%; MIMIC ARDS: mIoU: 86.20% vs. 72.07%, Sensitivity: 92.68% vs. 86.85%; JSRT: mIoU: 91.08% vs. 85.6%, Sensitivity: 97.62% vs. 95.04%). Our approach advances CXR imaging analysis, especially in investigating segmentation impacts through image translation techniques. • Unpaired translation removes lung opacities yet keeps key features in X-rays. • Adaptive masks highlight and constrain opacity changes for better interpretability. • Cross-domain alignment reduces artefacts and preserves real diagnostic features. • Experiments show improved image fidelity, segmentation, and lesion classification. Junzhi Ning, Dominic C. Marshall, Yijian Gao, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Matthieu Komorowski, Guang Yang 0006 |
Pattern Recognit. Lett. | 9 |
| 2025 | Can Rumor Detection Enhance Fact Verification? Unraveling Cross-Task Synergies Between Rumor Detection and Fact VerificationabstractRecently, rumor detection (fake news detection) has seen a surge in research interest, and fact verification (fake news checking) has simultaneously become a significant research aspect. Despite the inherent distinction between fact verification and rumor detection – the former being a three-category task and the latter a binary one – there has yet to be in-depth exploration into the synergies between these two tasks. Furthermore, given the severe scarcity and the time-consuming and costly construction nature of fact verification datasets, few-shot/zero-shot fact verification methods are particularly favored. To tackle these challenges, we conduct a series of studies around “How can rumor detection enhance few-shot fact verification, and to what extent?”. Specifically, we systematically investigate the knowledge transferability between the two tasks, proposing a framework, Det2Ver, that is applicable to both rumor detection and fact verification. Through the construction of adaptive prompt templates and prompt-tuned LLMs like T5, Det2Ver structural-level synchronizes the two tasks and utilizes the external knowledge from rumor detection to reinforce fact verification task. We demonstrate the significance and effectiveness of Det2Ver. Through the few-shot/zero-shot experiments on three widely-used datasets, compared to other LLMs prompt-tuning baselines, the Det2Ver for cross-task knowledge augmentation brings a significant improvement in macro-F1 for fact verification. Weiqiang Jin, Mengying Jiang, Tao Tao 0005, Hao Zhou 0038, Biao Zhao 0003, Guang Yang 0006 |
IEEE Trans. Big Data | 7 |
| 2025 | Editorial Emerging Horizons: The Rise of Large Language Models and Cross-Modal Generative AI
Guang Yang 0006, Jing Zhang 0037, Giorgos Papanastasiou, Ge Wang 0001, Dacheng Tao |
IEEE Trans. Big Data | 1 |
| 2025 | Enhancing Visual Reasoning With LLM-Powered Knowledge Graphs for Visual Question Localized-Answering in Robotic SurgeryabstractExpert surgeons often have heavy workloads and cannot promptly respond to queries from medical students and junior doctors about surgical procedures. Thus, research on Visual Question Localized-Answering in Surgery (Surgical-VQLA) is essential to assist medical students and junior doctors in understanding surgical scenarios. Surgical-VQLA aims to generate accurate answers and locate relevant areas in the surgical scene, requiring models to identify and understand surgical instruments, operative organs, and procedures. A key issue is the model's ability to accurately distinguish surgical instruments. Current Surgical-VQLA models rely primarily on sparse textual information, limiting their visual reasoning capabilities. To address this issue, we propose a framework called Enhancing Visual Reasoning with LLM-Powered Knowledge Graphs (EnVR-LPKG) for the Surgical-VQLA task. This framework enhances the model's understanding of the surgical scenario by utilizing knowledge graphs of surgical instruments constructed by the Large Language Model (LLM). Specifically, we design a Fine-grained Knowledge Extractor (FKE) to extract the most relevant information from knowledge graphs and perform contrastive learning with the extracted knowledge graphs and local image. Furthermore, we design a Multi-attention-based Surgical Instrument Enhancer (MSIE) module, which employs knowledge graphs to obtain an enhanced representation of the corresponding surgical instrument in the global scene. Through the MSIE module, the model can learn how to fuse visual features with knowledge graph text features, thereby strengthening the understanding of surgical instruments and further improving visual reasoning capabilities. Extensive experimental results on the EndoVis-17-VQLA and EndoVis-18-VQLA datasets demonstrate that our proposed method outperforms other state-of-the-art methods. We will release our code for future research. Pengfei Hao, Hongqiu Wang, Guang Yang 0006, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Multi-Level Noise Sampling From Single Image for Low-Dose Tomography ReconstructionabstractLow-dose digital radiography (DR) and computed tomography (CT) become increasingly popular due to reduced radiation dose. However, they often result in degraded images with lower signal-to-noise ratios, creating an urgent need for effective denoising techniques. The recent advancement of the single-image-based denoising approach provides a promising solution without requirement of pairwise training data, which are scarce in medical imaging. These methods typically rely on sampling image pairs from a noisy image for inter-supervised denoising. Although enjoying simplicity, the generated image pairs are at the same noise level and only include partial information about the input images. This study argues that generating image pairs at different noise levels while fully using the information of the input image is preferable since it could provide richer multi-perspective clues to guide the denoising process. To this end, we present a novel Multi-Level Noise Sampling (MNS) method for low-dose tomography denoising. Specifically, MNS method generates multi-level noisy sub-images by partitioning the high-dimensional input space into multiple low-dimensional sub-spaces with a simple yet effective strategy. The superiority of the MNS method in single-image-based denoising over the competing methods has been investigated and verified theoretically. Moreover, to bridge the gap between self-supervised and supervised denoising networks, we introduce an optimization function that leverages prior knowledge of multi-level noisy sub-images to guide the training process. Through extensive quantitative and qualitative experiments conducted on large-scale clinical low-dose CT and DR datasets, we validate the effectiveness and superiority of our MNS approach over other state-of-the-art supervised and self-supervised methods. Weiwen Wu, Yifei Long, Zhifan Gao, Guang Yang 0006, Fangxiao Cheng, Jianjia Zhang |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Feedback Attention to Enhance Unsupervised Deep Learning Image Registration in 3D EchocardiographyabstractCardiac motion estimation is important for assessing the contractile health of the heart, and performing this in 3D can provide advantages due to the complex 3D geometry and motions of the heart. Deep learning image registration (DLIR) is a robust way to achieve cardiac motion estimation in echocardiography, providing speed and precision benefits, but DLIR in 3D echo remains challenging. Successful unsupervised 2D DLIR strategies are often not effective in 3D, and there have been few 3D echo DLIR implementations. Here, we propose a new spatial feedback attention (FBA) module to enhance unsupervised 3D DLIR and enable it. The module uses the results of initial registration to generate a co-attention map that describes remaining registration errors spatially and feeds this back to the DLIR to minimize such errors and improve self-supervision. We show that FBA improves a range of promising 3D DLIR designs, including networks with and without transformer enhancements, and that it can be applied to both fetal and adult 3D echo, suggesting that it can be widely and flexibly applied. We further find that the optimal 3D DLIR configuration is when FBA is combined with a spatial transformer and a DLIR backbone modified with spatial and channel attention, which outperforms existing 3D DLIR approaches. FBA's good performance suggests that spatial attention is a good way to enable scaling up from 2D DLIR to 3D and that a focus on the quality of the image after registration warping is a good way to enhance DLIR performance. Codes and data are available at: https://github.com/kamruleee51/Feedback_DLIR. Md. Kamrul Hasan 0002, Yihao Luo, Guang Yang 0006, Choon Hwai Yap |
IEEE Trans. Medical Imaging | 3 |
| 2025 | AMVLM: Alignment-Multiplicity Aware Vision-Language Model for Semi-Supervised Medical Image SegmentationabstractLow-quality pseudo labels pose a significant obstacle in semi-supervised medical image segmentation (SSMIS), impeding consistency learning on unlabeled data. Leveraging vision-language model (VLM) holds promise in ameliorating pseudo label quality by employing textual prompts to delineate segmentation regions, but it faces the challenge of cross-modal alignment uncertainty due to multiple correspondences (multiple images/texts tend to correspond to one text/image). Existing VLMs address this challenge by modeling semantics as distributions but such distributions lead to semantic degradation. To address these problems, we propose Alignment-Multiplicity Aware Vision-Language Model (AMVLM), a new VLM pretraining paradigm with two novel similarity metric strategies. (i) Cross-modal Similarity Supervision (CSS) proposes a probability distribution transformer to supervise similarity scores across fine-granularity semantics through measuring cross-modal distribution disparities, thus learning cross-modal multiple alignments. (ii) Intra-modal Contrastive Learning (ICL) takes into account the similarity metric of coarse-fine granularity information within each modality to encourage cross-modal semantic consistency. Furthermore, using the pretrained AMVLM, we propose a pioneering text-guided SSMIS network to compensate for the quality deficiencies of pseudo-labels. This network incorporates a text mask generator to produce multimodal supervision information, enhancing pseudo label quality and the model's consistency learning. Extensive experimentation validates the efficacy of our AMVLM-driven SSMIS, showcasing superior performance across four publicly available datasets. The code will be available at: https://github.com/QingtaoPan/AMVLM. Qingtao Pan, Zhengrong Li, Wenhao Qiao, Jingjiao Lou, Guang Yang 0006, Bing Ji 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Serp-Mamba: Advancing High-Resolution Retinal Vessel Segmentation With Selective State-Space ModelabstractUltra-Wide-Field Scanning Laser Ophthalmoscopy (UWF-SLO) images capture high-resolution views of the retina with typically spanning 200 degrees. Accurate segmentation of vessels in UWF-SLO images is essential for detecting and diagnosing fundus disease. Recent studies highlight that Mamba's selective State Space Model (SSM) excels in modeling long-range dependencies with linear computational complexity, making it highly suitable for preserving the continuity of elongated vessel structures, especially for high-resolution UWF images. Inspired by this, we propose the Serpentine Mamba (Serp-Mamba) network to address this challenging task. Specifically, we recognize the intricate, varied, and delicate nature of the tubular structure of vessels. Furthermore, the high-resolution of UWF-SLO images exacerbates the imbalance between the vessel and background categories. Based on the above observations, we first devise a Serpentine Interwoven Adaptive (SIA) scan mechanism, which scans UWF-SLO images along curved vessel structures in a snake-like crawling manner. This approach, consistent with vascular texture transformations, ensures the effective and continuous capture of curved vascular structure features. Second, we propose an Ambiguity-Driven Dual Recalibration (ADDR) module to address the category imbalance problem intensified by high-resolution images. Our ADDR module delineates pixels by two learnable thresholds and refines ambiguous pixels through a dual-driven strategy, thereby accurately distinguishing vessels and background regions. Experiment results on three datasets demonstrate the superior performance of our Serp-Mamba on high-resolution vessel segmentation. We also conduct a series of ablation studies to verify the impact of our designs. Our code will be released upon publication (https://github.com/whq-xxh/Serp-Mamba). Hongqiu Wang, Bin Sheng 0001, Huazhu Fu, Guang Yang 0006, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Enhanced DTCMR With Cascaded Alignment and Adaptive DiffusionabstractDiffusion tensor cardiovascular magnetic resonance (DTCMR) is the only non-invasive method for visualizing myocardial microstructure, but it is challenged by inconsistent breath-holds and imperfect cardiac triggering, causing in-plane shifts and through-plane warping with an inadequate tensor fitting. While rigid registration corrects in-plane shifts, deformable registration risks distorting the diffusion distribution, and selecting a reference frame among low SNR frames is challenging. Existing pairwise deep learning and iterative methods are unsuitable for DTCMR due to their inability to handle the drastic in-plane motion and disentangle the diffusion contrast distortion with through-plane motions on low SNR frames, which compromises the accuracy of clinical biomarker tensor estimation. Our study introduces a novel deep learning framework incorporating tensor information for groupwise deformable registration, effectively correcting intra-subject inter-frame motion. This framework features a cascaded registration branch for addressing in-plane and through-plane motions and a parallel branch for generating pseudo-frames with diffusion contrasts and template updates to guide registration with a refined loss function and denoising. We evaluated our method on four DTCMR-specific metrics using data from over 900 cases from 2012 to 2023. Our method outperformed three traditional and two deep learning-based methods, achieving reduced fitting errors, the lowest percentage of negative eigenvalues at 0.446%, the highest R2 of HA line profiles at 0.911, no negative Jacobian Determinant, and the shortest reference time of 0.06 seconds per case. In conclusion, our deep learning framework significantly improves DTCMR imaging by effectively correcting inter-frame motion and surpassing existing methods across multiple metrics, demonstrating substantial clinical potential. Fanwen Wang, Yihao Luo, Camila Munoz, Yaqing Luo, Yinzhe Wu 0001, Zohya Khalique, Maria Molto, Ramyah Rajakulasingam, Ranil De Silva, Dudley Pennell, Pedro F. Ferreira, Andrew D. Scott, Sonia Nielles-Vallespin, Guang Yang 0006 |
IEEE Trans. Medical Imaging | 16 |
| 2025 | CT-SDM: A Sampling Diffusion Model for Sparse-View CT Reconstruction Across Various Sampling RatesabstractSparse views X-ray computed tomography has emerged as a contemporary technique to mitigate radiation dose. Because of the reduced number of projection views, traditional reconstruction methods can lead to severe artifacts. Recently, research studies utilizing deep learning methods has made promising progress in removing artifacts for Sparse-View Computed Tomography (SVCT). However, given the limitations on the generalization capability of deep learning models, current methods usually train models on fixed sampling rates, affecting the usability and flexibility of model deployment in real clinical settings. To address this issue, our study proposes a adaptive reconstruction method to achieve high-performance SVCT reconstruction at various sampling rate. Specifically, we design a novel imaging degradation operator in the proposed sampling diffusion model for SVCT (CT-SDM) to simulate the projection process in the sinogram domain. Thus, the CT-SDM can gradually add projection views to highly undersampled measurements to generalize the full-view sinograms. By choosing an appropriate starting point in diffusion inference, the proposed model can recover the full-view sinograms from various sampling rate with only one trained model. Experiments on several datasets have verified the effectiveness and robustness of our approach, demonstrating its superiority in reconstructing high-quality images from sparse-view CT scans across various sampling rates. Liutao Yang, Guang Yang 0006, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Beyond the Hype: A Dispassionate Look at Vision-Language Models in Medical ScenarioabstractRecent advancements in large vision-language models (LVLMs) have demonstrated remarkable capabilities across diverse tasks, garnering significant attention in AI communities. However, their performance and reliability in specialized domains such as medicine remain insufficiently assessed. In particular, most assessments overconcentrate on evaluating VLMs based on simple visual question answering (VQA) on multimodality data while ignoring the in-depth characteristics of LVLMs. In this study, we introduce RadVUQA, a novel radiological visual understanding and question answering benchmark, to comprehensively evaluate existing LVLMs. RadVUQA mainly validates LVLMs across five dimensions: 1) anatomical understanding, assessing the models' ability to visually identify biological structures; 2) multimodal comprehension, which involves the capability of interpreting linguistic and visual instructions to produce desired outcomes; 3) quantitative and spatial reasoning, evaluating the models' spatial awareness and proficiency in combining quantitative analysis with visual and linguistic information; 4) physiological knowledge, measuring the models' capability to comprehend functions and mechanisms of organs and systems; and 5) robustness, which assesses the models' capabilities against unharmonized and synthetic data. The results indicate that both generalized LVLMs and medical-specific LVLMs have critical deficiencies with weak multimodal comprehension and quantitative reasoning capabilities. Our findings reveal the large gap between existing LVLMs and clinicians, highlighting the urgent need for more robust and intelligent LVLMs. The code is available at https://github.com/Nandayang/RadVUQA. Yang Nan 0002, Huichi Zhou, Xiaodan Xing, Guang Yang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | SRViT: Self-Supervised Relation-Aware Vision Transformer for Hyperspectral UnmixingabstractVision transformer (ViT) has recently been a popular topic in the foundation model field, taking advantage of its strong scalability and outstanding representation capabilities. As a deep model, ViT introduces a new architecture for achieving hyperspectral image (HSI) unmixing. However, traditional ViTs overlook pixel-level spatial continuity by partitioning the input image into nonoverlapping fixed-size patches. This approach disrupts local structural relationships and hinders the model's ability to capture fine-grained spatial dependencies, resulting in suboptimal feature representation for dense prediction tasks in unmixing. To address these challenges, this article proposes the development of a self-supervised relation-aware ViT (SRViT). SRViT incorporates a self-embedded module comprising encoders, a pixel-level position encoder (PLPE), a self-supervised contrastive mechanism (SCM), and a decoder. The self-embedded module and PLPE preserve local correlations in HSI across different views, facilitating cross-view learning through SCM to ensure generalization. In addition, the decoder incorporates Kronecker-factored approximate curvature (K-FAC) to capture the local geometric structure of spectral information. Ultimately, SRViT learns endmembers and fractional abundance as the unmixing result. The effectiveness and competitiveness of SRViT have been systematically validated through comparative experiments, demonstrating its superior performance. The source code is available at the following link: https://github.com/yuanchaosu/TNNLS-SRViT. Yuanchao Su, Lianru Gao, Antonio Plaza, Xu Sun 0005, Mengying Jiang, Guang Yang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | HAMLET: Graph Transformer Neural Operator for Partial Differential EquationsabstractWe present a novel graph transformer framework, HAMLET, designed to address the challenges in solving partial differential equations (PDEs) using neural networks. The framework uses graph transformers with modular input encoders to directly incorporate differential equation information into the solution process. This modularity enhances parameter correspondence control, making HAMLET adaptable to PDEs of arbitrary geometries and varied input formats. Notably, HAMLET scales effectively with increasing data complexity and noise, showcasing its robustness. HAMLET is not just tailored to a single type of physical simulation, but can be applied across various domains. Moreover, it boosts model resilience and performance, especially in scenarios with limited data. We demonstrate, through extensive experiments, that our framework is capable of outperforming current techniques for PDEs. Andrey Bryutkin, Zhongying Deng, Guang Yang 0006, Carola-Bibiane Schönlieb, Angelica I. Avilés-Rivero |
ICML | 4 |
| 2024 | DiffExplainer: Unveiling Black Box Models Via Counterfactual Generation
Yingying Fang, Shuang Wu 0002, Zihao Jin, Caiwen Xu, Simon Walsh, Guang Yang 0006 |
MICCAI (10) | 7 |
| 2024 | Diff3Dformer: Leveraging Slice Sequence Diffusion for Enhanced 3D CT Classification with Transformer Networks
Zihao Jin, Yingying Fang, Caiwen Xu, Simon Walsh, Guang Yang 0006 |
MICCAI (1) | 6 |
| 2024 | Groupwise Deformable Registration of Diffusion Tensor Cardiovascular Magnetic Resonance: Disentangling Diffusion Contrast, Respiratory and Cardiac Motions
Fanwen Wang, Yihao Luo, Pedro F. Ferreira, Yaqing Luo, Yinzhe Wu 0001, Camila Munoz, Dudley Pennell, Andrew D. Scott, Sonia Nielles-Vallespin, Guang Yang 0006 |
MICCAI (2) | 12 |
| 2024 | LGRNet: Local-Global Reciprocal Network for Uterine Fibroid Segmentation in Ultrasound Videos
Angelica I. Avilés-Rivero, Guang Yang 0006, Harry Qin, Lei Zhu 0003 |
MICCAI (4) | 4 |
| 2024 | Variational Field Constraint Learning for Degree of Coronary Artery Ischemia Assessment
Qi Zhang 0078, Xiujian Liu, Heye Zhang, Chenchu Xu, Guang Yang 0006, Yixuan Yuan, Tao Tan 0002, Zhifan Gao |
MICCAI (3) | 5 |
| 2024 | Fuzzy Attention-Based Border Rendering Network for Lung Organ Segmentation
Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Xiaodan Xing, Zhifan Gao, Guang Yang 0006 |
MICCAI (9) | 7 |
| 2024 | Dynamic Multimodal Information Bottleneck for Multimodality ClassificationabstractEffectively leveraging multimodal data such as various images, laboratory tests and clinical information is becoming increasingly attractive in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on enhancing their performance by leveraging the differences or shared features from various modalities and fusing feature across different modalities. These approaches are generally not optimal for clinical settings, which pose the additional challenges of limited training data, as well as being rife with redundant data or noisy modality channels, leading to subpar performance. To address this gap, we study the robustness of existing methods to data redundancy and noise and propose a generalized dynamic multimodal information bottleneck framework for attaining a robust fused feature representation. Specifically, our information bottleneck module serves to filter out the task-irrelevant information and noises in the fused feature, and we further introduce a sufficiency loss to prevent dropping of task-relevant information, thus explicitly preserving the sufficiency of prediction information in the distilled feature. We validate our model on an in-house and a public COVID19 dataset for mortality prediction as well as two public biomedical datasets for diagnostic tasks. Extensive experiments show that our method surpasses the state-of-the-art and is significantly more robust, being the only method to remain performance when large-scale noisy channels exist. Our code is publicly available at https://github.com/ayanglab/DMIB. Yingying Fang, Shuang Wu 0002, Sheng Zhang 0024, Chaoyan Huang, Tieyong Zeng, Xiaodan Xing, Simon Walsh, Guang Yang 0006 |
WACV | 8 |
| 2024 | DisCo-FEND: Social Context Veracity Dissemination Consistency-Guided Case Reasoning for Few-Shot Fake News Detection
Weiqiang Jin, Ningwei Wang, Tao Tao 0005, Mengying Jiang, Biao Zhao 0003, Haibin Duan, Guang Yang 0006 |
WISE (5) | 9 |
| 2024 | Probing perfection: The relentless art of meddling for pulmonary airway segmentation from HRCT via a human-AI collaboration based active learning methodabstractIn the realm of pulmonary tracheal segmentation, the scarcity of annotated data stands as a prevalent pain point in most medical segmentation endeavors. Concurrently, most Deep Learning (DL) methodologies employed in this domain invariably grapple with other dual challenges: the inherent opacity of 'black box' models and the ongoing pursuit of performance enhancement. In response to these intertwined challenges, the core concept of our Human-Computer Interaction (HCI) based learning models (RS_UNet, LC_UNet, UUNet and WD_UNet) hinge on the versatile combination of diverse query strategies and an array of deep learning models. We train four HCI models based on the initial training dataset and sequentially repeat the following steps 1-4: (1) Query Strategy: Our proposed HCI models selects those samples which contribute the most additional representative information when labeled in each iteration of the query strategy (showing the names and sequence numbers of the samples to be annotated). Additionally, in this phase, the model selects the unlabeled samples with the greatest predictive disparity by calculating the Wasserstein Distance, Least Confidence, Entropy Sampling, and Random Sampling. (2) Central line correction: The selected samples in previous stage are then used for domain expert correction of the system-generated tracheal central lines in each training round. (3) Update training dataset: When domain experts are involved in each epoch of the DL model's training iterations, they update the training dataset with greater precision after each epoch, thereby enhancing the trustworthiness of the 'black box' DL model and improving the performance of models. (4) Model training: Proposed HCI model is trained using the updated training dataset and an enhanced version of existing UNet. Experimental results validate the effectiveness of this Human-Computer Interaction-based approaches, demonstrating that our proposed WD-UNet, LC-UNet, UUNet, RS-UNet achieve comparable or even superior performance than the state-of-the-art DL models, such as WD-UNet with only 15 %-35 % of the training data, leading to substantial reductions (65 %-85 % reduction of annotation effort) in physician annotation time. Yang Nan 0002, Sheng Zhang 0024, Federico Felder, Xiaodan Xing, Yingying Fang, Javier Del Ser, Simon Walsh, Guang Yang 0006 |
Artif. Intell. Medicine | 9 |
| 2024 | MLC: Multi-level consistency learning for semi-supervised left atrium segmentation
Zhebin Shi, Mingfeng Jiang, Yang Li 0097, Bo Wei 0004, Yongquan Wu, Tao Tan 0002, Guang Yang 0006 |
Expert Syst. Appl. | 8 |
| 2024 | Guest Editorial: Special Issue on the British Machine Vision Conference 2022
Guang Yang 0006, Angelica I. Avilés-Rivero, Yingying Fang, Zhenhua Feng 0001, Gianluigi Ciocca, Yulia Hicks, Constantino Carlos Reyes-Aldasoro |
Int. J. Comput. Vis. | 1 |
| 2024 | RCAR-UNet: Retinal vessel segmentation network algorithm via novel rough attention mechanism
Weiping Ding 0001, Jiashuang Huang, Hengrong Ju, Chongsheng Zhang, Guang Yang 0006, Chin-Teng Lin |
Inf. Sci. | 6 |
| 2024 | Adaptive dynamic inference for few-shot left atrium segmentation
Jun Chen 0030, Heye Zhang, Yongwon Cho, Sung Ho Hwang, Zhifan Gao, Guang Yang 0006 |
Medical Image Anal. | 7 |
| 2024 | Deep learning based synthesis of MRI, CT and PET: Review and analysisabstractMedical image synthesis represents a critical area of research in clinical decision-making, aiming to overcome the challenges associated with acquiring multiple image modalities for an accurate clinical workflow. This approach proves beneficial in estimating an image of a desired modality from a given source modality among the most common medical imaging contrasts, such as Computed Tomography (CT), Magnetic Resonance Imaging (MRI), and Positron Emission Tomography (PET). However, translating between two image modalities presents difficulties due to the complex and non-linear domain mappings. Deep learning-based generative modelling has exhibited superior performance in synthetic image contrast applications compared to conventional image synthesis methods. This survey comprehensively reviews deep learning-based medical imaging translation from 2018 to 2023 on pseudo-CT, synthetic MR, and synthetic PET. We provide an overview of synthetic contrasts in medical imaging and the most frequently employed deep learning networks for medical image synthesis. Additionally, we conduct a detailed analysis of each synthesis method, focusing on their diverse model designs based on input domains and network architectures. We also analyse novel network architectures, ranging from conventional CNNs to the recent Transformer and Diffusion models. This analysis includes comparing loss functions, available datasets and anatomical regions, and image quality assessments and performance in other downstream tasks. Finally, we discuss the challenges and identify solutions within the literature, suggesting possible future directions. We hope that the insights offered in this survey paper will serve as a valuable roadmap for researchers in the field of medical image synthesis. Sanuwani Dayarathna, Kh Tohidul Islam, Sergio Uribe, Guang Yang 0006, Munawar Hayat, Zhaolin Chen |
Medical Image Anal. | 4 |
| 2024 | Labelling with dynamics: A data-efficient learning paradigm for medical image segmentationabstractThe success of deep learning on image classification and recognition tasks has led to new applications in diverse contexts, including the field of medical imaging. However, two properties of deep neural networks (DNNs) may limit their future use in medical applications. The first is that DNNs require a large amount of labeled training data, and the second is that the deep learning-based models lack interpretability. In this paper, we propose and investigate a data-efficient framework for the task of general medical image segmentation. We address the two aforementioned challenges by introducing domain knowledge in the form of a strong prior into a deep learning framework. This prior is expressed by a customized dynamical system. We performed experiments on two different datasets, namely JSRT and ISIC2016 (heart and lungs segmentation on chest X-ray images and skin lesion segmentation on dermoscopy images). We have achieved competitive results using the same amount of training data compared to the state-of-the-art methods. More importantly, we demonstrate that our framework is extremely data-efficient, and it can achieve reliable results using extremely limited training data. Furthermore, the proposed method is rotationally invariant and insensitive to initialization. Yuanhan Mo, Fangde Liu, Guang Yang 0006, Shuo Wang 0011, Jian-Qing Zheng, Fuping Wu, Bartlomiej Wladyslaw Papiez, Douglas McIlwraith, Taigang He, Yike Guo |
Medical Image Anal. | 3 |
| 2024 | Hunting imaging biomarkers in pulmonary fibrosis: Benchmarks of the AIIB23 challengeabstract• This paper investigates the capacity of AI models for airway modelling on national datasets with paired clinical metadata. • We evaluated AI models against unharmonised, noisy, and out-of-distribution data, as well as the prognostication for FLD. • We found a new biomarker for mortality prediction, outperforming existing clinical measurements (FVC% and fibrosis scores). • In-depth analysis of AI models on airway modelling and prognosis, highlighting challenges and future research directions. Airway-related quantitative imaging biomarkers are crucial for examination, diagnosis, and prognosis in pulmonary diseases. However, the manual delineation of airway structures remains prohibitively time-consuming. While significant efforts have been made towards enhancing automatic airway modelling, current public-available datasets predominantly concentrate on lung diseases with moderate morphological variations. The intricate honeycombing patterns present in the lung tissues of fibrotic lung disease patients exacerbate the challenges, often leading to various prediction errors. To address this issue, the 'Airway-Informed Quantitative CT Imaging Biomarker for Fibrotic Lung Disease 2023′ (AIIB23) competition was organized in conjunction with the official 2023 International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI). The airway structures were meticulously annotated by three experienced radiologists. Competitors were encouraged to develop automatic airway segmentation models with high robustness and generalization abilities, followed by exploring the most correlated QIB of mortality prediction. A training set of 120 high-resolution computerised tomography (HRCT) scans were publicly released with expert annotations and mortality status. The online validation set incorporated 52 HRCT scans from patients with fibrotic lung disease and the offline test set included 140 cases from fibrosis and COVID-19 patients. The results have shown that the capacity of extracting airway trees from patients with fibrotic lung disease could be enhanced by introducing voxel-wise weighted general union loss and continuity loss. In addition to the competitive image biomarkers for mortality prediction, a strong airway-derived biomarker (Hazard ratio>1.5, p < 0.0001) was revealed for survival prognostication compared with existing clinical measurements, clinician assessment and AI-based biomarkers. Yang Nan 0002, Xiaodan Xing, Zeyu Tang 0001, Federico Felder, Sheng Zhang 0024, Roberta Eufrasia Ledda, Xiaoliu Ding, Feng Shi 0001, Tianyang Sun, Zehong Cao, Yun Gu, Pingyu Wang, Wen Tang 0005, Pengxin Yu, Han Kang, Junqiang Chen, Michail Mamalakis, Francesco Prinzi, Gianluca Carlini, Lisa Cuneo, Abhirup Banerjee, Zhaohu Xing, Lei Zhu 0003, Zacharia Mesbah, Dhruv Jain, Tsiry Mayet, Hongyu Yuan, Qing Lyu 0009, Abdul Qayyum 0002, Moona Mazher, Athol Wells, Simon Walsh, Guang Yang 0006 |
Medical Image Anal. | 41 |
| 2024 | Fuzzy Attention-Based Border Rendering Orthogonal Network for Lung Organ SegmentationabstractAutomatic lung organ segmentation on computerized tomography images is crucial for lung disease diagnosis. However, the unlimited voxel values and class imbalance of lung organs can lead to false-negative/positive and leakage issues in numerous state-of-the-art methods. In addition, some lung organs are easily lost during therecycleddown/up-sample procedure, e.g., bronchioles and arterioles, which can cause severe discontinuity issue. Inspired by these, this article introduces an effective lung organ segmentation method called fuzzy attention-based border rendering feature orthogonal network, which 1) integrates an efficient transformer-like fuzzy-attention module into deep networks to cope with the uncertainty in feature representations; 2) decouples and depicts the lung organ regions as cube-trees by focusing only onrecycle-sampling border vulnerable points, rendering the severely discontinuous, false-negative/positive organ regions with two novel global-local cube-tree fusion and sparse patched feature orthogonal modules; 3) develops a multiscale self-knowledge guidance module to improve model performance and robustness. We have demonstrated the efficacy of proposed method on five challenging datasets of lung organ segmentation, i.e., airway and artery. All experimental results demonstrate that our method can achieve the favorable performance significantly. Sheng Zhang 0024, Yingying Fang, Yang Nan 0002, Weiping Ding 0001, Yew-Soon Ong, Alejandro F. Frangi, Witold Pedrycz, Simon Walsh, Guang Yang 0006 |
IEEE Trans. Fuzzy Syst. | 10 |
| 2024 | Is Attention all You Need in Medical Image Analysis? A ReviewabstractMedical imaging is a key component in clinical diagnosis, treatment planning and clinical trial design, accounting for almost 90% of all healthcare data. CNNs achieved performance gains in medical image analysis (MIA) over the last years. CNNs can efficiently model local pixel interactions and be trained on small-scale MI data. Despite their important advances, typical CNN have relatively limited capabilities in modelling "global" pixel interactions, which restricts their generalisation ability to understand out-of-distribution data with different "global" information. The recent progress of Artificial Intelligence gave rise to Transformers, which can learn global relationships from data. However, full Transformer models need to be trained on large-scale data and involve tremendous computational complexity. Attention and Transformer compartments ("Transf/Attention") which can well maintain properties for modelling global relationships, have been proposed as lighter alternatives of full Transformers. Recently, there is an increasing trend to co-pollinate complementary local-global properties from CNN and Transf/Attention architectures, which led to a new era of hybrid models. The past years have witnessed substantial growth in hybrid CNN-Transf/Attention models across diverse MIA problems. In this systematic review, we survey existing hybrid CNN-Transf/Attention models, review and unravel key architectural designs, analyse breakthroughs, and evaluate current and future opportunities as well as challenges. We also introduced an analysis framework on generalisation opportunities of scientific and clinical impact, based on which new data-driven domain generalisation and adaptation methods can be stimulated. Giorgos Papanastasiou, Nikolaos Dikaios, Chengjia Wang, Guang Yang 0006 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Real-Time Non-Invasive Imaging and Detection of Spreading Depolarizations through EEG: An Ultra-Light Explainable Deep Learning ApproachabstractA core aim of neurocritical care is to prevent secondary brain injury. Spreading depolarizations (SDs) have been identified as an important independent cause of secondary brain injury. SDs are usually detected using invasive electrocorticography recorded at high sampling frequency. Recent pilot studies suggest a possible utility of scalp electrodes generated electroencephalogram (EEG) for non-invasive SD detection. However, noise and attenuation of EEG signals makes this detection task extremely challenging. Previous methods focus on detecting temporal power change of EEG over a fixed high-density map of scalp electrodes, which is not always clinically feasible. Having a specialized spectrogram as an input to the automatic SD detection model, this study is the first to transform SD identification problem from a detection task on a 1-D time-series wave to a task on a sequential 2-D rendered imaging. This study presented a novel ultra-light-weight multi-modal deep-learning network to fuse EEG spectrogram imaging and temporal power vectors to enhance SD identification accuracy over each single electrode, allowing flexible EEG map and paving the way for SD detection on ultra-low-density EEG with variable electrode positioning. Our proposed model has an ultra-fast processing speed (<0.3 sec). Compared to the conventional methods (2 hours), this is a huge advancement towards early SD detection and to facilitate instant brain injury prognosis. Seeing SDs with a new dimension - frequency on spectrograms, we demonstrated that such additional dimension could improve SD detection accuracy, providing preliminary evidence to support the hypothesis that SDs may show implicit features over the frequency profile. Yinzhe Wu 0001, Sharon Jewell, Xiaodan Xing, Yang Nan 0002, Anthony J. Strong, Guang Yang 0006, Martyn G. Boutelle |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Dual-Domain Collaborative Diffusion Sampling for Multi-Source Stationary Computed Tomography ReconstructionabstractThe multi-source stationary CT, where both the detector and X-ray source are fixed, represents a novel imaging system with high temporal resolution that has garnered significant interest. Limited space within the system restricts the number of X-ray sources, leading to sparse-view CT imaging challenges. Recent diffusion models for reconstructing sparse-view CT have generally focused separately on sinogram or image domains. Sinogram-centric models effectively estimate missing projections but may introduce artifacts, lacking mechanisms to ensure image correctness. Conversely, image-domain models, while capturing detailed image features, often struggle with complex data distribution, leading to inaccuracies in projections. Addressing these issues, the Dual-domain Collaborative Diffusion Sampling (DCDS) model integrates sinogram and image domain diffusion processes for enhanced sparse-view reconstruction. This model combines the strengths of both domains in an optimized mathematical framework. A collaborative diffusion mechanism underpins this model, improving sinogram recovery and image generative capabilities. This mechanism facilitates feedback-driven image generation from the sinogram domain and uses image domain results to complete missing projections. Optimization of the DCDS model is further achieved through the alternative direction iteration method, focusing on data consistency updates. Extensive testing, including numerical simulations, real phantoms, and clinical cardiac datasets, demonstrates the DCDS model's effectiveness. It consistently outperforms various state-of-the-art benchmarks, delivering exceptional reconstruction quality and precise sinogram. Zirong Li, Dingyue Chang, Fulin Luo, Qiegen Liu, Jianjia Zhang, Guang Yang 0006, Weiwen Wu |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Video-Instrument Synergistic Network for Referring Video Instrument Segmentation in Robotic SurgeryabstractSurgical instrument segmentation is fundamentally important for facilitating cognitive intelligence in robot-assisted surgery. Although existing methods have achieved accurate instrument segmentation results, they simultaneously generate segmentation masks of all instruments, which lack the capability to specify a target object and allow an interactive experience. This paper focuses on a novel and essential task in robotic surgery, i.e., Referring Surgical Video Instrument Segmentation (RSVIS), which aims to automatically identify and segment the target surgical instruments from each video frame, referred by a given language expression. This interactive feature offers enhanced user engagement and customized experiences, greatly benefiting the development of the next generation of surgical education systems. To achieve this, this paper constructs two surgery video datasets to promote the RSVIS research. Then, we devise a novel Video-Instrument Synergistic Network (VIS-Net) to learn both video-level and instrument-level knowledge to boost performance, while previous work only utilized video-level information. Meanwhile, we design a Graph-based Relation-aware Module (GRM) to model the correlation between multi-modal information (i.e., textual description and video frame) to facilitate the extraction of instrument-level information. Extensive experimental results on two RSVIS datasets exhibit that the VIS-Net can significantly outperform existing state-of-the-art referring segmentation methods. We will release our code and dataset for future research (https://github.com/whq-xxh/RSVIS). Hongqiu Wang, Guang Yang 0006, Harry Qin, Yike Guo, Yueming Jin, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Constraint-Aware Learning for Fractional Flow Reserve Pullback Curve Estimation From Invasive Coronary ImagingabstractEstimation of the fractional flow reserve (FFR) pullback curve from invasive coronary imaging is important for the intraoperative guidance of coronary intervention. Machine/deep learning has been proven effective in FFR pullback curve estimation. However, the existing methods suffer from inadequate incorporation of intrinsic geometry associations and physics knowledge. In this paper, we propose a constraint-aware learning framework to improve the estimation of the FFR pullback curve from invasive coronary imaging. It incorporates both geometrical and physical constraints to approximate the relationships between the geometric structure and FFR values along the coronary artery centerline. Our method also leverages the power of synthetic data in model training to reduce the collection costs of clinical data. Moreover, to bridge the domain gap between synthetic and real data distributions when testing on real-world imaging data, we also employ a diffusion-driven test-time data adaptation method that preserves the knowledge learned in synthetic data. Specifically, this method learns a diffusion model of the synthetic data distribution and then projects real data to the synthetic data distribution at test time. Extensive experimental studies on a synthetic dataset and a real-world dataset of 382 patients covering three imaging modalities have shown the better performance of our method for FFR estimation of stenotic coronary arteries, compared with other machine/deep learning-based FFR estimation models and computational fluid dynamics-based model. The results also provide high agreement and correlation between the FFR predictions of our method and the invasively measured FFR values. The plausibility of FFR predictions along the coronary artery centerline is also validated. Dong Zhang 0012, Xiujian Liu, Anbang Wang, Guang Yang 0006, Heye Zhang, Zhifan Gao |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Fuzzy Attention Neural Network to Tackle Discontinuity in Airway SegmentationabstractAirway segmentation is crucial for the examination, diagnosis, and prognosis of lung diseases, while its manual delineation is unduly burdensome. To alleviate this time-consuming and potentially subjective manual procedure, researchers have proposed methods to automatically segment airways from computerized tomography (CT) images. However, some small-sized airway branches (e.g., bronchus and terminal bronchioles) significantly aggravate the difficulty of automatic segmentation by machine learning models. In particular, the variance of voxel values and the severe data imbalance in airway branches make the computational module prone to discontinuous and false-negative predictions, especially for cohorts with different lung diseases. The attention mechanism has shown the capacity to segment complex structures, while fuzzy logic can reduce the uncertainty in feature representations. Therefore, the integration of deep attention networks and fuzzy theory, given by the fuzzy attention layer, should be an escalated solution for better generalization and robustness. This article presents an efficient method for airway segmentation, comprising a novel fuzzy attention neural network (FANN) and a comprehensive loss function to enhance the spatial continuity of airway segmentation. The deep fuzzy set is formulated by a set of voxels in the feature map and a learnable Gaussian membership function. Different from the existing attention mechanism, the proposed channel-specific fuzzy attention addresses the issue of heterogeneous features in different channels. Furthermore, a novel evaluation metric is proposed to assess both the continuity and completeness of airway structures. The efficiency, generalization, and robustness of the proposed method have been proved by training on normal lung disease while testing on datasets of lung cancer, COVID-19, and pulmonary fibrosis. Yang Nan 0002, Javier Del Ser, Zeyu Tang 0001, Peng Tang 0004, Xiaodan Xing, Yingying Fang, Francisco Herrera, Witold Pedrycz, Simon Walsh, Guang Yang 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2023 | The Beauty or the Beast: Which Aspect of Synthetic Medical Images Deserves Our Focus?abstractTraining medical AI algorithms requires large volumes of accurately labeled datasets, which are difficult to obtain in the real world. Synthetic images generated from deep generative models can help alleviate the data scarcity problem, but their effectiveness relies on their fidelity to real-world images. Typically, researchers select synthesis models based on image quality measurements, prioritizing synthetic images that appear realistic. However, our empirical analysis shows that high-fidelity and visually appealing synthetic images are not necessarily superior. In fact, we present a case where low-fidelity synthetic images outperformed their high-fidelity counterparts in downstream tasks. Our findings highlight the importance of comprehensive analysis before incorporating synthetic data into real-world applications. We hope our results will raise awareness among the research community of the value of low-fidelity synthetic images in medical AI algorithm training. Xiaodan Xing, Yang Nan 0002, Federico Felder, Simon Walsh, Guang Yang 0006 |
CBMS | 5 |
| 2023 | CDiffMR: Can We Replace the Gaussian Noise with K-Space Undersampling for Fast MRI?
Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Guang Yang 0006 |
MICCAI (10) | 4 |
| 2023 | Conditional Physics-Informed Graph Neural Network for Fractional Flow Reserve Assessment
Baihong Xie, Xiujian Liu, Heye Zhang, Chenchu Xu, Tieyong Zeng, Yixuan Yuan, Guang Yang 0006, Zhifan Gao |
MICCAI (7) | 7 |
| 2023 | You Don't Have to Be Perfect to Be Amazing: Unveil the Utility of Synthetic Images
Xiaodan Xing, Federico Felder, Yang Nan 0002, Giorgos Papanastasiou, Simon Walsh, Guang Yang 0006 |
MICCAI (5) | 6 |
| 2023 | CellFusion: Multipath Vehicle-to-Cloud Video Streaming with Network Coding in the WildabstractThis paper presents CellFusion, a system designed for high-quality, real-time video streaming from vehicles to the cloud. It leverages an innovative blend of multipath QUIC transport and network coding. Surpassing the limitations of individual cellular carriers, CellFusion uses a unique last-mile overlay that integrates multiple cellular networks into a single, unified cloud connection. This integration is made possible through the use of in-vehicle Customer Premises Equipment (CPEs) and edge-cloud proxy servers. Yunzhe Ni, Zhilong Zheng, Xianshang Lin, Fengyu Gao, Xuan Zeng 0002, Yirui Liu 0001, Senlang Du, Guang Yang 0006, Yuanchao Su, Dennis Cai, Hongqiang Harry Liu, Chenren Xu, Ennan Zhai |
SIGCOMM | 11 |
| 2023 | Mutually aided uncertainty incorporated dual consistency regularization with pseudo label for semi-supervised medical image segmentation
Shanfu Lu, Zijian Zhang 0004, Ziye Yan, Rongrong Zhou, Guang Yang 0006 |
Neurocomputing | 7 |
| 2023 | ChatAgri: Exploring potentials of ChatGPT on cross-linguistic agricultural text classificationabstractIn the era of sustainable smart agriculture, a vast amount of agricultural news text is posted online, accumulating significant agricultural knowledge. To efficiently access this knowledge, effective text classification techniques are urgently needed. Deep learning approaches, such as fine-tuning strategies on pre-trained language models (PLMs), have shown remarkable performance gains. Nonetheless, these methods face several complex challenges, including limited agricultural training data, poor domain transferability (especially across languages), and complex and expensive deployment of large models. Inspired by the success of recent ChatGPT models (e.g., GPT-3.5, GPT-4), this work explores the potential of applying ChatGPT in the field of agricultural informatization. Various crucial factors, such as prompt construction, answer parsing, and different ChatGPT variants, are thoroughly investigated to maximize its capabilities. A preliminary comparative study is conducted, comparing ChatGPT with PLMs-based fine-tuning methods and PLMs-based prompt-tuning methods. Empirical results demonstrate that ChatGPT effectively addresses the mentioned research challenges and bottlenecks, making it an ideal solution for agricultural text classification. Moreover, ChatGPT achieves comparable performance to existing PLM-based fine-tuning methods, even without fine-tuning on agricultural data samples. We hope this preliminary study could inspire the emergence of a general-purpose AI paradigm for agricultural text processing. Biao Zhao 0003, Weiqiang Jin, Javier Del Ser, Guang Yang 0006 |
Neurocomputing | 4 |
| 2023 | Prompt learning for metonymy resolution: Enhancing performance with internal prior knowledge of pre-trained language modelsabstractLinguistic metonymy is a common type of figurative language in natural language processing (NLP), where a concept is represented by a closely associated word or phrase, for example “business executives suits”. As a result, metonymy resolution has become an important NLP task aimed at correctly identifying metonymic expressions within sentences. Previous approaches to this task have typically relied on pre-trained language models (PLMs) using a fine-tuning process. However, this can be time-consuming and resource-intensive, and may lead to a loss of factual prior knowledge. The emergence of a novel learning paradigm termed “prompt learning” or “prompt-tuning” has recently sparked widespread interest and captured considerable attention, as it has proven to yield remarkable results and surpass previous benchmarks. This approach uses a “pre-train→prompt→predict” paradigm and has been shown to better utilize the internal prior knowledge of a PLM, especially in situations with limited supervised resources. Inspired by this success, we investigated how prompt learning could improve metonymy resolution. We have developed a series of prompt learning approaches, called PromptMR, for metonymy resolution, and applied them to several widely-used metonymy resolution datasets. We also designed additional prompt-tuning augmentation strategies to further enhance the potential of prompt learning. Our experiments demonstrated that our method achieved state-of-the-art performance over multiple competitive baselines in both data-sufficient and data-scarce scenarios. The code implementations for PromptMR are accessible on GitHub via the URL: https://github.com/albert-jin/PromptTuning2MetonymyResolution. Biao Zhao 0003, Weiqiang Jin, Yu Zhang 0205, Subin Huang, Guang Yang 0006 |
Knowl. Based Syst. | 5 |
| 2023 | Multi-site, Multi-domain Airway Tree Modeling
Yangqian Wu, Yulei Qin, Hao Zheng 0008, Wen Tang 0005, Corey W. Arnold, Chenhao Pei, Pengxin Yu, Yang Nan 0002, Guang Yang 0006, Simon Walsh, Dominic C. Marshall, Matthieu Komorowski, Puyang Wang, Dazhou Guo, Dakai Jin, Shuiqing Zhao, Runsheng Chang, Abdul Qayyum 0002, Moona Mazher, Yonghuang Wu, Ying'ao Liu, Jiancheng Yang, Ashkan Pakzad, Bojidar Rangelov, Raúl San José Estépar, Carlos Cano-Espinosa, Jiayuan Sun, Guang-Zhong Yang, Yun Gu |
Medical Image Anal. | 11 |
| 2023 | Region-based evidential deep learning to quantify uncertainty and improve robustness of brain tumor segmentationabstractDespite recent advances in the accuracy of brain tumor segmentation, the results still suffer from low reliability and robustness. Uncertainty estimation is an efficient solution to this problem, as it provides a measure of confidence in the segmentation results. The current uncertainty estimation methods based on quantile regression, Bayesian neural network, ensemble, and Monte Carlo dropout are limited by their high computational cost and inconsistency. In order to overcome these challenges, Evidential Deep Learning (EDL) was developed in recent work but primarily for natural image classification and showed inferior segmentation results. In this paper, we proposed a region-based EDL segmentation framework that can generate reliable uncertainty maps and accurate segmentation results, which is robust to noise and image corruption. We used the Theory of Evidence to interpret the output of a neural network as evidence values gathered from input features. Following Subjective Logic, evidence was parameterized as a Dirichlet distribution, and predicted probabilities were treated as subjective opinions. To evaluate the performance of our model on segmentation and uncertainty estimation, we conducted quantitative and qualitative experiments on the BraTS 2020 dataset. The results demonstrated the top performance of the proposed method in quantifying segmentation uncertainty and robustly segmenting tumors. Furthermore, our proposed new framework maintained the advantages of low computational cost and easy implementation and showed the potential for clinical application. Hao Li 0082, Yang Nan 0002, Javier Del Ser, Guang Yang 0006 |
Neural Comput. Appl. | 4 |
| 2023 | AI-based medical e-diagnosis for fast and automatic ventricular volume measurement in patients with normal pressure hydrocephalusabstractBased on CT and MRI images acquired from normal pressure hydrocephalus (NPH) patients, using machine learning methods, we aim to establish a multimodal and high-performance automatic ventricle segmentation method to achieve an efficient and accurate automatic measurement of the ventricular volume. First, we extract the brain CT and MRI images of 143 definite NPH patients. Second, we manually label the ventricular volume (VV) and intracranial volume (ICV). Then, we use the machine learning method to extract features and establish automatic ventricle segmentation model. Finally, we verify the reliability of the model and achieved automatic measurement of VV and ICV. In CT images, the Dice similarity coefficient (DSC), intraclass correlation coefficient (ICC), Pearson correlation, and Bland-Altman analysis of the automatic and manual segmentation result of the VV were 0.95, 0.99, 0.99, and 4.2 ± 2.6, respectively. The results of ICV were 0.96, 0.99, 0.99, and 6.0 ± 3.8, respectively. The whole process takes 3.4 ± 0.3 s. In MRI images, the DSC, ICC, Pearson correlation, and Bland-Altman analysis of the automatic and manual segmentation result of the VV were 0.94, 0.99, 0.99, and 2.0 ± 0.6, respectively. The results of ICV were 0.93, 0.99, 0.99, and 7.9 ± 3.8, respectively. The whole process took 1.9 ± 0.1 s. We have established a multimodal and high-performance automatic ventricle segmentation method to achieve efficient and accurate automatic measurement of the ventricular volume of NPH patients. This can help clinicians quickly and accurately understand the situation of NPH patient's ventricles. Xi Zhou 0003, Qinghao Ye, Jiakun Chen, Haiqin Ma, Jun Xia 0002, Javier Del Ser, Guang Yang 0006 |
Neural Comput. Appl. | 8 |
| 2023 | Global Transformer and Dual Local Attention Network via Deep-Shallow Hierarchical Feature Fusion for Retinal Vessel SegmentationabstractClinically, retinal vessel segmentation is a significant step in the diagnosis of fundus diseases. However, recent methods generally neglect the difference of semantic information between deep and shallow features, which fail to capture the global and local characterizations in fundus images simultaneously, resulting in the limited segmentation performance for fine vessels. In this article, a global transformer (GT) and dual local attention (DLA) network via deep-shallow hierarchical feature fusion (GT-DLA-dsHFF) are investigated to solve the above limitations. First, the GT is developed to integrate the global information in the retinal image, which effectively captures the long-distance dependence between pixels, alleviating the discontinuity of blood vessels in the segmentation results. Second, DLA, which is constructed using dilated convolutions with varied dilation rates, unsupervised edge detection, and squeeze-excitation block, is proposed to extract local vessel information, consolidating the edge details in the segmentation result. Finally, a novel deep-shallow hierarchical feature fusion (dsHFF) algorithm is studied to fuse the features in different scales in the deep learning framework, respectively, which can mitigate the attenuation of valid information in the process of feature fusion. We verified the GT-DLA-dsHFF on four typical fundus image datasets. The experimental results demonstrate our GT-DLA-dsHFF achieves superior performance against the current methods and detailed discussions verify the efficacy of the proposed three modules. Segmentation results of diseased images show the robustness of our proposed GT-DLA-dsHFF. Implementation codes will be available on https://github.com/YangLibuaa/GT-DLA-dsHFF. Yang Li 0010, Yue Zhang 0045, Jingyu Liu 0002, Kang Wang 0017, Gen-Sheng Zhang, Xiaofeng Liao 0001, Guang Yang 0006 |
IEEE Trans. Cybern. | 8 |
| 2023 | Adversarial Transformer for Repairing Human Airway SegmentationabstractAutomated airway segmentation models often suffer from discontinuities in peripheral bronchioles, which limits their clinical applicability. Furthermore, data heterogeneity across different centres and pathological abnormalities pose significant challenges to achieving accurate and robust segmentation in distal small airways. Accurate segmentation of airway structures is essential for the diagnosis and prognosis of lung diseases. To address these issues, we propose a patch-scale adversarial-based refinement network that takes in preliminary segmentation and original CT images and outputs a refined mask of the airway structure. Our method is validated on three datasets, including healthy cases, pulmonary fibrosis, and COVID-19 cases, and quantitatively evaluated using seven metrics. Our method achieves more than a 15% increase in the detected length ratio and detected branch ratio compared to previously proposed models, demonstrating its promising performance. The visual results show that our refinement approach, guided by a patch-scale discriminator and centreline objective functions, effectively detects discontinuities and missing bronchioles. We also demonstrate the generalizability of our refinement pipeline on three previous models, significantly improving their segmentation completeness. Our method provides a robust and accurate airway segmentation tool that can help improve diagnosis and treatment planning for lung diseases. Zeyu Tang 0001, Yang Nan 0002, Simon Walsh, Guang Yang 0006 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | HDL: Hybrid Deep Learning for the Synthesis of Myocardial Velocity Maps in Digital Twins for Cardiac AnalysisabstractSynthetic digital twins based on medical data accelerate the acquisition, labelling and decision making procedure in digital healthcare. A core part of digital healthcare twins is model-based data synthesis, which permits the generation of realistic medical signals without requiring to cope with the modelling complexity of anatomical and biochemical phenomena producing them in reality. Unfortunately, algorithms for cardiac data synthesis have been so far scarcely studied in the literature. An important imaging modality in the cardiac examination is three-directional CINE multi-slice myocardial velocity mapping (3Dir MVM), which provides a quantitative assessment of cardiac motion in three orthogonal directions of the left ventricle. The long acquisition time and complex acquisition produce make it more urgent to produce synthetic digital twins of this imaging modality. In this study, we propose a hybrid deep learning (HDL) network, especially for synthetic 3Dir MVM data. Our algorithm is featured by a hybrid UNet and a Generative Adversarial Network with a foreground-background generation scheme. The experimental results show that from temporally down-sampled magnitude CINE images (six times), our proposed algorithm can still successfully synthesise high temporal resolution 3Dir MVM CMR data (PSNR=42.32) with precise left ventricle segmentation (DICE=0.92). These performance scores indicate that our proposed HDL algorithm can be implemented in real-world digital twins for myocardial velocity mapping data simulation. To the best of our knowledge, this work is the first one investigating digital twins of the 3Dir MVM CMR, which has shown great potential for improving the efficiency of clinical studies via synthesised cardiac data. Xiaodan Xing, Javier Del Ser, Yinzhe Wu 0001, Yang Li 0010, Jun Xia 0002, Lei Xu 0037, David N. Firmin, Peter Gatehouse, Guang Yang 0006 |
IEEE J. Biomed. Health Informatics | 9 |
| 2023 | Multiple Adversarial Learning Based Angiography Reconstruction for Ultra-Low-Dose Contrast Medium CTabstractIodinated contrast medium (ICM) dose reduction is beneficial for decreasing potential health risk to renal-insufficiency patients in CT scanning. Due to the low-intensity vessel in ultra-low-dose-ICM CT angiography, it cannot provide clinical diagnosis of vascular diseases. Angiography reconstruction for ultra-low-dose-ICM CT can enhance vascular intensity for directly vascular diseases diagnosis. However, the angiography reconstruction is challenging since patient individual differences and vascular disease diversity. In this paper, we propose a Multiple Adversarial Learning based Angiography Reconstruction (i.e., MALAR) framework to enhance vascular intensity. Specifically, a bilateral learning mechanism is developed for mapping a relationship between source and target domains rather than the image-to-image mapping. Then, a dual correlation constraint is introduced to characterize both distribution uniformity from across-domain features and sample inconsistency within domain simultaneously. Finally, an adaptive fusion module by combining multi-scale information and long-range interactive dependency is explored to alleviate the interference of high-noise metal. Experiments are performed on CT sequences with different ICM doses. Quantitative results based on multiple metrics demonstrate the effectiveness of our MALAR on angiography reconstruction. Qualitative assessments by radiographers confirm the potential of our MALAR for the clinical diagnosis of vascular diseases. Weiwei Zhang 0006, Zhifan Gao, Guang Yang 0006, Lei Xu 0037, Weiwen Wu, Heye Zhang |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Vehicular Abandoned Object Detection Based on VANET and Edge AI in Road ScenesabstractRapid processing of abandoned objects is one of the most important tasks in road maintenance. Abandoned object detection heavily relies on traditional object detection approaches at a fixed location. However, detection accuracy and range are still far from satisfactory. This study proposes an abandoned object detection approach based on vehicular ad-hoc networks (VANETs) and edge artificial intelligence (AI) in road scenes. We propose a vehicular detection architecture for abandoned objects to achieve task-based AI technology for large-scale road maintenance in mobile computing circumstances. To improve detection accuracy and reduce repeated detection rates in mobile computing, we propose a detection algorithm that combines a deep learning network and a deduplication module for high-frequency detection. Finally, we propose a location estimation approach for abandoned objects based on the World Geodetic System 1984 (WGS84) coordinate system and an affine projection model to accurately compute the positions of abandoned objects. Experimental results show that our proposed algorithm achieves an average accuracy of 99.57% and 53.11% on the two datasets, respectively. Additionally, our whole system achieves real-time detection and high-precision localization performance on real roads. Gang Wang 0023, Mingliang Zhou 0001, Xuekai Wei, Guang Yang 0006 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Hierarchical Perception Adversarial Learning Framework for Compressed Sensing MRIabstractThe long acquisition time has limited the accessibility of magnetic resonance imaging (MRI) because it leads to patient discomfort and motion artifacts. Although several MRI techniques have been proposed to reduce the acquisition time, compressed sensing in magnetic resonance imaging (CS-MRI) enables fast acquisition without compromising SNR and resolution. However, existing CS-MRI methods suffer from the challenge of aliasing artifacts. This challenge results in the noise-like textures and missing the fine details, thus leading to unsatisfactory reconstruction performance. To tackle this challenge, we propose a hierarchical perception adversarial learning framework (HP-ALF). HP-ALF can perceive the image information in the hierarchical mechanism: image-level perception and patch-level perception. The former can reduce the visual perception difference in the entire image, and thus achieve aliasing artifact removal. The latter can reduce this difference in the regions of the image, and thus recover fine details. Specifically, HP-ALF achieves the hierarchical mechanism by utilizing multilevel perspective discrimination. This discrimination can provide the information from two perspectives (overall and regional) for adversarial learning. It also utilizes a global and local coherent discriminator to provide structure information to the generator during training. In addition, HP-ALF contains a context-aware learning block to effectively exploit the slice information between individual images for better reconstruction performance. The experiments validated on three datasets demonstrate the effectiveness of HP-ALF and its superiority to the comparative methods. Zhifan Gao, Yifeng Guo, Tieyong Zeng, Guang Yang 0006 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Less Is More: Unsupervised Mask-Guided Annotated CT Image Synthesis With Minimum Manual SegmentationsabstractAs a pragmatic data augmentation tool, data synthesis has generally returned dividends in performance for deep learning based medical image analysis. However, generating corresponding segmentation masks for synthetic medical images is laborious and subjective. To obtain paired synthetic medical images and segmentations, conditional generative models that use segmentation masks as synthesis conditions were proposed. However, these segmentation mask-conditioned generative models still relied on large, varied, and labeled training datasets, and they could only provide limited constraints on human anatomical structures, leading to unrealistic image features. Moreover, the invariant pixel-level conditions could reduce the variety of synthetic lesions and thus reduce the efficacy of data augmentation. To address these issues, in this work, we propose a novel strategy for medical image synthesis, namely Unsupervised Mask (UM)-guided synthesis, to obtain both synthetic images and segmentations using limited manual segmentation labels. We first develop a superpixel based algorithm to generate unsupervised structural guidance and then design a conditional generative model to synthesize images and annotations simultaneously from those unsupervised masks in a semi-supervised multi-task setting. In addition, we devise a multi-scale multi-task Fréchet Inception Distance (MM-FID) and multi-scale multi-task standard deviation (MM-STD) to harness both fidelity and variety evaluations of synthetic CT images. With multiple analyses on different scales, we could produce stable image quality measurements with high reproducibility. Compared with the segmentation mask guided synthesis, our UM-guided synthesis provided high-quality synthetic images with significantly higher fidelity, variety, and utility ( by Wilcoxon Signed Ranked test). Xiaodan Xing, Giorgos Papanastasiou, Simon Walsh, Guang Yang 0006 |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Explainable AI (XAI) In Biomedical Signal and Image Processing: Promises and ChallengesabstractArtificial intelligence has become pervasive across disciplines and fields, and biomedical image and signal processing is no exception. The growing and widespread interest on the topic has triggered a vast research activity that is reflected in an exponential research effort. Through study of massive and diverse biomedical data, machine and deep learning models have revolutionized various tasks such as modeling, segmentation, registration, classification and synthesis, outperforming traditional techniques. However, the difficulty in translating the results into biologically/clinically interpretable information is preventing their full exploitation in the field. Explainable AI (XAI) attempts to fill this translational gap by providing means to make the models interpretable and providing explanations. Different solutions have been proposed so far and are gaining increasing interest from the community. This paper aims at providing an overview on XAI in biomedical data processing and points to an upcoming Special Issue on Deep Learning in Biomedical Image and Signal Processing of the IEEE Signal Processing Magazine that is going to appear in March 2022. Guang Yang 0006, Arvind Rao, Christine Fernandez-Maloigne, Vince D. Calhoun, Gloria Menegaz |
ICIP | 1 |
| 2022 | A Novel Automated Classification and Segmentation for COVID-19 using 3D CT ScansabstractMedical image classification and segmentation based on deep learning (DL) are emergency research topics for diagnosing variant viruses of the current COVID-19 situation. In COVID-19 computed tomography (CT) images of the lungs, ground glass turbidity is the most common finding that requires specialist diagnosis. Based on this situation, some researchers propose the relevant DL models which can replace professional diagnostic specialists in clinics when lacking expertise. However, although DL methods have a stunning performance in medical image processing, the limited datasets can be a challenge in developing the accuracy of diagnosis at the human level. In addition, deep learning algorithms face the challenge of classifying and segmenting medical images in three or even multiple dimensions and maintaining high accuracy rates. Consequently, with a guaranteed high level of accuracy, our model can classify the patients' CT images into three types: Normal, Pneumonia and COVID. Subsequently, two datasets are used for segmentation, one of the datasets even has only a limited amount of data (20 cases). Our system combined the classification model and the segmentation model together, a fully integrated diagnostic model was built on the basis of ResNet50 and 3D U-Net algorithm. By feeding with different datasets, the COVID image segmentation of the infected area will be carried out according to classification results. Our model achieves 94.52% accuracy in the classification of lung lesions by 3 types: COVID, Pneumonia and Normal. For 2 labels (ground truth, lung lesions) segmentation, the model gets 99.57% of accuracy, 0.2191 of train loss and$0.78\pm 0.03$of MeanDice±Std, while the 4 labels (ground truth, left lung, right lung, lung lesions) segmentation achieves 98.89% of accuracy, 0.1132 of train loss and$0.83\pm 0.13$of MeanDice±Std. For future medical use, embedding the model into the medical facilities might be an efficient way of assisting or substituting doctors with diagnoses, therefore, a broader range of the problem of variant viruses in the COVID-19 situation may also be successfully solved. Guang Yang 0006 |
IPAS | 2 |
| 2022 | Swin Deformable Attention U-Net Transformer (SDAUT) for Explainable Fast MRI
Xiaodan Xing, Zhifan Gao, Guang Yang 0006 |
MICCAI (6) | 4 |
| 2022 | CS2: A Controllable and Simultaneous Synthesizer of Images and Annotations with Minimal Human Intervention
Xiaodan Xing, Yang Nan 0002, Yinzhe Wu 0001, Chengjia Wang, Zhifan Gao, Simon Walsh, Guang Yang 0006 |
MICCAI (8) | 8 |
| 2022 | Edge-enhanced dual discriminator generative adversarial network for fast MRI with parallel imaging using multi-view informationabstract-space. In recent years, most MRI reconstruction methods proposed in the literature focus on holistic image reconstruction rather than enhancing the edge information. This work steps aside this general trend by elaborating on the enhancement of edge information. Specifically, we introduce a novel parallel imaging coupled dual discriminator generative adversarial network (PIDD-GAN) for fast multi-channel MRI reconstruction by incorporating multi-view information. The dual discriminator design aims to improve the edge information in MRI reconstruction. One discriminator is used for holistic image reconstruction, whereas the other one is responsible for enhancing edge information. An improved U-Net with local and global residual learning is proposed for the generator. Frequency channel attention blocks (FCA Blocks) are embedded in the generator for incorporating attention mechanisms. Content loss is introduced to train the generator for better reconstruction quality. We performed comprehensive experiments on Calgary-Campinas public brain MR dataset and compared our method with state-of-the-art MRI reconstruction methods. Ablation studies of residual learning were conducted on the MICCAI13 dataset to validate the proposed modules. Results show that our PIDD-GAN provides high-quality reconstructed MR images, with well-preserved edge information. The time of single-image reconstruction is below 5ms, which meets the demand of faster processing. Weiping Ding 0001, Hao Dong 0003, Javier Del Ser, Jun Xia 0002, Tiaojuan Ren, Guang Yang 0006 |
Appl. Intell. | 10 |
| 2022 | Swin transformer for fast MRIabstractMagnetic resonance imaging (MRI) is an important non-invasive clinical tool that can produce high-resolution and reproducible images. However, a long scanning time is required for high-quality MR images, which leads to exhaustion and discomfort of patients, inducing more artefacts due to voluntary movements of the patients and involuntary physiological movements. To accelerate the scanning process, methods by k-space undersampling and deep learning based reconstruction have been popularised. This work introduced SwinMR, a novel Swin transformer based method for fast MRI reconstruction. The whole network consisted of an input module (IM), a feature extraction module (FEM) and an output module (OM). The IM and OM were 2D convolutional layers and the FEM was composed of a cascaded of residual Swin transformer blocks (RSTBs) and 2D convolutional layers. The RSTB consisted of a series of Swin transformer layers (STLs). The shifted windows multi-head self-attention (W-MSA/SW-MSA) of STL was performed in shifted windows rather than the multi-head self-attention (MSA) of the original transformer in the whole image space. A novel multi-channel loss was proposed by using the sensitivity maps, which was proved to reserve more textures and details. We performed a series of comparative studies and ablation studies in the Calgary-Campinas public brain MR dataset and conducted a downstream segmentation experiment in the Multi-modal Brain Tumour Segmentation Challenge 2017 dataset. The results demonstrate our SwinMR achieved high-quality reconstruction compared with other benchmark methods, and it shows great robustness with different undersampling masks, under noise interruption and on different datasets. The code is publicly available at https://github.com/ayanglab/SwinMR. Yingying Fang, Yinzhe Wu 0001, Huanjun Wu, Zhifan Gao, Yang Li 0010, Javier Del Ser, Jun Xia 0002, Guang Yang 0006 |
Neurocomputing | 9 |
| 2022 | AI-Based Reconstruction for Fast MRI - A Systematic Review and Meta-AnalysisabstractCompressed sensing (CS) has been playing a key role in accelerating the magnetic resonance imaging (MRI) acquisition process. With the resurgence of artificial intelligence, deep neural networks and CS algorithms are being integrated to redefine the state of the art of fast MRI. The past several years have witnessed substantial growth in the complexity, diversity, and performance of deep-learning-based CS techniques that are dedicated to fast MRI. In this meta-analysis, we systematically review the deep-learning-based CS techniques for fast MRI, describe key model designs, highlight breakthroughs, and discuss promising directions. We have also introduced a comprehensive analysis framework and a classification system to assess the pivotal role of deep learning in CS-based acceleration for MRI. Carola-Bibiane Schönlieb, Pietro Liò, Tim Leiner, Pier Luigi Dragotti, Ge Wang 0001, Daniel Rueckert, David N. Firmin, Guang Yang 0006 |
Proc. IEEE | 9 |
| 2022 | Automatic fine-grained glomerular lesion recognition in kidney pathologyabstractRecognition of glomeruli lesions is the key for diagnosis and treatment planning in kidney pathology; however, the coexisting glomerular structures such as mesangial regions exacerbate the difficulties of this task. In this paper, we introduce a scheme to recognize fine-grained glomeruli lesions from whole slide images. First, a focal instance structural similarity loss is proposed to drive the model to locate all types of glomeruli precisely. Then an Uncertainty Aided Apportionment Network is designed to carry out the fine-grained visual classification without bounding-box annotations. This double branch-shaped structure extracts common features of the child class from the parent class and produces the uncertainty factor for reconstituting the training dataset. Results of slide-wise evaluation illustrate the effectiveness of the entire scheme, with an 8–22% improvement of the mean Average Precision compared with remarkable detection methods. The comprehensive results clearly demonstrate the effectiveness of the proposed method. Yang Nan 0002, Fengyi Li, Peng Tang 0004, Guyue Zhang, Caihong Zeng, Guo Tong Xie, Guang Yang 0006 |
Pattern Recognit. | 8 |
| 2022 | JAS-GAN: Generative Adversarial Network Based Joint Atrium and Scar Segmentations on Unbalanced Atrial TargetsabstractAutomated and accurate segmentations of left atrium (LA) and atrial scars from late gadolinium-enhanced cardiac magnetic resonance (LGE CMR) images are in high demand for quantifying atrial scars. The previous quantification of atrial scars relies on a two-phase segmentation for LA and atrial scars due to their large volume difference (unbalanced atrial targets). In this paper, we propose an inter-cascade generative adversarial network, namely JAS-GAN, to segment the unbalanced atrial targets from LGE CMR images automatically and accurately in an end-to-end way. Firstly, JAS-GAN investigates an adaptive attention cascade to automatically correlate the segmentation tasks of the unbalanced atrial targets. The adaptive attention cascade mainly models the inclusion relationship of the two unbalanced atrial targets, where the estimated LA acts as the attention map to adaptively focus on the small atrial scars roughly. Then, an adversarial regularization is applied to the segmentation tasks of the unbalanced atrial targets for making a consistent optimization. It mainly forces the estimated joint distribution of LA and atrial scars to match the real ones. We evaluated the performance of our JAS-GAN on a 3D LGE CMR dataset with 192 scans. Compared with the state-of-the-art methods, our proposed approach yielded better segmentation performance (Average Dice Similarity Coefficient (DSC) values of 0.946 and 0.821 for LA and atrial scars, respectively), which indicated the effectiveness of our proposed approach for segmenting unbalanced atrial targets. Jun Chen 0030, Guang Yang 0006, Habib Khan, Heye Zhang, Yanping Zhang 0001, Shu Zhao 0005, Raad Mohiaddin, Tom Wong, David N. Firmin, Jennifer Keegan |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Adaptive Hierarchical Dual Consistency for Semi-Supervised Left Atrium Segmentation on Cross-Domain DataabstractSemi-supervised learning provides great significance in left atrium (LA) segmentation model learning with insufficient labelled data. Generalising semi-supervised learning to cross-domain data is of high importance to further improve model robustness. However, the widely existing distribution difference and sample mismatch between different data domains hinder the generalisation of semi-supervised learning. In this study, we alleviate these problems by proposing anAdaptive Hierarchical Dual Consistency(AHDC) for the semi-supervised LA segmentation on cross-domain data. The AHDC mainly consists of a Bidirectional Adversarial Inference module (BAI) and a Hierarchical Dual Consistency learning module (HDC). The BAI overcomes the difference of distributions and the sample mismatch between two different domains. It mainly learns two mapping networks adversarially to obtain two matched domains through mutual adaptation. The HDC investigates a hierarchical dual learning paradigm for cross-domain semi-supervised segmentation based on the obtained matched domains. It mainly builds two dual-modelling networks for mining the complementary information in both intra-domain and inter-domain. For the intra-domain learning, a consistency constraint is applied to the dual-modelling targets to exploit the complementary modelling information. For the inter-domain learning, a consistency constraint is applied to the LAs modelled by two dual-modelling networks to exploit the complementary knowledge among different data domains. We demonstrated the performance of our proposed AHDC on four 3D late gadolinium enhancement cardiac MR (LGE-CMR) datasets from different centres and a 3D CT dataset. Compared to other state-of-the-art methods, our proposed AHDC achieved higher segmentation accuracy, which indicated its capability in the cross-domain semi-supervised LA segmentation. Jun Chen 0030, Heye Zhang, Raad Mohiaddin, Tom Wong, David N. Firmin, Jennifer Keegan, Guang Yang 0006 |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Unsupervised Tissue Segmentation via Deep Constrained Gaussian NetworkabstractTissue segmentation is the mainstay of pathological examination, whereas the manual delineation is unduly burdensome. To assist this time-consuming and subjective manual step, researchers have devised methods to automatically segment structures in pathological images. Recently, automated machine and deep learning based methods dominate tissue segmentation research studies. However, most machine and deep learning based approaches are supervised and developed using a large number of training samples, in which the pixel-wise annotations are expensive and sometimes can be impossible to obtain. This paper introduces a novel unsupervised learning paradigm by integrating an end-to-end deep mixture model with a constrained indicator to acquire accurate semantic tissue segmentation. This constraint aims to centralise the components of deep mixture models during the calculation of the optimisation function. In so doing, the redundant or empty class issues, which are common in current unsupervised learning methods, can be greatly reduced. By validation on both public and in-house datasets, the proposed deep constrained Gaussian network achieves significantly (Wilcoxon signed-rank test) better performance (with the average Dice scores of 0.737 and 0.735, respectively) on tissue segmentation with improved stability and robustness, compared to other existing unsupervised segmentation approaches. Furthermore, the proposed method presents a similar performance (p-value >0.05) compared to the fully supervised U-Net. Yang Nan 0002, Peng Tang 0004, Guyue Zhang, Caihong Zeng, Zhifan Gao, Heye Zhang, Guang Yang 0006 |
IEEE Trans. Medical Imaging | 8 |
| 2022 | Annealing Genetic GAN for Imbalanced Web Data LearningabstractClass imbalance is one of the most basic and important problems of web data. The key to overcoming the class imbalance problems is to increase the effective instances of the minority, that is, data augmentation. Generative Adversarial Networks (GANs), which have recently been successfully applied in the field of image generation, can be used for data augmentation because they can learn the data distribution given ample training data instances and generate more data. However, learning the distributions from the imbalanced data can make GANs easily get stuck in a local optimum. In this work, we propose a new training strategy called Annealing Genetic GAN (AGGAN), which incorporates simulated annealing genetic algorithm into the training process of GANs. And this can help GANs avoid the local optimum trapping problem, which easily occurs when the training set is imbalanced. Unlike existing GANs, which use a fixed adversarial learning objective alternately training a generator, we use multiple adversarial learning objectives to train a set of generators and use the Metropolis criterion in simulated annealing to decide whether the generator should update. More specifically, the Metropolis criterion accepts worse solutions with a certain probability, so it can make our AGGAN escape from the local optimum and find a better solution. Theory and mathematical analysis provide strong theoretical support for the proposed training strategy. And experiments on several datasets demonstrate that AGGAN achieves convincing ability to solve the class imbalanced problem and reduces the training problems inherent in existing GANs. Jingyu Hao, Chengjia Wang, Guang Yang 0006, Zhifan Gao, Jinglin Zhang 0003, Heye Zhang |
IEEE Trans. Multim. | 3 |
| 2021 | FIRE: Unsupervised bi-directional inter- and intra-modality registration using deep networksabstractMagnetic resonance imaging (MRI) benefits from the acquisition of multiple sequences (thereafter, referred to as “modalities”) under a single imaging session. Each modality offers different complementary spatial and functional information in the clinical setting. Inter- and intra (across MR sequence slices)-modality image registration is an important pre-processing step across multiple applications in routine clinical workflows, such as when visual or quantitative imaging biomarkers need to be assessed across multi-sequence/multi-slice MRI data. This paper presents an unsupervised deep learning-based registration network that can learn affine and non-rigid transformations, simultaneously. Inverse-consistency is an important property that is commonly ignored in recent deep learning-based inter-modality registration algorithms. We address this issue through our proposed multi-task, cross-domain image synthesis architecture, in which we incorporated a new comprehensive transformation network. The proposed model learns a modality-independent latent representation to perform cycle-consistent cross-modality synthesis and uses an inverse-consistency loss to learn paired transformations, to align the synthesized with the target image. We name this proposed framework as “FIRE” due to the shape of its structure and we focus on interpreting model components to enhance model interpretability for clinical MR applications. Our method shows comparable and better performances against a well-established baseline method in experiments on multi-sequence brain MR data and intra-modality 4D cardiac Cine-MR data. Chengjia Wang, Guang Yang 0006, Giorgos Papanastasiou |
CBMS | 2 |
| 2021 | Explainable AI for COVID-19 CT Classifiers: An Initial Comparison StudyabstractArtificial Intelligence (AI) has made leapfrogs in development across all the industrial sectors especially when deep learning has been introduced. Deep learning helps to learn the behaviour of an entity through methods of recognising and interpreting patterns. Despite its limitless potential, the mystery is how deep learning algorithms make a decision in the first place. Explainable AI (XAI) is the key to unlocking AI and the black-box for deep learning. XAI is an AI model that is programmed to explain its goals, logic, and decision making so that the end users can understand. The end users can be domain experts, regulatory agencies, managers and executive board members, data scientists, users that use AI, with or without awareness, or someone who is affected by the decisions of an AI model. Chest CT has emerged as a valuable tool for the clinical diagnostic and treatment management of the lung diseases associated with COVID-19. AI can support rapid evaluation of CT scans to differentiate COVID-19 findings from other lung diseases. However, how these AI tools or deep learning algorithms reach such a decision and which are the most influential features derived from these neural networks with typically deep layers are not clear. The aim of this study is to propose and develop XAI strategies for COVID-19 classification models with an investigation of comparison. The results demonstrate promising quantification and qualitative visualisations that can further enhance the clinician's understanding and decision making with more granular information from the results given by the learned XAI models. Qinghao Ye, Jun Xia 0002, Guang Yang 0006 |
CBMS | 3 |
| 2021 | Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual FusionabstractVideo highlight detection plays an increasingly important role in social media content filtering, however, it remains highly challenging to develop automated video highlight detection methods because of the lack of temporal annotations (i.e., where the highlight moments are in long videos) for supervised learning. In this paper, we propose a novel weakly supervised method that can learn to detect highlights by mining video characteristics with video level annotations (topic tags) only. Particularly, we exploit audio-visual features to enhance video representation and take temporal cues into account for improving detection performance. Our contributions are threefold: 1) we propose an audio-visual tensor fusion mechanism that efficiently models the complex association between two modalities while reducing the gap of the heterogeneity between the two modalities; 2) we introduce a novel hierarchical temporal context encoder to embed local temporal clues in between neighboring segments; 3) finally, we alleviate the gradient vanishing problem theoretically during model optimization with attention-gated instance aggregation. Extensive experiments on two benchmark datasets (YouTube Highlights and TVSum) have demonstrated our method outperforms other state-of-the-art methods with remarkable improvements. Qinghao Ye, Xiyue Shen, Yuan Gao 0017, Qi Bi, Ping Li 0006, Guang Yang 0006 |
ICCV | 7 |
| 2021 | DHQN: a Stable Approach to Remove Target Network from Deep Q-learning NetworkabstractAs the first successful attempt to combine deep neural network and reinforcement learning, Deep Q-learning Network (DQN) draws a lot of attention from reinforcement learning researchers. One of the most important components of DQN is target network, which is used to stabilize learning process. When confront complex network structure, the existence of target network means extra memory resource to preserve the neural network weights and high computing cost to calculate target. Thus, we propose a Deep Hybrid Q-learning Network (DHQN) algorithm, which introduces an alternative approach, Random Hybrid Optimization (RHO), that can simplify DQN and attain a more stable and faster learning without a target network. We illustrate that RHO can decelerate divergence in the classical off-policy counterexample θ → 2θ problem. We also testify the effectiveness of DHQN in several control and Atari domains, which shows DHQN outperforms DQN without a target network and original DQN. Guang Yang 0006, Di'an Fei, Tian Huang, Qingyun Li, Xingguo Chen |
ICTAI | 1 |
| 2021 | Arbitrary Scale Super-Resolution for Medical ImagesabstractSingle image super-resolution (SISR) aims to obtain a high-resolution output from one low-resolution image. Currently, deep learning-based SISR approaches have been widely discussed in medical image processing, because of their potential to achieve high-quality, high spatial resolution images without the cost of additional scans. However, most existing methods are designed for scale-specific SR tasks and are unable to generalize over magnification scales. In this paper, we propose an approach for medical image arbitrary-scale super-resolution (MIASSR), in which we couple meta-learning with generative adversarial networks (GANs) to super-resolve medical images at any scale of magnification in [Formula: see text]. Compared to state-of-the-art SISR algorithms on single-modal magnetic resonance (MR) brain images (OASIS-brains) and multi-modal MR brain images (BraTS), MIASSR achieves comparable fidelity performance and the best perceptual quality with the smallest model size. We also employ transfer learning to enable MIASSR to tackle SR tasks of new medical modalities, such as cardiac MR images (ACDC) and chest computed tomography images (COVID-CT). The source code of our work is also public. Thus, MIASSR has the potential to become a new foundational pre-/post-processing step in clinical image analysis tasks such as reconstruction, image quality enhancement, and segmentation. Chuan Tan, Guang Yang 0006, Pietro Liò |
Int. J. Neural Syst. | 4 |
| 2021 | Industrial Cyber-Physical Systems-Based Cloud IoT Edge for Federated Heterogeneous DistillationabstractDeep convoloutional networks have been widely deployed in modern cyber-physical systems performing different visual classification tasks. As the fog and edge devices have different computing capacity and perform different subtasks, models trained for one device may not be deployable on another. Knowledge distillation technique can effectively compress well trained convolutional neural networks into light-weight models suitable to different devices. However, due to privacy issue and transmission cost, manually annotated data for training the deep learning models are usually gradually collected and archived in different sites. Simply training a model on powerful cloud servers and compressing them for particular edge devices failed to use the distributed data stored at different sites. This offline training approach is also inefficient to deal with new data collected from the edge devices. To overcome these obstacles, in this article, we propose the heterogeneous brain storming (HBS) method for object recognition tasks in real-world Internet of Things (IoT) scenarios. Our method enables flexible bidirectional federated learning of heterogeneous models trained on distributed datasets with a new “brain storming” mechanism and optimizable temperature parameters. In our comparison experiments, this HBS method outperformed multiple state-of-the-art single-model compression methods, as well as the newest multinetwork knowledge distillation methods with both homogeneous and heterogeneous classifiers. The ablation experiment results proved that the trainable temperature parameter into the conventional knowledge distillation loss can effectively ease the learning process of student networks in different methods. To the best of authors' knowledge, this is the first IoT-oriented method that allows asynchronous bidirectional heterogeneous knowledge distillation in deep networks. Chengjia Wang, Guang Yang 0006, Giorgos Papanastasiou, Heye Zhang, Joel J. P. C. Rodrigues, Victor Hugo C. de Albuquerque |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Multitask Learning for Estimating Multitype Cardiac Indices in MRI and CT Based on Adversarial Reverse MappingabstractThe estimation of multitype cardiac indices from cardiac magnetic resonance imaging (MRI) and computed tomography (CT) images attracts great attention because of its clinical potential for comprehensive function assessment. However, the most exiting model can only work in one imaging modality (MRI or CT) without transferable capability. In this article, we propose the multitask learning method with the reverse inferring for estimating multitype cardiac indices in MRI and CT. Different from the existing forward inferring methods, our method builds a reverse mapping network that maps the multitype cardiac indices to cardiac images. The task dependencies are then learned and shared to multitask learning networks using an adversarial training approach. Finally, we transfer the parameters learned from MRI to CT. A series of experiments were conducted in which we first optimized the performance of our framework via ten-fold cross-validation of over 2900 cardiac MRI images. Then, the fine-tuned network was run on an independent data set with 2360 cardiac CT images. The results of all the experiments conducted on the proposed adversarial reverse mapping show excellent performance in estimating multitype cardiac indices. Chengjin Yu, Zhifan Gao, Weiwei Zhang 0006, Guang Yang 0006, Shu Zhao 0005, Heye Zhang, Yanping Zhang 0001, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Annealing Genetic GAN for Minority Oversampling
Jingyu Hao, Chengjia Wang, Heye Zhang, Guang Yang 0006 |
BMVC | 4 |
| 2020 | Deep Attentive Wasserstein Generative Adversarial Networks for MRI Reconstruction with Recurrent Context-Awareness
Yifeng Guo, Chengjia Wang, Heye Zhang, Guang Yang 0006 |
MICCAI (2) | 4 |
| 2020 | Simultaneous left atrium anatomy and scar segmentations via deep learning in multiview information with attentionabstractThree-dimensional late gadolinium enhanced (LGE) cardiac MR (CMR) of left atrial scar in patients with atrial fibrillation (AF) has recently emerged as a promising technique to stratify patients, to guide ablation therapy and to predict treatment success. This requires a segmentation of the high intensity scar tissue and also a segmentation of the left atrium (LA) anatomy, the latter usually being derived from a separate bright-blood acquisition. Performing both segmentations automatically from a single 3D LGE CMR acquisition would eliminate the need for an additional acquisition and avoid subsequent registration issues. In this paper, we propose a joint segmentation method based on multiview two-task (MVTT) recursive attention model working directly on 3D LGE CMR images to segment the LA (and proximal pulmonary veins) and to delineate the scar on the same dataset. Using our MVTT recursive attention model, both the LA anatomy and scar can be segmented accurately (mean Dice score of 93% for the LA anatomy and 87% for the scar segmentations) and efficiently (∼0.27 s to simultaneously segment the LA anatomy and scars directly from the 3D LGE CMR dataset with 60–68 2D slices). Compared to conventional unsupervised learning and other state-of-the-art deep learning based methods, the proposed MVTT model achieved excellent results, leading to an automatic generation of a patient-specific anatomical model combined with scar segmentation for patients in AF. Guang Yang 0006, Jun Chen 0030, Zhifan Gao, Shuo Li 0001, Hao Ni 0001, Elsa D. Angelini, Tom Wong, Raad Mohiaddin, Eva Nyktari, Rick Wage, Lei Xu 0037, Yanping Zhang 0001, Xiuquan Du, Heye Zhang, David N. Firmin, Jennifer Keegan |
Future Gener. Comput. Syst. | 1 |
| 2020 | Atrial scar quantification via multi-scale CNN in the graph-cuts frameworkabstractLate gadolinium enhancement magnetic resonance imaging (LGE MRI) appears to be a promising alternative for scar assessment in patients with atrial fibrillation (AF). Automating the quantification and analysis of atrial scars can be challenging due to the low image quality. In this work, we propose a fully automated method based on the graph-cuts framework, where the potentials of the graph are learned on a surface mesh of the left atrium (LA) using a multi-scale convolutional neural network (MS-CNN). For validation, we have included fifty-eight images with manual delineations. MS-CNN, which can efficiently incorporate both the local and global texture information of the images, has been shown to evidently improve the segmentation accuracy of the proposed graph-cuts based method. The segmentation could be further improved when the contribution between the t-link and n-link weights of the graph is balanced. The proposed method achieves a mean accuracy of 0.856 ± 0.033 and mean Dice score of 0.702 ± 0.071 for LA scar quantification. Compared to the conventional methods, which are based on the manual delineation of LA for initialization, our method is fully automatic and has demonstrated significantly better Dice score and accuracy (p < 0.01). The method is promising and can be potentially useful in diagnosis and prognosis of AF. Lei Li 0020, Fuping Wu, Guang Yang 0006, Lingchao Xu, Tom Wong, Raad Mohiaddin, David N. Firmin, Jennifer Keegan, Xiahai Zhuang |
Medical Image Anal. | 3 |
| 2020 | SaliencyGAN: Deep Learning Semisupervised Salient Object Detection in the Fog of IoTabstractIn modern Internet of Things (IoT), visual analysis and predictions are often performed by deep learning models. Salient object detection (SOD) is a fundamental preprocessing for these applications. Executing SOD on the fog devices is a challenging task due to the diversity of data and fog devices. To adopt convolutional neural networks (CNN) on fog-cloud infrastructures for SOD-based applications, we introduce a semisupervised adversarial learning method in this article. The proposed model, named as SaliencyGAN, is empowered by a novel concatenated generative adversarial network (GAN) framework with partially shared parameters. The backbone CNN can be chosen flexibly based on the specific devices and applications. In the meanwhile, our method uses both the labeled and unlabeled data from different problem domains for training. Using multiple popular benchmark datasets, we compared state-of-the-art baseline methods to our SaliencyGAN obtained with 10-100% labeled training data. SaliencyGAN gained performance comparable to the supervised baselines when the percentage of labeled data reached 30%, and outperformed the weakly supervised and unsupervised baselines. Furthermore, our ablation study shows that SaliencyGAN were more robust to the common “mode missing” (or “mode collapse”) issue compared to the selected popular GAN models. The visualized ablation results have proved that SaliencyGAN learned a better estimation of data distributions. To the best of our knowledge, this is the first IoT-oriented semisupervised SOD method. Chengjia Wang, Shizhou Dong, Giorgos Papanastasiou, Heye Zhang, Guang Yang 0006 |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | Direct Quantification of Coronary Artery Stenosis Through Hierarchical Attentive Multi-View LearningabstractQuantification of coronary artery stenosis on X-ray angiography (XRA) images is of great importance during the intraoperative treatment of coronary artery disease. It serves to quantify the coronary artery stenosis by estimating the clinical morphological indices, which are essential in clinical decision making. However, stenosis quantification is still a challenging task due to the overlapping, diversity and small-size region of the stenosis in the XRA images. While efforts have been devoted to stenosis quantification through low-level features, these methods have difficulty in learning the real mapping from these features to the stenosis indices. These methods are still cumbersome and unreliable for the intraoperative procedures due to their two-phase quantification, which depends on the results of segmentation or reconstruction of the coronary artery. In this work, we are proposing a hierarchical attentive multi-view learning model (HEAL) to achieve a direct quantification of coronary artery stenosis, without the intermediate segmentation or reconstruction. We have designed a multi-view learning model to learn more complementary information of the stenosis from different views. For this purpose, an intra-view hierarchical attentive block is proposed to learn the discriminative information of stenosis. Additionally, a stenosis representation learning module is developed to extract the multi-scale features from the keyframe perspective for considering the clinical workflow. Finally, the morphological indices are directly estimated based on the multi-view feature embedding. Extensive experiment studies on clinical multi-manufacturer dataset consisting of 228 subjects show the superiority of our HEAL against nine comparing methods, including direct quantification methods and multi-view learning methods. The experimental results demonstrate the better clinical agreement between the ground truth and the prediction, which endows our proposed method with a great potential for the efficient intraoperative treatment of coronary artery disease. Dong Zhang 0012, Guang Yang 0006, Shu Zhao 0005, Yanping Zhang 0001, Dhanjoo N. Ghista, Heye Zhang, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2019 | A Deep Learning Based Approach to Skin Lesion Border Extraction With a Novel Edge Detector in Dermoscopy ImagesabstractLesion border detection is considered a crucial step in diagnosing skin cancer. However, performing such a task automatically is challenging due to the low contrast between the surrounding skin and lesion, ambiguous lesion borders, and the presence of artifacts such as hair. In this paper we propose a two-stage approach for skin lesion border detection: (i) segmenting the skin lesion dermoscopy image using U-Net, and (ii) extracting the edges from the segmented image using a novel approach we call FuzzEdge. The proposed approach is compared with another published skin lesion border detection approach, and the results show that our approach performs better in detecting the main borders of the lesion and is more robust to artifacts that might be present in the image. The approach is also compared with the manual border drawings of a dermatologist, resulting in an average Dice similarity of 87.7%. Abder-Rahman Ali, Jingpeng Li 0001, Sally Jane O'Shea, Guang Yang 0006, Thomas Trappenberg, Xujiong Ye |
IJCNN | 4 |
| 2019 | Discriminative Consistent Domain Generation for Semi-supervised Learning
Jun Chen 0030, Heye Zhang, Yanping Zhang 0001, Shu Zhao 0005, Raad Mohiaddin, Tom Wong, David N. Firmin, Guang Yang 0006, Jennifer Keegan |
MICCAI (2) | 8 |
| 2019 | Recurrent Aggregation Learning for Multi-view Echocardiographic Sequences Segmentation
Ming Li 0005, Weiwei Zhang 0006, Guang Yang 0006, Chengjia Wang, Heye Zhang, Huafeng Liu 0003, Shuo Li 0001 |
MICCAI (2) | 3 |
| 2019 | Direct Quantification for Coronary Artery Stenosis Using Multiview Learning
Dong Zhang 0012, Guang Yang 0006, Shu Zhao 0005, Yanping Zhang 0001, Heye Zhang, Shuo Li 0001 |
MICCAI (2) | 2 |
| 2019 | Evaluation of algorithms for Multi-Modality Whole Heart Segmentation: An open-access grand challengeabstractKnowledge of whole heart anatomy is a prerequisite for many clinical applications. Whole heart segmentation (WHS), which delineates substructures of the heart, can be very valuable for modeling and analysis of the anatomy and functions of the heart. However, automating this segmentation can be challenging due to the large variation of the heart shape, and different image qualities of the clinical data. To achieve this goal, an initial set of training data is generally needed for constructing priors or for training. Furthermore, it is difficult to perform comparisons between different methods, largely due to differences in the datasets and evaluation metrics used. This manuscript presents the methodologies and evaluation results for the WHS algorithms selected from the submissions to the Multi-Modality Whole Heart Segmentation (MM-WHS) challenge, in conjunction with MICCAI 2017. The challenge provided 120 three-dimensional cardiac images covering the whole heart, including 60 CT and 60 MRI volumes, all acquired in clinical environments with manual delineation. Ten algorithms for CT data and eleven algorithms for MRI data, submitted from twelve groups, have been evaluated. The results showed that the performance of CT WHS was generally better than that of MRI WHS. The segmentation of the substructures for different categories of patients could present different levels of challenge due to the difference in imaging and variations of heart shapes. The deep learning (DL)-based methods demonstrated great potential, though several of them reported poor results in the blinded evaluation. Their performance could vary greatly across different network structures and training strategies. The conventional algorithms, mainly based on multi-atlas segmentation, demonstrated good performance, though the accuracy and computational efficiency could be limited. The challenge, including provision of the annotated training data and the blinded evaluation for submitted algorithms on the test data, continues as an ongoing benchmarking resource via its homepage (www.sdspeople.fudan.edu.cn/zhuangxiahai/0/mmwhs/). Xiahai Zhuang, Lei Li 0020, Christian Payer, Darko Stern, Martin Urschler, Mattias P. Heinrich, Julien Oster, Chunliang Wang, Örjan Smedby, Cheng Bian, Xin Yang 0009, Pheng-Ann Heng, Aliasghar Mortazi, Ulas Bagci, Guanyu Yang 0001, Chenchen Sun, Gaetan Galisot, Jean-Yves Ramel, Guang Yang 0006 |
Medical Image Anal. | 19 |
| 2018 | Holistic and Deep Feature Pyramids for Saliency Detection
Shizhong Dong, Zhifan Gao, Shanhui Sun, Xin Wang 0045, Ming Li 0005, Heye Zhang, Guang Yang 0006, Huafeng Liu 0003, Shuo Li 0001 |
BMVC | 7 |
| 2018 | Deep Learning intra-image and inter-images features for Co-saliency detection
Shizhong Dong, Zhifan Gao, Xi Wu 0004, Heye Zhang, Guang Yang 0006, Shuo Li 0001 |
BMVC | 7 |
| 2018 | Multiview Two-Task Recursive Attention Model for Left Atrium and Atrial Scars Segmentation
Jun Chen 0030, Guang Yang 0006, Zhifan Gao, Hao Ni 0001, Elsa D. Angelini, Raad Mohiaddin, Tom Wong, Yanping Zhang 0001, Xiuquan Du, Heye Zhang, Jennifer Keegan, David N. Firmin |
MICCAI (2) | 2 |
| 2018 | The Deep Poincaré Map: A Novel Approach for Left Ventricle Segmentation
Yuanhan Mo, Fangde Liu, Douglas McIlwraith, Guang Yang 0006, Jingqing Zhang, Taigang He, Yike Guo |
MICCAI (4) | 4 |
| 2018 | Stochastic Deep Compressive Sensing for the Reconstruction of Diffusion Tensor Cardiac MRI
Jo Schlemper, Guang Yang 0006, Pedro F. Ferreira, Andrew D. Scott, Laura-Ann McGill, Zohya Khalique, Margarita Gorodezky, Malte Roehl, Jennifer Keegan, Dudley Pennell, David N. Firmin, Daniel Rueckert |
MICCAI (1) | 2 |
| 2018 | Adversarial and Perceptual Refinement for Compressed Sensing MRI Reconstruction
Maximilian Seitzer, Guang Yang 0006, Jo Schlemper, Ozan Oktay, Tobias Würfl, Vincent Christlein, Tom Wong, Raad Mohiaddin, David N. Firmin, Jennifer Keegan, Daniel Rueckert, Andreas K. Maier |
MICCAI (1) | 2 |
| 2018 | Bayesian VoxDRN: A Probabilistic Deep Voxelwise Dilated Residual Network for Whole Heart Segmentation from 3D MR Images
Zenglin Shi, Guodong Zeng, Le Zhang 0001, Xiahai Zhuang, Lei Li 0020, Guang Yang 0006, Guoyan Zheng |
MICCAI (4) | 6 |
| 2018 | Atrial Fibrosis Quantification Based on Maximum Likelihood Estimator of Multivariate Images
Fuping Wu, Lei Li 0020, Guang Yang 0006, Tom Wong, Raad Mohiaddin, David N. Firmin, Jennifer Keegan, Lingchao Xu, Xiahai Zhuang |
MICCAI (4) | 3 |
| 2018 | DAGAN: Deep De-Aliasing Generative Adversarial Networks for Fast Compressed Sensing MRI ReconstructionabstractCompressed sensing magnetic resonance imaging (CS-MRI) enables fast acquisition, which is highly desirable for numerous clinical applications. This can not only reduce the scanning cost and ease patient burden, but also potentially reduce motion artefacts and the effect of contrast washout, thus yielding better image quality. Different from parallel imaging-based fast MRI, which utilizes multiple coils to simultaneously receive MR signals, CS-MRI breaks the Nyquist-Shannon sampling barrier to reconstruct MRI images with much less required raw data. This paper provides a deep learning-based strategy for reconstruction of CS-MRI, and bridges a substantial gap between conventional non-learning methods working only on data from a single image, and prior knowledge from large training data sets. In particular, a novel conditional Generative Adversarial Networks-based model (DAGAN)-based model is proposed to reconstruct CS-MRI. In our DAGAN architecture, we have designed a refinement learning method to stabilize our U-Net based generator, which provides an end-to-end network to reduce aliasing artefacts. To better preserve texture and edges in the reconstruction, we have coupled the adversarial loss with an innovative content loss. In addition, we incorporate frequency-domain information to enforce similarity in both the image and frequency domains. We have performed comprehensive comparison studies with both conventional CS-MRI reconstruction methods and newly investigated deep learning approaches. Compared with these methods, our DAGAN method provides superior reconstruction with preserved perceptual image details. Furthermore, each image is reconstructed in about 5 ms, which is suitable for real-time processing. Guang Yang 0006, Simiao Yu, Hao Dong 0003, Gregory Slabaugh, Pier Luigi Dragotti, Xujiong Ye, Fangde Liu, Simon R. Arridge, Jennifer Keegan, Yike Guo, David N. Firmin |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Super-Resolved Enhancement of a Single Image and Its Application in Cardiac MRI
Guang Yang 0006, Xujiong Ye, Gregory Slabaugh, Jennifer Keegan, Raad Mohiaddin, David N. Firmin |
ICISP | 1 |