EDBT 2026 Demo / reviewers in the wild / expert
Changmiao Wang
dblp:175/0741
· DBLP profile ↗
50ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0003-2466-5990ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 20 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WDT-MD: Wavelet Diffusion Transformers for Microaneurysm Detection in Fundus ImagesabstractMicroaneurysms (MAs), the earliest pathognomonic signs of Diabetic Retinopathy (DR), present as sub-60 μm lesions in fundus images with highly variable photometric and morphological characteristics, rendering manual screening not only labor-intensive but inherently error-prone. While diffusion-based anomaly detection has emerged as a promising approach for automated MA screening, its clinical application is hindered by three fundamental limitations. First, these models often fall prey to "identity mapping", where they inadvertently replicate the input image. Second, they struggle to distinguish MAs from other anomalies, leading to high false positives. Third, their suboptimal reconstruction of normal features hampers overall performance. To address these challenges, we propose a Wavelet Diffusion Transformer framework for MA Detection (WDT-MD), which features three key innovations: a noise-encoded image conditioning mechanism to avoid "identity mapping" by perturbing image conditions during training; pseudo-normal pattern synthesis via inpainting to introduce pixel-level supervision, enabling discrimination between MAs and other anomalies; and a wavelet diffusion Transformer architecture that combines the global modeling capability of diffusion Transformers with multi-scale wavelet analysis to enhance reconstruction of normal retinal features. Comprehensive experiments on the IDRiD and e-ophtha MA datasets demonstrate that WDT-MD outperforms state-of-the-art methods in both pixel-level and image-level MA detection. This advancement holds significant promise for improving early DR screening. Yifei Sun 0005, Yuzhi He, Junhao Jia, Ruiquan Ge, Changmiao Wang |
AAAI | 6 |
| 2026 | LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung NodulesabstractDiagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been made in using multimodal large language models for analyzing lung CT scans, challenges remain in accurately describing nodule morphology and incorporating medical expertise. These limitations affect the reliability and effectiveness of these models in clinical settings. Collaborative multi-agent systems offer a promising strategy for achieving a balance between generality and precision in medical applications, yet their potential in pathology has not been thoroughly explored. To bridge these gaps, we introduce LungNoduleAgent, an innovative collaborative multi-agent system specifically designed for analyzing lung CT scans. LungNoduleAgent streamlines the diagnostic process into sequential components, improving precision in describing nodules and grading malignancy through three primary modules. The first module, the Nodule Spotter, coordinates clinical detection models to accurately identify nodules. The second module, the Radiologist, integrates localized image description techniques to produce comprehensive CT reports. Finally, the Doctor Agent System performs malignancy reasoning by using images and CT reports, supported by a pathology knowledge base and a multi-agent system framework. Extensive testing on two private datasets and the public LIDC-IDRI dataset indicates that LungNoduleAgent surpasses mainstream vision-language models, agent systems, and advanced expert models such as GPT-4o, Claude 3.7 Sonnet, LLaMA-3.2 Vision, Qwen2.5-VL, Med-R1, MedGemma, MedAgent-Pro, MedAgents, MDAgent and LLaVA-Med. These results highlight the importance of region-level semantic alignment and multi-agent collaboration in diagnosing nodules. LungNoduleAgent stands out as a promising foundational tool for supporting clinical analyses of lung nodules. Yaoqun Liu, Fenglei Fan, Dajiang Lei, Gangyong Jia, Changmiao Wang, Ruiquan Ge |
AAAI | 9 |
| 2026 | DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy PredictionabstractLung cancer continues to be the leading cause of cancer-related deaths globally. Early detection and diagnosis of pulmonary nodules are essential for improving patient survival rates. Although previous research has integrated multimodal and multi-temporal information, outperforming single modality and single time point, the fusion methods are limited to inefficient vector concatenation and simple mutual attention, highlighting the need for more effective multimodal information fusion. To address these challenges, we introduce a Dual-Graph Spatiotemporal Attention Network, which leverages temporal variations and multimodal data to enhance the accuracy of predictions. Our methodology involves developing a Global-Local Feature Encoder to better capture the local, global, and fused characteristics of pulmonary nodules. Additionally, a Dual-Graph Construction method organizes multimodal features into inter-modal and intra-modal graphs. Furthermore, a Hierarchical Cross-Modal Graph Fusion Module is introduced to refine feature integration. We also compiled a novel multimodal dataset named the NLST-cmst dataset as a comprehensive source of support for related research. Our extensive experiments, conducted on both the NLST-cmst and curated CSTL-derived datasets, demonstrate that our DGSAN significantly outperforms state-of-the-art methods in classifying pulmonary nodules with exceptional computational efficiency. Zhaojie Fang, Guanyu Zhou, Yin Shen, Huoling Luo, Ahmed El-Azab, Ruiquan Ge, Changmiao Wang |
AAAI | 10 |
| 2026 | CervNet: A Hybrid Deep Learning Model for Cervical Cell Multi-Classification Using CNN and Swin Transformer
Khadija Idaissa, Fei-wei Qin, Changmiao Wang, Yanming Zhu 0001 |
ICIC (27) | 4 |
| 2026 | MAPTab: Missing-Aware Progressive Masking for Self-Supervised Tabular Imputation
Jinyi Xu, Chenlei Li, Mengping Zhong, Changmiao Wang |
ICIC (26) | 6 |
| 2026 | A cross-scale interaction framework combining Mamba and Convolutional Neural Networks for Arbitrary-Scale Super-Resolution of infrared images
Fei-wei Qin, Changmiao Wang, Kai Zhang 0008, Yong Peng 0001, Jing Bai 0004 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | No modality left behind: Adapting to missing modalities via knowledge distillation for brain tumor segmentation
Shenghao Zhu, Yifei Chen 0019, Guanyu Zhou, Yuanhan Wang, Fei-wei Qin, Changmiao Wang, Qiyuan Tian |
Medical Image Anal. | 8 |
| 2026 | Progressive Fusion of Multi-Scale Mamba Context and Local Detail Priors for Infrared Small Target DetectionabstractInfrared Small Target Detection (IRSTD) requires strong target-level detection capability, which depends on effective modeling of long-range global dependencies. This demand has driven the transition from CNN-based approaches to Transformer-based architectures. Although Transformers improve global context modeling, their high computational cost limits practical deployment. Recent advances in Mamba enable efficient long-range dependency modeling with reduced complexity, offering a promising alternative that alleviates the efficiency limitations of Transformers while preserving target-level detection performance. However, Mamba is not inherently tailored for IRSTD, as it lacks explicit mechanisms for capturing fine-grained local details and modeling background variations across multiple spatial scales. To address these limitations, we propose MCFNet, an encoder-decoder framework that integrates Mamba to enhance target-level detection performance with moderate computational cost. MCFNet introduces a Detail-Capturable Convolution Block to strengthen local detail perception and a Multi-scale Contextual Mamba Block to improve background modeling across different scales. While the resulting dual-branch design enhances both global semantics and local details, it also introduces challenges in feature fusion. To this end, a Feature Fusion Decoding Module is further proposed to enable effective collaboration between global and local representations. Extensive experiments on multiple public IRSTD benchmark datasets demonstrate that MCFNet consistently outperforms existing methods in both pixel-level and target-level metrics, achieving higher detection accuracy with reduced false alarms. The code of our model is available at: https://github.com/Fihven/MCFNet. Xiangjun Zhu, Fei-wei Qin, Changmiao Wang, Jin Fan 0003, Fei Lin 0006, Jing Bai 0004, Chenglong Zhang 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | TC-KANRecon: High-Quality and Accelerated MRI Reconstruction via Adaptive KAN Mechanisms and Intelligent Feature ScalingabstractMRI has become essential in clinical diagnosis due to its high resolution and multiple contrast mechanisms. However, the relatively long acquisition time limits its broader application. To address this issue, this study presents an innovative conditional guided diffusion model, named TC-KANRecon, which incorporates the Multi-Free U-KAN module and a dynamic clipping strategy. TC-KANRecon model aims to accelerate the MRI reconstruction process through deep learning methods while maintaining the reconstruction quality. The MF-UKAN module can effectively balance the tradeoff between image denoising and structure preservation. Specifically, it presents the multi-head attention mechanisms and scalar modulation factors, which significantly enhance the model's robustness and structure preservation capabilities in complex noise environments. Moreover, the dynamic clipping strategy in TC-KANRecon adjusts the cropping interval according to the sampling steps, thereby mitigating image detail loss while preserving the visual features of the images. Furthermore, the Conditional Guidance Model incorporates full-sampling k-space information, realizing efficient fusion of conditional information, enhancing the model's ability to process complex data, and improving the realism and detail richness of reconstructed images. Experimental results demonstrate that the proposed method outperforms other MRI reconstruction methods in both qualitative and quantitative evaluations. Notably, TC-KANRecon method exhibits excellent reconstruction results when processing high-noise, low-sampling-rate MRI data. Ruiquan Ge, Yifei Chen 0019, Shenghao Zhu, Dong Zeng, Changmiao Wang, Qiegen Liu, Shanzhou Niu |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | BS-LDM: Effective Bone Suppression in High-Resolution Chest X-Ray Images With Conditional Latent Diffusion ModelsabstractLung diseases represent a significant global health challenge, with Chest X-Ray (CXR) being a key diagnostic tool due to its accessibility and affordability. Nonetheless, the detection of pulmonary lesions is often hindered by overlapping bone structures in CXR images, leading to potential misdiagnoses. To address this issue, we develop an end-to-end framework called BS-LDM, designed to effectively suppress bone in high-resolution CXR images. This framework is based on conditional latent diffusion models and incorporates a multi-level hybrid loss-constrained vector-quantized generative adversarial network which is crafted for perceptual compression, ensuring the preservation of details. To further enhance the framework's performance, we utilize offset noise in the forward process, and a temporal adaptive thresholding strategy in the reverse process. These additions help minimize discrepancies in generating low-frequency information of soft tissue images. Additionally, we have compiled a high-quality bone suppression dataset named SZCH-X-Rays. This dataset includes 818 pairs of high-resolution CXR and soft tissue images collected from our partner hospital. Moreover, we processed 241 data pairs from the JSRT dataset into negative images, which are more commonly used in clinical practice. Our comprehensive experiments and downstream evaluations reveal that BS-LDM excels in bone suppression, underscoring its clinical value. Yifei Sun 0005, Zhanghao Chen, Wenming Deng, Jin Liu 0012, Wenwen Min, Ahmed El-Azab, Changmiao Wang, Ruiquan Ge |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | ICH-PFNet: Prompt-Free Intracerebral Hemorrhage Segmentation via Convolutional Sparse Embeddings and Contrastive Semantic ConsistencyabstractIntracerebral hemorrhage (ICH) necessitates precise and efficient segmentation of hemorrhagic regions in head computed tomography (CT) scans to facilitate timely clinical decisions. To address challenges such as irregular shapes of hematomas, unclear lesion boundaries, and the scarcity of annotated data, we introduce the ICH-PFNet, a text-guided segmentation framework specifically designed for ICH imaging that operates without prompts. The Mamba Pyramid Downsampling module ensures robust multi-scale feature extraction, while the GCS-CLIP fusion mechanism enhances semantic consistency through batch-level contrastive similarity. The Enhanced SAM module provides automatic spatial guidance and convolution-based sparse embeddings to eliminate manual input. Furthermore, a Feature Pyramid Network combined with a Group Aggregation Bridge enhances multi-scale feature fusion and refines boundaries. Our model showed superior performance in segmenting small and structurally complex hemorrhages by using a private CT dataset. These results highlight its potential for integration into automated ICH assessment workflows. The code is available at https://github.com/Hzchzc123/ICH-CMNet. Chenxin Di, Qiwei Yang, Yaoqun Liu, Haoxuan Sun, Ahmed El-Azab, Changmiao Wang |
BIBM | 10 |
| 2025 | RTGMFF: Enhanced fMRI-Based Brain Disorder Diagnosis via ROI-Driven Text Generation and Multimodal Feature FusionabstractFunctional magnetic resonance imaging (fMRI) is a powerful tool for probing brain function, yet reliable clinical diagnosis is hampered by low signal-to-noise ratios, inter-subject variability, and the limited frequency awareness of prevailing CNN- and Transformer-based models. Moreover, most fMRI datasets lack textual annotations that could contextualize regional activation and connectivity patterns. We introduce RTGMFF, a framework that unifies automatic ROI-level text generation with multimodal feature fusion for brain-disorder diagnosis. RTGMFF consists of three components: (i) ROI-driven fMRI text generation deterministically condenses each subject's activation, connectivity, age, and sex into reproducible text tokens; (ii) Hybrid frequency-spatial encoder fuses a hierarchical waveletmamba branch with a cross-scale Transformer encoder to capture frequency-domain structure alongside long-range spatial dependencies; and (iii) Adaptive semantic alignment module embeds the ROI token sequence and visual features in a shared space, using a regularized cosine-similarity loss to narrow the modality gap. Extensive experiments on the ADHD-200 and ABIDE benchmarks show that RTGMFF surpasses current methods in diagnostic accuracy, achieving notable gains in sensitivity, specificity, and area under the ROC curve. Code is available at https://github.com/BeistMedAI/RTGMFF. Junhao Jia, Yifei Sun 0005, Yunyou Liu, Changmiao Wang, Fei-wei Qin, Yong Peng 0001, Wenwen Min |
BIBM | 5 |
| 2025 | DR-TTA: Dynamic and Robust Test-Time Adaptation Under Low-Quality Mri Conditions for Brain Tumor SegmentationabstractBrain tumor segmentation from low-quality MRI scans poses significant challenges, particularly in sub-Saharan Africa, where the scans frequently suffer from low resolution and artifacts. Such degradations introduce substantial domain shifts that hinder the effectiveness of existing test-time adaptation (TTA) methods, largely due to catastrophic forgetting and the unreliability of pseudo-labels. In response, we introduce DRTTA, a dynamic and robust framework designed for effective test-time adaptation. This method maintains essential knowledge from the source domain by freezing certain parameters and utilizing adaptive BatchNorm, allowing for successful alignment with the target domain. During inference, DR-TTA employs a learnable augmentation strategy that is optimized to simulate distortions specific to the target domain. Additionally, a hybrid loss function incorporating geometric constraints is used to filter out unreliable pseudo-labels, thus stabilizing the training process. Our extensive experiments on the BraTS-SSA and BraTS-SIM datasets demonstrate that DR-TTA significantly surpasses existing state-of-the-art methods across key performance metrics. This advancement provides a viable solution for deploying brain tumor segmentation technology in real-world scenarios, particularly within resource-limited environments. Our source code is available at https://github.com/baiyou1234/DR-TTA. Yuanhan Wang, Yifei Chen 0019, Wenjing Yu, Mingxuan Liu 0001, Beining Wu, Shenghao Zhu, Fei-wei Qin, Jin Fan 0003, Changmiao Wang |
BIBM | 10 |
| 2025 | RE-SAM2: Boosting Few-Shot Medical Image Segmentation via Reinforcement Learning and Ensemble LearningabstractDeep learning models for medical image segmentation often encounter difficulties when there is a lack of annotated data. While current few-shot segmentation methods have reduced these challenges, they frequently fail to fully utilize the information in the limited samples available. Additionally, they typically depend on large quantities of unlabeled data with pseudo-labels for domain adaptation. In response to these issues, we propose RE-SAM2, a novel framework for fewshot medical image segmentation that combines reinforcement learning with ensemble learning. The central concept involves retaining the reward model from reinforcement learning after training and integrating it into the model through ensemble learning techniques. Unlike previous methods, RE-SAM2 does not require extra unlabeled data and achieves notable improvements in segmentation accuracy with limited supervision. Experiments conducted on benchmark datasets reveal that RE-SAM2 surpasses current leading approaches. The code is available at the link https://github.com/zzzzz37/RE-SAM2. Shougan Teng, Wenwen Min, Changmiao Wang, Zhenbing Liu |
BIBM | 4 |
| 2025 | Toward Robust Early Detection of Alzheimer's Disease via an Integrated Multimodal Learning ApproachabstractAlzheimer’s Disease (AD) is a complex neurodegenerative disorder marked by memory loss, executive dysfunction, and personality changes. Early diagnosis is challenging due to subtle symptoms and varied presentations, often leading to misdiagnosis with traditional unimodal diagnostic methods due to their limited scope. This study introduces an advanced multimodal classification model that integrates clinical, cognitive, neuroimaging, and EEG data to enhance diagnostic accuracy. The model incorporates a feature tagger with a tabular data coding architecture and utilizes the TimesBlock module to capture intricate temporal patterns in Electroencephalograms (EEG) data. By employing Cross-modal Attention Aggregation module, the model effectively fuses Magnetic Resonance Imaging (MRI) spatial information with EEG temporal data, significantly improving the distinction between AD, Mild Cognitive Impairment, and Normal Cognition. Simultaneously, we have constructed the first AD classification dataset that includes three modalities: EEG, MRI, and tabular data. Our innovative approach aims to facilitate early diagnosis and intervention, potentially slowing the progression of AD. The source code and our private ADMC dataset are available at https://github.com/JustlfC03/MSTNet. Yifei Chen 0019, Shenghao Zhu, Zhaojie Fang, Chang Liu 0090, Binfeng Zou, Linwei Qiu, Shuo Chang, Fei-wei Qin, Jin Fan 0003, Yong Peng 0001, Changmiao Wang |
ICASSP | 13 |
| 2025 | 3D-Telepathy: Reconstructing 3D Objects from EEG Signals
Yuxiang Ge, Jionghao Cheng, Ruiquan Ge, Zhaojie Fang, Gangyong Jia, Nannan Li 0001, Ahmed El-Azab, Changmiao Wang |
IJCNN | 9 |
| 2025 | Clinical Prior Guided Cross-Modal Hierarchical Fusion for Histological Subtyping of Lung Cancer in CT Scans
Ahmed El-Azab, Songqi Zhang, Qinghua Liang, Danna Li, Ying Xiang, Changmiao Wang |
MICCAI (15) | 10 |
| 2025 | Endo-GSMT: Endoscopic Monocular Scene Reconstruction with Dynamic Gaussian Splatting and Motion Tracking
Hao Gou, Changmiao Wang, Yaoqun Liu, Fucang Jia, Deqiang Xiao, Fei-wei Qin, Huoling Luo |
MICCAI (9) | 2 |
| 2025 | GL-LCM: Global-Local Latent Consistency Models for Fast High-Resolution Bone Suppression in Chest X-Ray Images
Yifei Sun 0005, Zhanghao Chen, Yuqing Lu, Lixin Duan, Fenglei Fan, Ahmed El-Azab, Changmiao Wang, Ruiquan Ge |
MICCAI (13) | 9 |
| 2025 | Inferring Super-Resolved Gene Expression by Integrating Histology Images and Spatial Transcriptomics with HISTEX
Shuailin Xue, Changmiao Wang, Xiaomao Fan, Wenwen Min |
MICCAI (13) | 2 |
| 2025 | Small Lesions-aware Bidirectional Multimodal Multiscale Fusion Network for Lung Disease Classification
Jianxun Yu, Ruiquan Ge, Chenyu Lin, Xianjun Fu, Jikui Liu, Ahmed El-Azab, Changmiao Wang |
MICCAI (1) | 9 |
| 2025 | Bridging the Gap in Missing Modalities: Leveraging Knowledge Distillation and Style Matching for Brain Tumor Segmentation
Shenghao Zhu, Yifei Chen 0019, Yuanhan Wang, Chang Liu 0090, Fei-wei Qin, Changmiao Wang |
MICCAI (8) | 8 |
| 2025 | MT-WilmsNet: A Multi-level Transformer Fusion Network for Wilms' Tumor Segmentation and Metastasis Prediction
Wenjing Yu, Changmiao Wang |
MICCAI (4) | 7 |
| 2025 | Drawing2CAD: Sequence-to-Sequence Learning for CAD Generation from Vector DrawingsabstractComputer-Aided Design (CAD) generative modeling is driving significant innovations across industrial applications. Recent works have shown remarkable progress in creating solid models from various inputs such as point clouds, meshes, and text descriptions. However, these methods fundamentally diverge from traditional industrial workflows that begin with 2D engineering drawings. The automatic generation of parametric CAD models from these 2D vector drawings remains underexplored despite being a critical step in engineering design. To address this gap, our key insight is to reframe CAD generation as a sequence-to-sequence learning problem where vector drawing primitives directly inform the generation of parametric CAD operations, preserving geometric precision and design intent throughout the transformation process. We propose Drawing2CAD, a framework with three key technical components: a network-friendly vector primitive representation that preserves precise geometric information, a dual-decoder transformer architecture that decouples command type and parameter generation while maintaining precise correspondence, and a soft target distribution loss function accommodating inherent flexibility in CAD parameters. To train and evaluate Drawing2CAD, we create CAD-VGDrawing, a dataset of paired engineering drawings and parametric CAD models, and conduct thorough experiments to demonstrate the effectiveness of our method. Code and dataset are available at https://github.com/lllssc/Drawing2CAD. Fei-wei Qin, Shichao Lu, Junhao Hou, Changmiao Wang, Meie Fang, Ligang Liu 0001 |
ACM Multimedia | 4 |
| 2025 | CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ SegmentationabstractMulti-organ medical segmentation is a crucial component of medical image processing, essential for doctors to make accurate diagnoses and develop effective treatment plans. Despite significant progress in this field, current multi-organ segmentation models often suffer from inaccurate details, dependence on geometric prompts and loss of spatial information. Addressing these challenges, we introduce a novel model named CRISP-SAM2 with CR oss-modal Interaction and Semantic Prompting based on SAM2. This model represents a promising approach to multi-organ medical segmentation guided by textual descriptions of organs. Our method begins by converting visual and textual inputs into cross-modal contextualized semantics using a progressive cross-attention interaction mechanism. These semantics are then injected into the image encoder to enhance the detailed understanding of visual information. To eliminate reliance on geometric prompts, we use a semantic prompting strategy, replacing the original prompt encoder to sharpen the perception of challenging targets. In addition, a similarity-sorting self-updating strategy for memory and a mask-refining process is applied to further adapt to medical imaging and enhance localized details. Comparative experiments conducted on seven public datasets indicate that CRISP-SAM2 outperforms existing models. Extensive analysis also demonstrates the effectiveness of our method, thereby confirming its superior performance, especially in addressing the limitations mentioned earlier. Our code is available at: https://github.com/YU-deep/CRISP_SAM2.git. Changmiao Wang, Ahmed El-Azab, Gangyong Jia, Changqing Zou, Ruiquan Ge |
ACM Multimedia | 2 |
| 2025 | LPUWF-LDM: Enhanced latent diffusion model for precise late-phase UWF-FA generation on limited dataset
Zhaojie Fang, Guanyu Zhou, Ke Zhuang, Yifei Chen 0019, Ruiquan Ge, Changmiao Wang, Gangyong Jia, Qing Wu 0008, Juan Ye, Maimaiti Nuliqiman, Peifang Xu, Ahmed El-Azab |
Expert Syst. Appl. | 7 |
| 2025 | InfraFFN: A Feature Fusion Network leveraging dual-path convolution and self-attention for infrared image super-resolution
Fei-wei Qin, Ruiquan Ge, Kai Zhang 0008, Fei Lin 0006, Yeru Wang, Juan Manuel Górriz, Ahmed El-Azab, Changmiao Wang |
Knowl. Based Syst. | 9 |
| 2025 | ICH-PRNet: a cross-modal intracerebral haemorrhage prognostic prediction method using joint-attention interaction mechanism
Ahmed El-Azab, Ruiquan Ge, Jichao Zhu, Gangyong Jia, Qing Wu 0008, Changmiao Wang |
Neural Networks | 10 |
| 2025 | Frequency-Domain Convolutional Network With Historical Data Fusion Module for Regional Streamflow PredictionabstractAccurate runoff prediction is essential for effective water resource management, particularly in addressing flood control and monitoring drought conditions. However, the diverse nature of land types and varying climate conditions often complicate this task, requiring frequent adaptations to prediction models for local applications. Existing methods primarily focus on modeling for individual regions, while regional runoff prediction models cannot often learn long-term patterns, limiting their regional adaptability. To overcome this challenge, we present the temporal fusion runoff network (TFRN), a new framework designed to enhance long short-term memory (LSTM) models by enabling them to incorporate distant historical information. This innovation offers a promising framework for regional runoff prediction by enhancing model performance and minimizing computational demands. In this study, the proposed TFRN utilizes convolutional networks to extract and integrate both long-term and short-term trends from input sequences, and by merging the strengths of LSTM and Transformer architectures, TFRN achieves a thorough integration of historical data. Specifically, our method employs convolutional networks across both time and frequency domains to capture multi-scale features. Within the Transformer component, we introduce an adaptive fusion module to improve the integration of historical information. We validated the effectiveness of our model using two extensive hydrological datasets for a 7-day runoff prediction task. The results underscore the superiority of our approach, demonstrating its advantages over several leading methods. The source code is available at https://github.com/redtea-code/TFRN. Yuanhao Chen, Haoqi Yu, Jingrong Dai, Nannan Li 0001, Changmiao Wang, Ahmed El-Azab |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | SCKansformer: Fine-Grained Classification of Bone Marrow Cells via Kansformer Backbone and Hierarchical Attention MechanismsabstractThe incidence and mortality rates of malignant tumors, such as acute leukemia, have risen significantly. Clinically, hospitals rely on cytological examination of peripheral blood and bone marrow smears to diagnose malignant tumors, with accurate blood cell counting being crucial. Existing automated methods face challenges such as low feature expression capability, poor interpretability, and redundant feature extraction when processing high-dimensional microimage data. We propose a novel fine-grained classification model, SCKansformer, for bone marrow blood cells, which addresses these challenges and enhances classification accuracy and efficiency. The model integrates the Kansformer Encoder, SCConv Encoder, and Global-Local Attention Encoder. The Kansformer Encoder replaces the traditional MLP layer with the KAN, improving nonlinear feature representation and interpretability. The SCConv Encoder, with its Spatial and Channel Reconstruction Units, enhances feature representation and reduces redundancy. The Global-Local Attention Encoder combines Multi-head Self-Attention with a Local Part module to capture both global and local features. We validated our model using the Bone Marrow Blood Cell Fine-Grained Classification Dataset (BMCD-FGCD), comprising over 10,000 samples and nearly 40 classifications, developed with a partner hospital. Comparative experiments on our private dataset, as well as the publicly available PBC and ALL-IDB datasets, demonstrate that SCKansformer outperforms both typical and advanced microcell classification methods across all datasets. Yifei Chen 0019, Shenghao Zhu, Linwei Qiu, Binfeng Zou, Chenyan Zhang, Zhaojie Fang, Fei-wei Qin, Jin Fan 0003, Changmiao Wang |
IEEE J. Biomed. Health Informatics | 12 |
| 2024 | Masked Conditional Diffusion Model with GNN for Spatial Transcriptomics Data ImputationabstractSpatially resolved transcriptomics represents a significant advancement in single-cell analysis by offering both gene expression data and their corresponding physical locations. However, this high degree of spatial resolution entails a drawback, as the resulting spatial transcriptomic data at the cellular level is notably plagued by a high incidence of missing values. Furthermore, most existing imputation methods either overlook the spatial information between spots or compromise the overall gene expression data distribution. To address these challenges, our primary focus is on effectively utilizing the spatial location information within spatial transcriptomic data to impute missing values, while preserving the overall data distribution. We introduce stMCDI, a masked conditional diffusion model for spatial transcriptomics data imputation, which employs a denoising network trained using randomly masked data portions as guidance, with the unmasked data serving as conditions. Additionally, it utilizes a GNN encoder to integrate the spatial position information, thereby enhancing model performance. Compared with baseline methods, our model achieves state-of-the-art performance in all evaluation metrics on six real-world datasets. The results obtained from spatial transcriptomics datasets elucidate the performance of our methods relative to existing approaches. Our code can be accessed at https://github.com/wenwenmin/stMCDI. Wenwen Min, Shunfang Wang, Changmiao Wang, Taosheng Xu |
BIBM | 4 |
| 2024 | PGP: Prior-Guided Pretraining for Small-sample Esophageal Cancer SegmentationabstractTransformer-based models have demonstrated substantial potential in medical image segmentation tasks due to their exceptional ability to capture long-range dependencies. To further enhance segmentation performance, various effective methods have been proposed, including pretraining methods (weakly supervised or self-supervised pretraining schemes), contrastive learning schemes, and knowledge distillation methods. However, segmenting esophageal cancer (EC) from CT images remains a significant challenge, partly due to the complex anatomy of EC, such as variable shapes, extensive extents, and often blurred boundaries with adjacent anatomical structures. In this study, we propose a prior-guided pretraining (PGP) regimen based on bounding boxes, which enhances the model’s ability to discern textural differences between EC and the surrounding tissues. Using Swin UNITR as the backbone, our proposed pretraining scheme demonstrates superior performance in EC segmentation compared to other schemes. To further improve the segmentation accuracy of EC, we also addressed the class imbalance and long-tail problems inherent in EC segmentation, thereby further enhancing segmentation performance. Qinglei Shi, Wenhan Duan, Haochen Lu, Kecan Wu, Junxi Zhu, Juefei Yuan, Qiyan Ke, Andu Zhang, Changmiao Wang, Renzhi Wang 0002 |
BIBM | 12 |
| 2024 | CCLNet: Causal and Contrastive Learning Framework for Enhanced Pulmonary Embolism DetectionabstractThe fusion of multimodal medical data is crucial for helping doctors make accurate treatment decisions. For example, combining Computed Tomography Pulmonary Angiography (CTPA) with Electronic Health Records (EHR) can significantly improve the accuracy of Pulmonary Embolism (PE) detection, thereby increasing patient survival rates. Although multimodal learning has advantages in PE diagnosis, the heterogeneity of multimodal data poses a significant challenge to accurate diagnosis. The natural semantic and structural differences between data modalities make it difficult to effectively integrate their information. In addition, within a single modality, the existence of redundant and irrelevant information introduces unnecessary variability, making the data more complex, and making stable diagnosis challenging. To address these issues, we propose a new framework called CCLNet, which includes a contrastive learning component for addressing inter-modality heterogeneity and a causal learning component for handling intra-modality heterogeneity. Specifically, we achieve precise alignment between visual and tabular modalities by using global-level information to soften labels during contrastive learning. In addition, by using causal intervention methods to eliminate the influence of heterogeneous factors within the modality, we can accurately reveal the causal relationship between features and targets, thereby improving the accuracy and stability of the model. Experimental results demonstrate that our method performs excellently, achieving the best results. Our code is available at https://github.com/LeavingStarW/CLPE. Ruiquan Ge, Jianxun Yu, Fei-wei Qin, Nannan Li 0001, Wenwen Min, Ahmed El-Azab, Changmiao Wang |
BIBM | 9 |
| 2024 | ICH-SCNet: Intracerebral Hemorrhage Segmentation and Prognosis Classification Network Using CLIP-guided SAM mechanismabstractIntracerebral hemorrhage (ICH) is the most fatal subtype of stroke and is characterized by a high incidence of disability. Accurate segmentation of the ICH region and prognosis prediction are critically important for developing and refining treatment plans for post-ICH patients. However, existing approaches address these two tasks independently and predominantly focus on imaging data alone, thereby neglecting the intrinsic correlation between the tasks and modalities. This paper introduces a multi-task network, ICH-SCNet, designed for both ICH segmentation and prognosis classification. Specifically, we integrate a SAM-CLIP cross-modal interaction mechanism that combines medical text and segmentation auxiliary information with neuroimaging data to enhance cross-modal feature recognition. Additionally, we develop an effective feature fusion module and a multi-task loss function to improve performance further. Extensive experiments on an ICH dataset reveal that our approach surpasses other state-of-the-art methods. It excels in the overall performance of classification tasks and outperforms competing models in all segmentation task metrics. Ahmed El-Azab, Ruiquan Ge, Xinchen Jiang, Gangyong Jia, Qing Wu 0008, Qinglei Shi, Changmiao Wang |
BIBM | 9 |
| 2024 | Infrared Image Super-Resolution via Lightweight Information Split Network
Fei-wei Qin, Changmiao Wang, Ruiquan Ge, Kai Zhang 0008, Yong Peng 0001 |
ICIC (8) | 4 |
| 2024 | Make an Image Move: Few-Shot Based Video Generation Guided by CLIP
Yonglong Huang, Nannan Li 0001, Fuqin Deng, Ruiquan Ge, Changmiao Wang |
ICPR (6) | 6 |
| 2024 | CircMAN: Multi-channel Attention Networks Based on Feature Fusion for CircRNA-Binding Protein Site Prediction
Huiliang Luo, Guojian Deng, Riqian Hu, Ruiquan Ge, Fei-wei Qin, Changmiao Wang |
ISBRA (1) | 6 |
| 2024 | Spatial Gene Expression Prediction from Histology Images with STco
Zhiceng Shi, Changmiao Wang, Wenwen Min |
ISBRA (1) | 3 |
| 2024 | stEnTrans: Transformer-Based Deep Learning for Spatial Transcriptomics Enhancement
Shuailin Xue, Changmiao Wang, Wenwen Min |
ISBRA (1) | 3 |
| 2024 | Cache-Driven Spatial Test-Time Adaptation for Cross-Modality Medical Image Segmentation
Xiang Li 0115, Huihui Fang, Changmiao Wang, Mingsi Liu, Lixin Duan, Yanwu Xu 0001 |
MICCAI (11) | 3 |
| 2024 | SCUNet++: Swin-UNet and CNN Bottleneck Hybrid Architecture with Multi-Fusion Dense Skip Connection for Pulmonary Embolism CT Image SegmentationabstractPulmonary embolism (PE) is a prevalent lung disease that can lead to right ventricular hypertrophy and failure in severe cases, ranking second in severity only to myocardial infarction and sudden death. Pulmonary artery CT angiography (CTPA) is a widely used diagnostic method for PE. However, PE detection presents challenges in clinical practice due to limitations in imaging technology. CTPA can produce noises similar to PE, making confirmation of its presence time-consuming and prone to overdiagnosis. Nevertheless, the traditional segmentation method of PE can not fully consider the hierarchical structure of features, local and global spatial features of PE CT images. In this paper, we propose an automatic PE segmentation method called SCUNet++ (Swin Conv UNet++). This method incorporates multiple fusion dense skip connections between the encoder and decoder, utilizing the Swin Transformer as the encoder. And fuses features of different scales in the decoder subnetwork to compensate for spatial information loss caused by the inevitable downsampling in Swin-UNet or other state-of-the-art methods, effectively solving the above problem. We provide a theoretical analysis of this method in detail and validate it on publicly available PE CT image datasets FUMPE and CAD-PE. The experimental results indicate that our proposed method achieved a Dice similarity coefficient (DSC) of 83.47% and a Hausdorff distance 95th percentile (HD95) of 3.83 on the FUMPE dataset, as well as a DSC of 83.42% and an HD95 of 5.10 on the CAD-PE dataset. These findings demonstrate that our method exhibits strong performance in PE segmentation tasks, potentially enhancing the accuracy of automatic segmentation of PE and providing a powerful diagnostic tool for clinical physicians. Our source code and new FUMPE dataset are available at https://github.com/JustlfC03/SCUNet-plusplus. Yifei Chen 0019, Binfeng Zou, Zhaoxin Guo, Yiyu Huang, Fei-wei Qin, Qinhai Li, Changmiao Wang |
WACV | 8 |
| 2024 | Multimodal contrastive learning for spatial gene expression prediction using histology imagesabstractIn recent years, the advent of spatial transcriptomics (ST) technology has unlocked unprecedented opportunities for delving into the complexities of gene expression patterns within intricate biological systems. Despite its transformative potential, the prohibitive cost of ST technology remains a significant barrier to its widespread adoption in large-scale studies. An alternative, more cost-effective strategy involves employing artificial intelligence to predict gene expression levels using readily accessible whole-slide images stained with Hematoxylin and Eosin (H&E). However, existing methods have yet to fully capitalize on multimodal information provided by H&E images and ST data with spatial location. In this paper, we propose mclSTExp, a multimodal contrastive learning with Transformer and Densenet-121 encoder for Spatial Transcriptomics Expression prediction. We conceptualize each spot as a "word", integrating its intrinsic features with spatial context through the self-attention mechanism of a Transformer encoder. This integration is further enriched by incorporating image features via contrastive learning, thereby enhancing the predictive capability of our model. We conducted an extensive evaluation of highly variable genes in two breast cancer datasets and a skin squamous cell carcinoma dataset, and the results demonstrate that mclSTExp exhibits superior performance in predicting spatial gene expression. Moreover, mclSTExp has shown promise in interpreting cancer-specific overexpressed genes, elucidating immune-related genes, and identifying specialized spatial domains annotated by pathologists. Our source code is available at https://github.com/shizhiceng/mclSTExp. Wenwen Min, Zhiceng Shi, Jun Wan 0005, Changmiao Wang |
Briefings Bioinform. | 5 |
| 2024 | Alzheimer's disease diagnosis from single and multimodal data using machine and deep learning models: Achievements and future directions
Ahmed El-Azab, Changmiao Wang, Mohammed Abdelaziz, Jason Gu, Juan Manuel Górriz, Yudong Zhang 0001, Chunqi Chang |
Expert Syst. Appl. | 2 |
| 2024 | LKFormer: large kernel transformer for infrared image super-resolution
Fei-wei Qin, Changmiao Wang, Ruiquan Ge, Yong Peng 0001, Kai Zhang 0008 |
Multim. Tools Appl. | 3 |
| 2024 | TDFFM: Transformer and Deep Forest Fusion Model for Predicting Coronavirus 3C-Like Protease Cleavage SitesabstractCOVID-19, caused by the highly contagious SARS-CoV-2 virus, is distinguished by its positive-sense, single-stranded RNA genome. A thorough understanding of SARS-CoV-2 pathogenesis is crucial for halting its proliferation. Notably, the 3C-like protease of the coronavirus (denoted as$3CL^{pro}$) is instrumental in the viral replication process. Precise delineation of$3CL^{pro}$cleavage sites is imperative for elucidating the transmission dynamics of SARS-CoV-2. While machine learning tools have been deployed to identify potential$3CL^{pro}$cleavage sites, these existing methods often fall short in terms of accuracy. To improve the performances of these predictions, we propose a novel analytical framework, the Transformer and Deep Forest Fusion Model (TDFFM). Within TDFFM, we utilize the AAindex and the BLOSUM62 matrix to encode protein sequences. These encoded features are subsequently input into two distinct components: a Deep Forest, which is an effective decision tree ensemble methodology, and a Transformer equipped with a Multi-Level Attention Model (TMLAM). The integration of the attention mechanism allows our model to more accurately identify positive samples, thus enhancing the overall predictive performance. Evaluation on a test set demonstrates that our TDFFM achieves an accuracy of 0.955, an AUC of 0.980, and an F1-score of 0.367, substantiating the model's superior prediction capabilities. Ruiquan Ge, Changmiao Wang, Ahmed El-Azab, Qiming Fang, Renfeng Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | UWAFA-GAN: Ultra-Wide-Angle Fluorescein Angiography Transformation via Multi-Scale Generation and Registration EnhancementabstractFundus photography, in combination with the ultra-wide-angle fundus (UWF) techniques, becomes an indispensable diagnostic tool in clinical settings by offering a more comprehensive view of the retina. Nonetheless, UWF fluorescein angiography (UWF-FA) necessitates the administration of a fluorescent dye via injection into the patient's hand or elbow unlike UWF scanning laser ophthalmoscopy (UWF-SLO). To mitigate potential adverse effects associated with injections, researchers have proposed the development of cross-modality medical image generation algorithms capable of converting UWF-SLO images into their UWF-FA counterparts. Current image generation techniques applied to fundus photography encounter difficulties in producing high-resolution retinal images, particularly in capturing minute vascular lesions. To address these issues, we introduce a novel conditional generative adversarial network (UWAFA-GAN) to synthesize UWF-FA from UWF-SLO. This approach employs multi-scale generators and an attention transmit module to efficiently extract both global structures and local lesions. Additionally, to counteract the image blurriness issue that arises from training with misaligned data, a registration module is integrated within this framework. Our method performs non-trivially on inception scores and details generation. Clinical user studies further indicate that the UWF-FA images generated by UWAFA-GAN are clinically comparable to authentic images in terms of diagnostic reliability. Empirical evaluations on our proprietary UWF image datasets elucidate that UWAFA-GAN outperforms extant methodologies. Ruiquan Ge, Zhaojie Fang, Pengxue Wei, Zhanghao Chen, Hongyang Jiang 0001, Ahmed El-Azab, Wangting Li, Shaochong Zhang, Changmiao Wang |
IEEE J. Biomed. Health Informatics | 10 |
| 2023 | GCS-ICHNet: Assessment of Intracerebral Hemorrhage Prognosis using Self-Attention with Domain Knowledge IntegrationabstractIntracerebral Hemorrhage (ICH) is a severe condition resulting from damaged brain blood vessel ruptures, often leading to complications and fatalities. Timely and accurate prognosis and management are essential due to its high mortality rate. However, conventional methods heavily rely on subjective clinician expertise, which can lead to inaccurate diagnoses and delays in treatment. Artificial intelligence (AI) models have been explored to assist clinicians, but many prior studies focused on model modification without considering domain knowledge. This paper introduces a novel deep learning algorithm, GCS-ICHNet, which integrates multimodal brain CT image data and the Glasgow Coma Scale (GCS) score to improve ICH prognosis. The algorithm utilizes a transformer-based fusion module for assessment. GCS-ICHNet demonstrates high sensitivity 81.03% and specificity 91.59%, outperforming average clinicians and other state-of-the-art methods. The code is available at https://github.com/Windbelll/Prognosis-analysis-of-cerebral-hemorrhage. Xuhao Shan, Ruiquan Ge, Shibin Wu, Ahmed El-Azab, Jichao Zhu, Gangyong Jia, Qingying Xiao, Changmiao Wang |
BIBM | 11 |
| 2023 | UWAT-GAN: Fundus Fluorescein Angiography Synthesis via Ultra-Wide-Angle Transformation Multi-scale GAN
Zhaojie Fang, Zhanghao Chen, Pengxue Wei, Wangting Li, Shaochong Zhang, Ahmed El-Azab, Gangyong Jia, Ruiquan Ge, Changmiao Wang |
MICCAI (7) | 9 |
| 2021 | Hepatocellular Carcinoma Segmentation from Digital Subtraction Angiography Videos Using Learnable Temporal Difference
Wenting Jiang, Lu Zhang 0051, Changmiao Wang, Xiaoguang Han 0001, Shuixing Zhang, Shuguang Cui |
MICCAI (5) | 4 |
| 2020 | GP-GAN: Brain tumor growth prediction using stacked 3D generative adversarial networks from longitudinal MR Images
Ahmed El-Azab, Changmiao Wang, Syed Jamal Safdar Gardezi, Hongmin Bai, Qingmao Hu, Tianfu Wang 0001, Chunqi Chang, Bai Ying Lei |
Neural Networks | 2 |