Wei Wang 0169

dblp:35/7092-169 · DBLP profile ↗
← Back
60ranked-venue papers
3as first author
51since 2021 · last 2026
0000-0002-1874-9947ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 3 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 19 since 2021Artificial intelligence and machine learning · 19 · 17 since 2021
YearPublicationVenuePosition
2026 Ambiguity-aware Truncated Flow Matching for Ambiguous Medical Image Segmentation
abstract
A simultaneous enhancement of accuracy and diversity of predictions remains a challenge in ambiguous medical image segmentation (AMIS) due to the inherent trade-offs. While truncated diffusion probabilistic models (TDPMs) hold strong potential with a paradigm optimization, existing TDPMs suffer from entangled accuracy and diversity of predictions with insufficient fidelity and plausibility. To address the aforementioned challenges, we propose Ambiguity-aware Truncated Flow Matching (ATFM), which introduces a novel inference paradigm and dedicated model components. Firstly, we propose Data-Hierarchical Inference, a redefinition of AMIS-specific inference paradigm, which enhances accuracy and diversity at data-distribution and data-sample level, respectively, for an effective disentanglement. Secondly, Gaussian Truncation Representation (GTR) is introduced to enhance both fidelity of predictions and reliability of truncation distribution, by explicitly modeling it as a Gaussian distribution at Ttrunc instead of using sampling-based approximations. Thirdly, Segmentation Flow Matching (SFM) is proposed to enhance the plausibility of diverse predictions by extending semantic-aware flow transformation in Flow Matching (FM). Comprehensive evaluations on LIDC and ISIC3 datasets demonstrate that ATFM outperforms SOTA methods and simultaneously achieves a more efficient inference. ATFM improves GED and HM-IoU by up to 12% and 7.3% compared to advanced methods.
Fanding Li, Xiangyu Li 0004, Xianghe Su, Xingyu Qiu, Suyu Dong, Wei Wang 0169, Kuanquan Wang, Gongning Luo, Shuo Li 0001
AAAI6
2026 Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease Diagnosis
abstract
Multimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness.
Haoran Li 0024, Haoyu Cao 0002, Yongting Hu, Qihao Xu, Chengliang Liu 0003, Xiaoling Luo 0001, Zhihao Wu 0002, Yong Xu 0001, Wei Wang 0169
AAAI10
2026 SCULPT: Semantic-aware causal prompt tuning for out-of-distribution detection of whole slide images
Pengzhong Sun, Xiangyu Li 0004, Dong Liang 0001, Jun Liu 0080, Zhanshi Zhu, Xiaokun Li, Suyu Dong, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
Knowl. Based Syst.9
2026 A Novel Multi-Perspective Framework for Molecule Pretraining: From Atom to Motif Views
abstract
Predicting molecular properties is vital for drug discovery, but experimental measurement is costly and limited by scarce labeled data. Self-supervised molecular pretraining can leverage large unlabeled datasets, reducing dependence on extensive annotations. However, most methods struggle to preserve domain-specific chemical knowledge, especially clinically relevant substructures such as motifs. Random masking and generic graph augmentations often degrade critical chemical information and harm interpretability. Many approaches also work at a single scale-either atom or motif-missing opportunities for cross-scale integration. We propose A2M-Mol, a multi-perspective molecular pretraining framework that combines atom-level and motif-level views through four parallel graph constructions. This design enables cross-view alignment and multiscale fusion, explicitly encoding chemical knowledge. A2M-Mol employs a suite of self-supervised tasks, including cross-view correspondence, atomic reconstruction, global topology modeling, and property constraint enforcement, all coordinated via tailored contrastive learning. Extensive experiments across benchmarks and backbone architectures show consistent improvements over state-of-the-art methods. Ablation studies confirm strong synergies among the tasks. A2M-Mol maintains robust predictive accuracy across data scales, demonstrating effectiveness for real-world molecular property prediction and potential to accelerate drug discovery.
Wei Wang 0169, Dengzhen Lu, Suyu Dong, Gongning Luo, Kuanquan Wang, Shanzhuo Zhang
IEEE J. Biomed. Health Informatics1
2026 Causality-Adjusted Data Augmentation for Domain Continual Medical Image Segmentation
abstract
In domain continual medical image segmentation, distillation-based methods mitigate catastrophic forgetting by continuously reviewing old knowledge. However, these approaches often exhibit biases towards both new and old knowledge simultaneously due to confounding factors, which can undermine segmentation performance. To address these biases, we propose the Causality-Adjusted Data Augmentation (CauAug) framework, introducing a novel causal intervention strategy called the Texture-Domain Adjustment Hybrid-Scheme (TDAHS) alongside two causality-targeted data augmentation approaches: the Cross Kernel Network (CKNet) and the Fourier Transformer Generator (FTGen). (1) TDAHS establishes a domain-continual causal model that accounts for two types of knowledge biases by identifying irrelevant local textures (L) and domain-specific features (D) as confounders. It introduces a hybrid causal intervention that combines traditional confounder elimination with a proposed replacement approach to better adapt to domain shifts, thereby promoting causal segmentation. (2) CKNet eliminates confounder L to reduce biases in new knowledge absorption. It decreases reliance on local textures in input images, forcing the model to focus on relevant anatomical structures and thus improving generalization. (3) FTGen causally intervenes on confounder D by selectively replacing it to alleviate biases that impact old knowledge retention. It restores domain-specific features in images, aiding in the comprehensive distillation of old knowledge. Our experiments show that CauAug significantly mitigates catastrophic forgetting and surpasses existing methods in various medical image segmentation tasks.
Zhanshi Zhu, Gongning Luo, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Guohua Wang 0001, Shuo Li 0001
IEEE J. Biomed. Health Informatics4
2026 TKRL: Targeted Knowledge Rectification Learning Against Teacher-Originated Defects in Domain Continual Segmentation
abstract
Knowledge distillation can mitigate catastrophic forgetting in domain continual segmentation by transferring knowledge from the older model to the newer model. However, existing distillation-based methods primarily emphasize knowledge retention while overlooking inherent defects in the older teacher models. As a result, these teacher-originated defects, such as knowledge gaps or biases, are propagated and exacerbate forgetting. To address this challenge, we propose a Targeted Knowledge Rectification Learning framework (TKRL) to probe and correct teacher-originated defects. TKRL consists of two modules: 1) Probe-augmented Class Distillation, which generates gradient-driven "probes" to uncover underrepresented features in the older model, thereby bridging knowledge gaps by distilling hidden information into the new model; 2) Variance-guided Masked Autoencoder, which selectively masks and reconstructs critical high-uncertainty patches across multi-level semantic regions, thereby correcting biases inherited from the older model. Our experimental results show that TKRL effectively rectifies knowledge gaps and biases, thereby mitigating catastrophic forgetting and enhancing performance in domain continual segmentation.
Zhanshi Zhu, Wenjian Gu, Xiangyu Li 0004, Qince Li, Yongfeng Yuan, Wei Wang 0169, Kuanquan Wang, Suyu Dong, Shuo Li 0001
IEEE J. Biomed. Health Informatics6
2025 Multi-view Evidential Learning-based Medical Image Segmentation
abstract
Medical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limited embedded knowledge scope. Vision foundation models have been demonstrated to be effective in extracting generalizable knowledge, but they cannot extract domain-specific knowledge without fine-tuning. In this work, we propose a novel multi-view evidential learning-based framework, which can extract both domain-specific and generalizable knowledge from multi-view features by combining the advantages of traditional and vision foundation models. Specifically, a novel multi-view state space model (MV-SSM) is designed to extract task-related knowledge while removing redundant information within multi-view features. The proposed MV-SSM utilizes Mamba, a state space model, to model cross-view contextual dependencies between domain-specific and generalizable features. Additionally, evidential learning is adopted to quantify the segmentation uncertainty of the model for boundary. In special, variational Dirichlet is introduced to characterize the distribution of the result probabilities, parameterized with collected evidence to quantify uncertainty. As a result, the model can reduce the segmentation uncertainties of boundaries by optimizing the parameters of the Dirichlet distribution. Experimental results on three datasets show that our method obtains superior segmentation performance.
Chao Huang 0008, Yushu Shi, Wai Keung Wong, Chengliang Liu 0003, Wei Wang 0169, Zhihua Wang 0002, Jie Wen 0001
AAAI5
2025 Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus Diseases
abstract
With the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper addresses these challenges by conceptualizing them as multi-level feature fusion and self-supervised disease-indicative feature learning problems. We decode fundus images at various levels of granularity to delineate scenarios wherein multiple diseases and lesions co-occur. To effectively integrate these features, we introduce a hierarchical vision transformer (HVT) that adeptly captures both inter-level and intra-level dependencies. A novel forward-attention module is proposed to enhance the integration of lower-level semantic information into higher semantic layers, thereby enriching the representation of complex features. Additionally, we introduce a novel self-supervised mask-consistent feature learner (MCFL). Unlike traditional mask-autoencoders that reconstruct original images using encoder-decoder structures, MCFL utilizes a teacher-student framework to reconstruct mask-consistent feature maps. In this setup, exponential moving averaging is employed to derive classification-guided features, serving as labels for reconstruction rather than merely reconstructing the original images. This innovative approach facilitates the extraction of disease-indicative features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art models.
Wei Wang 0169, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001
AAAI2
2025 A Trusted Lesion-assessment Network for Interpretable Diagnosis of Coronary Artery Disease in Coronary CT Angiography
abstract
Coronary Artery Disease (CAD) poses a significant threat to cardiovascular patients worldwide, underscoring the critical importance of automated CAD diagnostic technologies in clinical practice. Previous technologies for lesion assessment in Coronary CT Angiography (CCTA) images have been insufficient in terms of interpretability, resulting in solutions that lack clinical reliability in both network architecture and prediction outcomes, even when diagnoses are accurate. To address the limitation of interpretability, we introduce the Trusted Lesion-Assessment Network (TLA-Net), which provides a clinically reliable solution for multi-view CAD diagnosis: (1) The causality-informed evidence collection constructs a causal graph for the diagnostic process and implements causal interventions, preventing confounders' interference and enhancing the transparency of the network architecture. (2) The clinically-aligned uncertainty integration hierarchically combines Dirichlet distributions from various views based on clinical priors, offering confidence coefficients for prediction outcomes that align with physicians' image analysis procedures. Experimental results on a dataset of 2,618 lesions demonstrate that TLA-Net, supported by its interpretable methodological design, exhibits superior performance with outstanding generalization, domain adaptability, and robustness.
Xinghua Ma, Xinyan Fang, Mingye Zou, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Zhaowen Qiu, Xin Gao 0001, Shuo Li 0001
AAAI5
2025 Federated Weakly Supervised Video Anomaly Detection with Multimodal Prompt
abstract
Video anomaly detection (VAD) aims at locating the abnormal events in videos. Recently, the Weakly Supervised VAD has made great progress, which only requires video-level annotations when training. In practical applications, different institutions may have different types of abnormal videos. However, the abnormal videos cannot be circulated on the internet due to privacy protection. To train a more generalized anomaly detector that can identify various anomalies, it is reasonable to introduce federated learning into WSVAD. In this paper, we propose Global and Local Context-driven Federated Learning, a new paradigm for privacy protected weakly supervised video anomaly detection. Specifically, we utilize the vision-language association of CLIP to detect whether the video frame is abnormal. Instead of leveraging handcrafted text prompts for CLIP, we propose a text prompt generator. The generated prompt is simultaneously influenced by text and visual. On the one hand, the text provides global context related to anomaly, which improves the model's ability of generalization. On the other hand, the visual provides personalized local context because different clients may have videos with different types of anomalies or scenes. The generated prompt ensures global generalization while processing personalized data from different clients. Extensive experiments show that the proposed method achieves remarkable performance.
Benfeng Wang, Chao Huang 0008, Jie Wen 0001, Wei Wang 0169, Yong Xu 0001
AAAI4
2025 Finding Local Diffusion Schrodinger Bridge using Kolmogorov-Arnold Network
abstract
In image generation, Schrödinger Bridge (SB)-based methods theoretically enhance the efficiency and quality compared to the diffusion models by finding the least costly path between two distributions. However, they are computationally expensive and time-consuming when applied to complex image data. The reason is that they focus on fitting globally optimal paths in high-dimensional spaces, directly generating images as next step on the path using complex networks through self-supervised training, which typically results in a gap with the global optimum. Meanwhile, most diffusion models are in the same path subspace generated by weights fA(t) and fB(t), as they follow the paradigm (xt= fA(t)xImg+ fB(t)ϵ). To address the limitations of SB-based methods, this paper proposes for the first time to find local Diffusion Schrödinger Bridges (LDSB) in the diffusion path subspace, which strengthens the connection between the SB problem and diffusion models. Specifically, our method optimizes the diffusion paths using Kolmogorov-Arnold Network (KAN), which has the advantage of resistance to forgetting and continuous output. The experiment shows that our LDSB significantly improves the quality and efficiency of image generation using the same pretrained denoising network and the KAN for optimising is only less than 0.1MB. The FID metric is reduced by more than 15%, especially with a reduction of 48.50% when NFE of DDIM is 5 for the CelebA dataset. Code is available at https://github.com/PerceptionComputingLab/LDSB.
Xingyu Qiu, Mengying Yang, Xinghua Ma, Fanding Li, Dong Liang 0001, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
CVPR7
2025 Ex-VAD: Explainable Fine-grained Video Anomaly Detection Based on Visual-Language Models
abstract
With advancements in visual language models (VLMs) and large language models (LLMs), video anomaly detection (VAD) has progressed beyond binary classification to fine-grained categorization and multidimensional analysis. However, existing methods focus mainly on coarse-grained detection, lacking anomaly explanations. To address these challenges, we propose Ex-VAD, an Explainable Fine-grained Video Anomaly Detection approach that combines fine-grained classification with detailed explanations of anomalies. First, we use a VLM to extract frame-level captions, and an LLM converts them to video-level explanations, enhancing the model's explainability. Second, integrating textual explanations of anomalies with visual information greatly enhances the model's anomaly detection capability. Finally, we apply label-enhanced alignment to optimize feature fusion, enabling precise fine-grained detection. Extensive experimental results on the UCF-Crime and XD-Violence datasets demonstrate that Ex-VAD significantly outperforms existing State-of-The-Art methods.
Chao Huang 0008, Yushu Shi, Jie Wen 0001, Wei Wang 0169, Yong Xu 0001, Xiaochun Cao
ICML4
2025 Deep Opinion-Unaware Blind Image Quality Assessment by Learning and Adapting from Multiple Annotators
abstract
Existing deep neural network (DNN)-based blind image quality assessment (BIQA) methods primarily rely on human-rated datasets for training. However, collecting human labels is extremely time-consuming and labor-intensive, posing a significant bottleneck for practical applications. To address this challenge, we propose a Deep opinion-Unaware BIQA model by learning and adapting from Multiple Annotators, termed DUBMA, thereby eliminating the need for human annotations. Specifically, we first generate a large-scale set of distorted image pairs and then assign relative quality rankings using existing full-reference IQA models. The resulting dataset is subsequently employed for training our DUBMA. Due to the inherent discrepancies between synthetic and real-world distortions, a domain shift may occur. To address this, we propose an outlier-robust unsupervised domain adaptation approach leveraging optimal transport. This strategy effectively reduces the gap between synthetic and real-world distortion domains, thereby boosting the model’s adaptability and overall performance. Extensive experiments show that DUBMA outperforms existing opinion-unaware BIQA methods in terms of prediction accuracy across multiple datasets.
Zhihua Wang 0002, Xuelin Liu, Jiebin Yan, Jie Wen 0001, Wei Wang 0169, Chao Huang 0008
IJCAI5
2025 Structure and Smoothness Constrained Dual Networks for MR Bias Field Correction
Dong Liang 0001, Xingyu Qiu, Wei Wang 0169, Kuanquan Wang, Suyu Dong, Gongning Luo
MICCAI (13)4
2025 Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
abstract
Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLMs to reason about anomalies step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks.
Chao Huang 0008, Benfeng Wang, Wei Wang 0169, Jie Wen 0001, Chengliang Liu 0003, Li Shen 0008, Xiaochun Cao
NeurIPS3
2025 TCRdesign: an antigen-specific generative language model for de novo design of T-cell receptors
abstract
T-cell receptors (TCR), which are heterodimers of $\alpha $ and $\beta $ chains that recognize foreign antigens, are of great significance to current immunotherapy. Although artificial intelligence (AI) has explosively accelerated de novo protein design, the challenge of therapeutic TCR design has been overlooked by most researchers. Existing TCR engineering relies heavily on isolating antigen-specific TCRs from tumor tissues, which requires a large amount of labor resources and wet experimental verification. To mitigate this issue, we present TCRdesign, a pretrained generative protein language model (PLM) for the de novo design of artificial TCR $\beta $-chain complementarity-determining region 3 sequences conditioned on antigen-binding specificity (BS). In parallel, we develop a high-accuracy binding predictor (TCRBinder) that couples paired $\alpha $/$\beta $ chain information with antigen sequences to assess BS. Our in silico comparisons demonstrate that (i) TCRdesign surpasses state-of-the-art baselines in generating antigen-specific TCR sequences. The model leverages paired-chain coherence to refine amino-acid level interaction patterns. (ii) TCRdesign-generated TCR sequences exhibit better antigen binding capability to diverse oncogenic hotspots compared with natural counterparts. (iii) TCRdesign inherits the intrinsic properties of large PLMs, enabling effectively identify the determinant residues in TCR-antigen binding, which enhances its interpretability. These results highlight the significant capability of TCRdesign in understanding and generating TCR sequences with an antigen-specific interaction pattern, charting a versatile path toward AI-driven T-cell engineering for precision immunotherapy.
Xiaokun Li, Qiang Yang 0015, Weihe Dong, Kuanquan Wang, Suyu Dong, Wei Wang 0169, Gongning Luo, Xianyu Zhang 0004, Tiansong Yang, Xin Gao 0001, Guohua Wang 0001
Briefings Bioinform.7
2025 TransDiffECG: Semantically controllable ECG synthesis via transformer-based diffusion modeling
Suyu Dong, Chaoyu Sun, Wanting Cong, Kuanquan Wang, Gongning Luo, Wei Wang 0169
J. Biomed. Informatics8
2025 MeMGB-Diff: Memory-Efficient Multivariate Gaussian Bias Diffusion Model for 3D bias field correction
abstract
Bias fields inevitably degrade MRI that seriously interferes the diagnosis of physicians for accurate analysis, and removing it is a crucial image analysis task. Generative models (such as GANs) are used for bias field correction, and outperform traditional methods, however are hindered by the high cost of data annotation and instability during training. Recently, the diffusion-based methods have excelled over GANs in many applications, and they are powerful in removing noise from images, while the bias field can be regarded as a smooth noise. However, it is a challenge to directly apply to 3D bias field correction due to sampling inefficiency, the heavy computational demand, and implicit correction process. We propose a Memory-Efficient Multivariate Gaussian Bias Diffusion Model (MeMGB-Diff) that is an explicit, sampling, and memory both efficient diffusion model for 3D bias field correction without using clinical labels. MeMGB-Diff extends the diffusion models to multivariate Gaussian and models the bias field as a multivariate Gaussian variable, allowing direct diffusion and removal of the 3D bias fields without Gaussian noise. For memory efficiency, MeMGB-Diff performs diffusion model in smaller readable image domain at the expense of a negligible accuracy loss, based on the strong correlation among adjacent voxels of bias field. We also propose a loss function to mainly learn the intensity trend, which mainly causes the inhomogeneity of MRI, and effectively increases the correction accuracy. For comprehensive performance comparison, we propose a synthetic method for generating more varied bias fields during testing. Both quantitative and qualitative assessments on synthetic and clinical data confirm the high fidelity and uniform intensity of our results. MeMGB-Diff reduces data size by 64 times to use less memory, improves sampling efficiency by more than 10 times compared to other diffusion-based methods, and achieves optimal metrics, including SSIM, PSNR, COCO, and CV for various tissues. Hence, our MeMGB-Diff is a state-of-the-art (SOTA) method for 3D bias field correction.
Xingyu Qiu, Dong Liang 0001, Gongning Luo, Xiangyu Li 0004, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
Medical Image Anal.5
2025 Pixel is All You Need: Adversarial Spatio-Temporal Ensemble Active Learning for Salient Object Detection
abstract
Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial spatio-temporal ensemble active learning. Our contributions are four-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed spatio-temporal ensemble strategy not only achieves outstanding performance but significantly reduces the model's computational cost. 3) Our proposed relationship-aware diversity sampling can conquer oversampling while boosting model performance. 4) We provide theoretical proof for the existence of such a point-labeled dataset. Experimental results show that our approach can find such a point-labeled dataset, where a saliency model trained on it obtained 98%-99% performance of its fully-supervised version with only ten annotated points per image.
Wei Wang 0169, Yacong Li, Fengmao Lv, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Multi-view diabetic retinopathy grading via cross-view spatial alignment and adaptive vessel reinforcing
Xiaoyan Dou, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Tianyi Luo, Jie Wen 0001, Bingo Wing-Kuen Ling, Yong Xu 0001, Wei Wang 0169
Pattern Recognit.10
2025 Adjacency-Aware Fuzzy Label Learning for Skin Disease Diagnosis
abstract
Automatic acne severity grading is crucial for the accurate diagnosis and effective treatment of skin diseases. However, the acne severity grading process is often ambiguous due to the similar appearance of acne with close severity, making it challenging to achieve reliable acne severity grading. Following the idea of fuzzy logic for handling uncertainty in decision-making, we transforms the acne severity grading task into a fuzzy label learning (FLL) problem, and propose a novel adjacency-aware fuzzy label learning (AFLL) framework to handle uncertainties in this task. The AFLL framework makes four significant contributions, each demonstrated to be highly effective in extensive experiments. First, we introduce a novel adjacency-aware decision sequence generation method that enhances sequence tree construction by reducing bias and improving discriminative power. Second, we present a consistency-guided decision sequence prediction method that mitigates error propagation in hierarchical decision-making through a novel selective masking decision strategy. Third, our proposed sequential conjoint distribution loss innovatively captures the differences for both high and low fuzzy memberships across the entire fuzzy label set while modeling the internal temporal order among different acne severity labels with a cumulative distribution, leading to substantial improvements in FLL. Fourth, to the best of our knowledge, AFLL is the first approach to explicitly address the challenge of distinguishing adjacent categories in acne severity grading tasks. Experimental results on the public ACNE04 dataset demonstrate that AFLL significantly outperforms existing methods, establishing a new state-of-the-art in acne severity grading.
Murong Zhou, Baifu Zuo, Guohua Wang 0001, Gongning Luo, Fanding Li, Suyu Dong, Wei Wang 0169, Kuanquan Wang, Xiangyu Li 0004, Lifeng Xu
IEEE Trans. Fuzzy Syst.7
2025 MedFILIP: Medical Fine-Grained Language-Image Pre-Training
abstract
Medical vision-language pretraining (VLP) that leverages naturally-paired medical image-report data is crucial for medical image analysis. However, existing methods struggle to accurately characterize associations between images and diseases, leading to inaccurate or incomplete diagnostic results. In this work, we propose MedFILIP, a fine-grained VLP model, introduces medical image-specific knowledge through contrastive learning, specifically: 1) An information extractor based on a large language model is proposed to decouple comprehensive disease details from reports, which excels in extracting disease deals through flexible prompt engineering, thereby effectively reducing text complexity while retaining rich information at a tiny cost. 2) A knowledge injector is proposed to construct relationships between categories and visual attributes, which help the model to make judgments based on image features, and fosters knowledge extrapolation to unfamiliar disease categories. 3) A semantic similarity matrix based on fine-grained annotations is proposed, providing smoother, information-richer labels, thus allowing fine-grained image-text alignment. 4) We validate MedFILIP on numerous datasets, e.g., RSNA-Pneumonia, NIH ChestX-ray14, VinBigData, and COVID-19. For single-label, multi-label, and fine-grained classification, our model achieves state-of-the-art performance, the classification accuracy has increased by a maximum of 6.69%.
Xinjie Liang, Xiangyu Li 0004, Fanding Li, Wei Wang 0169, Kuanquan Wang, Suyu Dong, Gongning Luo, Shuo Li 0001
IEEE J. Biomed. Health Informatics6
2025 A Benchmark Framework for the Right Atrium Cavity Segmentation From LGE-MRIs
abstract
The right atrium (RA) is critical for cardiac hemodynamics but is often overlooked in clinical diagnostics. This study presents a benchmark framework for RA cavity segmentation from late gadolinium-enhanced magnetic resonance imaging (LGE-MRIs), leveraging a two-stage strategy and a novel 3D deep learning network, RASnet. The architecture addresses challenges in class imbalance and anatomical variability by incorporating multi-path input, multi-scale feature fusion modules, Vision Transformers, context interaction mechanisms, and deep supervision. Evaluated on datasets comprising 354 LGE-MRIs, RASnet achieves SOTA performance with a Dice score of 92.19% on a primary dataset and demonstrates robust generalizability on an independent dataset. The proposed framework establishes a benchmark for RA cavity segmentation, enabling accurate and efficient analysis for cardiac imaging applications. Open-source code (https://github.com/zjinw/RAS) and data (https://zenodo.org/records/15524472) are provided to facilitate further research and clinical adoption.
Jieyun Bai, Jinwen Zhu, Zhiting Chen, Ziduo Yang, Yaosheng Lu, Lei Li 0020, Qince Li, Wei Wang 0169, Henggui Zhang, Kuanquan Wang, Jichao Zhao, Hua Lu 0022, Suining Li, Xiaoshen Zhang, Xiaowei Xu 0004, Yanfeng Tian, Víctor M. Campello, Karim Lekadir
IEEE Trans. Medical Imaging8
2024 Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering Structures
abstract
Incomplete multi-view clustering (IMVC) aims to reveal shared clustering structures within multi-view data, where only partial views of the samples are available. Existing IMVC methods primarily suffer from two issues: 1) Imputation-based methods inevitably introduce inaccurate imputations, which in turn degrade clustering performance; 2) Imputation-free methods are susceptible to unbalanced information among views and fail to fully exploit shared information. To address these issues, we propose a novel method based on variational autoencoders. Specifically, we adopt multiple view-specific encoders to extract information from each view and utilize the Product-of-Experts approach to efficiently aggregate information to obtain the common representation. To enhance the shared information in the common representation, we introduce a coherence objective to mitigate the influence of information imbalance. By incorporating the Mixture-of-Gaussians prior information into the latent representation, our proposed method is able to learn the common representation with clustering-friendly structures. Extensive experiments on four datasets show that our method achieves competitive clustering performance compared with state-of-the-art methods.
Gehui Xu, Jie Wen 0001, Chengliang Liu 0003, Lunke Fei, Wei Wang 0169
AAAI7
2024 Att-EMD-Unet: A Novel weakly supervised perspective for ECG segmentation
abstract
Automatic ECG segmentation has gained significant attention due to its critical role in cardiac analysis and diagnosis. However, current automatic ECG segmentation methods are hindered by the need for labor-intensive and expert-level annotations. To alleviate the annotation burden, we explore a weakly supervised perspective for ECG QT segmentation. Specifically, we employ annotator-friendly and less expert-intensive casual annotations as supervision signals for model training. In this paper, we propose a novel model called Att-EMD-Unet, which employs U-Net as base network structure and incorporates channel/ temporal attention mechanisms to predict the QT segments from original signals. Recognizing the challenges posed by casual and incomplete labels in our weakly supervised learning framework, we have innovatively incorporated an empirical mode decomposition (EMD) based R and T peaks attentive loss function during the training phase. This function is specifically designed to address and rectify potential inaccuracies or omissions in the estimation of R and T peaks within ECGs. An expert clinician conducted an evaluation of our proposed casual annotation method. The findings from this assessment indicate that it cuts down the time needed for labeling by about 46.58% for each beat in various ECG signals, compared to the usual detailed methods. And the experimental results reveal that our Att-EMD-Unet model surpasses conventional unsupervised methods and achieves comparable performance to state-of-the-art fully supervised learning methods. Our approach effectively overcomes the limitations of extensive labeling required in fully supervised learning, presenting an efficient and accurate solution for ECG QT segmentation
Wei Wang 0169
BIBM2
2024 Biomedically Informed ECG Synthesis: Customizing Cardiac Cycle Phases with Diffusion Model
abstract
Cardiovascular diseases are a major global health challenge, with electrocardiography (ECG) being critical for diagnosis and monitoring. As artificial intelligence and automated ECG diagnostic technologies rapidly advance, the demand for large-scale ECG databases continues to grow. Generative ECG has become a mainstream method to enhance database size and diversity. However, existing methods typically generate ECG randomly or focus on limited physiological categories, lacking the ability to synthesize ECG with varying physiological features and cardiac cycles, which is crucial for various practical applications. In response to this need, we propose a novel approach introducing a diffusion model called DIFF-ECG to generate precisely customized ECG that accurately reflect diverse cardiac conditions. Segmentation-based quality assessments confirmed that the synthesized ECG accurately followed the specified cardiac cycle information, with our model significantly outperforming baseline diffusion and GAN-based methods. Therefore, our approach addresses the critical need for generating clinically relevant and customizable ECG, contributing significantly to the field of automated cardiac disease diagnosis. By enabling fine-tuning of cardiac cycle phases, our method significantly expands the application range of generative ECG, potentially improving the diagnostic accuracy for rare diseases and advancing personalized medicine.
Wei Wang 0169, Zhihao Wu 0002, Suyu Dong, Gongning Luo, Kuanquan Wang
BIBM3
2024 EdgeReg: Edge-assisted Unsupervised Medical Image Registration
abstract
Medical image registration (MIR) is essential for various clinical diagnoses and treatments. Despite the rapid progress in deep learning-based MIR techniques, most methods focus on directly optimizing the raw image intensity information. In this paper, we explore the usage of edge information of anatomical structures associated with the spatial location of image intensities to assist in registration, termed EdgeReg. The intuition is that the edge information can provide additional rich boundary information to the raw images, enhancing the network’s feature representation. Additionally, as the edge images are strictly spatially consistent with the raw images, additional supervised information can be added to network training. Specifically, we first extract the edge images from the raw moving and fixed images using the Sobel operator and feed these images into a lightweight feature extractor to merge the image intensity and edge information. The enriched features are subsequently input into established registration networks. Finally, similarity loss is applied to both the raw and edge images. Extensive experiments show that EdgeReg is compatible with various networks across diverse datasets and dimensions (2D and 3D), achieving superior registration performance. In particular, EdgeReg does not rely on segmentation labels and is trained in an unsupervised paradigm. Therefore, edge information is a beneficial assistance for unsupervised MIR. The code is available at https://github.com/PerceptionComputingLab/EdgeReg.
Jun Liu 0080, Wei Wang 0169, Gongning Luo, Yacong Li, Kuanquan Wang
BIBM3
2024 Hierarchical Retrieval of High-Resolution Fingerprints Based on Pore Feature
abstract
Faced with an escalating number of fingerprint images, most existing retrieval approachs suffer from a common problem: diminishing computational efficiency. This paper presents a hierarchical retrieval system tailored for high-resolution fingerprint images that utilizes abundant pore features and robust recognizability to improve retrieval performance. The framework comprises two core components. Firstly, a CNN-based feature extraction network is established, incorporating an attention mechanism to capture pore features in fingerprint images comprehensively. Subsequently, a hierarchical fingerprint retrieval approach is introduced, involving connection graph construction and a hierarchy of jump table structures for efficient retrieval of query pores. Empirical experiments conducted on high-resolution fingerprint image datasets underscore the system’s effectiveness. Compared with other advanced pore-based fingerprint retrieval methods, the proposed method exhibits a notable rise in the hit rate with reduced penetration rates, significantly reducing the retrieval time.
Yuanrong Xu, Suyu Dong, Wei Wang 0169
BIBM4
2024 3D Electromechanical Coupling Simulation Under Heart Failure: Exploring Reentry Phenomena and Arrhythmogenesis
abstract
Heart failure alters the electrophysiological properties of cardiomyocytes, leading to changes in the overall mechanical function of the heart, with significant implications for human health. Current research on heart failure predominantly focuses on two-dimensional electrophysiological models, which limits their ability to fully capture the complexity of heart failure. In contrast, our work integrates both electrophysio-logical and mechanical aspects by developing a cardiomyocyte model incorporating heart failure remodeling within a three-dimensional electromechanical coupling model, providing a more comprehensive understanding of heart failure dynamics. We investigated the variations in electromechanical coupling properties of the left ventricle under heart failure conditions, focusing on key indicators such as ventricular action potentials, myocardial contractility, ventricular volume, and pressure. The simulation results successfully reproduced the impact of heart failure on both electrophysiological and mechanical characteristics. Additionally, the simulation observed reentrant waves under heart failure conditions, further revealing that heart failure predisposes the heart to reentrant arrhythmias. In conclusion, the 3D electromechanical coupling model presented in this study offers crucial insights into the mechanisms of reentry phenomena and arrhythmogenesis under heart failure conditions.
Wei Wang 0169, Xianda Bu, Qince Li, Kuanquan Wang
BIBM1
2024 Mutualreg: Mutual Learning for Unsupervised Medical Image Registration
abstract
Recently, self-training strategies have shown outstanding performance in the unsupervised medical image registration field. These strategies use their own network to generate pseudo-displacement fields (PFs) to supervise network training. However, limited diversity and accuracy of these PFs hinder their effectiveness. To address these limitations, we propose a novel mutual learning registration paradigm (MutualReg), where knowledge is distilled mutually between teacher and student networks for alternate improvement via recursive training. This involves two fundamental challenges: 1) how to generate more diverse and accurate PFs; and 2) how to effectively integrate knowledge distillation from the teacher network and learning from the student network. For the former, we employ a different and powerful teacher network thanks to the decoupling nature of MutualReg. For the latter, we introduce a Voxel-wise Reliability Criterion (VRC) module to retain reliable voxel locations of knowledge distillation. In the abdominal CT registration task, MutualReg outperforms state-of-the-art competitors, demonstrating its effectiveness. Code is available from https://github.com/PerceptionComputingLab/MutualReg/.
Jun Liu 0080, Nuo Shen, Wei Wang 0169, Kuanquan Wang, Qince Li, Yongfeng Yuan, Henggui Zhang, Gongning Luo
ICASSP4
2024 Batch Singular Value Polarization and Weighted Semantic Augmentation for Universal Domain Adaptation
abstract
As a more challenging domain adaptation setting, universal domain adaptation (UniDA) introduces category shift on top of domain shift, which needs to identify unknown category in the target domain and avoid misclassifying target samples into source private categories. To this end, we propose a novel UniDA approach named Batch Singular value Polarization and Weighted Semantic Augmentation (BSP-WSA). Specifically, we adopt an adversarial classifier to identify the target unknown category and align feature distributions between the two domains. Then, we propose to perform SVD on the classifier's outputs to maximize larger singular values while minimizing those smaller ones, which could prevent target samples from being wrongly assigned to source private classes. To better bridge the domain gap, we propose a weighted semantic augmentation approach for UniDA to generate data on common categories between the two domains. Extensive experiments on three benchmarks demonstrate that BSP-WSA could outperform existing state-of-the-art UniDA approaches.
Wangzi Qi, Wei Wang 0169, Chao Huang 0008, Jie Wen 0001, Cong Wang 0018
ICML2
2024 Spatio-Temporal Contrast Network for Data-Efficient Learning of Coronary Artery Disease in Coronary CT Angiography
Xinghua Ma, Mingye Zou, Xinyan Fang, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Zhaowen Qiu, Xin Gao 0001, Shuo Li 0001
MICCAI (11)6
2024 DrugMGR: a deep bioactive molecule binding method to identify compounds targeting proteins
abstract
MOTIVATION: Understanding the intermolecular interactions of ligand-target pairs is key to guiding the optimization of drug research on cancers, which can greatly mitigate overburden workloads for wet labs. Several improved computational methods have been introduced and exhibit promising performance for these identification tasks, but some pitfalls restrict their practical applications: (i) first, existing methods do not sufficiently consider how multigranular molecule representations influence interaction patterns between proteins and compounds; and (ii) second, existing methods seldom explicitly model the binding sites when an interaction occurs to enable better prediction and interpretation, which may lead to unexpected obstacles to biological researchers. RESULTS: To address these issues, we here present DrugMGR, a deep multigranular drug representation model capable of predicting binding affinities and regions for each ligand-target pair. We conduct consistent experiments on three benchmark datasets using existing methods and introduce a new specific dataset to better validate the prediction of binding sites. For practical application, target-specific compound identification tasks are also carried out to validate the capability of real-world compound screen. Moreover, the visualization of some practical interaction scenarios provides interpretable insights from the results of the predictions. The proposed DrugMGR achieves excellent overall performance in these datasets, exhibiting its advantages and merits against state-of-the-art methods. Thus, the downstream task of DrugMGR can be fine-tuned for identifying the potential compounds that target proteins for clinical treatment. AVAILABILITY AND IMPLEMENTATION: https://github.com/lixiaokun2020/DrugMGR.
Xiaokun Li, Qiang Yang 0015, Weihe Dong, Gongning Luo, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Ping Xuan, Xianyu Zhang 0004, Xin Gao 0001
Bioinform.6
2024 Boosting knowledge diversity, accuracy, and stability via tri-enhanced distillation for domain continual medical image segmentation
Zhanshi Zhu, Xinghua Ma, Wei Wang 0169, Suyu Dong, Kuanquan Wang, Lianming Wu, Gongning Luo, Guohua Wang 0001, Shuo Li 0001
Medical Image Anal.3
2024 A simulation study on the antiarrhythmic mechanisms of established agents in myocardial ischemia and infarction
abstract
Patients with myocardial ischemia and infarction are at increased risk of arrhythmias, which in turn, can exacerbate the overall risk of mortality. Despite the observed reduction in recurrent arrhythmias through antiarrhythmic drug therapy, the precise mechanisms underlying their effectiveness in treating ischemic heart disease remain unclear. Moreover, there is a lack of specialized drugs designed explicitly for the treatment of myocardial ischemic arrhythmia. This study employs an electrophysiological simulation approach to investigate the potential antiarrhythmic effects and underlying mechanisms of various pharmacological agents in the context of ischemia and myocardial infarction (MI). Based on physiological experimental data, computational models are developed to simulate the effects of a series of pharmacological agents (amiodarone, telmisartan, E-4031, chromanol 293B, and glibenclamide) on cellular electrophysiology and utilized to further evaluate their antiarrhythmic effectiveness during ischemia. On 2D and 3D tissues with multiple pathological conditions, the simulation results indicate that the antiarrhythmic effect of glibenclamide is primarily attributed to the suppression of efflux of potassium ion to facilitate the restitution of [K+]o, as opposed to recovery of IKATP during myocardial ischemia. This discovery implies that, during acute cardiac ischemia, pro-arrhythmogenic alterations in cardiac tissue's excitability and conduction properties are more significantly influenced by electrophysiological changes in the depolarization rate, as opposed to variations in the action potential duration (APD). These findings offer specific insights into potentially effective targets for investigating ischemic arrhythmias, providing significant guidance for clinical interventions in acute coronary syndrome.
Qince Li, Cuiping Liang, Xiqian Wang, Xianghu Wu, Wei Wang 0169, Yongfeng Yuan, Kuanquan Wang
PLoS Comput. Biol.7
2024 Graph Regularized and Feature Aware Matrix Factorization for Robust Incomplete Multi-View Clustering
abstract
In recent years, many incomplete multi-view clustering methods have been proposed to address the challenging and new clustering task on incomplete multi-view data whose part of view representations are not fully collected for some samples. Although extensive experiments have validated the effectiveness of these methods for handling the incomplete learning issue, a common issue exists, i.e., these methods all ignore the discriminative/important difference of discriminative features and noisy features. In this paper, to address the above issue, a new incomplete multi-view clustering model, called Graph Regularized and fEature Aware maTrix Factorization (GreatF), is proposed. Different from the existing methods, we introduce an adaptive feature weighting constraint to the matrix factorization-based multi-view representation learning model. With this weighting constraint, the effect of the discriminative features can be enhanced while the negative effect caused by the redundant and noisy features can be eliminated for the model optimization; thus, the robustness of the model can be enhanced. In addition, in this work, we designed a new graph-embedded consensus representation learning term in which consensus representation learning and structure information preservation are integrated into a joint model with one term. In particular, this term provides a more concise approach to obtain the structured consensus representation from incomplete multi-view data. Experimental results on four well-known datasets demonstrate that GreatF performs better than the state-of-the-art incomplete multi-view clustering methods.
Jie Wen 0001, Gehui Xu, Zhanyan Tang, Wei Wang 0169, Lunke Fei, Yong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection
abstract
Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image.
Wei Wang 0169, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001
AAAI3
2023 Synergistically Learning Class-specific Tokens for Multi-class Whole Slide Image Classification
abstract
The application of transformer architecture in analyzing whole slide images (WSIs) has become increasingly popular due to its remarkable ability to learn complex associations. Nevertheless, a significant drawback emerges in the multiclass analysis of WSIs. The majority of the transformer-based methods available currently rely primarily on a single, class-agnostic token. This approach might not ideally capture the subtleties of class-discriminative information. To address this challenge, we present an innovative approach tailored for multi-class WSI analysis that harnesses the power of class-specific tokens. Central to our method is a novel attention mechanism designed to foster a synergistic learning relationship between patch and class tokens, enhancing the granularity of information captured and ensuring a more comprehensive representation of the WSI. Complementing this, we introduce a dynamic class-centric training strategy designed to optimize token representation learning, ensuring each token is informatively aligned with its corresponding class. Through extensive experimentation on three challenging multi-class WSI analysis datasets, our method consistently demonstrates superior performance, underscoring its potential as a robust solution for multi-class WSI analysis tasks.
Pengzhong Sun, Wei Wang 0169, Xiangyu Li 0004, Suyu Dong, Shuo Li 0001, Kuanquan Wang, Gongning Luo
BIBM2
2023 Vision Transformers(ViT) Pretraining on 3D ABUS Image and Dual-CapsViT: Enhancing ViT Decoding via Dual-Channel Dynamic Routing
abstract
Breast cancer continues to be a pressing global health concern, emphasizing the essential need for effective diagnostic techniques. Automated Breast Ultrasound Systems (ABUS) provide a promising advance in breast tumor detection, yet they require significant expertise in interpreting 3D ABUS images, a task fraught with distinctive challenges. Although Vision Transformers (ViT) display remarkable potential for image processing, their low inductive bias and significant data requirements pose obstacles, particularly in the data-constrained medical field. To mitigate these issues, we introduce a Mask-Recover strategy for pretraining Transformer models on 3D ABUS images, enhancing model adaptability and reducing the data demands of the ViT model. Moreover, recognizing the risk that ViTs’ average pooling approach may unintentionally mask small but vital features, we propose Dual-CapsViT, an inventive model combining Transformers and Capsule Networks. This integration affords efficient token routing while preserving fine-grained details. To reconcile potential inconsistencies between capsules and tokens, we engineer a novel dual-channel routing algorithm, strengthening the decoder’s performance. We benchmarked our models against well-known standards such as ResNet and ViT for classifying breast tumors in ABUS images. Our models exhibited superior performance, as evidenced by improved accuracy, specificity, and Area Under the Receiver Operating Characteristic Curve (AUC) metrics, thereby affirming Dual-CapsViT’s potential to enhance breast cancer diagnostics.
Mingwang Xu, Wei Wang 0169, Kuanquan Wang, Suyu Dong, Pengzhong Sun, Jinwei Sun, Gongning Luo
BIBM2
2023 Localized and Balanced Efficient Incomplete Multi-view Clustering
abstract
In recent years, many incomplete multi-view clustering methods have been proposed to address the challenging unsupervised clustering issue on the multi-view data with missing views. However, most of the existing works are inapplicable to large-scale clustering task and their clustering results are unstable since these methods have high computational complexities and their results are produced by kmeans rather than their designed learning models. In this paper, we propose a new one-step incomplete multi-view clustering model, called Localized and Balanced Incomplete Multi-view Clustering (LBIMVC), to address these issues. Specifically, LBIMVC develops a new graph regularized incomplete multi-matrix-factorization model to obtain the unique clustering result by learning a consensus probability representation, where each element of the consensus representation can directly reflect the probability of the corresponding sample to the class. In addition, the proposed graph regularized model integrates geometric preserving and consensus representation learning into one term without introducing any extra constraint terms and parameters to explore the structure of data. Moreover, to avoid that samples are over divided into a few clusters, a balanced constraint is introduced to the model. Experimental results on four databases demonstrate that our method not only obtains competitive clustering performance, but also performs faster than some state-of-the-art methods.
Jie Wen 0001, Gehui Xu, Chengliang Liu 0003, Lunke Fei, Chao Huang 0008, Wei Wang 0169, Yong Xu 0001
ACM Multimedia6
2023 Ambiguity-aware breast tumor cellularity estimation via self-ensemble label distribution learning
Xiangyu Li 0004, Xinjie Liang, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
Medical Image Anal.4
2023 Curriculum label distribution learning for imbalanced medical image segmentation
Xiangyu Li 0004, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
Medical Image Anal.3
2023 Trajectory-Aware Adaptive Imaging Clue Analysis for Guidewire Artifact Removal in Intravascular Optical Coherence Tomography
abstract
Guidewire Artifact Removal (GAR) involves restoring missing imaging signals in areas of IntraVascular Optical Coherence Tomography (IVOCT) videos affected by guidewire artifacts. GAR helps overcome imaging defects and minimizes the impact of missing signals on the diagnosis of CardioVascular Diseases (CVDs). To restore the actual vascular and lesion information within the artifact area, we propose a reliable Trajectory-aware Adaptive imaging Clue analysis Network (TAC-Net) that includes two innovative designs: (i) Adaptive clue aggregation, which considers both texture-focused original (ORI) videos and structure-focused relative total variation (RTV) videos, and suppresses texture-structure imbalance with an active weight-adaptation mechanism; (ii) Trajectory-aware Transformer, which uses a novel attention calculation to perceive the attention distribution of artifact trajectories and avoid the interference of irregular and non-uniform artifacts. We provide a detailed formulation for the procedure and evaluation of the GAR task and conduct comprehensive quantitative and qualitative experiments. The experimental results demonstrate that TAC-Net reliably restores the texture and structure of guidewire artifact areas as expected by experienced physicians (e.g., SSIM: 97.23%). We also discuss the value and potential of the GAR task for clinical applications and computer-aided diagnosis of CVDs.
Gongning Luo, Xinghua Ma, Jinwen Guo, Mingye Zou, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
IEEE J. Biomed. Health Informatics5
2022 ULTRA: Uncertainty-Aware Label Distribution Learning for Breast Tumor Cellularity Assessment
Xiangyu Li 0004, Xinjie Liang, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Shuo Li 0001
MICCAI (3)4
2022 Position-Prior Clustering-Based Self-attention Module for Knee Cartilage Segmentation
Dong Liang 0012, Jun Liu 0080, Kuanquan Wang, Gongning Luo, Wei Wang 0169, Shuo Li 0001
MICCAI (5)5
2022 Synthetic Data Supervised Salient Object Detection
abstract
Although deep salient object detection (SOD) has achieved remarkable progress, deep SOD models are extremely data-hungry, requiring large-scale pixel-wise annotations to deliver such promising results. In this paper, we propose a novel yet effective method for SOD, coined SODGAN, which can generate infinite high-quality image-mask pairs requiring only a few labeled data, and these synthesized pairs can replace the human-labeled DUTS-TR to train any off-the-shelf SOD model. Its contribution is three-fold. 1) Our proposed diffusion embedding network can address the manifold mismatch and is tractable for the latent code generation, better matching with the ImageNet latent space. 2) For the first time, our proposed few-shot saliency mask generator can synthesize infinite accurate image synchronized saliency masks with a few labeled data. 3) Our proposed quality-aware discriminator can select highquality synthesized image-mask pairs from noisy synthetic data pool, improving the quality of synthetic data. For the first time, our SODGAN tackles SOD with synthetic data directly generated from the generative model, which opens up a new research paradigm for SOD. Extensive experimental results show that the saliency model trained on synthetic data can achieve $98.4%$ F-measure of the saliency model trained on the DUTS-TR. Moreover, our approach achieves a new SOTA performance in semi/weakly-supervised methods, and even outperforms several fully-supervised SOTA methods. Code is available at https://github.com/wuzhenyubuaa/SODGAN
Wei Wang 0169, Tengfei Shi, Chenglizhao Chen, Aimin Hao, Shuo Li 0001
ACM Multimedia3
2022 Mechanisms of ventricular arrhythmias elicited by coexistence of multiple electrophysiological remodeling in ischemia: A simulation study
abstract
Myocardial ischemia, injury and infarction (MI) are the three stages of acute coronary syndrome (ACS). In the past two decades, a great number of studies focused on myocardial ischemia and MI individually, and showed that the occurrence of reentrant arrhythmias is often associated with myocardial ischemia or MI. However, arrhythmogenic mechanisms in the tissue with various degrees of remodeling in the ischemic heart have not been fully understood. In this study, biophysical detailed single-cell models of ischemia 1a, 1b, and MI were developed to mimic the electrophysiological remodeling at different stages of ACS. 2D tissue models with different distributions of ischemia and MI areas were constructed to investigate the mechanisms of the initiation of reentrant waves during the progression of ischemia. Simulation results in 2D tissues showed that the vulnerable windows (VWs) in simultaneous presence of multiple ischemic conditions were associated with the dynamics of wave propagation in the tissues with each single pathological condition. In the tissue with multiple pathological conditions, reentrant waves were mainly induced by two different mechanisms: one is the heterogeneity along the excitation wavefront, especially the abrupt variation in conduction velocity (CV) across the border of ischemia 1b and MI, and the other is the decreased safe factor (SF) for conduction at the edge of the tissue in MI region which is attributed to the increased excitation threshold of MI region. Finally, the reentrant wave was observed in a 3D model with a scar reconstructed from MRI images of a MI patient. These comprehensive findings provide novel insights for understanding the arrhythmic risk during the progression of myocardial ischemia and highlight the importance of the multiple pathological stages in designing medical therapies for arrhythmias in ischemia.
Cuiping Liang, Qince Li, Kuanquan Wang, Yimei Du, Wei Wang 0169, Henggui Zhang
PLoS Comput. Biol.5
2022 Hematoma Expansion Context Guided Intracranial Hemorrhage Segmentation and Uncertainty Estimation
abstract
Accurate segmentation of the Intracranial Hemorrhage (ICH) in non-contrast CT images is significant for computer-aided diagnosis. Although existing methods have achieved remarkable 1 1 The code will be available from https://github.com/JohnleeHIT/SLEX-Net. results, none of them incorporated ICH's prior information in their methods. In this work, for the first time, we proposed a novel SLice EXpansion Network (SLEX-Net), which incorporated hematoma expansion in the segmentation architecture by directly modeling the hematoma variation among adjacent slices. Firstly, a new module named Slice Expansion Module (SEM) was built, which can effectively transfer contextual information between two adjacent slices by mapping predictions from one slice to another. Secondly, to perceive contextual information from both upper and lower slices, we designed two information transmission paths: forward and backward slice expansion, and aggregated results from those paths with a novel weighing strategy. By further exploiting intra-slice and inter-slice context with the information paths, the network significantly improved the accuracy and continuity of segmentation results. Moreover, the proposed SLEX-Net enables us to conduct an uncertainty estimation with one-time inference, which is much more efficient than existing methods. We evaluated the proposed SLEX-Net and compared it with some state-of-the-art methods. Experimental results demonstrate that our method makes significant improvements in all metrics on segmentation performance and outperforms other existing uncertainty estimation methods in terms of several metrics.
Xiangyu Li 0004, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Yue Gao 0002, Shuo Li 0001
IEEE J. Biomed. Health Informatics3
2021 Transformer Network for Significant Stenosis Detection in CCTA of Coronary Arteries
Xinghua Ma, Gongning Luo, Wei Wang 0169, Kuanquan Wang
MICCAI (6)3
2021 Mining a stroke knowledge graph from literature
abstract
BACKGROUND: Stroke has an acute onset and a high mortality rate, making it one of the most fatal diseases worldwide. Its underlying biology and treatments have been widely studied both in the "Western" biomedicine and the Traditional Chinese Medicine (TCM). However, these two approaches are often studied and reported in insolation, both in the literature and associated databases. RESULTS: To aid research in finding effective prevention methods and treatments, we integrated knowledge from the literature and a number of databases (e.g. CID, TCMID, ETCM). We employed a suite of biomedical text mining (i.e. named-entity) approaches to identify mentions of genes, diseases, drugs, chemicals, symptoms, Chinese herbs and patent medicines, etc. in a large set of stroke papers from both biomedical and TCM domains. Then, using a combination of a rule-based approach with a pre-trained BioBERT model, we extracted and classified links and relationships among stroke-related entities as expressed in the literature. We construct StrokeKG, a knowledge graph includes almost 46 k nodes of nine types, and 157 k links of 30 types, connecting diseases, genes, symptoms, drugs, pathways, herbs, chemical, ingredients and patent medicine. CONCLUSIONS: Our Stroke-KG can provide practical and reliable stroke-related knowledge to help with stroke-related research like exploring new directions for stroke research and ideas for drug repurposing and discovery. We make StrokeKG freely available at http://114.115.208.144:7474/browser/ (Please click "Connect" directly) and the source structured data for stroke at https://github.com/yangxi1016/Stroke.
Xi Yang 0020, Chengkun Wu, Goran Nenadic, Wei Wang 0169, Kai Lu 0001
BMC Bioinform.4
2021 Correction to: Mining a stroke knowledge graph from literature
Xi Yang 0020, Chengkun Wu, Goran Nenadic, Wei Wang 0169, Kai Lu 0001
BMC Bioinform.4
2020 Branch-Aware Double DQN for Centerline Extraction in Coronary CT Angiography
Gongning Luo, Wei Wang 0169, Kuanquan Wang
MICCAI (6)3
2020 Generating electrocardiogram signals by deep learning
Naren Wulan, Wei Wang 0169, Pengzhong Sun, Kuanquan Wang, Yong Xia 0005, Henggui Zhang
Neurocomputing2
2020 Deep Atlas Network for Efficient 3D Left Ventricle Segmentation on Echocardiography
Suyu Dong, Gongning Luo, Clara M. Tam, Wei Wang 0169, Kuanquan Wang, Shaodong Cao, Bo Chen 0013, Henggui Zhang, Shuo Li 0001
Medical Image Anal.4
2020 Commensal correlation network between segmentation and direct area estimation for bi-ventricle quantification
Gongning Luo, Suyu Dong, Wei Wang 0169, Kuanquan Wang, Shaodong Cao, Clara M. Tam, Henggui Zhang, Joanne Howey, Pavlo Ohorodnyk, Shuo Li 0001
Medical Image Anal.3
2020 Dynamically constructed network with error correction for accurate ventricle volume estimation
Gongning Luo, Wei Wang 0169, Clara M. Tam, Kuanquan Wang, Shaodong Cao, Henggui Zhang, Bo Chen 0013, Shuo Li 0001
Medical Image Anal.2
2019 Grassroots VS elites: Which ones are better candidates for influence maximization in social networks?
Dong Li 0052, Wei Wang 0169, Jiming Liu 0001
Neurocomputing2
2018 Mechanistic insight into spontaneous transition from cellular alternans to arrhythmia - A simulation study
abstract
Cardiac electrical alternans (CEA), manifested as T-wave alternans in ECG, is a clinical biomarker for predicting cardiac arrhythmias and sudden death. However, the mechanism underlying the spontaneous transition from CEA to arrhythmias remains incompletely elucidated. In this study, multiscale rabbit ventricular models were used to study the transition and a potential role of INa in perpetuating such a transition. It was shown CEA evolved into either concordant or discordant action potential (AP) conduction alternans in a homogeneous one-dimensional tissue model, depending on tissue AP duration and conduction velocity (CV) restitution properties. Discordant alternans was able to cause conduction failure in the model, which was promoted by impaired sodium channel with either a reduced or increased channel current. In a two-dimensional homogeneous tissue model, a combined effect of rate- and curvature-dependent CV broke-up alternating wavefronts at localised points, facilitating a spontaneous transition from CEA to re-entry. Tissue inhomogeneity or anisotropy further promoted break-up of re-entry, leading to multiple wavelets. Similar observations have also been seen in human atrial cellular and tissue models. In conclusion, our results identify a mechanism by which CEA spontaneously evolves into re-entry without a requirement for premature ventricular complexes or pre-existing tissue heterogeneities, and demonstrated the important pro-arrhythmic role of impaired sodium channel activity. These findings are model-independent and have potential human relevance.
Wei Wang 0169, Shanzhuo Zhang, Haibo Ni, Clifford J. Garratt, Mark R. Boyett, Jules C. Hancox, Henggui Zhang
PLoS Comput. Biol.1
2015 Simulation of effects of TBX18 on the pacemaker activity of human ventricular cells
abstract
Transcription factor TBX18 could reduce the electrical coupling of ventricular myocytes and slow the electrical propagation, leading to pacemaker activity. In this article, the effect of TBX18 was analyzed by modulating coupling conductance (diffusion coefficient) and we found that with the decreasing of coupling, the pacemaker activity of ventricle increased. The first pacing time decreased with the reduction of coupling. However, when coupling conductance was lower than a critical value, the automatic excitation could not propagate, although the pacemaker worked robustly. Once the working myocytes could be driven, the pacemakers with different coupling conductance made no significant difference. Action potentials (APs) of pacemaker cells and normal cardiac myocytes at the same coordinates were similar for different coupling.
Yue Zhang 0015, Kuanquan Wang, Henggui Zhang, Wei Wang 0169
BIBM4
2014 Simulation of ventricular automaticity induced by reducing inward-rectifier K+ current
abstract
Turning non-autonomic ventricular cells into pacemaking cells is believed to hold the key for making a bio-pacemaker that could potentially treat patients with cardiac conduction diseases. In this article, we analyze the effects of various membrane ion channel currents on ventricular automaticity induced by reducing the inward-rectifier K+current (IK1). It was found that the L-type calcium current (ICaL), rather than the fast sodium current (INa), plays a major role in the rapid depolarization phase of the action potential. With a small ICaL, the automaticity of cells failed due to incompletion of the rapid depolarization. However, during the slow depolarization phase of the action potential, the background sodium current (IbNa), background calcium current (IbCa) and Na+/Ca2+exchanger current (INaCa) were playing more important roles. In 2D simulations, the automatic ventricular excitations arising from IK1reduction only couldn't propagate; it required other currents to be modulated at the same time for driving the surrounding cardiac tissues.
Yue Zhang 0015, Kuanquan Wang, Henggui Zhang, Yongfeng Yuan, Wei Wang 0169
BIBM5