EDBT 2026 Demo / reviewers in the wild / expert
Yanda Meng
dblp:275/6822
· DBLP profile ↗
44ranked-venue papers
11as first author
41since 2021 · last 2026
0000-0001-7344-2174ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 7 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 22 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FRMF-Net: Feature rectification and adaptive modality fusion guided multi-modal brain tumor segmentation networkabstractBrain tumor segmentation from multi-modal magnetic resonance imaging (MRI) is crucial for computer-assisted diagnosis and treatment planning. However, this task remains highly challenging due to substantial image heterogeneity, modality-inherent variability, and severe class imbalance among tumor sub-regions. To address these issues, we propose FRMF-Net , a F eature R ectification and adaptive M odality F usion guided multi-modal brain tumor segmentation Net work, which consists of three key components: a Modality-Specific Feature Rectification (MSFR) module, an Adaptive Modality Fusion (AMF) module, and a Region-Adaptive Loss (RAL). Specifically, MSFR enhances modality-specific representations by jointly modeling shared and private information, thereby mitigating inter-modality noise and reducing feature discrepancies across modalities. Building on this, AMF performs voxel-wise adaptive fusion through modality-, channel-, and spatial-wise attention, enabling the network to dynamically emphasize the most informative features for accurate tumor delineation. In addition, RAL alleviates the class imbalance issue by adaptively reweighting the contribution of each tumor sub-region according to its spatial extent in each sample. Extensive experiments on the BraTS 2019 and BraTS 2020 datasets demonstrate that FRMF-Net consistently outperforms the state-of-the-art methods, achieving superior Dice score and lower Hausdorff distance, particularly in small and challenging tumor regions. These results confirm that FRMF-Net provides a robust and effective solution for multi-modal brain tumor segmentation. Tongxue Zhou, Su Ruan, Jinming Duan 0001, Yanda Meng, Zhiwei Ji, Bangli Liu, Maël Balluet, Bai Ying Lei |
Expert Syst. Appl. | 5 |
| 2026 | A hierarchical teacher-student learning framework with adaptive cross-modal fusion for brain tumor segmentationabstractAccurate brain tumor segmentation plays an important role in clinical diagnosis, treatment planning, and therapeutic response monitoring. Multi-modal MRI provides complementary structural and functional information, but existing methods remain limited by their inadequate exploitation of cross-modal complementarity and their inability to effectively handle modality-specific disparities and redundant information. To address these challenges, this paper proposes a novel hierarchical teacher-student learning framework with adaptive cross-modal fusion. MRI modalities are grouped into teacher modalities (Flair and T1c) and student modalities (T2 and T1) based on their intrinsic tumor-related characteristics. Central to this framework is the Modality Guidance Module (MGM), which consists of two key components designed to achieve multi-modal feature distillation. Within MGM, the Modality Enhancement Module (MEM) extracts highly discriminative features from teacher modalities. While the Modality Fusion Module (MFM) leverages these features to guide and refine the learning of student modalities. To further capture inter-modal dependencies, a Cross-Modal Fusion Module (CMFM) is introduced to adaptively integrate complementary information across all modalities. Extensive experiments on the BraTS 2018, 2019 and 2020 datasets demonstrate that the proposed method achieves superior performance compared with state-of-the-art approaches. Beyond brain tumor segmentation, the hierarchical teacher-student paradigm and adaptive fusion strategy also hold potential for broader multi-modal image analysis tasks. Tongxue Zhou, Su Ruan, Jinming Duan 0001, Haigen Hu, Yanda Meng, Ling Huang 0003, Defu Yang, Bingbing Jiang 0001, Tingjin Luo, Zhiwei Ji, Bai Ying Lei |
Expert Syst. Appl. | 5 |
| 2026 | UTriGate-Net : Uncertainty-aware brain tumor segmentation via triaxial context encoding and gated modality fusionabstractAccurate segmentation of brain tumors from multi-modal MRI is crucial for diagnosis and treatment planning. However, challenges such as severe class imbalance, modality-specific feature heterogeneity, and predictive uncertainty hinder reliable performance. In this work, we propose UTriGate-Net, a novel uncertainty-aware multi-modal brain tumor segmentation framework. First, we design a Triaxial Context Encoding (TCE) block that extracts anisotropic spatial features by applying directional convolutions along the axial, coronal, and sagittal planes, thereby enhancing 3D contextual representation. Second, we introduce a Gated Modality Fusion (GMF) module, which adaptively integrates complementary information across modalities through modality-specific gating weights that suppress redundancy while retaining salient features. Finally, to improve segmentation reliability, we develop an Uncertainty-Regularized Weighted Loss (URWL) that combines dynamic class-specific weighting to mitigate class imbalance with an entropy-based uncertainty penalty to encourage well-calibrated predictions. Experiments on the BraTS 2019 and 2020 datasets demonstrate that UTriGate-Net achieves superior segmentation accuracy and robustness, particularly in challenging subregions. Overall, the proposed framework offers a promising solution for reliable and precise brain tumor delineation in clinical practice. Tongxue Zhou, Su Ruan, Yanda Meng, Jinming Duan 0001, Haigen Hu, Bingbing Jiang 0001, Zhiwei Ji, Bangli Liu, Tingjin Luo, Bai Ying Lei |
Expert Syst. Appl. | 3 |
| 2026 | Artifact-suppressed 3D retinal microvascular segmentation via multi-scale topology regulation
Ting Luo 0001, Jinxian Zhang, Tao Chen 0003, Zhouyan He, Yanda Meng, Jiong Zhang 0004, Dan Zhang 0026 |
Medical Image Anal. | 5 |
| 2026 | DFuse-Net: Disentangled feature fusion with uncertainty-aware learning for reliable multi-modal brain tumor segmentation
Tongxue Zhou, Su Ruan, Yanda Meng, Jinming Duan 0001, Bai Ying Lei |
Medical Image Anal. | 4 |
| 2026 | Multimodal human video generation with uncertainty-aware pose guidance
Kun Yang 0010, Yuanyuan Meng, Juncen Guo, Yanda Meng, Songwen Pei, Jing Liu 0050, Yang Liu 0246 |
Pattern Recognit. | 6 |
| 2026 | StealthMark: Harmless and Stealthy Ownership Verification for Medical Segmentation via Uncertainty-Guided BackdoorsabstractAnnotating medical data for training AI models is often costly and limited due to the shortage of specialists with relevant clinical expertise. This challenge is further compounded by privacy and ethical concerns associated with sensitive patient information. As a result, well-trained medical segmentation models on private datasets constitute valuable intellectual property requiring robust protection mechanisms. Existing model protection techniques primarily focus on classification and generative tasks, while segmentation models-crucial to medical image analysis-remain largely underexplored. In this paper, we propose a novel, stealthy, and harmless method, StealthMark, for verifying the ownership of medical segmentation models under closed-box conditions. Our approach subtly modulates model uncertainty without altering the final segmentation outputs, thereby preserving the model's performance. To enable ownership verification, we incorporate model-agnostic explanation methods, e.g. LIME, to extract feature attributions from the model outputs. Under specific triggering conditions, these explanations reveal a distinct and verifiable watermark. We further design the watermark as a QR code to facilitate robust and recognizable ownership claims. We conducted extensive experiments across four medical imaging datasets (CMR dataset from UK Biobank, the SEG fundus dataset, the EchoNet echocardiography dataset, and the PraNet colonoscopy dataset) and five mainstream segmentation models. The results demonstrate the effectiveness, stealthiness, and harmlessness of our method on the original model's segmentation performance. For example, when applied to the SAM model, StealthMark consistently achieved attack success rates (ASR) above 95% across various datasets while maintaining less than a 1% drop in Dice and AUC scores-significantly outperforming backdoor-based watermarking methods and highlighting its strong potential for practical deployment. Our implementation code is made available at https://github.com/Qinkaiyu/StealthMark. Qinkai Yu, Chong Zhang 0006, Gaojie Jin, Tianjin Huang, Wei Zhou 0021, Xiao-Bo Jin, Bo Huang 0012, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng, Aline Villavicencio, Yanda Meng |
IEEE Trans. Image Process. | 14 |
| 2026 | Super-Resolution Reconstruction of OCTA via Multi-Field-of-View Representation LearningabstractHigh-resolution Optical Coherence Tomography Angiography (OCTA) images are essential for morphological analysis and biomarker measurement of the retinal vasculature. They can also provide underlying biomarkers for the accurate analysis of eye-related diseases. The trade-off between the high resolution (HR) and large scanning field-of-view (FOV) is a long-standing problem for OCTA image instrument. A large FOV image provides more retinal information with shorter acquisition time but often suffers from low resolution (LR), high scatter noise, and poor vascular contrast. In order to obtain HR OCTA images with larger FOV, we propose a novel self-similar dynamic domain adaptation network based on cross-field-of-view representation learning. The network enables LR images (i.e., $6\times \text{6}\,\text{mm}^{2}$) to learn HR image (i.e., $3\times \text{3}\,\text{mm}^{2}$) feature representations specialized for OCTA by constructing feature mapping relations for cross-field-of-view OCTA scans. To be specific, a multiple random degradation model is proposed on HR images to generate various synthetic LR images. Further, we propose a dynamic domain adaptation framework that prompts feature dynamic alignment of the LR image reconstruction results with those of synthetic LR images. Finally, a novel self-similar supervision loss is proposed to optimize the reconstruction results from LR to HR by exploiting the similarity between vessels in different regions. Experimental results on three OCTA datasets show that the proposed method surpasses existing state-of-the-art ones, significantly enhancing retinal structure segmentation and disease classification. Our OCTA dataset (the first dataset in this research area with paired $3\times 3$ and $6\times \text{6}\,\text{mm}^{2}$ OCTA images) and code are publicly available. Huaying Hao, Shaoyi Leng, Yanda Meng, Yonghuai Liu, Yalin Zheng, Huazhu Fu, Jiong Zhang 0004, Quanyong Yi, Yue Liu 0005, Jingfeng Zhang, Yitian Zhao |
IEEE J. Biomed. Health Informatics | 3 |
| 2026 | Beyond Correlation: Causal Intervention for Multi-Label Medical Image DiagnosisabstractThis paper addresses the challenge of multi-disease diagnosis by integrating causal reasoning into the diagnostic framework. In clinical practice, multiple conditions often co-occur, making multi-disease diagnosis more relevant than isolated single-disease cases. However, most deep learning methods focus on single-disease detection and fail to capture the complexity of diagnosing concurrent conditions. Even in multi-label settings, existing approaches mainly rely on correlation-based inference, capturing statistical associations rather than true causal relationships. This can lead to spurious feature-disease associations, where features linked to one disease are mistakenly attributed to another due to frequent co-occurrence, ultimately undermines diagnostic accuracy and interpretability. To address this challenge, we propose a novel framework that incorporates causal intervention into multi-label medical image diagnosis, enabling the model to identify true causal signals rather than misleading correlations arising from co-occurring diseases. Specifically, we model latent disease-related confounders and apply backdoor adjustment to disentangle genuine causal effects from spurious associations. This is achieved by implicitly learning shared feature representations that serve as confounding variables, which are then used to refine image-derived features during prediction. The resulting causal adjustment allows the model to focus on disease-specific cues, improving accuracy and interpretability. Extensive experiments on four diverse medical imaging datasets: ODIR (color fundus photography), LID-FFA (fundus fluorescein angiography), Endo (colonoscopy), and Chestpert (X-ray) demonstrate that our method consistently outperforms existing approaches. Furthermore, our model also effectively separates the diagnosis of co-occurring diseases, highlighting the potential of causal reasoning to enhance the reliability and clinical applicability of AI-assisted diagnosis. The source code is publicly available at https://github.com/davelailai/BankCausal.git. Jianyang Xie, Yitian Zhao, Xiuju Chen, Yanda Meng, He Zhao 0002, Uazman Alam, Yalin Zheng |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and DiagnosisabstractOphthalmologists typically require multimodal data sources to improve diagnostic accuracy in clinical decisions. However, due to medical device shortages, low-quality data and data privacy concerns, missing data modalities are common in real-world scenarios. Existing deep learning methods tend to address it by learning an implicit latent subspace representation for different modality combinations. We identify two significant limitations of these methods: (1) implicit representation constraints that hinder the model's ability to capture modality-specific information and (2) modality heterogeneity, causing distribution gaps and redundancy in feature representations. To address these, we propose an Incomplete Modality Disentangled Representation (IMDR) strategy, which disentangles features into explicit independent modal-common and modal-specific features by guidance of mutual information, distilling informative knowledge and enabling it to reconstruct valuable missing semantics and produce robust multimodal representations. Furthermore, we introduce a joint proxy learning module that assists IMDR in eliminating intra-modality redundancy by exploiting the extracted proxies from each class. Experiments on four ophthalmology multimodal datasets demonstrate that the proposed IMDR outperforms the state-of-the-art methods significantly. Zile Huang, Zhongxing Xu, Zihong Luo, Yalin Zheng, Yanda Meng |
AAAI | 9 |
| 2025 | DFuse-Net: Disentangled Multi-Modal Fusion Via Contrastive and Consistency-Aware Learning for Reliable Brain Tumor SegmentationabstractAccurate brain tumor segmentation from multimodal MRI is critical for clinical diagnosis and treatment planning. However, effectively leveraging the complementary information across different modalities remains a significant challenge due to modality-specific noise, information redundancy and inherent model uncertainty. To tackle these challenges, we propose a Disentangled Fusion Network (DFuse-Net) that integrates disentangled feature fusion with contrastive and consistency-aware learning to enable reliable multi-modal brain tumor segmentation. Our method first explicitly disentangles modality-shared and modality-specific feature representations. Then, a Disentangled Feature Fusion Module (DFFM) is proposed to effectively integrate modality-shared and modalityspecific feature representations. In addition, a contrastive-aware learning scheme is employed to enhance feature discriminability, while a consistency-aware learning strategy is applied to enforce structural coherence across modalities. Moreover, Monte Carlo dropout is applied during inference to generate voxelwise aleatoric and epistemic uncertainty maps, enhancing the robustness of segmentation. Extensive experiments on the BraTS datasets demonstrate that DFuse-Net achieves superior segmentation accuracy and reliability compared to the state-of-the-art methods. Tongxue Zhou, Nan Zhang 0014, Huiling Chen 0001, Yanda Meng, Zhiwei Ji |
BIBM | 6 |
| 2025 | Uncertainty Quantification for Multiple-Choice Questions is Just One-Token DeepabstractMultiple-choice question (MCQ) benchmarks such as MMLU and GPQA are widely used to assess the capabilities of large language models (LLMs). While accuracy remains the standard evaluation metric, recent work has introduced uncertainty quantification (UQ) methods, such as entropy, conformal prediction, and verbalized confidence, as complementary measures of model reliability and calibration. However, we find that these UQ methods, when applied to MCQ tasks, are unexpectedly fragile. Specifically, we show that fine-tuning a model on just 1,000 examples to adjust the probability of the first generated token, under the common prompting setup where the model is instructed to output only a single answer choice, can systematically distort a broad range of UQ methods across models, prompts, and domains, all while leaving answer accuracy unchanged. We validate this phenomenon through extensive experiments on five instruction-tuned LLMs, tested under standard prompting, zero-shot chain-of-thought reasoning, and a biomedical question answering setting. In all cases, models retain similar accuracy but exhibit significantly degraded calibration. These results suggest that current UQ practices for MCQs are ''one-token deep'', driven more by first-token decoding behavior than by any deeper representation of uncertainty, and are easily manipulated through minimal interventions. Our findings call for more robust and interpretable approaches to uncertainty estimation, particularly in structured formats like MCQs, where confidence signals are often reduced to token-level heuristics. Qingcheng Zeng, Mingyu Jin, Qinkai Yu, Zhenting Wang, Wenyue Hua, Guangyan Sun, Yanda Meng, Shiqing Ma, Qifan Wang 0001, Felix Juefei-Xu, Fan Yang 0023, Kaize Ding, Ruixiang Tang, Yongfeng Zhang 0003 |
CIKM | 7 |
| 2025 | Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers?abstractLarge language models (LLMs) have shown remarkable performances across a wide range of tasks. However, the mechanisms by which these models encode tasks of varying complexities remain poorly understood. In this paper, we explore the hypothesis that LLMs process concepts of varying complexities in different layers, introducing the idea of “Concept Depth” to suggest that more complex concepts are typically acquired in deeper layers. Specifically, we categorize concepts based on their level of abstraction, defining them in the order of increasing complexity within factual, emotional, and inferential tasks. We conduct extensive probing experiments using layer-wise representations across various LLM families (Gemma, LLaMA, Qwen) on various datasets spanning the three domains of tasks. Our findings reveal that models could efficiently conduct probing for simpler tasks in shallow layers, and more complex tasks typically necessitate deeper layers for accurate understanding. Additionally, we examine how external factors, such as adding noise to the input and quantizing the model weights, might affect layer-wise representations. Our findings suggest that these factors can impede the development of a conceptual understanding of LLMs until deeper layers are explored. We hope that our proposed concept and experimental insights will enhance the understanding of the mechanisms underlying LLMs. Our codes are available at https://github.com/Luckfort/CD. Mingyu Jin, Qinkai Yu, Qingcheng Zeng, Zhenting Wang, Wenyue Hua, Haiyan Zhao 0003, Kai Mei, Yanda Meng, Kaize Ding, Fan Yang 0023, Mengnan Du, Yongfeng Zhang 0003 |
COLING | 9 |
| 2025 | Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?abstractSpatial-temporal graph convolutional networks (ST-GCNs) showcase impressive performance in skeleton-based human action recognition (HAR). However, despite the development of numerous models, their recognition performance does not differ significantly after aligning the input settings. With this observation, we hypothesize that ST-GCNs are over-parameterized for HAR, a conjecture subsequently confirmed through experiments employing the lottery ticket hypothesis. Additionally, a novel sparse ST-GCNs generator is proposed, which trains a sparse architecture from a randomly initialized dense network while maintaining comparable performance levels to the dense components. Moreover, we generate multi-level sparsity ST-GCNs by integrating sparse structures at various sparsity levels and demonstrate that the assembled model yields a significant enhancement in HAR performance. Thorough experiments on four datasets, including NTU-RGB+D 60(120), Kinetics-400, and FineGYM, demonstrate that the proposed sparse ST-GCNs can achieve comparable performance to their dense components. Even with 95% fewer parameters, the sparse ST-GCNs exhibit a degradation of1% in top-1 accuracy. The code is available at https://github.com/davelailai/Sparse-ST-GCN. Jianyang Xie, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng |
CVPR | 3 |
| 2025 | tHPM-LDM: Integrating Individual Historical Record with Population Memory in Latent Diffusion-Based Glaucoma Forecasting
Jianyang Xie, Yimin Luo, Yanda Meng, Savita Madhusudhan, Gregory Yoke Hong Lip, Li Cheng 0001, Yalin Zheng, He Zhao 0002 |
MICCAI (1) | 4 |
| 2025 | Self-adaptive Vision-Language Model for 3D Segmentation of Pulmonary Artery and Vein
Deqian Yang, Zhilin Sui, Yanda Meng |
MICCAI (9) | 11 |
| 2025 | Fairness-Aware vCDR-Controlled Generation for Glaucoma Diagnosis
Shuran Yang, Feixiang Zhou, Meng Wang 0038, Yitian Zhao, Yalin Zheng, Yanda Meng |
MICCAI (9) | 11 |
| 2025 | Memory-Augmented SAM2 for Training-Free Surgical Video Segmentation
Xujiong Ye, Yanda Meng, Zeyu Fu |
MICCAI (10) | 4 |
| 2025 | Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport
Qinkai Yu, Jianyang Xie, Yitian Zhao, Cheng Chen 0013, Jun Cheng 0003, Lu Liu 0001, Yalin Zheng, Yanda Meng |
MICCAI (15) | 10 |
| 2025 | Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng |
MICCAI (15) | 10 |
| 2025 | GLCP: Global-to-Local Connectivity Preservation for Tubular Structure Segmentation
Feixiang Zhou, Zhuangzhi Gao, He Zhao 0002, Jianyang Xie, Yanda Meng, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng |
MICCAI (16) | 5 |
| 2025 | $\text{MR}^{2}$-Net: Retinal OCTA Image Stitching via Multi-Scale Representation Learning and Dynamic Location GuidanceabstractOptical coherence tomography angiography (OCTA) plays a crucial role in quantifying and analyzing retinal vascular diseases. However, the limited field of view (FOV) inherent in most commercial OCTA imaging systems poses a significant challenge for clinicians, restricting the possibility to analyze larger retinal regions of high resolution. Automatic stitching of OCTA scans in adjacent regions may provide a promising solution to extend the region of interest. However, commonly-used stitching algorithms face difficulties in achieving effective alignment due to noise, artifacts and dense vasculature present in OCTA images. To address these challenges, we propose a novel retinal OCTA image stitching network, named -Net, which integrates multi-scale representation learning and dynamic location guidance. In the first stage, an image registration network with a progressive multi-resolution feature fusion is proposed to derive deep semantic information effectively. Additionally, we introduce a dynamic guidance strategy to locate the foveal avascular zone (FAZ) and constrain registration errors in overlapping vascular regions. In the second stage, an image fusion network based on multiple mask constraints and adjacent image aggregation (AIA) strategies is developed to further eliminate the artifacts in the overlapping areas of stitched images, thereby achieving precise vessel alignment. To validate the effectiveness of our method, we conduct a series of experiments on two delicately constructed datasets, i.e., OPTOVUE-OCTA and SVision-OCTA. Experimental results demonstrate that our method outperforms other image stitching methods and effectively generates high-quality wide-field OCTA images, achieving a structural similarity index (SSIM) score of 0.8264 and 0.8014 on the two datasets, respectively. Haiting Mao, Yuhui Ma, Dan Zhang 0026, Yanda Meng, Shaodong Ma, Yuchuan Qiao, Huazhu Fu, Caifeng Shan, Da Chen 0002, Yitian Zhao, Jiong Zhang 0004 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Randomness-Restricted Diffusion Model for Ocular Surface Structure SegmentationabstractOcular surface diseases affect a significant portion of the population worldwide. Accurate segmentation and quantification of different ocular surface structures are crucial for the understanding of these diseases and clinical decision-making. However, the automated segmentation of the ocular surface structure is relatively unexplored and faces several challenges. Ocular surface structure boundaries are often inconspicuous and obscured by glare from reflections. In addition, the segmentation of different ocular structures always requires training of multiple individual models. Thus, developing a one-model-fits-all segmentation approach is desirable. In this paper, we introduce a randomness-restricted diffusion model for multiple ocular surface structure segmentation. First, a time-controlled fusion-attention module (TFM) is proposed to dynamically adjust the information flow within the diffusion model, based on the temporal relationships between the network's input and time. TFM enables the network to effectively utilize image features to constrain the randomness of the generation process. We further propose a low-frequency consistency filter and a new loss to alleviate model uncertainty and error accumulation caused by the multi-step denoising process. Extensive experiments have shown that our approach can segment seven different ocular surface structures. Our method performs better than both dedicated ocular surface segmentation methods and general medical image segmentation methods. We further validated the proposed method over two clinical datasets, and the results demonstrated that it is beneficial to clinical applications, such as the meibomian gland dysfunction grading and aqueous deficient dry eye diagnosis. Huaying Hao, Yifan Zhao 0001, Yanda Meng, Jiang Liu 0001, Yalin Zheng, Wei Chen 0089, Yitian Zhao |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action RecognitionabstractGraph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasible to represent the semantic information. In this paper, we proposed a dynamic semantic-based graph convolution network (DS-GCN) for skeleton-based human action recognition, where the joints and edge types were encoded in the skeleton topology in an implicit way. Specifically, two semantic modules, the joints type-aware adaptive topology and the edge type-aware adaptive topology, were proposed. Combining proposed semantics modules with temporal convolution, a powerful framework named DS-GCN was developed for skeleton-based action recognition. Extensive experiments in two datasets, NTU-RGB+D and Kinetics-400 show that the proposed semantic modules were generalized enough to be utilized in various backbones for boosting recognition accuracy. Meanwhile, the proposed DS-GCN notably outperformed state-of-the-art methods. The code is released here https://github.com/davelailai/DS-GCN Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng |
AAAI | 2 |
| 2024 | Clinical Insight-Augmented Multi-View Learning for Alzheimer's Detection in Retinal OCTA ImagesabstractAlzheimer’s disease (AD) poses a significant global challenge, with a notable absence of accessible and cost-effective diagnostic tools for widespread AD detection. The retina, mirroring the brain in anatomy and physiology, has emerged as a potential avenue for rapid AD identification through retinal imaging. The current retinal image-based AD detection methods usually focus primarily on the macular area, but ignore the potential value that the optic disc region may have for the detection task. In this study, we leverage both macular- and disc-centered OCTA images and propose a multi-region fusion framework for AD detection. Based on clinical evidence, we integrate handcrafted features into the framework to improve model performance and interpretability. Specifically, vascular morphological parameters extracted from the macular and disc regions are used as input to a revalued KNN model to improve predictive capabilities. Furthermore, recognizing the significance of extracting and utilizing complementary information from the macular and optic disc regions, we propose an uncertainty-guided strategy based on Dempster-Shefer Theory (DST) to fuse knowledge from different regions. This approach considers each region’s forecast quality and significantly improves the effectiveness and robustness of the model. Through comparative analysis with existing methods, we have demonstrated that our method outperforms the state-of-the-art ones and provides more valuable pathological evidence for the association between retinal vascular changes and AD. Yuandi Zhang, Jinkui Hao, Botian Zheng, Yonghuai Liu, Yanda Meng, Jiong Zhang 0004, Yang Chen 0008, Yitian Zhao |
BIBM | 5 |
| 2024 | Multi-disease Detection in Retinal Images Guided by Disease Causal Estimation
Jianyang Xie, Xiuju Chen, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng |
MICCAI (1) | 4 |
| 2024 | CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-Aware Prompting
Qinkai Yu, Jianyang Xie, Anh Nguyen 0003, He Zhao 0002, Jiong Zhang 0004, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng |
MICCAI (1) | 9 |
| 2024 | Multi-granularity learning of explicit geometric constraint and contrast for label-efficient medical image segmentation and differentiable clinical function assessmentabstractAutomated segmentation is a challenging task in medical image analysis that usually requires a large amount of manually labeled data. However, most current supervised learning based algorithms suffer from insufficient manual annotations, posing a significant difficulty for accurate and robust segmentation. In addition, most current semi-supervised methods lack explicit representations of geometric structure and semantic information, restricting segmentation accuracy. In this work, we propose a hybrid framework to learn polygon vertices, region masks, and their boundaries in a weakly/semi-supervised manner that significantly advances geometric and semantic representations. Firstly, we propose multi-granularity learning of explicit geometric structure constraints via polygon vertices (PolyV) and pixel-wise region (PixelR) segmentation masks in a semi-supervised manner. Secondly, we propose eliminating boundary ambiguity by using an explicit contrastive objective to learn a discriminative feature space of boundary contours at the pixel level with limited annotations. Thirdly, we exploit the task-specific clinical domain knowledge to differentiate the clinical function assessment end-to-end. The ground truth of clinical function assessment, on the other hand, can serve as auxiliary weak supervision for PolyV and PixelR learning. We evaluate the proposed framework on two tasks, including optic disc (OD) and cup (OC) segmentation along with vertical cup-to-disc ratio (vCDR) estimation in fundus images; left ventricle (LV) segmentation at end-diastolic and end-systolic frames along with ejection fraction (LVEF) estimation in two-dimensional echocardiography images. Experiments on nine large-scale datasets of the two tasks under different label settings demonstrate our model’s superior performance on segmentation and clinical function assessment. Yanda Meng, Jianyang Xie, Jinming Duan 0001, Martha Joddrell, Savita Madhusudhan, Tunde Peto, Yitian Zhao, Yalin Zheng |
Medical Image Anal. | 1 |
| 2024 | Dynamic Semantic-Based Spatial-Temporal Graph Convolution Network for Skeleton-Based Human Action RecognitionabstractHuman action recognition is an essential topic in computer vision and image processing. Graph convolutional networks (GCNs) have attracted significant attention and achieved noteworthy performance in skeleton-based human action recognition tasks. However, most of the previous graph-based works are designed to refine skeleton topology without considering the types of different joints and edges and the occurrence order of the frames. Such a limitation makes them insufficient to represent intrinsic semantic information. Differently, we proposed a dynamic semantic-based spatial-temporal graph convolution network (DS-STGCN) to address the challenge. DS-STGCN has two dynamic semantic modules for spatial and temporal contexts respectively. Specifically, the joints and edge types were encoded in the spatial module implicitly, and the occurrence order of frames was encoded in the temporal module implicitly. Extensive experiments on four datasets including NTU-RGB+D 60(120), Kinetics-400, and FineGYM show that our proposed two semantic modules can bring consistent recognition performance improvement with various backbones. Meanwhile, the proposed DS-STGCN notably surpassed state-of-the-art methods on these datasets. Notably, in the more challenging dataset, such as Kinetics-400, our model significantly outperformed other state-of-the-art GCN-based methods by a large margin. The code has been released at https://github.com/davelailai/DS-STGCN. Jianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 0003, Xiaoyun Yang, Yalin Zheng |
IEEE Trans. Image Process. | 2 |
| 2023 | Weakly Supervised Segmentation with Point Annotations for Histopathology Images via Contrast-Based Variational ModelabstractImage segmentation is a fundamental task in the field of imaging and vision. Supervised deep learning for segmentation has achieved unparalleled success when sufficient training data with annotated labels are available. However, annotation is known to be expensive to obtain, especially for histopathology images where the target regions are usually with high morphology variations and irregular shapes. Thus, weakly supervised learning with sparse annotations of points is promising to reduce the annotation workload. In this work, we propose a contrast-based variational model to generate segmentation results, which serve as reliable complementary supervision to train a deep segmentation model for histopathology images. The proposed method considers the common characteristics of target regions in histopathology images and can be trained in an end-to-end manner. It can generate more regionally consistent and smoother boundary segmentation, and is more robust to unlabeled ‘novel’ regions. Experiments on two different histology datasets demonstrate its effectiveness and efficiency in comparison to previous models. Code is available at: https://github.com/hrzhang1123/CVM_WS_Segmentation. Hongrun Zhang, Liam Burrows, Yanda Meng, Declan Sculthorpe, Abhik Mukherjee, Sarah E. Coupland, Ke Chen 0002, Yalin Zheng |
CVPR | 3 |
| 2023 | Weakly/Semi-supervised Left Ventricle Segmentation in 2D Echocardiography with Uncertain Region-Aware Contrastive Learning
Yanda Meng, Jianyang Xie, Jinming Duan 0001, Yitian Zhao, Yalin Zheng |
PRCV (13) | 1 |
| 2023 | Bilateral adaptive graph convolutional network on CT based Covid-19 diagnosis with uncertainty-aware consensus-assisted multiple instance learningabstractCoronavirus disease (COVID-19) has caused a worldwide pandemic, putting millions of people's health and lives in jeopardy. Detecting infected patients early on chest computed tomography (CT) is critical in combating COVID-19. Harnessing uncertainty-aware consensus-assisted multiple instance learning (UC-MIL), we propose to diagnose COVID-19 using a new bilateral adaptive graph-based (BA-GCN) model that can use both 2D and 3D discriminative information in 3D CT volumes with arbitrary number of slices. Given the importance of lung segmentation for this task, we have created the largest manual annotation dataset so far with 7,768 slices from COVID-19 patients, and have used it to train a 2D segmentation model to segment the lungs from individual slices and mask the lungs as the regions of interest for the subsequent analyses. We then used the UC-MIL model to estimate the uncertainty of each prediction and the consensus between multiple predictions on each CT slice to automatically select a fixed number of CT slices with reliable predictions for the subsequent model reasoning. Finally, we adaptively constructed a BA-GCN with vertices from different granularity levels (2D and 3D) to aggregate multi-level features for the final diagnosis with the benefits of the graph convolution network's superiority to tackle cross-granularity relationships. Experimental results on three largest COVID-19 CT datasets demonstrated that our model can produce reliable and accurate COVID-19 predictions using CT volumes with any number of slices, which outperforms existing approaches in terms of learning and generalisation ability. To promote reproducible research, we have made the datasets, including the manual annotations and cleaned CT dataset, as well as the implementation code, available at https://doi.org/10.5281/zenodo.6361963. Yanda Meng, Joshua Bridge, Cliff Addison, Manhui Wang, Cristin Merritt, Stu Franks, Maria Mackey, Steve Messenger, Renrong Sun, Thomas Fitzmaurice, Caroline McCann, Yitian Zhao, Yalin Zheng |
Medical Image Anal. | 1 |
| 2023 | Transportation Object Counting With Graph-Based Adaptive Auxiliary LearningabstractThis paper proposes an adaptive auxiliary task learning-based approach for transport object counting problems such as humans and vehicles. These problems are essential in many real-world tasks such as video surveillance, traffic monitoring, public security, and urban planning, to aid intelligent transportation systems. Unlike existing auxiliary task learning-based methods, we develop an attention-enhanced adaptively shared backbone network to enable both task-shared and task-tailored features that are learned in an end-to-end manner. The network seamlessly combines a standard Convolution Neural Network (CNN) and a Graph Convolution Network (GCN) for feature extraction and feature reasoning among different domains of tasks. Our approach gains enriched contextual information by iteratively and hierarchically fusing features across different task branches of the adaptive CNN backbone. The whole framework pays special attention to objects’ spatial locations and varied density levels, informed by object (or crowd) segmentation and density level segmentation auxiliary tasks. In particular, thanks to the proposed dilated contrastive density loss function, our network benefits from individual and regional context supervision, along with strengthened robustness. Experiments on six challenging multi-domain datasets demonstrate that our method achieves superior performance compared with state-of-the-art auxiliary task learning-based counting methods. Our code is publicly available. Yanda Meng, Joshua Bridge, Yitian Zhao, Martha Joddrell, Yihong Qiao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Dual Consistency Enabled Weakly and Semi-Supervised Optic Disc and Cup Segmentation With Dual Adaptive Graph Convolutional NetworksabstractGlaucoma is a progressive eye disease that results in permanent vision loss, and the vertical cup to disc ratio (vCDR) in colour fundus images is essential in glaucoma screening and assessment. Previous fully supervised convolution neural networks segment the optic disc (OD) and optic cup (OC) from color fundus images and then calculate the vCDR offline. However, they rely on a large set of labeled masks for training, which is expensive and time-consuming to acquire. To address this, we propose a weakly and semi-supervised graph-based network that investigates geometric associations and domain knowledge between segmentation probability maps (PM), modified signed distance function representations (mSDF), and boundary region of interest characteristics (B-ROI) in three aspects. Firstly, we propose a novel Dual Adaptive Graph Convolutional Network (DAGCN) to reason the long-range features of the PM and the mSDF w.r.t. the regional uniformity. Secondly, we propose a dual consistency regularization-based semi-supervised learning paradigm. The regional consistency between the PM and the mSDF, and the marginal consistency between the derived B-ROI from each of them boost the proposed model's performance due to the inherent geometric associations. Thirdly, we exploit the task-specific domain knowledge via the oval shapes of OD & OC, where a differentiable vCDR estimating layer is proposed. Furthermore, without additional annotations, the supervision on vCDR serves as weakly-supervisions for segmentation tasks. Experiments on six large-scale datasets demonstrate our model's superior performance on OD & OC segmentation and vCDR estimation. The implementation code has been made available.https://github.com/smallmax00/Dual_Adaptive_Graph_Reasoning. Yanda Meng, Hongrun Zhang, Yitian Zhao, Dongxu Gao, Barbra Hamill, Godhuli Patri, Tunde Peto, Savita Madhusudhan, Yalin Zheng |
IEEE Trans. Medical Imaging | 1 |
| 2023 | 3D Human Pose and Shape Reconstruction From Videos via Confidence-Aware Temporal Feature AggregationabstractEstimating 3D human body shapes and poses from videos is a challenging computer vision task. The intrinsic temporal information embedded in adjacent frames is helpful in making accurate estimations. Existing approaches learn temporal features of the target frames simply by aggregating features of their adjacent frames, using off-the-shelf deep neural networks. Consequently these approaches cannot explicitly and effectively use the correlations between adjacent frames to help infer the parameters of the target frames. In this paper, we propose a novel framework that can measure the correlations amongst adjacent frames in the form of an estimated confidence metric. The confidence value will indicate to what extent the adjacent frames can help predict the target frames’ 3D shapes and poses. Based on the estimated confidence values, temporally aggregated features are then obtained by adaptively allocating different weights to the temporal predicted features from the adjacent frames. The final 3D shapes and poses are estimated by regressing from the temporally aggregated features. Experimental results on three benchmark datasets show that the proposed method outperforms state-of-the-art approaches (even without the motion priors involved in training). In particular, the proposed method is more robust against corrupted frames. Hongrun Zhang, Yanda Meng, Yitian Zhao, Xuesheng Qian, Yihong Qiao, Xiaoyun Yang, Yalin Zheng |
IEEE Trans. Multim. | 2 |
| 2022 | DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image ClassificationabstractMultiple instance learning (MIL) has been increasingly used in the classification of histopathology whole slide images (WSIs). However, MIL approaches for this specific classification problem still face unique challenges, particularly those related to small sample cohorts. In these, there are limited number of WSI slides (bags), while the resolution of a single WSI is huge, which leads to a large number of patches (instances) cropped from this slide. To address this issue, we propose to virtually enlarge the number of bags by introducing the concept of pseudo-bags, on which a double-tier MIL framework is built to effectively use the intrinsic features. Besides, we also contribute to deriving the instance probability under the framework of attentionbased MIL, and utilize the derivation to help construct and analyze the proposed framework. The proposed method outperforms other latest methods on the CAMELYON-16 by substantially large margins, and is also better in performance on the TCGA lung cancer dataset. The proposed framework is ready to be extended for wider MIL applications. The code is available at: https://github. com/hrzhang1123/DTFD-MIL. Hongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao, Xiaoyun Yang, Sarah E. Coupland, Yalin Zheng |
CVPR | 2 |
| 2022 | Shape-Aware Weakly/Semi-Supervised Optic Disc and Cup Segmentation with Regional/Marginal Consistency
Yanda Meng, Xu Chen 0030, Hongrun Zhang, Yitian Zhao, Dongxu Gao, Barbra Hamill, Godhuli Patri, Tunde Peto, Savita Madhusudhan, Yalin Zheng |
MICCAI (4) | 1 |
| 2022 | Graph-Based Region and Boundary Aggregation for Biomedical Image SegmentationabstractSegmentation is a fundamental task in biomedical image analysis. Unlike the existing region-based dense pixel classification methods or boundary-based polygon regression methods, we build a novel graph neural network (GNN) based deep learning framework with multiple graph reasoning modules to explicitly leverage both region and boundary features in an end-to-end manner. The mechanism extracts discriminative region and boundary features, referred to as initialized region and boundary node embeddings, using a proposed Attention Enhancement Module (AEM). The weighted links between cross-domain nodes (region and boundary feature domains) in each graph are defined in a data-dependent way, which retains both global and local cross-node relationships. The iterative message aggregation and node update mechanism can enhance the interaction between each graph reasoning module's global semantic information and local spatial characteristics. Our model, in particular, is capable of concurrently addressing region and boundary feature reasoning and aggregation at several different feature levels due to the proposed multi-level feature node embeddings in different parallel graph reasoning modules. Experiments on two types of challenging datasets demonstrate that our method outperforms state-of-the-art approaches for segmentation of polyps in colonoscopy images and of the optic disc and optic cup in colour fundus images. The trained models will be made available at: https://github.com/smallmax00/Graph_Region_Boudnary. Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Yihong Qiao, Ian J. C. MacCormick, Xiaowei Huang 0001, Yalin Zheng |
IEEE Trans. Medical Imaging | 1 |
| 2021 | BI-GCN: Boundary-Aware Input-Dependent Graph Convolution Network for Biomedical Image Segmentation
Yanda Meng, Hongrun Zhang, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng |
BMVC | 1 |
| 2021 | Spatial Uncertainty-Aware Semi-Supervised Crowd CountingabstractSemi-supervised approaches for crowd counting attract attention, as the fully supervised paradigm is expensive and laborious due to its request for a large number of images of dense crowd scenarios and their annotations. This paper proposes a spatial uncertainty-aware semi-supervised approach via regularized surrogate task (binary segmentation) for crowd counting problems. Different from existing semi-supervised learning-based crowd counting methods, to exploit the unlabeled data, our proposed spatial uncertainty-aware teacher-student framework focuses on high confident regions’ information while addressing the noisy supervision from the unlabeled data in an end-to-end manner. Specifically, we estimate the spatial uncertainty maps from the teacher model’s surrogate task to guide the feature learning of the main task (density regression) and the surrogate task of the student model at the same time. Besides, we introduce a simple yet effective differential transformation layer to enforce the inherent spatial consistency regularization between the main task and the surrogate task in the student model, which helps the surrogate task to yield more reliable predictions and generates high-quality uncertainty maps. Thus, our model can also address the task-level perturbation problems that occur spatial inconsistency between the primary and surrogate tasks in the student model. Experimental results on four challenging crowd counting datasets demonstrate that our method achieves superior performance to the state-of-the-art semi-supervised methods. Code is available at : https://github.com/smallmax00/SUA_crowd_counting Yanda Meng, Hongrun Zhang, Yitian Zhao, Xiaoyun Yang, Xuesheng Qian, Xiaowei Huang 0001, Yalin Zheng |
ICCV | 1 |
| 2021 | Learning Unsupervised Parameter-Specific Affine Transformation for Medical Images Registration
Xu Chen 0030, Yanda Meng, Yitian Zhao, Rachel Williams, Srinivasa R. Vallabhaneni, Yalin Zheng |
MICCAI (4) | 2 |
| 2020 | Regression of Instance Boundary by Aggregated CNN and GCN
Yanda Meng, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng |
ECCV (8) | 1 |
| 2020 | CNN-GCN Aggregation Enabled Boundary Regression for Biomedical Image Segmentation
Yanda Meng, Dongxu Gao, Yitian Zhao, Xiaoyun Yang, Xiaowei Huang 0001, Yalin Zheng |
MICCAI (4) | 1 |
| 2020 | Introducing the GEV Activation Function for Highly Unbalanced Data to Develop COVID-19 Diagnostic ModelsabstractFast and accurate diagnosis is essential for the efficient and effective control of the COVID-19 pandemic that is currently disrupting the whole world. Despite the prevalence of the COVID-19 outbreak, relatively few diagnostic images are openly available to develop automatic diagnosis algorithms. Traditional deep learning methods often struggle when data is highly unbalanced with many cases in one class and only a few cases in another; new methods must be developed to overcome this challenge. We propose a novel activation function based on the generalized extreme value (GEV) distribution from extreme value theory, which improves performance over the traditional sigmoid activation function when one class significantly outweighs the other. We demonstrate the proposed activation function on a publicly available dataset and externally validate on a dataset consisting of 1,909 healthy chest X-rays and 84 COVID-19 X-rays. The proposed method achieves an improved area under the receiver operating characteristic (DeLong's p-value < 0.05) compared to the sigmoid activation. Our method is also demonstrated on a dataset of healthy and pneumonia vs. COVID-19 X-rays and a set of computerized tomography images, achieving improved sensitivity. The proposed GEV activation function significantly improves upon the previously used sigmoid activation for binary classification. This new paradigm is expected to play a significant role in the fight against COVID-19 and other diseases, with relatively few training cases available. Joshua Bridge, Yanda Meng, Yitian Zhao, Mingfeng Zhao, Renrong Sun, Yalin Zheng |
IEEE J. Biomed. Health Informatics | 2 |