EDBT 2026 Demo / reviewers in the wild / expert
Wenting Chen
dblp:135/7011 · also Wen-Ting Chen
· DBLP profile ↗
46ranked-venue papers
13as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse AttentionabstractMedical vision-language pre-training (VLP) offers significant potential for advancing medical image understanding by leveraging paired image-report data. However, existing methods are limited by False Negatives (FaNe) induced by semantically similar texts and insufficient fine-grained cross-modal alignment. To address these limitations, we propose FaNe, a semantic-enhanced VLP framework. To mitigate false negatives, we introduce a semantic-aware positive pair mining strategy based on text-text similarity with adaptive normalization. Furthermore, we design a text-conditioned sparse attention pooling module to enable fine-grained image-text alignment through localized visual representations guided by textual cues. To strengthen intra-modal discrimination, we develop a hard-negative aware contrastive loss that adaptively reweights semantically similar negatives. Extensive experiments on five downstream medical imaging benchmarks demonstrate that FaNe achieves state-of-the-art performance across image classification, object detection, and semantic segmentation, validating the effectiveness of our framework. Peng Zhang 0057, Zhihui Lai 0001, Wenting Chen, Xu Wu 0001, Heng Kong |
AAAI | 3 |
| 2026 | MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential DiagnosisabstractDespite achieving high accuracy on medical benchmarks, LLMs exhibit the Einstellung Effect in clinical diagnosis—relying on statistical shortcuts rather than patient-specific evidence, causing misdiagnosis in atypical cases. Existing benchmarks fail to detect this critical failure mode. We introduce MedEinst, a counterfactual benchmark with 5,383 paired clinical cases across 49 diseases. Each pair contains a control case and a “trap” case with altered discriminative evidence that flips the diagnosis. We measure susceptibility via Bias Trap Rate—probability of misdiagnosing traps despite correctly diagnosing controls. Evaluation shows frontier models achieve high baseline accuracy but severe bias trap rates. Thus, we propose ECR-Agent, aligning LLM reasoning with Evidence-Based Medicine via two components: (1) Dynamic Causal Inference (DCI) performs structured reasoning through dual-pathway perception, dynamic causal graph reasoning across three levels (association, intervention, counterfactual), and evidence audit for final diagnosis; (2) Critic-Driven Graph Memory Evolution (CGME) iteratively refines the system by storing validated reasoning paths in an exemplar base and consolidating disease-specific knowledge into evolving illness graphs. Source code is to be released. Wenting Chen, Guolin Huang, Wenxuan Wang 0001, Zhongrui Zhu |
ACL (1) | 1 |
| 2026 | Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language ModelsabstractShaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaonan Liu, Xiaoling Luo 0001, Shiyi Zheng, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 6 |
| 2026 | Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsabstractWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung, Meidan Ding, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenxuan Wang 0001, Zizhan Ma, Yiu-Fai Cheung, Meidan Ding, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 7 |
| 2026 | Enhancing mutation testing for deep neural networks: a novel approach to generating high-quality mutants
Yongming Yao, Wenting Chen |
Autom. Softw. Eng. | 5 |
| 2026 | MNSeg: Mamba-based 3D neuron segmentation integrated with bidirectional attention mechanism and topological loss
Xinle Dai, Qiufu Li, LinLin Shen, Wenting Chen, Weijia Fan |
Neurocomputing | 4 |
| 2026 | FProtoSeg: Fine-grained prototype alignment for Weakly Supervised Semantic Segmentation of histopathology images
Meidan Ding, Wenting Chen, Xiaoling Luo 0001, Haiqin Zhong, LinLin Shen |
Pattern Recognit. | 2 |
| 2026 | Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography
Leilei Zeng, Jie Liu 0044, Wenting Chen, Chenyang Lyu, Wenxi Li, Shaonan Liu, Xiande Zhou, LinLin Shen |
Pattern Recognit. | 3 |
| 2026 | ACGM: Attribute-Centric Graph Modeling Network for Concurrent Missing Tabular Data Imputation and COVID-19 PrognosisabstractCOVID-19 prognosis using clinical tabular data faces significant challenges due to missing values and class imbalance issues. Existing methods often overlook the complex high-order interrelationship among clinical attributes and struggle with training stability on imbalanced datasets. We propose ACGM, an attribute-centric graph modeling network that simultaneously addresses missing data imputation and COVID-19 prognosis. ACGM consists of three key modules: an attributes preprocessing module (APM) for coarse-grained imputation initialization, a graph-enhanced attributes imputation module (GEAIM) that models high-order inter-attribute relationships through graph structures, and a graph-enhanced disease prognosis module (GEDPM) that leverages these complex attribute interactions for final prediction. GEAIM and GEDPM employ a mean-teacher strategy with attributes graph matching to preserve high-order relationships, enhance training stability, and maintain structural integrity of attribute interactions. Extensive experiments are conducted on four public COVID-19 tabular datasets, demonstrating the superiority of our ACGM over existing methods. Through comprehensive interpretability analysis, we identify that attributes such as LDH, Difficulty In Breathing, and SaO2 significantly impact COVID-19 prognosis, aligning well with clinical insights and radiologist assessments. Zhuoru Wu, Wenting Chen, Xuechen Li 0001, Filippo Ruffini, Shaonan Liu, Lorenzo Tronchin, Domenico Albano, Eliodoro Faiella, Deborah Fazzini, Domiziana Santucci, Xiaoling Luo 0001, Valerio Guarrasi, Paolo Soda, LinLin Shen |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph MatchingabstractMedical report generation is crucial for clinical diagnosis and patient management, summarizing diagnoses and recommendations based on medical imaging. However, existing work often overlook the clinical pipeline involved in report writing, where physicians typically conduct an initial quick review followed by a detailed examination. Moreover, current alignment methods may lead to misaligned relationships. To address these issues, we propose DAMPER, a dual-stage framework for medical report generation that mimics the clinical pipeline of report writing in two stages. In the first stage, a MeSH-Guided Coarse-Grained Alignment (MCG) stage that aligns chest X-ray (CXR) image features with medical subject headings (MeSH) features to generate a rough keyphrase representation of the overall impression. In the second stage, a Hypergraph-Enhanced Fine-Grained Alignment (HFG) stage that constructs hypergraphs for image patches and report annotations, modeling high-order relationships within each modality and performing hypergraph matching to capture semantic correlations between image regions and textual phrases. Finally,the coarse-grained visual features, generated MeSH representations, and visual hypergraph features are fed into a report decoder to produce the final medical report. Extensive experiments on public datasets demonstrate the effectiveness of DAMPER in generating comprehensive and accurate medical reports, outperforming state-of-the-art methods across various evaluation metrics. Wenting Chen, Jie Liu 0044, Qisheng Lu, Xiaoling Luo 0001, LinLin Shen |
AAAI | 2 |
| 2025 | S³-Mamba: Small-Size-Sensitive Mamba for Lesion SegmentationabstractSmall lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small lesions. To tackle the challenges, we propose a Small-Size-Sensitive Mamba (S³-Mamba), which promotes the sensitivity to small lesions across three dimensions: channel, spatial, and training strategy. Specifically, an Enhanced Visual State Space block is designed to focus on small lesions through multiple residual connections to preserve local features, and selectively amplify important details while suppressing irrelevant ones through channel-wise attention. A Tensor-based Cross-feature Multi-scale Attention is designed to integrate input image features and intermediate-layer features with edge features and exploit the attentive support of features across multiple scales, thereby retaining spatial details of small lesions at various granularities. Finally, we introduce a novel regularized curriculum learning to automatically assess lesion size and sample difficulty, and gradually focus from easy samples to hard ones like small lesions. Extensive experiments on three medical image segmentation datasets show the superiority of our S³-Mamba, especially in segmenting small lesions. Gui Wang, Yuexiang Li, Wenting Chen, Meidan Ding, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, LinLin Shen |
AAAI | 3 |
| 2025 | EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language ModelabstractRecent advancements in Large Vision Language Models (LVLMs) show promise for pathological diagnosis, yet their application in clinical settings faces critical challenges of multimodal hallucination and biased responses. While preference alignment methods have proven effective in general domains, acquiring high-quality preference data for pathology remains challenging due to limited expert resources and domain complexity. In this paper, we propose EAGLE (Expert-guided self-enhancement for preference Alignment in patholoGy Large vision-languagE model), a novel framework that systematically integrates medical expertise into preference alignment. EAGLE consists of three key stages: initialization through supervised fine-tuning, self-preference creation leveraging expert prompting and medical entity recognition, and iterative preference following-tuning. The self-preference creation stage uniquely combines expert-verified chosen sampling with expert-guided rejected sampling to generate high-quality preference data, while the iterative tuning process continuously refines both data quality and model performance. Extensive experiments demonstrate that EAGLE significantly outperforms existing pathological LVLMs, effectively reducing hallucination and bias while maintaining pathological accuracy. The source code is available at https://github.com/meidandz/EAGLE. © 2025 Association for Computational Linguistics. Meidan Ding, Wenxuan Wang 0001, Haiqin Zhong, Xinheng Lyu, Wenting Chen, LinLin Shen |
ACL (1) | 7 |
| 2025 | Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsabstractThe significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gas-troenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments. © 2025 Association for Computational Linguistics. Jie Liu 0044, Wenxuan Wang 0001, Yihang Su, Yudi Zhang 0005, Cheng-Yi Li, Wenting Chen, Xiaohan Xing, Kao-Jung Chang, LinLin Shen, Michael R. Lyu |
ACL (1) | 7 |
| 2025 | FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual questionanswering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench Xusen Ma, Xianxu Hou, Meidan Ding, Yudong Li 0001, Junliang Chen 0002, Wenting Chen, Xiaoyang Peng, LinLin Shen |
CVPR | 7 |
| 2025 | WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageabstractRecent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy. Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Xiangjian He, Xiaohan Xing, Sen Yang 0006, LinLin Shen |
ICCV | 3 |
| 2025 | Metamorphic Testing for Audio Content Moderation SoftwareabstractThe rapid growth of audio-centric platforms and applications such as Whatsapp and Twitter has transformed the way people communicate and share audio content in modern society. However, these platforms are increasingly misused to disseminate harmful audio content, such as hate speech, deceptive advertisements, and explicit material, which can have significant negative consequences (e.g., detrimental effects on mental health). In response, researchers and practitioners have been actively developing and deploying audio content moderation tools to tackle this issue. Despite these efforts, malicious actors can bypass moderation systems by making subtle alterations to audio content, such as modifying pitch or inserting noise. Moreover, the effectiveness of modern audio moderation tools against such adversarial inputs remains insufficiently studied. To address these challenges, we propose MTAM, a Metamorphic Testing framework for Audio content Moderation software. Specifically, we conduct a pilot study on 2000 audio clips and define 14 metamorphic relations across two perturbation categories: Audio Features-Based and Heuristic perturbations. MTAM applies these metamorphic relations to toxic audio content to generate test cases that remain harmful while being more likely to evade detection. In our evaluation, we employ MTAM to test five commercial textual content moderation software and an academic model against three kinds of toxic content. The results show that MTAM achieves up to 38.6%, 18.3%, 35.1%, 16.7%, and 51.1% error finding rates (EFR) when testing commercial moderation software provided by Gladia, Assembly AI, Baidu, Nextdata, and Tencent respectively, and it obtains up to 45.7% EFR when testing the state-of-the-art algorithms from the academy. In addition, we leverage the test cases generated by MTAM to retrain the model we explored, which largely improves model robustness (nearly 0% EFR) while maintaining the accuracy on the original test set. We release the code and experiment data to facilitate future research1. Wenxuan Wang 0001, Yongjiang Wu, Junyuan Zhang, Shuqing Li 0001, Yun Peng 0003, Wenting Chen, Shuai Wang 0011, Michael R. Lyu |
ASE | 6 |
| 2025 | 🤖 WSI-Agents: A Collaborative Multi-agent System for Multi-modal Whole Slide Image Analysis
Xinheng Lyu, Yuci Liang, Wenting Chen, Meidan Ding, Guolin Huang, Daokun Zhang, Xiangjian He, LinLin Shen |
MICCAI (5) | 3 |
| 2025 | MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive SequenceabstractClinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive benchmark that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12,163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MedCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches. The relevant dataset and code will be released upon acceptance of this paper. Jie Liu 0044, Wenxuan Wang 0001, Zizhan Ma, Guolin Huang, Yihang Su, Kao-Jung Chang, Haoliang Li, LinLin Shen, Michael R. Lyu, Wenting Chen |
NeurIPS | 10 |
| 2025 | EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy AnalysisabstractEndoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are limited, as they typically cover specific endoscopic scenarios and a small set of clinical tasks, failing to capture the real-world diversity of endoscopic scenarios and the full range of skills needed in clinical workflows. To address these issues, we introduce EndoBench, the first comprehensive benchmark specifically designed to assess MLLMs across the full spectrum of endoscopic practice with multi-dimensional capacities. EndoBench encompasses 4 distinct endoscopic scenarios, 12 specialized clinical tasks with 12 secondary subtasks, and 5 levels of visual prompting granularities, resulting in 6,832 rigorously validated VQA pairs from 21 diverse datasets. Our multi-dimensional evaluation framework mirrors the clinical workflow—spanning anatomical recognition, lesion analysis, spatial localization, and surgical operations—to holistically gauge the perceptual and diagnostic abilities of MLLMs in realistic scenarios. We benchmark 23 state-of-the-art models, including general-purpose, medical-specialized, and proprietary MLLMs, and establish human clinician performance as a reference standard. Our extensive experiments reveal: (1) proprietary MLLMs outperform open-source and medical-specialized models overall, but still trail human experts; (2) medical-domain supervised fine-tuning substantially boosts task-specific accuracy; and (3) model performance remains sensitive to prompt format and clinical task complexity. EndoBench establishes a new standard for evaluating and advancing MLLMs in endoscopy, highlighting both progress and persistent gaps between current models and expert clinical reasoning. We publicly release our benchmark and code. Boyun Zheng, Wenting Chen, Zhihao Peng 0002, Zhenfei Yin, Jiancong Hu, Yixuan Yuan |
NeurIPS | 3 |
| 2025 | Bi-VLGM: Bi-Level Class-Severity-Aware Vision-Language Graph Matching for Text Guided Medical Image SegmentationabstractAbstract Medical reports containing specific diagnostic results and additional information not present in medical images can be effectively employed to assist image understanding tasks, and the modality gap between vision and language can be bridged by vision-language matching (VLM). However, current vision-language models distort the intra-model relation and only include class information in reports that is insufficient for segmentation task. In this paper, we introduce a novel Bi-level class-severity-aware Vision-Language Graph Matching (Bi-VLGM) for text guided medical image segmentation, composed of a word-level VLGM module and a sentence-level VLGM module, to exploit the class-severity-aware relation among visual-textual features. In word-level VLGM, to mitigate the distorted intra-modal relation during VLM, we reformulate VLM as graph matching problem and introduce a vision-language graph matching (VLGM) to exploit the high-order relation among visual-textual features. Then, we perform VLGM between the local features for each class region and class-aware prompts to bridge their gap. In sentence-level VLGM, to provide disease severity information for segmentation task, we introduce a severity-aware prompting to quantify the severity level of disease lesion, and perform VLGM between the global features and the severity-aware prompts. By exploiting the relation between the local (global) and class (severity) features, the segmentation model can include the class-aware and severity-aware information to promote segmentation performance. Extensive experiments proved the effectiveness of our method and its superiority to existing methods. The source code will be released. Wenting Chen, Jie Liu 0044, Tianming Liu 0001, Yixuan Yuan |
Int. J. Comput. Vis. | 1 |
| 2024 | Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report GenerationabstractFine-grained vision-language models (VLM) have been widely used for inter-modality local alignment between the predefined fixed patches and textual words. However, in medical analysis, lesions exhibit varying sizes and positions, and using fixed patches may cause incomplete representations of lesions. Moreover, these methods provide explainability by using heatmaps to show the general image areas potentially associated with texts rather than specific regions, making their explanations not explicit and specific enough. To address these issues, we propose a novel Adaptive patch-word Matching (AdaMatch) model to correlate chest X-ray (CXR) image regions with words in medical reports and apply it to CXR-report generation to provide explainability for the generation process. AdaMatch exploits the fine-grained relation between adaptive patches and words to provide explanations of specific image regions with corresponding words. To capture the abnormal regions of varying sizes and positions, we introduce an Adaptive Patch extraction (AdaPatch) module to acquire adaptive patches for these regions adaptively. Aiming to provide explicit explainability for the CXR-report generation task, we propose an AdaMatch-based bidirectional LLM for Cyclic CXR-report generation (AdaMatch-Cyclic). It employs AdaMatch to obtain the keywords for CXR images and 'keypatches' for medical reports as hints to guide CXR-report generation. Extensive experiments on two publicly available CXR datasets validate the effectiveness of our method and its superior performance over existing methods. © 2024 Association for Computational Linguistics. Wenting Chen, LinLin Shen, Jiebo Luo 0001, Xiang Li 0001, Yixuan Yuan |
ACL (1) | 1 |
| 2024 | Medical Image Synthesis via Fine-Grained Image-Text Alignment and Anatomy-Pathology Prompting
Wenting Chen, Pengyu Wang 0005, Hui Ren 0001, Lichao Sun 0001, Quanzheng Li, Yixuan Yuan, Xiang Li 0001 |
MICCAI (12) | 1 |
| 2024 | 💎 GEM: Context-Aware Gaze EstiMation with Visual Search Behavior Matching for Chest Radiograph
Shaonan Liu, Wenting Chen, Jie Liu 0044, Xiaoling Luo 0001, LinLin Shen |
MICCAI (1) | 2 |
| 2024 | Multi-Dataset Multi-Task Learning for COVID-19 Prognosis
Filippo Ruffini, Lorenzo Tronchin, Zhuoru Wu, Wenting Chen, Paolo Soda, LinLin Shen, Valerio Guarrasi |
MICCAI (12) | 4 |
| 2024 | Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningabstractIn the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework. Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
NeurIPS | 3 |
| 2024 | Two-stream regression network for dental implant position prediction
Xinquan Yang, Xuechen Li 0001, Wenting Chen, LinLin Shen, Xin Li 0196, Yongqiang Deng |
Expert Syst. Appl. | 4 |
| 2024 | SCPMan: Shape context and prior constrained multi-scale attention network for pancreatic segmentation
Leilei Zeng, Xuechen Li 0001, Xinquan Yang, Wenting Chen, Jingxin Liu 0005, LinLin Shen |
Expert Syst. Appl. | 4 |
| 2024 | Mask-aware transformer with structure invariant loss for CT translation
Wenting Chen, Wei Zhao 0040, Zhen Chen 0013, Tianming Liu 0001, Li Liu 0017, Jun Liu 0007, Yixuan Yuan |
Medical Image Anal. | 1 |
| 2024 | STAR-RL: Spatial-Temporal Hierarchical Reinforcement Learning for Interpretable Pathology Image Super-ResolutionabstractPathology image are essential for accurately interpreting lesion cells in cytopathology screening, but acquiring high-resolution digital slides requires specialized equipment and long scanning times. Though super-resolution (SR) techniques can alleviate this problem, existing deep learning models recover pathology image in a black-box manner, which can lead to untruthful biological details and misdiagnosis. Additionally, current methods allocate the same computational resources to recover each pixel of pathology image, leading to the sub-optimal recovery issue due to the large variation of pathology image. In this paper, we propose the first hierarchical reinforcement learning framework named Spatial-Temporal hierARchical Reinforcement Learning (STAR-RL), mainly for addressing the aforementioned issues in pathology image super-resolution problem. We reformulate the SR problem as a Markov decision process of interpretable operations and adopt the hierarchical recovery mechanism in patch level, to avoid sub-optimal recovery. Specifically, the higher-level spatial manager is proposed to pick out the most corrupted patch for the lower-level patch worker. Moreover, the higher-level temporal manager is advanced to evaluate the selected patch and determine whether the optimization should be stopped earlier, thereby avoiding the over-processed problem. Under the guidance of spatial-temporal managers, the lower-level patch worker processes the selected patch with pixel-wise interpretable actions at each time step. Experimental results on medical images degraded by different kernels show the effectiveness of STAR-RL. Furthermore, STAR-RL validates the promotion in tumor diagnosis with a large margin and shows generalizability under various degradations. The source code is available at https://github.com/CUHK-AIM-Group/STAR-RL. Wenting Chen, Jie Liu 0044, Tommy W. S. Chow, Yixuan Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Multi-scale Contrastive Learning for Gastroenteroscopy ClassificationabstractIn gastroenteroscopy image analysis, numerous CADs demonstrate that deep learning aids doctors' diagnosis. The shapes and sizes of the lesions are varied. And in the clinic, the dataset appears to be data imbalanced. However, existing methods directly classify by texture and ignore lesions with various shapes and sizes. To address the issue above, we propose a deep neural network, which consists of multi-scale feature extraction, contrastive feature learning and a multi-scale feature fusion module. We train the contrastive feature learning module and multi-scale feature fusion module simultaneously to alleviate the issue of data distribution differences. Thus, the proposed network can better identify various categories. Extensive experiments on the Hyper Kvasir dataset show that the proposed Hybrid-M2CL outperforms the benchmark proposed by the dataset with 5.0% Macro Precision, 3.3% Macro Recall, 3.4% Macro F1-score, 3.3% Micro Precision, 3.6% MCC. In addition, it outperforms the SOTA by 1.1% Macro F1-score, 2.6% MCC, and 2.0% B-ACC. Xuechen Li 0001, Zhibin Peng, Wenting Chen, LinLin Shen, Guangyao Wu |
CBMS | 4 |
| 2023 | Adversarial Keyword Extraction and Semantic-Spatial Feature Aggregation for Clinical Report Guided Thyroid Nodule Segmentation
Yudi Zhang 0005, Wenting Chen, Xuechen Li 0001, LinLin Shen, Zhihui Lai 0001, Heng Kong |
PRCV (13) | 2 |
| 2022 | RI-GCN: Review-aware Interactive Graph Convolutional Network for Review-based Item RecommendationabstractA wealth of semantic features exist in the reviews written by users, such as rich information on item features and implicit preferences of users. Existing review-based recommendation models usually employ Convolutional Neural Networks (CNNs) to learn representations of users and items from reviews. However, these CNNs-based models suffer from two main problems: (1) they only consider the information of the word itself during the convolution, ignoring the high-order contextual semantic information of the word; (2) they model user/item attributes in a static and independent way, ignoring the potential feature interaction between them. Therefore, we propose a novel Review-aware Interactive Graph Convolutional Network (RI-GCN) for review-based item recommendation. Specifically, we design a Review-aware GCN component to model the message propagation of graphs constructed from reviews, capturing the contextual features of words. A feature interactive GCN component is then proposed to capture the user/item high-order collaborative features in the user-item graph, enabling the model to further complement and refine u ser/item a ttributes. Finally, we adopt a Factorization Machine model for the recommendation task. Experimental results demonstrate that the proposed model is superior to state-of-the-art models. Yijin Cai, Weijin Wang, Wenting Chen |
IEEE Big Data | 4 |
| 2022 | TW-GAN: Topology and width aware GAN for retinal artery/vein classification
Wenting Chen, Kai Ma 0002, Wei Ji 0011, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
Medical Image Anal. | 1 |
| 2022 | Dynamic Depth-Aware Network for Endoscopy Super-ResolutionabstractEndoscopy super-resolution (SR) plays an important role in improving diagnostic results and reducing the misdiagnosis rate. Even though recent studies have investigated the SR for endoscopy, these methods apply equal importance to the whole image and do not consider the relationship among pixels, especially the depth information, which can provide diagnosis-related information for clinicians. To address this problem, we propose a dynamic depth-aware network for endoscopy super-resolution, which represents the first effort to comprehensively integrate the depth information to the SR task for endoscopic images. It includes a depth-wise feature extracting branch (DW-B) and a depth-guided SR branch (DGSR-B). The DW-B aims to extract the representative feature for each depth level (i.e. depth matrix) further to provide auxiliary information and guide the super-resolution of texture under different depth levels. In DGSR-B, a depth-guided block (DGB) consisting of depth-focus normalization (DFN) is introduced to inject both the depth matrix and depth map into the LR image feature, so as to guide the image generation for each depth region. To adaptively super-resolve the regions under different depth levels, we devise a dynamic depth-aware loss to assign different trainable weights to each region for SR optimization. Extensive experiments have been conducted on two main publicly available datasets, i.e., the Kvasir dataset and the EndoScene dataset, and the superior performance verifies the effectiveness of our method for SR task and polyp segmentation. Source code is to be released. Wenting Chen, Yifan Liu 0010, Jiancong Hu, Yixuan Yuan |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Gated SwitchGAN for Multi-Domain Facial Image TranslationabstractRecent studies on multi-domain facial image translation have achieved impressive results. The existing methods generally provide a discriminator with an auxiliary classifier to impose domain translation. However, these methods neglect important information regarding domain distribution matching. To solve this problem, we propose a switch generative adversarial network (SwitchGAN) with a more adaptive discriminator structure and a matched generator to perform delicate image translation among multiple domains. A feature-switching operation is proposed to achieve feature selection and fusion in our conditional modules. We demonstrate the effectiveness of our model. Furthermore, we also introduce a new capability of our generator that represents attribute intensity control and extracts content information without tailored training. Experiments on the Morph, RaFD and CelebA databases visually and quantitatively show that our extended SwitchGAN (i.e., Gated SwitchGAN) can achieve better translation results than StarGAN, AttGAN and STGAN. The attribute classification accuracy achieved using the trained ResNet-18 model and the FID score obtained using the ImageNet pretrained Inception-v3 model also quantitatively demonstrate the superior performance of our models. Yuanlue Zhu, Wenting Chen, Wenshuang Liu, LinLin Shen |
IEEE Trans. Multim. | 3 |
| 2021 | Translate the Facial Regions You Like Using Self-Adaptive Region TranslationabstractWith the progression of Generative Adversarial Networks (GANs), image translation methods has achieved increasingly remarkable performance. However, most available methods can only achieve image level translation, which is unable to precisely control the regions to be translated. In this paper, we propose a novel self-adaptive region translation network (SART) for region-level translation, which uses region-adaptive instance normalization (RIN) and a region matching loss (RML) for this task. We first encode the style and content image for each region with style and content encoder. To translate both shape and texture of the target region, we inject region-adaptive style features into the decoder by RIN. To ensure independent translation among different regions, RML is proposed to measure the similarity between the non-translated/translated regions of content and translated images. Extensive experiments on three publicly available datasets, i.e. Morph, RaFD and CelebAMask-HQ, suggest that our approach demonstrate obvious improvement over state-of-the-art methods like StarGAN, SEAN and FUNIT. Our approach has further advantages in precise control of the regions to be translated. As a result, region level expression changes and step-by-step make-up can be achieved. The video demo is available at (https://youtu.be/DvIdmcR2LEc). Wenshuang Liu, Wenting Chen, Zhanjia Yang, LinLin Shen |
AAAI | 2 |
| 2021 | Surrogate network-based sparseness hyper-parameter optimization for deep expression recognition
Weicheng Xie 0001, Wenting Chen, LinLin Shen, Jinming Duan 0001, Meng Yang 0001 |
Pattern Recognit. | 2 |
| 2020 | SATGAN: Augmenting Age Biased Dataset for Cross-Age Face RecognitionabstractIn this paper, we propose a Stable Age Translation GAN (SATGAN) to generate fake face images at different ages to augment age biased face datasets for Cross-Age Face Recognition (CAFR). The proposed SATGAN consists of both generator and discriminator. As a part of the generator, a novel Mask Attention Module (MAM) is introduced to make the generator focus on the face area. In addition, the generator employs a Uniform Distribution Discriminator (UDD) to supervise the learning of latent feature map and enforce the uniform distribution. Besides, the discriminator employs a Feature Separation Module (FSM) to disentangle identity information from the age information. The quantitative and qualitative evaluations on Morph dataset prove that SATGAN achieves much better performance than existing methods. The face recognition model trained using dataset (VGGFace2 and MS-Celeb-lM) augmented using our SATGAN achieves better accuracy on cross age dataset like Cross-Age LFW and AgeDB-30. Wenshuang Liu, Wenting Chen, Yuanlue Zhu, LinLin Shen |
ICPR | 2 |
| 2020 | TR-GAN: Topology Ranking GAN with Triplet Loss for Retinal Artery/Vein Classification
Wenting Chen, Kai Ma 0002, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (5) | 1 |
| 2020 | Leveraging Undiagnosed Data for Glaucoma Classification with Teacher-Student Learning
Wenting Chen, Kai Ma 0002, Hanruo Liu, Xiaoguang Di, Yefeng Zheng 0001 |
MICCAI (1) | 3 |
| 2019 | Introspective Gan for Meshface RecognitionabstractMajority of face recognition systems can only retrieve protected ID photos from the government agency in China. These facial images covered with mesh-like curves are termed meshface. Meshface could significantly affect the performance of face recognition systems. Although some GAN based methods are proposed to address this issue by translating meshface images to clean face images, their capacities are limited. In this paper, we introduce the introspective modules to encourage the generator to reconstruct face images and accelerate the learning process of the discriminator. Both a private meshface dataset and the public LFW dataset are used for experiments. The quantitative evaluation on both datasets proves that the introspective GAN can recover face image with better quality. Additionally, face recognition performance is also significantly improved. Wenting Chen, LinLin Shen, Zhihui Lai 0001 |
ICIP | 1 |
| 2019 | Texture Deformation Based Generative Adversarial Networks for Multi-domain Face Editing
Wenting Chen, Xinpeng Xie, Xi Jia, LinLin Shen |
PRICAI (1) | 1 |
| 2019 | Research on online consumer behavior and psychology under the background of big dataabstractSummary The emergence of emerging services and technologies such as cloud computing and social networks has driven the variety and scale of data in human society. It can increase and expand at an unprecedented rate. Data has changed from a simple object of processing to a basic resource. How can it be better? Managing and utilizing big data has become a topic of common concern. At the same time, as people's living standards improve, the scale of consumption continues to expand, and consumer demand for personalization becomes more apparent. The online consumer is a new consumer group and has distinct characteristics from the traditional market consumer groups. This article begins with the consumer's point of view. Through the analysis of psychological characteristics, behavioral characteristics, shopping needs, shopping motivation, price perception and risk perception, it will deepen consumer awareness. As a result, they can help network marketing become more economic and effective, and promote virtuous cycle of consumption and sales. Wenting Chen, Qian Zhang 0062, Maozhu Jin |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Ship classification in TerraSAR-X SAR images based on classifier combinationabstractShip classification is an important step in maritime surveillance utilizing synthetic aperture radar images. In this paper, we focus on the classifier architecture. The paper investigates three individual classifiers, i.e., the K nearest neighbor classifier, the Bayes classifier, and the back-propagation neural network classifier from the viewpoint of discrimination measurements firstly. Then, we propose a SVM combination strategy to fuse the results of individual classifiers. Extensive experiments conducted on the TerraSAR-X SAR images validate the effectiveness of the proposed method. Kefeng Ji, Xiangwei Xing, Wenting Chen, Huanxin Zou, Junli Chen |
IGARSS | 3 |
| 2013 | Ship Classification in TerraSAR-X Images With Feature Space Based Sparse RepresentationabstractShip classification is the key step in maritime surveillance using synthetic aperture radar (SAR) imagery. In this letter, we develop a new ship classification method in TerraSAR-X images based on sparse representation in feature space, in which the sparse representation classification (SRC) method is exploited. In particular, to describe the ship more accurately and to reduce the dimension of the dictionary in SRC, we propose to employ a representative feature vector to construct the dictionary instead of utilizing the image pixels directly. By testing on a ship data set collected from TerraSAR-X images, we show that the proposed method is superior to traditional methods such as the template matching (TM), K-nearest neighbor (K-NN), Bayes and Support Vector Machines (SVM). Xiangwei Xing, Kefeng Ji, Huanxin Zou, Wenting Chen, Jixiang Sun |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2005 | Integration of Transfer of Learning to the Adaptive Learning EnvironmentabstractThe instructional activity model (IAM) is a general purpose model to generate an adaptive learning course which is compatible with the SCORM standard. IAM is composed of related activity tree (AT) nodes and capability nodes. Prerequisites are capabilities supposed to possess before learning an AT while contributions are capabilities after learning an AT. IAM model supports the adaptive learning sequencing by considering the relationships between AT and capability nodes. However, the IAM model does not take the transfer of learning into consideration. In this paper, we propose the mechanism to integrate the concept of learning transfer to the IAM model. In our proposed mechanism, the relationships between capabilities are considered based on the similarity measure between capabilities. The selection process of IAM model is also modified to reflect the relationships of capabilities. Wenting Chen, Jung-Chuan Yen, Man-Kwan Shan |
ICALT | 1 |