Dexuan Xu

dblp:349/1378 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-3100-5886ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models
abstract
Recent advances in multimodal large language models (MLLMs) have significantly improved medical AI, enabling it to unify the understanding of visual and textual information. However, as medical knowledge continues to evolve, it is critical to allow these models to efficiently update outdated or incorrect information without retraining from scratch. Although textual knowledge editing has been widely studied, there is still a lack of systematic benchmarks for multimodal medical knowledge editing involving image and text modalities. To fill this gap, we present MedMKEB, the first comprehensive benchmark designed to evaluate the reliability, generality, locality, portability, and robustness of knowledge editing in medical multimodal large language models. MedMKEB is built on a high-quality medical visual question-answering dataset and enriched with carefully constructed editing tasks, including counterfactual correction, semantic generalization, knowledge transfer, and adversarial robustness. We incorporate human expert validation to ensure the accuracy and reliability of the benchmark. Extensive experiments on state-of-the-art general and medical MLLMs demonstrate the limitations of existing knowledge editing methods in the medical domain, highlighting the need to develop specialized editing strategies.
Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao, Hanpin Wang, Huamin Zhang, Yu Huang 0004
AAAI1
2026 Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image Understanding
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in interpreting single medical images. However, real-world clinical diagnosis is intrinsically a multi-view process, requiring the synthesis of information across volumetric slices, temporal sequences, and comparative modalities. Existing benchmarks fail to capture this complexity, limiting the assessment of models in realistic clinical workflows. To bridge this gap, we introduce MedMultiBench, the first large-scale benchmark specifically designed for medical multi-image understanding. Comprising 11,392 expert-curated samples, MedMultiBench evaluates MLLMs across four distinct dimensions: Joint Reasoning, Comparative Analysis, Comprehensive Perception, and In-Context Learning. We benchmark 13 state-of-the-art MLLMs, revealing that while current models excel in single-view tasks, they struggle significantly with multi-image contexts. Our experiments identify a performance degradation in open-source models when processing increased visual loads, whereas closed-source models demonstrate better scalability. MedMultiBench provides a robust framework to facilitate the development of MLLMs capable of holistic clinical reasoning.
Dexuan Xu, Jiayin Yuan, Yanyuan Chen, Hanpin Wang, Yu Huang 0004
ACL (1)1
2025 STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling
abstract
Online psychological counseling dialogue systems are trending, offering a convenient and accessible alternative to traditional in-person therapy. However, existing psychological counseling dialogue systems mainly focus on basic empathetic dialogue or QA with minimal professional knowledge and without goal guidance. In many real-world counseling scenarios, clients often seek multi-type help, such as diagnosis, consultation, therapy, console, and common questions, but existing dialogue systems struggle to combine different dialogue types naturally. In this paper, we identify this challenge as how to construct mixed-type dialogue systems for psychological counseling that enable clients to clarify their goals before proceeding with counseling. To mitigate the challenge, we collect a mixed-type counseling dialogues corpus termed STAMPsy, covering five dialogue types, task-oriented dialogue for diagnosis, knowledge-grounded dialogue, conversational recommendation, empathetic dialogue, and question answering, over 5,000 conversations. Moreover, spatiotemporal-aware knowledge enables systems to have world awareness and has been proven to affect one's mental health. Therefore, we link dialogues in STAMPsy to spatiotemporal state and propose a spatiotemporal-aware mixed-type psychological counseling dataset. Additionally, we build baselines on STAMPsy and develop an iterative self-feedback psychological dialogue generation framework, named Self-STAMPsy. Results indicate that clarifying dialogue goals in advance and utilizing spatiotemporal states are effective.
Jieyi Wang, Zeming Liu, Dexuan Xu, Chuan Wang 0002, Ruiyuan Guan, Weihua Yue, Yu Huang 0004
AAAI4
2025 MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Output
abstract
Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lacks direct understanding of visual clues in the image; for output, the model only gives text answers and lacks connection with key areas in the image. To address these issues, we propose a unified medical vision language model MIMO, with visual referring Multimodal Input and pixel grounding Multimodal Output. MIMO can not only combine visual clues and textual instructions to understand complex medical images and semantics, but can also ground medical terminologies in textual output within the image. To overcome the scarcity of relevant data in the medical field, we propose MIMOSeg, a comprehensive medical multimodal dataset including 895K samples. MIMOSeg is constructed from four different perspectives, covering basic instruction following and complex question answering with multimodal input and multimodal output. We conduct experiments on several downstream medical multimodal tasks. Extensive experimental results verify that MIMO can uniquely combine visual referring and pixel grounding capabilities, which are not available in previous models. Our project can be found in https://github.com/pkusixspace/MIMO.
Yanyuan Chen, Dexuan Xu, Yu Huang 0004, Songkun Zhan, Hanpin Wang, Dongxue Chen, Meikang Qiu, Hang Li 0001
CVPR2
2025 DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automation
abstract
Existing large language model (LLM) agents for automating data science show promise, but they remain constrained by narrow task scopes, limited generalization across tasks and models, and over-reliance on state-of-the-art (SOTA) LLMs.We introduce DatawiseAgent 1 , a notebook-centric LLM agent framework for adaptive and robust data science automation.Inspired by how human data scientists work in computational notebooks, DatawiseAgent introduces a unified interaction representation and a multi-stage architecture based on finitestate transducers (FSTs).This design enables flexible long-horizon planning, progressive solution development, and robust recovery from execution failures.Extensive experiments across diverse data science scenarios and models show that DatawiseAgent consistently achieves SOTA performance by surpassing strong baselines such as AutoGen and TaskWeaver, demonstrating superior effectiveness and adaptability.Further evaluations reveal graceful performance degradation under weaker or smaller models, underscoring the robustness and scalability.
Ziming You, Yumiao Zhang, Dexuan Xu, Yiwei Lou, Yandong Yan, Huamin Zhang, Yu Huang 0004
EMNLP3
2025 Uncertainty-aware Masked Modeling in Medical Imaging
abstract
Denoising the corrupted images using an autoencoder is a promising approach to obtain a good performing encoder. This sub-field, known as MIM (Masked Image Modeling), has gained renewed attention with the advent of Masked Autoencoders (MAE). Although MAE and its subsequent work have revealed the significant impact of masking strategies on downstream task performance, existing masking strategy designs remain relatively basic. In this paper, we propose UAM3, a novel uncertainty-aware masked modeling framework for universal encoder pre-training on medical CT images. Our framework effectively leverages the model’s perceptual capacity during training by encouraging uncertainty prediction for different patches under controlled perturbations. Specifically, the framework adopts the mean teacher paradigm, consisting of a teacher model and a student model with the same architecture. During training, the teacher model generates challenging masks, while the student model completes the preset reconstruction task and updates the teacher model using exponential moving average (EMA). The core is an uncertainty-aware scheme that allows the student model to use predictive uncertainty information to progressively focus on high information entropy contents. Comprehensive experiments demonstrate that UAM3outperforms state-of-the-art MIM methods, highlighting the potential of our framework for addressing challenging medical downstream problems.
Dexuan Xu
ICASSP3
2025 Medical Vision-Language Pre-training with Multimodal Variational Masked Autoencoder for Robust Medical VQA
abstract
Medical Visual Question Answering (Medical VQA) plays an important role in medical informatics. However, the robustness of existing medical VQA models is severely challenged by adversarial attacks. Current methods (e.g. adversarial training and noise-based reasoning) heavily rely on additional data or complex procedures and often ignore model-level robustness. To address these issues, we propose Multimodal Variational Masked Autoencoder (MVMAE), a novel pre-training framework designed to enhance the robustness of the medical VQA task. MVMAE leverages masked modeling and variational inference to extract robust multimodal features. The framework introduces a low-cost multimodal bottleneck fusion module and employs reparameterization to sample robust latent representations, ensuring effective feature fusion and reconstruction. Extensive experiments on public medical VQA datasets demonstrate that MVMAE significantly improves resistance to various adversarial attacks and outperforms other medical multimodal pre-training methods.
Dexuan Xu, Yanyuan Chen, Yu Huang 0004, Shihao E, Yiwei Lou, Yongzhi Cao, Hanpin Wang, Meikang Qiu
ACM Multimedia1
2024 A Learnable Discrete-Prior Fusion Autoencoder with Contrastive Learning for Tabular Data Synthesis
abstract
The actual collection of tabular data for sharing involves confidentiality and privacy constraints, leaving the potential risks of machine learning for interventional data analysis unsafely averted. Synthetic data has emerged recently as a privacy-protecting solution to address this challenge. However, existing approaches regard discrete and continuous modal features as separate entities, thus falling short in properly capturing their inherent correlations. In this paper, we propose a novel contrastive learning guided Gaussian Transformer autoencoder, termed GTCoder, to synthesize photo-realistic multimodal tabular data for scientific research. Our approach introduces a transformer-based fusion module that seamlessly integrates multimodal features, permitting for mining more informative latent representations. The attention within the fusion module directs the integrated output features to focus on critical components that facilitate the task of generating latent embeddings. Moreover, we formulate a contrastive learning strategy to implicitly constrain the embeddings from discrete features in the latent feature space by encouraging the similar discrete feature distributions closer while pushing the dissimilar further away, in order to better enhance the representation of the latent embedding. Experimental results indicate that GTCoder is effective to generate photo-realistic synthetic data, with interactive interpretation of latent embedding, and performs favorably against some baselines on most real-world and simulated datasets.
Rongchao Zhang, Yiwei Lou, Dexuan Xu, Yongzhi Cao, Hanpin Wang, Yu Huang 0004
AAAI3
2024 MR Image Quality Assessment via Enhanced Mamba: A Hybrid Spatial-Frequency Approach
abstract
Magnetic resonance (MR) image quality assessment plays a crucial role in disease diagnosis and data analysis. Existing methods typically treat the data as images and process them with convolutional networks, thereby ignoring the sequential characteristics of MR data. In this paper, we propose a hybrid spatial-frequency network (HSFNet) for MR image quality assessment, which extracts MR image quality features in both spatial and frequency domains. Specifically, within each data domain, information from local images and global sequences is iteratively integrated by applying a Mamba-based cascading processing module for multiple times. Extensive experiments on both T1-weighting and T2-weighting MR datasets demonstrate the proposed method’s effectiveness and generalization ability in comparison with state-of-the-art MR image quality assessment methods.
Yiwei Lou, Dexuan Xu, Rongchao Zhang, Yongzhi Cao, Hanpin Wang, Yu Huang 0004
BIBM2
2024 Curriculum Learning for Self-Iterative Semi-Supervised Medical Image Segmentation
abstract
Self-training with data augmentation emerges as an efficacious strategy for harnessing unlabeled data in the realm of semi-supervised medical image segmentation. Within the synthetic domain, existing models make a deliberate trade-off, sacrificing some of its absolute performance on labeled data to bolster its generalization capabilities on the predominantly abundant unlabeled data encompassed within the entire dataset. In this study, we find out the essence of employing data augmentation techniques to create a proxy data domain that serves as a bridge between labeled and unlabeled data. To this end, we optimize the aforementioned approach by incorporating the concept of curriculum learning, which encompasses two primary components: Dynamic Copy-Paste strategies and the Self-Iterative Segmentation Model. Concerning the former, the dynamic scaling of the copy-paste box guides the model in acquiring shared semantics, progressing from easier (labeled data) to more challenging (unlabeled data). In order to facilitate this incremental learning process, we have devised models that supports progressive iterative evolution throughout the training phase. Our approach has demonstrated remarkable efficacy through a comprehensive series of benchmarks, consistently outperforming existing methods and achieving state-of-the-art performance.
Dexuan Xu, Yanyuan Chen, Yiwei Lou
BIBM2
2024 A Novel Multi-Atlas Fusion Model Based On Contrastive Learning For Functional Connectivity Graph Diagnosis
abstract
Functional connectivity (FC) graph analysis is an important method for diagnosing brain disorders using functional magnetic resonance imaging (fMRI). Existing FC graph diagnosis approaches preprocess the brain by dividing it into specific regions using atlases. However, relying on a single atlas exclusively for data preprocessing fails to fully harness the potential of the medical prior knowledge embedded within these atlases. To address this issue, this paper proposes a functional connectivity graph representation learning method that integrates multiple atlases and multiple views. This self-supervised approach to representation learning aligns features across different atlases through contrastive learning. Building upon this, we introduce an improved finite field- of-view multi-head attention mechanism for feature fusion. This mechanism is used to initialize a spectral graph convolutional network's nodes (GCN) with fusion features obtained through cross-atlas alignment. Meanwhile, edge initialization takes into account the inherent correlations among multiple independent atlas fusion features. The proposed method is validated on multi-site datasets and demonstrates superiority across various metrics. The experimental section provides visualizations of the learned feature alignment and consistency, evaluating the effectiveness of contrastive learning.
Dexuan Xu, Yiwei Lou, Yu Huang 0004
ICASSP2
2024 No-Reference MRI Quality Assessment via Contrastive Representation: Spatial and Frequency Domain Perspectives
abstract
No-reference image quality assessment for magnetic resonance images (MRI) aims to generate quality evaluations closely aligned with human perception without the reliance on high-quality reference images. This faces challenges due to the limited availability of large-scale, publicly accessible datasets and the absence of rigorous scientific evaluation. To address these challenges, we underscore the importance of six quality indicators that are closely related to spatial and frequency domain features. Based on these indicators, three experts are employed to provide comprehensive quality ratings across eight public MRI datasets. In addition, we propose a novel approach to extract and combine quality feature representations from spatial and frequency domains. Within the spatial domain, we emphasize capturing anatomical details and structural integrity by applying two sets of data augmentation transformations, which produce data variants that facilitate spatial feature extraction via contrastive learning. Meanwhile, a parallel strategy is carried out in the frequency domain, focusing on signal variations and artifacts. Experimental results demonstrate the robustness of our approach in MRI quality assessment, effectively integrating insights from both spatial and frequency domain perspectives.
Yiwei Lou, Dexuan Xu, Yongzhi Cao, Hanpin Wang, Yu Huang 0004
ICME3
2024 Decouple and Decorrelate: A Disentanglement Security Framework Combining Sample Weighting for Cross-Institution Biased Disease Diagnosis
abstract
There is an urgent need to address the effective diagnosis of multiple diseases across various medical institutions while ensuring the privacy of medical data in IoT environments. This requires the model to have the ability of zero-shot generalization, which can not be satisfied by existing models. To address this issue, we propose a two-stage model for medical image diagnosis, based on decoupling and decorrelating. An adversarial architecture is built using a gradient reversal discriminator to improve the model’s robustness. To further address the mixed correlation within domain-invariant features achieved by disentanglement, we propose to mitigate feature dependency through sample weighting. The effectiveness of the model is validated using both the diabetic retinopathy and the skin lesion datasets. For cross-dataset experiment, we select two datasets for symmetric decoupling and reserve the remaining dataset as the test set. The test dataset is analogous to real-world scenarios, where all the samples and labels are completely unknown to the model. The experiments show that the model achieves excellent performance and outperforms baselines in most metrics, which demonstrate the effectiveness of our approach to address the issue of multi-center data privacy in IoT, with a focus on enhancing diagnostic accuracy while ensuring data security.
Hang Li 0001, Dexuan Xu, Yiwei Lou, Menglong Ran, Zhi Jin 0001, Yu Huang 0004
IEEE Internet Things J.3
2023 Radiology Report Generation via Structured Knowledge-Enhanced Multi-modal Attention and Contrastive Learning
abstract
The automated generation of radiology reports has attracted significant attention in the field of bioinformatics. Currently, the main limitations of this task include insufficient utilization of prior medical knowledge, lack of efficient knowledge fusion algorithms, and less distinctiveness between different generated reports. To address these issues, we propose a novel algorithm for radiology report generation, which includes Structured Knowledge-Enhanced Multi-modal Attention (SKEMA) and Dual-Branch Contrastive Learning (DBCL) for the first time. SKEMA aims to effectively bridge the gap between visual and prior knowledge by leveraging the high-order adjacency matrix of the knowledge graph to weightedly fuse image features and knowledge features. We enhance both features through masking, and use the original features and augmented features as positive and negative samples in the dual-branch contrastive learning (DBCL). DBCL increases the differences between positive and negative samples to avoid generating templated results, and enhances the robustness of the model. Finally, we conducted experiments to demonstrate the effectiveness of our model on two public radiology datasets, IU-Xray and MIMIC-CXR. Our model outperformed previous baseline methods on both datasets and achieved excellent evaluation scores.
Dexuan Xu, Yanyuan Chen, Yiwei Lou, Hanpin Wang, Yu Huang 0004
BIBM1
2023 Refining the Unseen: Self-supervised Two-stream Feature Extraction for Image Quality Assessment
abstract
The inadequacy of labeled datasets for image quality assessment has led to the development and popularity of self-supervised approaches. However, most existing self-supervised methods primarily focus on content and fidelity features extracted with convolutional neural networks, overlooking the crucial importance of structural features in quality assessment. To address this problem, we present a novel self-supervised two-stream feature extraction and representation approach. In our approach, the first stream leverages a contrastive learning framework to extract image fidelity features, while the second stream emphasizes structural features by incorporating an attention mechanism. This innovative combination results in a comprehensive feature representation for quality assessment. Moreover, our proposed method facilitates transfer learning, allowing the pre-trained two-stream model in the source domain to be seamlessly applied to target domains for quality regression. This compatibility with transfer learning enhances the adaptability and generalization of the model. Extensive experiments are carried out on three synthetic distortion datasets to validate the effectiveness of our approach. The results demonstrate that our work not only competes with state-of-the-art self-supervised methods but also outperforms some supervised approaches.
Yiwei Lou, Yanyuan Chen, Dexuan Xu, Doudou Zhou, Yongzhi Cao, Hanpin Wang, Yu Huang 0004
ICDM3