EDBT 2026 Demo / reviewers in the wild / expert
Peng Xia 0005
dblp:44/6914-5
· DBLP profile ↗
16ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0003-2676-9128ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedVerse: Efficient and Reliable Medical Reasoning via DAG-Structured Parallel ExecutionabstractJianwen Chen, Xinyu Yang, Peng Xia, Arian Azarang, Yueh Z Lee, Gang Li, Hongtu Zhu, Yun Li, Beidi Chen, Huaxiu Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xinyu Yang 0002, Peng Xia 0005, Arian Azarang, Yueh Z. Lee, Gang Li 0001, Hongtu Zhu, Yun Li 0010, Beidi Chen, Huaxiu Yao |
ACL (1) | 3 |
| 2025 | Towards Realistic Semi-supervised Medical Image ClassificationabstractExisting semi-supervised learning (SSL) approaches follow the idealized closed-world assumption, neglecting the challenges present in realistic medical scenarios, such as open-set distribution and imbalanced class distribution. Although some methods in natural domains attempt to address the open-set problem, they are insufficient for medical domains, where intertwined challenges like class imbalance and small inter-class lesion discrepancies persist. Thus, this paper presents a novel self-recalibrated semantic training framework, which is tailored for SSL in medical imaging by ingeniously harvesting realistic unlabeled samples. Inspired by the observation that certain open-set samples share some similar disease-related representations with in-distribution samples, we first propose an informative sample selection strategy that identifies high-value samples to serve as augmentations, thereby effectively enriching the semantics of known categories. Furthermore, we adopt a compact semantic clustering strategy to address the semantic confusion raised by the above newly introduced open-set semantics. Moreover, to mitigate the interference of class imbalance in open-set SSL, we introduce a less biased dual-balanced classifier with similarity pseudo-label regularization and category-customized regularization. Extensive experiments on a variety of medical image datasets demonstrate the superior performance of our proposed method over state-of-the-art Closed-set and Open-set SSL methods. Wenxue Li 0003, Lie Ju, Peng Xia 0005, Xinyu Xiong, Lei Zhu 0002, ZongYuan Ge |
AAAI | 4 |
| 2025 | Neighbor Does Matter: Density-Aware Contrastive Learning for Medical Semi-supervised SegmentationabstractIn medical image analysis, multi-organ semi-supervised segmentation faces challenges such as insufficient labels and low contrast in soft tissues. To address these issues, existing studies typically employ semi-supervised segmentation techniques using pseudo-labeling and consistency regularization. However, these methods mainly rely on individual data samples for training, ignoring the rich neighborhood information present in the feature space. In this work, we argue that supervisory information can be directly extracted from the geometry of the feature space. Inspired by the density-based clustering hypothesis, we propose using feature density to locate sparse regions within feature clusters. Our goal is to increase intra-class compactness by addressing sparsity issues. To achieve this, we propose a Density-Aware Contrastive Learning (DACL) strategy, pushing anchored features in sparse regions towards cluster centers approximated by high-density positive samples, resulting in more compact clusters. Specifically, our method constructs density-aware neighbor graphs using labeled and unlabeled data samples to estimate feature density and locate sparse regions. We also combine label-guided co-training with density-guided geometric regularization to form complementary supervision for unlabeled data. Experiments on the Multi-Organ Segmentation Challenge dataset demonstrate that our proposed method outperforms state-of-the-art methods, highlighting its efficacy in medical image segmentation tasks. Zhongxing Xu, Wenxue Li 0003, Peng Xia 0005, Yiheng Zhong, Hanjun Wu, Jionglong Su, ZongYuan Ge |
AAAI | 5 |
| 2025 | HGCLIP: Exploring Vision-Language Models with Graph Representations for Hierarchical UnderstandingabstractObject categories are typically organized into a multi-granularity taxonomic hierarchy. When classifying categories at different hierarchy levels, traditional uni-modal approaches focus primarily on image features, revealing limitations in complex scenarios. Recent studies integrating Vision-Language Models (VLMs) with class hierarchies have shown promise, yet they fall short of fully exploiting the hierarchical relationships. These efforts are constrained by their inability to perform effectively across varied granularity of categories. To tackle this issue, we propose a novel framework (HGCLIP) that effectively combines CLIP with a deeper exploitation of the Hierarchical class structure via Graph representation learning. We explore constructing the class hierarchy into a graph, with its nodes representing the textual or image features of each category. After passing through a graph encoder, the textual features incorporate hierarchical structure information, while the image features emphasize class-aware features derived from prototypes through the attention mechanism. Our approach demonstrates significant improvements on 11 diverse visual recognition benchmarks. Our codes are fully available at https://github.com/richard-peng-xia/HGCLIP. Peng Xia 0005, Xingtong Yu, Lie Ju, Zhiyong Wang 0001, Peibo Duan, ZongYuan Ge |
COLING | 1 |
| 2025 | MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language ModelsabstractArtificial Intelligence (AI) has demonstrated significant potential in healthcare, particularly in disease diagnosis and treatment planning. Recent progress in Medical Large Vision-Language Models (Med-LVLMs) has opened up new possibilities for interactive diagnostic tools. However, these models often suffer from factual hallucination, which can lead to incorrect diagnoses. Fine-tuning and retrieval-augmented generation (RAG) have emerged as methods to address these issues. However, the amount of high-quality data and distribution shifts between training data and deployment data limit the application of fine-tuning methods. Although RAG is lightweight and effective, existing RAG-based approaches are not sufficiently general to different medical domains and can potentially cause misalignment issues, both between modalities and between the model and the ground truth. In this paper, we propose a versatile multimodal RAG system, MMed-RAG, designed to enhance the factuality of Med-LVLMs. Our approach introduces a domain-aware retrieval mechanism, an adaptive retrieved contexts selection, and a provable RAG-based preference fine-tuning strategy. These innovations make the RAG process sufficiently general and reliable, significantly improving alignment when introducing retrieved contexts. Experimental results across five medical datasets (involving radiology, ophthalmology, pathology) on medical VQA and report generation demonstrate that MMed-RAG can achieve an average improvement of 43.8% in factual accuracy in the factual accuracy of Med-LVLMs. Peng Xia 0005, Kangyu Zhu, Haoran Li 0011, Tianze Wang, Sheng Wang 0014, Linjun Zhang, James Zou 0001, Huaxiu Yao |
ICLR | 1 |
| 2025 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsabstractInterleaved multimodal comprehension and generation, enabling models to produce and interpret both images and text in arbitrary sequences, have become a pivotal area in multimodal learning. Despite significant advancements, the evaluation of this capability remains insufficient. Existing benchmarks suffer from limitations in data scale, scope, and evaluation depth, while current evaluation metrics are often costly or biased, lacking in reliability for practical applications. To address these challenges, we introduce MMIE, a large-scale knowledge-intensive benchmark for evaluating interleaved multimodal comprehension and generation in Large Vision-Language Models (LVLMs). MMIE comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. It supports both interleaved inputs and outputs, offering a mix of multiple-choice and open-ended question formats to evaluate diverse competencies. Moreover, we propose a reliable automated evaluation metric, leveraging a scoring model fine-tuned with human-annotated data and systematic evaluation criteria, aimed at reducing bias and improving evaluation accuracy. Extensive experiments demonstrate the effectiveness of our benchmark and metrics in providing a comprehensive evaluation of interleaved LVLMs. Specifically, we evaluate eight LVLMs, revealing that even the best models show significant room for improvement, with most achieving only moderate results. We believe MMIE will drive further advancements in the development of interleaved LVLMs. Peng Xia 0005, Siwei Han, Shi Qiu 0016, Yiyang Zhou, Zhaoyang Wang 0004, Zhaorun Chen, Chenhang Cui, Mingyu Ding, Huaxiu Yao |
ICLR | 1 |
| 2025 | Anyprefer: An Agentic Framework for Preference Data SynthesisabstractHigh-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding approach, where the target model generates and annotates its own preference data, but this can lead to inaccuracies since the reward model shares weights with the target model, thereby amplifying inherent biases. To address these issues, we propose Anyprefer, a framework designed to synthesize high-quality preference data for aligning the target model. Anyprefer frames the data synthesis process as a cooperative two-player Markov Game, where the target model and the judge model collaborate together. Here, a series of external tools are introduced to assist the judge model in accurately rewarding the target model’s responses, mitigating biases in the rewarding process. In addition, a feedback mechanism is introduced to optimize prompts for both models, enhancing collaboration and improving data quality.
The synthesized data is compiled into a new preference dataset, Anyprefer-V1, consisting of 58K high-quality preference pairs.
Extensive experiments show that Anyprefer significantly improves model alignment performance across four main applications, covering 21 datasets, achieving average improvements of 18.55% in five natural language generation datasets, 3.66% in nine vision-language understanding datasets, 30.05% in three medical image analysis datasets, and 16.00% in four visuo-motor control tasks. Yiyang Zhou, Zhaoyang Wang 0004, Tianle Wang 0009, Shangyu Xing, Peng Xia 0005, Bo Li 0026, Zijian Zhang 0010, Zhaorun Chen, Xuchao Zhang, Chetan Bansal, Mohit Bansal, Huaxiu Yao |
ICLR | 5 |
| 2025 | MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference OptimizationabstractThe advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due to modality misalignment, where the models prioritize textual knowledge over visual input, leading to hallucinations that contradict information in medical images. Previous attempts to enhance modality alignment in Med-LVLMs through preference optimization have inadequately addressed clinical relevance in preference data, making these samples easily distinguishable and reducing alignment effectiveness. In response, we propose MMedPO, a novel multimodal medical preference optimization approach that considers the clinical relevance of preference samples to enhance Med-LVLM alignment. MMedPO curates multimodal preference data by introducing two types of dispreference: (1) plausible hallucinations injected through target Med-LVLMs or GPT-4o to produce medically inaccurate responses, and (2) lesion region neglect achieved through local lesion-noising, disrupting visual understanding of critical areas. We then calculate clinical relevance for each sample based on scores from multiple Med-LLMs and visual tools, enabling effective alignment. Our experiments demonstrate that MMedPO significantly enhances factual accuracy in Med-LVLMs, achieving substantial improvements over existing preference optimization methods by 14.2% and 51.7% on the Med-VQA and report generation tasks, respectively. Our code are available in https://github.com/aiming-lab/MMedPO}{https://github.com/aiming-lab/MMedPO. Kangyu Zhu, Peng Xia 0005, Yun Li 0010, Hongtu Zhu, Sheng Wang 0014, Huaxiu Yao |
ICML | 2 |
| 2025 | MJ-Video: Benchmarking and Rewarding Video Generation with Fine-Grained Video PreferenceabstractRecent advancements in video generation have significantly improved the ability to synthesize videos from text instructions. However, existing models still struggle with key challenges such as instruction misalignment, content hallucination, safety concerns, and generation bias. To address these limitations, we introduce MJ-BENCH-VIDEO, a large-scale video preference benchmark designed to evaluate video generation across five critical aspects: Alignment, Safety, Fineness, Coherence & Consistency, and Bias & Fairness. This benchmark further incorporates 28 fine-grained criteria to provide a comprehensive evaluation of video preference. Building upon this dataset, we propose MJ-VIDEO, a Mixture-of-Experts (MoE)-based video reward model designed to deliver fine-grained reward. MJ-VIDEO can dynamically select relevant experts to accurately judge the preference based on the input text-video pair. This architecture enables more precise and adaptable preference judgments. Through extensive benchmarking on MJ-BENCH-VIDEO, we analyze the limitations of existing video reward models and demonstrate the superior performance of MJ-VIDEO in video preference assessment, achieving 17.58% and 15.87% improvements in overall and fine-grained preference judgments, respectively. Additionally, MJ-VIDEO is able to improve the alignment performance in video generation via preference fine-tuning. Haibo Tong, Zhaoyang Wang 0004, Zhaorun Chen, Haonian Ji, Shi Qiu 0016, Siwei Han, Kexin Geng, Zhongkai Xue, Yiyang Zhou, Peng Xia 0005, Mingyu Ding, Rafael Rafailov, Chelsea Finn, Huaxiu Yao |
NeurIPS | 10 |
| 2024 | OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
Peng Xia 0005, Lin Wang 0027, Siyuan Yan, Zhongxing Xu, Yimin Luo, Kaimin Song, Jürgen Leitner, Xuelian Cheng, Chi Liu 0002, Kaijing Zhou, ZongYuan Ge |
ECCV (4) | 2 |
| 2024 | RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language ModelsabstractThe recent emergence of Medical Large Vision Language Models (Med-LVLMs) has enhanced medical diagnosis.However, current Med-LVLMs frequently encounter factual issues, often generating responses that do not align with established medical facts.Retrieval-Augmented Generation (RAG), which utilizes external knowledge, can improve the factual accuracy of these models but introduces two major challenges.First, limited retrieved contexts might not cover all necessary information, while excessive retrieval can introduce irrelevant and inaccurate references, interfering with the model's generation.Second, in cases where the model originally responds correctly, applying RAG can lead to an over-reliance on retrieved contexts, resulting in incorrect answers.To address these issues, we propose RULE, which consists of two components.First, we introduce a provably effective strategy for controlling factuality risk through the calibrated selection of the number of retrieved contexts.Second, based on samples where over-reliance on retrieved contexts led to errors, we curate a preference dataset to fine-tune the model, balancing its dependence on inherent knowledge and retrieved contexts for generation.We demonstrate the effectiveness of RULE on medical VQA and report generation tasks across three datasets, achieving an average improvement of 47.4% in factual accuracy.We publicly release our benchmark and code in https: //github.com/richard-peng-xia/RULE. Peng Xia 0005, Kangyu Zhu, Haoran Li 0011, Hongtu Zhu, Yun Li 0010, Gang Li 0001, Linjun Zhang, Huaxiu Yao |
EMNLP | 1 |
| 2024 | TP-DRSeg: Improving Diabetic Retinopathy Lesion Segmentation with Explicit Text-Prompts Assisted SAM
Wenxue Li 0003, Xinyu Xiong, Peng Xia 0005, Lie Ju, ZongYuan Ge |
MICCAI (8) | 3 |
| 2024 | Generalizing to Unseen Domains in Diabetic Retinopathy with Disentangled Representations
Peng Xia 0005, Wenxue Li 0003, Lie Ju, Peibo Duan, Huaxiu Yao, ZongYuan Ge |
MICCAI (10) | 1 |
| 2024 | CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsabstractArtificial intelligence has significantly impacted medical applications, particularly with the advent of Medical Large Vision Language Models (Med-LVLMs), sparking optimism for the future of automated and personalized healthcare. However, the trustworthiness of Med-LVLMs remains unverified, posing significant risks for future model deployment. In this paper, we introduce CARES and aim to comprehensively evaluate the Trustworthiness of Med-LVLMs across the medical domain. We assess the trustworthiness of Med-LVLMs across five dimensions, including trustfulness, fairness, safety, privacy, and robustness. CARES comprises about 41K question-answer pairs in both closed and open-ended formats, covering 16 medical image modalities and 27 anatomical regions. Our analysis reveals that the models consistently exhibit concerns regarding trustworthiness, often displaying factual inaccuracies and failing to maintain fairness across different demographic groups. Furthermore, they are vulnerable to attacks and demonstrate a lack of privacy awareness. We publicly release our benchmark and code in https://github.com/richard-peng-xia/CARES. Peng Xia 0005, Juanxi Tian, Yangrui Gong, Ruibo Hou, Zhenbang Wu, Zhiyuan Fan, Yiyang Zhou, Kangyu Zhu, Zhaoyang Wang 0004, Xiao Wang 0044, Xuchao Zhang, Chetan Bansal, Marc Niethammer, Junzhou Huang, Hongtu Zhu, Yun Li 0010, Jimeng Sun 0001, ZongYuan Ge, Gang Li 0001, James Zou 0001, Huaxiu Yao |
NeurIPS | 1 |
| 2023 | NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity UnderstandingabstractThe application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education, improve quality control, and enable operational compliance monitoring. However, the development of automatic recognition systems in this field is currently hindered by the scarcity of appropriately labeled datasets. The existing video datasets pose several limitations: 1) these datasets are small-scale in size to support comprehensive investigations of nursing activity; 2) they primarily focus on single procedures, lacking expert-level annotations for various nursing procedures and action steps; and 3) they lack temporally localized annotations, which prevents the effective localization of targeted actions within longer video sequences. To mitigate these limitations, we propose NurViD, a large video dataset with expert-level annotation for nursing procedure activity understanding. NurViD consists of over 1.5k videos totaling 144 hours, making it approximately four times longer than the existing largest nursing activity datasets. Notably, it encompasses 51 distinct nursing procedures and 177 action steps, providing a much more comprehensive coverage compared to existing datasets that primarily focus on limited procedures. To evaluate the efficacy of current deep learning methods on nursing activity understanding, we establish three benchmarks on NurViD: procedure recognition on untrimmed videos, procedure and action recognition on trimmed videos, and action detection. Our benchmark and code will be available at https://github.com/minghu0830/NurViD-benchmark. Lin Wang 0027, Siyuan Yan, Don Ma, Qingli Ren, Peng Xia 0005, Wei Feng 0015, Peibo Duan, Lie Ju, ZongYuan Ge |
NeurIPS | 6 |
| 2019 | Topic model with incremental vocabulary based on Belief Propagation
Jianfeng Yan, Peng Xia 0005 |
Knowl. Based Syst. | 6 |