VLDB 2026 Research / reviewers in the wild / expert
Muxi Diao
dblp:368/7775
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0000-4423-4157ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 39% Vision and language · 24% Segmentation and scene understanding · 8% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 50% Visual content generation and editing · 50% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 77% Program synthesis and code generation · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
instruction tuning |
1.5 | 2 | 2024 | How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data · EMNLP 2024 DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › AI-generated content detection
AI-generated image detection |
1.0 | 1 | 2026 | Toward Generalizable Forgery Detection and Reasoning · IEEE Trans. Image Process. 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative reinforcement learning |
1.0 | 1 | 2026 | MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement Learning · ACL (1) 2026 |
Computer vision › Image recognition and object detection › image forensics
forgery detection |
1.0 | 1 | 2026 | Toward Generalizable Forgery Detection and Reasoning · IEEE Trans. Image Process. 2026 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
1.0 | 1 | 2026 | MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision · AAAI 2026 |
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning |
1.0 | 1 | 2026 | Toward Generalizable Forgery Detection and Reasoning · IEEE Trans. Image Process. 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model safety |
0.9 | 1 | 2025 | SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025 |
Visual content generation and editing
video generation |
0.9 | 1 | 2025 | CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation · NeurIPS 2025 |
Multimedia analysis and retrieval
video understanding |
0.9 | 1 | 2025 | CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation · NeurIPS 2025 |
Security and privacy of machine learning
adversarial attack |
0.9 | 1 | 2025 | SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025 |
Security and privacy of machine learning › adversarial attack › large language model attack
adversarial prompts |
0.9 | 1 | 2025 | SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025 |
Security and privacy of machine learning
red teaming |
0.9 | 1 | 2025 | SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.8 | 1 | 2024 | How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data · EMNLP 2024 |
Compilers and program optimization
code generation |
0.8 | 1 | 2024 | DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning · ACL (1) 2024 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2025 | CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation · NeurIPS 2025 |
Program synthesis and code generation
code generation with language models |
0.2 | 1 | 2024 | How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
multimodal large language model · 2.7supervised fine-tuning · 1.0semantic similarity rewards · 1.0reinforcement learning · 1.0feature fusion · 1.0cross-attention · 1.0citation feedback · 1.0alternating co-optimization · 1.0DINO · 1.0CLIP · 1.0video generation model · 0.9self-evolving optimization · 0.9multilingual evaluation · 0.9benchmark construction · 0.9adversarial training · 0.9multi-objective instruction tuning · 0.8large language model · 0.8fine-tuning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionabstractAccurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision–language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding. Zhonghao Yan, Muxi Diao, Ruoyan Jing, Jiayuan Xu, Kaizhou Zhang, Lele Yang, Yanxi Liu 0006, Kongming Liang, Zhanyu Ma |
AAAI | 2 |
| 2026 | MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement LearningabstractLarge Language Models (LLMs) are inherently constrained by their fixed-length context windows, which limits LLMs' ability to retain and utilize information across long-term interactions.To address this limitation, recent work has proposed external memory modules for LLMs.Using memory modules typically involves two stages: evidence retrieval and memory utilization.While prior work focuses on the architecture of memory modules and the retrieval stage, the equally critical memory utilization stage remains underexplored.Building on this, we propose MemCoRL, a two-stage alternating co-optimization reinforcement learning method.Stage 1 optimizes evidence retrieval using citation feedback and semantic accuracy from utilization as rewards.Stage 2 optimizes utilization with rewards combining semantic similarity and lexical overlap.Iterative co-optimization establishes a positive feedback loop: better retrieval improves memory utilization, which in turn refines retrieval rewards.Experimental results show our approach outperforms the leading baselines on both lexical overlap and semantic similarity metrics, confirming the co-optimization in memory retrieval and memory utilization. Yuewen Liu, Muxi Diao, Anyi Zhang |
ACL (1) | 3 |
| 2026 | Toward Generalizable Forgery Detection and ReasoningabstractAccurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a generalizable forgery detection model. Moreover, since every pixel in an AI-generated image is synthesized, traditional saliency-based forgery explanation methods are not well suited for this task. To address these challenges, we formulate detection and explanation as a unified Forgery Detection and Reasoning task (FDR-Task), leveraging Multi-Modal Large Language Models (MLLMs) to provide accurate detection through reliable reasoning over forgery attributes. To facilitate this task, we introduce the Multi-Modal Forgery Reasoning dataset (MMFR-Dataset), a large-scale dataset containing 120K images across 10 generative models, with 378K reasoning annotations on forgery attributes, enabling comprehensive evaluation of the FDR-Task. Furthermore, we propose FakeReasoning, a forgery detection and reasoning framework with three key components: 1) a dual-branch visual encoder that integrates CLIP and DINO to capture both high-level semantics and low-level artifacts; 2) a Forgery-Aware Feature Fusion Module that leverages DINO's attention maps and cross-attention mechanisms to guide MLLMs toward forgery-related clues; 3) a Classification Probability Mapper that couples language modeling and forgery detection, enhancing overall performance. Experiments across multiple generative models demonstrate that FakeReasoning not only achieves robust generalization but also outperforms state-of-the-art methods on both detection and reasoning tasks. The code is available at: https://github.com/PRIS-CV/FakeReasoning. Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin, Muxi Diao, Lei Chen 0069, Kongming Liang, Zhanyu Ma |
IEEE Trans. Image Process. | 5 |
| 2025 | SEAS: Self-Evolving Adversarial Safety Optimization for Large Language ModelsabstractAs Large Language Models (LLMs) continue to advance in capability and influence, ensuring their security and preventing harmful outputs has become crucial. A promising approach to address these concerns involves training models to automatically generate adversarial prompts for red teaming. However, the evolving subtlety of vulnerabilities in LLMs challenges the effectiveness of current adversarial methods, which struggle to generate diverse, complex prompts and dynamically explore the weaknesses of these models. To tackle these challenges, we introduce the Self-Evolving Adversarial Safety (SEAS) optimization framework, which includes both a SEAS dataset and a SEAS pipeline. The SEAS dataset comprises complex adversarial prompts, while the SEAS pipeline operates through three stages: Initialization, Attack, and Adversarial Optimization. This framework generates a diverse range of adversarial prompts and dynamically explores the model's vulnerabilities to enhance its security. Our contributions include a novel adversarial framework, a comprehensive safety dataset, and empirical evidence demonstrating the effectiveness of SEAS. Muxi Diao, Shiyang Liu, Guogang Liao, Jingang Wang, Weiran Xu |
AAAI | 1 |
| 2025 | We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?abstractRunqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, Honggang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Runqi Qiao, Qiuna Tan, Guanting Dong 0001, Minhui Wu, Xiaoshuai Song, Jiapeng Wang 0005, Zhuoma Gongque, Shanglin Lei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Peiqing Yang 0003, Zhimin Bao, Muxi Diao, Chen Li 0031, Honggang Zhang 0002 |
ACL (1) | 18 |
| 2025 | CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryabstractLarge language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g. mathematics and code generation), neglecting an all-round evaluation of the computer science field. To bridge this gap, we introduce CS-Bench, the first multilingual (English, Chinese, French, German) benchmark dedicated to evaluating the performance of LLMs in computer science. CS-Bench comprises approximately 10K meticulously curated test samples, covering 26 subfields across 4 key areas of computer science, encompassing various task forms and divisions of knowledge and reasoning. Utilizing CS-Bench, we conduct a comprehensive evaluation of over 30 mainstream LLMs, revealing the relationship between CS performance and model scales. We also quantitatively analyze the reasons for failures in existing LLMs and highlight directions for improvements, including knowledge supplementation and CS-specific reasoning. Further cross-capability experiments show a high correlation between LLMs' capabilities in computer science and their abilities in mathematics and coding. Moreover, expert LLMs specialized in mathematics and coding also demonstrate strong performances in several CS subfields. Looking ahead, we envision CS-Bench serving as a cornerstone for LLM applications in the CS field and paving new avenues in assessing LLMs' diverse reasoning capabilities. Our project homepage is available at https://csbench.github.io/. Xiaoshuai Song, Muxi Diao, Guanting Dong 0001, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Weihao Zeng 0003, Yejie Wang, Zhuoma Gongque, Jianing Yu 0001, Qiuna Tan, Weiran Xu |
ICLR | 2 |
| 2025 | CineTechBench: A Benchmark for Cinematographic Technique Understanding and GenerationabstractCinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects—shot scale, shot angle, composition, camera movement, lighting, color, and focal length—and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question–answer pairs and annotated descriptions to assess MLLMs’ ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatical film production and appreciation. The code and benchmark can be accessed at \url{https://github.com/PRIS-CV/CineTechBench}. Songyu Xu, Xiangxuan Shan, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma |
NeurIPS | 5 |
| 2024 | DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction TuningabstractYejie Wang, Keqing He, Guanting Dong, Pei Wang, Weihao Zeng, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang, Xunliang Cai. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yejie Wang, Keqing He 0001, Guanting Dong 0001, Weihao Zeng 0003, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang 0002 |
ACL (1) | 6 |
| 2024 | How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good DataabstractYejie Wang, Keqing He, Dayuan Fu, Zhuoma GongQue, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, Weiran Xu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yejie Wang, Keqing He 0001, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong 0001, Muxi Diao, Jingang Wang, Mengdi Zhang 0002, Weiran Xu |
EMNLP | 10 |