Muxi Diao

dblp:368/7775 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0000-4423-4157ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 39% Vision and language · 24% Segmentation and scene understanding · 8%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 50% Visual content generation and editing · 50%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 77% Program synthesis and code generation · 23%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction tuning
1.522024
How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data · EMNLP 2024
DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning · ACL (1) 2024
Machine learning › Trustworthy machine learning › AI-generated content detection
AI-generated image detection
1.012026
Toward Generalizable Forgery Detection and Reasoning · IEEE Trans. Image Process. 2026
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative reinforcement learning
1.012026
MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement Learning · ACL (1) 2026
Computer vision › Image recognition and object detection › image forensics
forgery detection
1.012026
Toward Generalizable Forgery Detection and Reasoning · IEEE Trans. Image Process. 2026
Computer vision › Segmentation and scene understanding
medical image segmentation
1.012026
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision · AAAI 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning
1.012026
Toward Generalizable Forgery Detection and Reasoning · IEEE Trans. Image Process. 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery · ICLR 2025
Natural language and speech › Language models and text generation
large language model safety
0.912025
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025
Computer vision › Vision and language
visual reasoning
0.912025
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? · ACL (1) 2025
Visual content generation and editing
video generation
0.912025
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation · NeurIPS 2025
Multimedia analysis and retrieval
video understanding
0.912025
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation · NeurIPS 2025
Security and privacy of machine learning
adversarial attack
0.912025
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025
Security and privacy of machine learning › adversarial attack › large language model attack
adversarial prompts
0.912025
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025
Security and privacy of machine learning
red teaming
0.912025
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models · AAAI 2025
Machine learning › Efficient and distributed learning
data selection
0.812024
How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data · EMNLP 2024
Compilers and program optimization
code generation
0.812024
DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning · ACL (1) 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation · NeurIPS 2025
Program synthesis and code generation
code generation with language models
0.212024
How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 2.7supervised fine-tuning · 1.0semantic similarity rewards · 1.0reinforcement learning · 1.0feature fusion · 1.0cross-attention · 1.0citation feedback · 1.0alternating co-optimization · 1.0DINO · 1.0CLIP · 1.0video generation model · 0.9self-evolving optimization · 0.9multilingual evaluation · 0.9benchmark construction · 0.9adversarial training · 0.9multi-objective instruction tuning · 0.8large language model · 0.8fine-tuning · 0.8
YearPublicationVenuePosition
2026 MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
abstract
Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision–language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding.
Zhonghao Yan, Muxi Diao, Ruoyan Jing, Jiayuan Xu, Kaizhou Zhang, Lele Yang, Yanxi Liu 0006, Kongming Liang, Zhanyu Ma
AAAI2
2026 MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement Learning
abstract
Large Language Models (LLMs) are inherently constrained by their fixed-length context windows, which limits LLMs' ability to retain and utilize information across long-term interactions.To address this limitation, recent work has proposed external memory modules for LLMs.Using memory modules typically involves two stages: evidence retrieval and memory utilization.While prior work focuses on the architecture of memory modules and the retrieval stage, the equally critical memory utilization stage remains underexplored.Building on this, we propose MemCoRL, a two-stage alternating co-optimization reinforcement learning method.Stage 1 optimizes evidence retrieval using citation feedback and semantic accuracy from utilization as rewards.Stage 2 optimizes utilization with rewards combining semantic similarity and lexical overlap.Iterative co-optimization establishes a positive feedback loop: better retrieval improves memory utilization, which in turn refines retrieval rewards.Experimental results show our approach outperforms the leading baselines on both lexical overlap and semantic similarity metrics, confirming the co-optimization in memory retrieval and memory utilization.
Yuewen Liu, Muxi Diao, Anyi Zhang
ACL (1)3
2026 Toward Generalizable Forgery Detection and Reasoning
abstract
Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a generalizable forgery detection model. Moreover, since every pixel in an AI-generated image is synthesized, traditional saliency-based forgery explanation methods are not well suited for this task. To address these challenges, we formulate detection and explanation as a unified Forgery Detection and Reasoning task (FDR-Task), leveraging Multi-Modal Large Language Models (MLLMs) to provide accurate detection through reliable reasoning over forgery attributes. To facilitate this task, we introduce the Multi-Modal Forgery Reasoning dataset (MMFR-Dataset), a large-scale dataset containing 120K images across 10 generative models, with 378K reasoning annotations on forgery attributes, enabling comprehensive evaluation of the FDR-Task. Furthermore, we propose FakeReasoning, a forgery detection and reasoning framework with three key components: 1) a dual-branch visual encoder that integrates CLIP and DINO to capture both high-level semantics and low-level artifacts; 2) a Forgery-Aware Feature Fusion Module that leverages DINO's attention maps and cross-attention mechanisms to guide MLLMs toward forgery-related clues; 3) a Classification Probability Mapper that couples language modeling and forgery detection, enhancing overall performance. Experiments across multiple generative models demonstrate that FakeReasoning not only achieves robust generalization but also outperforms state-of-the-art methods on both detection and reasoning tasks. The code is available at: https://github.com/PRIS-CV/FakeReasoning.
Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin, Muxi Diao, Lei Chen 0069, Kongming Liang, Zhanyu Ma
IEEE Trans. Image Process.5
2025 SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
abstract
As Large Language Models (LLMs) continue to advance in capability and influence, ensuring their security and preventing harmful outputs has become crucial. A promising approach to address these concerns involves training models to automatically generate adversarial prompts for red teaming. However, the evolving subtlety of vulnerabilities in LLMs challenges the effectiveness of current adversarial methods, which struggle to generate diverse, complex prompts and dynamically explore the weaknesses of these models. To tackle these challenges, we introduce the Self-Evolving Adversarial Safety (SEAS) optimization framework, which includes both a SEAS dataset and a SEAS pipeline. The SEAS dataset comprises complex adversarial prompts, while the SEAS pipeline operates through three stages: Initialization, Attack, and Adversarial Optimization. This framework generates a diverse range of adversarial prompts and dynamically explores the model's vulnerabilities to enhance its security. Our contributions include a novel adversarial framework, a comprehensive safety dataset, and empirical evidence demonstrating the effectiveness of SEAS.
Muxi Diao, Shiyang Liu, Guogang Liao, Jingang Wang, Weiran Xu
AAAI1
2025 We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
abstract
Runqi Qiao, Qiuna Tan, Guanting Dong, MinhuiWu MinhuiWu, Chong Sun, Xiaoshuai Song, Jiapeng Wang, Zhuoma GongQue, Shanglin Lei, YiFan Zhang, Zhe Wei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Yida Xu, Peiqing Yang, Zhimin Bao, Muxi Diao, Chen Li, Honggang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Runqi Qiao, Qiuna Tan, Guanting Dong 0001, Minhui Wu, Xiaoshuai Song, Jiapeng Wang 0005, Zhuoma Gongque, Shanglin Lei, Miaoxuan Zhang, Runfeng Qiao, Xiao Zong, Peiqing Yang 0003, Zhimin Bao, Muxi Diao, Chen Li 0031, Honggang Zhang 0002
ACL (1)18
2025 CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
abstract
Large language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g. mathematics and code generation), neglecting an all-round evaluation of the computer science field. To bridge this gap, we introduce CS-Bench, the first multilingual (English, Chinese, French, German) benchmark dedicated to evaluating the performance of LLMs in computer science. CS-Bench comprises approximately 10K meticulously curated test samples, covering 26 subfields across 4 key areas of computer science, encompassing various task forms and divisions of knowledge and reasoning. Utilizing CS-Bench, we conduct a comprehensive evaluation of over 30 mainstream LLMs, revealing the relationship between CS performance and model scales. We also quantitatively analyze the reasons for failures in existing LLMs and highlight directions for improvements, including knowledge supplementation and CS-specific reasoning. Further cross-capability experiments show a high correlation between LLMs' capabilities in computer science and their abilities in mathematics and coding. Moreover, expert LLMs specialized in mathematics and coding also demonstrate strong performances in several CS subfields. Looking ahead, we envision CS-Bench serving as a cornerstone for LLM applications in the CS field and paving new avenues in assessing LLMs' diverse reasoning capabilities. Our project homepage is available at https://csbench.github.io/.
Xiaoshuai Song, Muxi Diao, Guanting Dong 0001, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Weihao Zeng 0003, Yejie Wang, Zhuoma Gongque, Jianing Yu 0001, Qiuna Tan, Weiran Xu
ICLR2
2025 CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
abstract
Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects—shot scale, shot angle, composition, camera movement, lighting, color, and focal length—and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question–answer pairs and annotated descriptions to assess MLLMs’ ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatical film production and appreciation. The code and benchmark can be accessed at \url{https://github.com/PRIS-CV/CineTechBench}.
Songyu Xu, Xiangxuan Shan, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma
NeurIPS5
2024 DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning
abstract
Yejie Wang, Keqing He, Guanting Dong, Pei Wang, Weihao Zeng, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang, Xunliang Cai. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yejie Wang, Keqing He 0001, Guanting Dong 0001, Weihao Zeng 0003, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang 0002
ACL (1)6
2024 How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data
abstract
Yejie Wang, Keqing He, Dayuan Fu, Zhuoma GongQue, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, Weiran Xu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yejie Wang, Keqing He 0001, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong 0001, Muxi Diao, Jingang Wang, Mengdi Zhang 0002, Weiran Xu
EMNLP10