Dongfu Jiang

dblp:336/6970 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0007-9442-6721ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents
abstract
Zhuofeng Li, Yi Lu, Dongfu Jiang, Haoxiang Zhang, Yuyang Bai, Chuan Li, Yu Wang, Shuiwang Ji, Jianwen Xie, Yu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhuofeng Li, Dongfu Jiang, Yuyang Bai, Shuiwang Ji, Jianwen Xie, Yu Zhang 0044
ACL (1)3
2025 ACECODER: Acing Coder RL via Automated Test-Case Synthesis
abstract
Most progress in recent coder models has been driven by supervised fine-tuning (SFT), while the potential of reinforcement learning (RL) remains largely unexplored, primarily due to the lack of reliable reward data/model in the code domain.In this paper, we address this challenge by leveraging automated large-scale testcase synthesis to enhance code model training.Specifically, we design a pipeline that generates extensive (question, test-cases) pairs from existing code data.Using these test cases, we construct preference pairs based on pass rates over sampled programs to train reward models with Bradley-Terry loss.It shows an average of 10-point improvement for Llama-3.1-8B-Ins and 5-point improvement for Qwen2.5-Coder-7B-Insthrough best-of-32 sampling, making the 7B model on par with 236B DeepSeek-V2.5.Furthermore, we conduct reinforcement learning with both reward models and testcase pass rewards, leading to consistent improvements across HumanEval, MBPP, Big-CodeBench, and LiveCodeBench (V4).Notably, we follow the R1-style training to start from Qwen2.5-Coder-base directly and show that our RL training can improve model on HumanEval-plus by over 25% and MBPP-plus by 6% for merely 80 optimization steps.We believe our results highlight the huge potential of reinforcement learning in coder models.
Huaye Zeng, Dongfu Jiang, Haozhe Wang 0002, Ping Nie, Wenhu Chen
ACL (1)2
2025 MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
abstract
We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation. In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MM-Bench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks. Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions.
Tianhao Liang, Sherman Siu, Zhengqing Wang, Kai Wang 0068, Yubo Wang 0019, Yuansheng Ni, Ziyan Jiang, Wang Zhu 0001, Bohan Lyu 0001, Dongfu Jiang, Hexiang Hu, Xiang Yue, Wenhu Chen
ICLR11
2025 General-Reasoner: Advancing LLM Reasoning Across All Domains
abstract
Reinforcement learning (RL) has recently demonstrated strong potential in enhancing the reasoning capabilities of large language models (LLMs). Particularly, the "Zero" reinforcement learning introduced by Deepseek-R1-Zero, enables direct RL training of base LLMs without relying on an intermediate supervised fine-tuning stage. Despite these advancements, current works for LLM reasoning mainly focus on mathematical and coding domains, largely due to data abundance and the ease of answer verification. This limits the applicability and generalization of such models to broader domains, where questions often have diverse answer representations, and data is more scarce. In this paper, we propose General-Reasoner, a novel training framework designed to enhance LLM reasoning capabilities across diverse domains. Our key contributions include: (1) constructing a large-scale, high-quality dataset of questions with verifiable answers curated by web crawling, covering a wide range of disciplines; and (2) developing a generative model-based answer verifier, which replaces traditional rule-based verification with the capability of chain-of-thought and context-awareness. We train a series of models and evaluate them on a wide range of datasets covering wide domains like physics, chemistry, finance, electronics etc. Our comprehensive evaluation across these 12 benchmarks (e.g. MMLU-Pro, GPQA, SuperGPQA, TheoremQA, BBEH and MATH AMC) demonstrates that General-Reasoner outperforms existing baseline methods, achieving robust and generalizable reasoning performance while maintaining superior effectiveness in mathematical reasoning tasks.
Xueguang Ma, Qian Liu 0033, Dongfu Jiang, Ge Zhang 0009, Zejun Ma 0001, Wenhu Chen
NeurIPS3
2024 VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
abstract
In the rapidly advancing field of conditional image generation research, challenges such as limited explainability lie in effectively evaluating the performance and capabilities of various models.This paper introduces VIESCORE, a Visual Instruction-guided Explainable metric for evaluating any conditional image generation tasks.VIESCORE leverages general knowledge from Multimodal Large Language Models (MLLMs) as the backbone and does not require training or fine-tuning.We evaluate VIESCORE on seven prominent tasks in conditional image tasks and found: (1) VIESCORE (GPT4-o) achieves a high Spearman correlation of 0.4 with human evaluations, while the human-to-human correlation is 0.45.(2) VI-ESCORE (with open-source MLLM) is significantly weaker than GPT-4o and GPT-4v in evaluating synthetic images.(3) VIESCORE achieves a correlation on par with human ratings in the generation tasks but struggles in editing tasks.With these results, we believe VIESCORE shows its great potential to replace human judges in evaluating image synthesis tasks.
Max Ku, Dongfu Jiang, Cong Wei 0001, Xiang Yue, Wenhu Chen
ACL (1)2
2024 MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
abstract
We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and text-books, covering six core disciplines: Art & Design, Busi-ness, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly het-erogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 28 open-source LMMs as well as the propri-etary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.
Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 0033, Ruoqi Liu, Ge Zhang 0009, Samuel Stevens 0001, Dongfu Jiang, Weiming Ren, Yuxuan Sun 0002, Cong Wei 0001, Botao Yu, Ruibin Yuan, Renliang Sun, Boyuan Zheng 0001, Zhenzhu Yang, Wenhao Huang 0001, Huan Sun 0001, Yu Su 0001, Wenhu Chen
CVPR8
2024 VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
abstract
Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, Wenhu Chen. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Dongfu Jiang, Ge Zhang 0009, Max Ku, Achint Soni, Sherman Siu, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang 0068, Quy Duc Do, Yuansheng Ni, Bohan Lyu 0001, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Y. Lin, Wenhu Chen
EMNLP2
2024 GenAI Arena: An Open Evaluation Platform for Generative Models
abstract
Generative AI has made remarkable strides to revolutionize fields such as image and video generation. These advancements are driven by innovative algorithms, architecture, and data. However, the rapid proliferation of generative models has highlighted a critical gap: the absence of trustworthy evaluation metrics. Current automatic assessments such as FID, CLIP, FVD, etc often fail to capture the nuanced quality and user satisfaction associated with generative outputs. This paper proposes an open platform GenAI-Arena to evaluate different image and video generative models, where users can actively participate in evaluating these models. By leveraging collective user feedback and votes, GenAI-Arena aims to provide a more democratic and accurate measure of model performance. It covers three tasks of text-to-image generation, text-to-video generation, and image editing respectively. Currently, we cover a total of 35 open-source generative models. GenAI-Arena has been operating for seven months, amassing over 9000 votes from the community. We describe our platform, analyze the data, and explain the statistical methods for ranking the models. To further promote the research in building model-based evaluation metrics, we release a cleaned version of our preference data for the three tasks, namely GenAI-Bench. We prompt the existing multi-modal models like Gemini, and GPT-4o to mimic human voting. We compute the accuracy by comparing the model voting with the human voting to understand their judging abilities. Our results show existing multimodal models are still lagging in assessing the generated visual content, even the best model GPT-4o only achieves an average accuracy of $49.19\%$ across the three generative tasks. Open-source MLLMs perform even worse due to the lack of instruction-following and reasoning ability in complex vision scenarios.
Dongfu Jiang, Max Ku, Tianle Li, Yuansheng Ni, Shizhuo Sun, Rongqi Fan, Wenhu Chen
NeurIPS1
2024 WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
abstract
Recent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that collects human preferences to evaluate VLMs. We curated WV-Bench by selecting 500 high-quality samples from 8,000 user submissions in WV-Arena. WV-Bench uses GPT-4 as the judge to compare each VLM with Claude-3-Sonnet, achieving a Spearman correlation of 0.94 with the WV-Arena Elo. This significantly outperforms other benchmarks like MMVet, MMMU, and MMStar.Our comprehensive analysis of 20K real-world interactions reveals important insights into the failure cases of top-performing VLMs. For example, we find that although GPT-4V surpasses many other models like Reka-Flash, Opus, and Yi-VL-Plus in simple visual recognition and reasoning tasks, it still faces challenges with subtle contextual cues, spatial reasoning, visual imagination, and expert domain knowledge. Additionally, current VLMs exhibit issues with hallucinations and safety when intentionally provoked. We are releasing our chat and feedback data to further advance research in the field of VLMs.
Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi 0001, Bill Y. Lin
NeurIPS2
2023 LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
abstract
We present LLM-BL E N D E R, an ensembling framework designed to attain consistently superior performance by leveraging the diverse strengths of multiple open-source large language models (LLMs).Our framework consists of two modules: PAIRRANKER and GEN-FUSER, addressing the observation that optimal LLMs for different examples can significantly vary.PAIRRANKER employs a specialized pairwise comparison method to distinguish subtle differences between candidate outputs.It jointly encodes the input text and a pair of candidates, using cross-attention encoders to determine the superior one.Our results demonstrate that PAIRRANKER exhibits the highest correlation with ChatGPT-based ranking.Then, GENFUSER aims to merge the top-ranked candidates, generating an improved output by capitalizing on their strengths and mitigating their weaknesses.To facilitate largescale evaluation, we introduce a benchmark dataset, MixInstruct, which is a mixture of multiple instruction datasets featuring oracle pairwise comparisons.Our LLM-BL E N D E R significantly outperform individual LLMs and baseline methods across various metrics, establishing a substantial performance gap. 1 2 Open Assistant 12.61% Koala 6.71% Alpaca 11.61% Baize 11.61% StableLM 1.90% FLAN-T5 0.80% Vicuna 21.22% Dolly V2 4.50% MOSS 12.91% ChatGLM 8.51% MPT 7.61% Percentage of Examples Where Each Model Ranks First Which LLM should I use for my input?All!I can ensemble!
Dongfu Jiang, Xiang Ren 0001, Bill Y. Lin
ACL (1)1