VLDB 2026 Research / reviewers in the wild / expert
Qingfeng Sun
dblp:194/5100
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-4000-6717ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RubricBench: Aligning Model-Generated Rubrics with Human StandardsabstractJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junyi Zhou 0007, Qiyuan Zhang 0001, Yufei Wang 0005, Fuyuan Lyu, Yidong Ming, Can Xu 0002, Qingfeng Sun, Kai Zheng 0001, Xue (Steve) Liu, Chen Ma 0001 |
ACL (1) | 7 |
| 2025 | WarriorCoder: Learning from Expert Battles to Augment Code Large Language ModelsabstractHuawen Feng, Pu Zhao, Qingfeng Sun, Can Xu, Fangkai Yang, Lu Wang, Qianli Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Huawen Feng, Pu Zhao 0004, Qingfeng Sun, Can Xu 0002, Fangkai Yang, Lu Wang 0029, Qianli Ma 0001, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
ACL (1) | 3 |
| 2025 | WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructabstractLarge language models (LLMs), such as GPT-4, have shown remarkable performance in natural language processing (NLP) tasks, including challenging mathematical reasoning. However, most existing open-source models are only pre-trained on large-scale internet data and without math-related optimization. In this paper, we present WizardMath, which enhances the mathematical reasoning abilities of LLMs, by applying our proposed Reinforcement Learning from Evol-Instruct Feedback (RLEIF) method to the domain of math. Through extensive experiments on two mathematical reasoning benchmarks, namely GSM8k and MATH, we reveal the extraordinary capabilities of our model. Remarkably, WizardMath-Mistral 7B surpasses all other open-source LLMs by a substantial margin. Furthermore, WizardMath 70B even outperforms ChatGPT-3.5, Claude Instant, Gemini Pro and Mistral Medium. Additionally, our preliminary exploration highlights the pivotal role of instruction evolution and process supervision in achieving exceptional math performance. Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Jian-Guang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, Yansong Tang, Dongmei Zhang 0001 |
ICLR | 2 |
| 2025 | AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
Mengkang Hu, Pu Zhao 0004, Can Xu 0002, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo 0002, Saravan Rajmohan |
KDD (1) | 4 |
| 2024 | WizardCoder: Empowering Code Large Language Models with Evol-InstructabstractCode Large Language Models (Code LLMs), such as StarCoder, have demonstrated remarkable performance in various code-related tasks. However, different from their counterparts in the general language modeling field, the technique of instruction fine-tuning remains relatively under-researched in this domain. In this paper, we present Code Evol-Instruct, a novel approach that adapts the Evol-Instruct method to the realm of code, enhancing Code LLMs to create novel models, WizardCoder. Through comprehensive experiments on five prominent code generation benchmarks, namely HumanEval, HumanEval+, MBPP, DS-1000, and MultiPL-E, our models showcase outstanding performance. They consistently outperform all other open-source Code LLMs by a significant margin. Remarkably, WizardCoder 15B even surpasses the well-known closed-source LLMs, including Anthropic's Claude and Google's Bard, on the HumanEval and HumanEval+ benchmarks. Additionally, WizardCoder 34B not only achieves a HumanEval score comparable to GPT3.5 (ChatGPT) but also surpasses it on the HumanEval+ benchmark. Furthermore, our preliminary exploration highlights the pivotal role of instruction complexity in achieving exceptional coding performance. Can Xu 0002, Pu Zhao 0004, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma 0004, Qingwei Lin, Daxin Jiang |
ICLR | 4 |
| 2024 | WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsabstractTraining large language models (LLMs) with open-domain instruction following data brings colossal success. However, manually creating such instruction data is very time-consuming and labor-intensive. Moreover, humans may struggle to produce high-complexity instructions. In this paper, we show an avenue for creating large amounts of instruction data with varying levels of complexity using LLM instead of humans. Starting with an initial set of instructions, we use our proposed Evol-Instruct to rewrite them step by step into more complex instructions. Then, we mix all generated instruction data to fine-tune LLaMA. We call the resulting model WizardLM. Both automatic and human evaluations consistently indicate that WizardLM outperforms baselines such as Alpaca (trained from Self-Instruct) and Vicuna (trained from human-created instructions). The experimental results demonstrate that the quality of instruction-following dataset crafted by Evol-Instruct can significantly improve the performance of LLMs. Can Xu 0002, Qingfeng Sun, Kai Zheng 0021, Xiubo Geng, Pu Zhao 0004, Jiazhan Feng, Chongyang Tao, Qingwei Lin, Daxin Jiang |
ICLR | 2 |
| 2024 | WizardArena: Post-training Large Language Models via Simulated Offline Chatbot ArenaabstractRecent work demonstrates that, post-training large language models with open-domain instruction following data have achieved colossal success. Simultaneously, human Chatbot Arena has emerged as one of the most reasonable benchmarks for model evaluation and developmental guidance. However, the processes of manually curating high-quality training data and utilizing online human evaluation platforms are both expensive and limited. To mitigate the manual and temporal costs associated with post-training, this paper introduces a Simulated Chatbot Arena named WizardArena, which is fully based on and powered by open-source LLMs. For evaluation scenario, WizardArena can efficiently predict accurate performance rankings among different models based on offline test set. For training scenario, we simulate arena battles among various state-of-the-art models on a large scale of instruction data, subsequently leveraging the battle results to constantly enhance target model in both the supervised fine-tuning and reinforcement learning . Experimental results demonstrate that our WizardArena aligns closely with the online human arena rankings, and our models trained on offline extensive battle data exhibit significant performance improvements during SFT, DPO, and PPO stages. Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Qingwei Lin, Jian-Guang Lou, Shifeng Chen, Yansong Tang, Weizhu Chen |
NeurIPS | 2 |
| 2023 | MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationabstractJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, Qingwei Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiazhan Feng, Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Yaming Yang 0001, Chongyang Tao, Dongyan Zhao 0001, Qingwei Lin |
ACL (1) | 2 |
| 2022 | PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksabstractYufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yufei Wang 0003, Can Xu 0002, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, Daxin Jiang |
ACL (1) | 3 |
| 2022 | Multimodal Dialogue Response GenerationabstractQingfeng Sun, Yujing Wang, Can Xu, Kai Zheng, Yaming Yang, Huang Hu, Fei Xu, Jessica Zhang, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Qingfeng Sun, Can Xu 0002, Kai Zheng 0021, Yaming Yang 0001, Huang Hu, Xiubo Geng, Daxin Jiang |
ACL (1) | 1 |
| 2022 | Stylized Knowledge-Grounded Dialogue Generation via Disentangled Template RewritingabstractQingfeng Sun, Can Xu, Huang Hu, Yujing Wang, Jian Miao, Xiubo Geng, Yining Chen, Fei Xu, Daxin Jiang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Qingfeng Sun, Can Xu 0002, Huang Hu, Jian Miao, Xiubo Geng, Daxin Jiang |
NAACL-HLT | 1 |
| 2019 | Hierarchical Attention Prototypical Networks for Few-Shot Text ClassificationabstractShengli Sun, Qingfeng Sun, Kevin Zhou, Tengchao Lv. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shengli Sun, Qingfeng Sun, Kevin Zhou, Tengchao Lv |
EMNLP/IJCNLP (1) | 2 |