Yichang Zhang

dblp:165/9507 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
abstract
Xiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoyuan Li 0001, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang 0007, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu
ACL (1)7
2026 Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
abstract
Binghai Wang, Yantao Liu, Yuxuan Liu, Tianyi Tang, Shenzhi Wang, Chang Gao, Chujie Zheng, Yichang Zhang, Le Yu, Shixuan Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Bowen Yu, Fei Huang, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Binghai Wang, Yantao Liu, Shenzhi Wang, Chujie Zheng, Yichang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Bowen Yu 0002, Fei Huang 0002, Junyang Lin
ACL (1)8
2025 Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
abstract
Large Language Models (LLMs) can correct their self-generated responses, but a decline in accuracy after self-correction is also witnessed.To have a deeper understanding of selfcorrection, we endeavor to decompose, evaluate, and analyze the self-correction behaviors of LLMs.By enumerating and analyzing answer correctness before and after self-correction, we decompose the self-correction capability into confidence (being confident to correct answers) and critique (turning wrong answers to correct) capabilities, and propose two metrics from a probabilistic perspective to measure these 2 capabilities, along with another metric for overall self-correction capability evaluation.Based on our decomposition and evaluation metrics, we conduct extensive experiments and draw some empirical conclusions.For example, we find different models can exhibit distinct behaviors: some models are confident while others are more critical.We also find the trade-off between the two capabilities (i.e.improving one can lead to a decline in the other) when manipulating model self-correction behavior by prompts or in-context learning.Further, we find a simple yet efficient strategy to improve self-correction capability by transforming Supervision Fine-Tuning (SFT) data format, and our strategy outperforms vanilla SFT in both capabilities and achieves much higher accuracy after self-correction.Our code is publicly available on GitHub.
Zhe Yang 0013, Yichang Zhang, Yudong Wang 0005, Ziyao Xu 0001, Junyang Lin, Zhifang Sui
ACL (1)2
2025 CodeArena: Evaluating and Aligning CodeLLMs on Human Preference
abstract
Jian Yang, Jiaxi Yang, Wei Zhang, Jin Ke, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li, Binyuan Hui, Junyang Lin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jian Yang 0003, Jiaxi Yang 0004, Wei Zhang 0021, Yibo Miao, Lei Zhang 0201, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li 0001, Binyuan Hui, Junyang Lin
EMNLP9
2025 A Probabilistic Inference Scaling Theory for LLM Self-Correction
abstract
Large Language Models (LLMs) have demonstrated the capability to refine their generated answers through self-correction, enabling continuous performance improvement over multiple rounds. However, the mechanisms underlying how and why accuracy evolves during this iterative process remain unexplored. To fill this gap, we propose a probabilistic theory to model the dynamics of accuracy change and explain the performance improvements observed in multi-round self-correction. Through mathematical derivation, we establish that the accuracy after the t^{th} round of self-correction is given by: Acc_t = Upp - \alpha^t(Upp - Acc_0),where Acc_0 denotes the initial accuracy, Upp represents the upper bound of accuracy convergence, and \alpha determines the rate of convergence. Based on our theory, these parameters can be calculated and the predicted accuracy curve then can be obtained through only a single round of self-correction. Extensive experiments across diverse models and datasets demonstrate that our theoretical predictions align closely with empirical accuracy curves, validating the effectiveness of the theory. Our work provides a theoretical foundation for understanding LLM self-correction, thus paving the way for further explorations.
Yichang Zhang, Junyang Lin, Zhifang Sui
EMNLP2
2025 Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language Models
abstract
Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning.
Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang
ICLR17
2025 PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
abstract
In this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs.We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level.From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning:(1) Reasoning performance varies widely across languages for current LLMs;(2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance;(3) The thinking length differs significantly by language for current LLMs.Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs.
Yiming Wang 0011, Pei Zhang 0011, Jialong Tang, Baosong Yang, Rui Wang 0015, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang 0002, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001
NeurIPS12
2025 An explainable prediction of shale gas ultimate recovery based on Tree-Based ensemble machine learning and Shapley additive explanations
Min Pang, Zhaoming Zhou, Yichang Zhang
Appl. Intell.4
2024 Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
abstract
Large language models (LLMs) have demonstrated impressive capabilities, but still suffer from inconsistency issues (e.g.LLMs can react differently to disturbances like rephrasing or inconsequential order change).In addition to these inconsistencies, we also observe that LLMs, while capable of solving hard problems, can paradoxically fail at easier ones.To evaluate this hard-to-easy inconsistency, we develop the ConsisEval benchmark, where each entry comprises a pair of questions with a strict order of difficulty.Furthermore, we introduce the concept of consistency score to quantitatively measure this inconsistency and analyze the potential for improvement in consistency by relative consistency score.Based on comprehensive experiments across a variety of existing models, we find: (1) GPT-4 achieves the highest consistency score of 92.2% but is still inconsistent to specific questions due to distraction by redundant information, misinterpretation of questions, etc.; (2) models with stronger capabilities typically exhibit higher consistency, but exceptions also exist; (3) hard data enhances consistency for both fine-tuning and in-context learning.Our data and code will be publicly available on GitHub. 1
Zhe Yang 0013, Yichang Zhang, Tianyu Liu 0001, Jian Yang 0003, Junyang Lin, Chang Zhou 0005, Zhifang Sui
EMNLP2
2021 Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across Domains
abstract
Haojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang, Yaliang Li, Jun Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Haojie Pan, Chengyu Wang 0001, Minghui Qiu, Yichang Zhang, Yaliang Li, Jun Huang 0007
ACL/IJCNLP (1)4
2021 M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified Pretraining
abstract
Multimodal pretraining has demonstrated success in the downstream tasks of cross-modal representation learning. However, it is limited to the English data, and there is still a lack of large-scale dataset for multimodal pretraining in Chinese. In this work, we propose the largest dataset for pretraining in Chinese, which consists of over 1.9TB images and 292GB texts. The dataset has large coverage over domains, including encyclopedia, question answering, forum discussion, etc. Besides, we propose a method called M6, referring to Multi-Modality-to-Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. The model is pretrained with our proposed tasks, including text-to-text transfer, image-to-text transfer, as well as multi-modality-to-text transfer. The tasks endow the model with strong capability of understanding and generation. We scale the model to 10 billion parameters, and build the largest pretrained model in Chinese. Experimental results show that our proposed M6 outperforms the baseline in a number of downstream tasks concerning both single modality and multiple modalities, and the 10B-parameter pretrained model demonstrates strong potential in the setting of zero-shot learning.
Junyang Lin, Rui Men, An Yang, Chang Zhou 0005, Yichang Zhang, Peng Wang 0028, Jingren Zhou 0001, Jie Tang 0001, Hongxia Yang
KDD5
2019 Towards Knowledge-Based Recommender Dialog System
abstract
Qibin Chen, Junyang Lin, Yichang Zhang, Ming Ding, Yukuo Cen, Hongxia Yang, Jie Tang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junyang Lin, Yichang Zhang, Ming Ding 0004, Yukuo Cen, Hongxia Yang, Jie Tang 0001
EMNLP/IJCNLP (1)3
2019 Towards Knowledge-Based Personalized Product Description Generation in E-commerce
abstract
Quality product descriptions are critical for providing competitive customer experience in an E-commerce platform. An accurate and attractive description not only helps customers make an informed decision but also improves the likelihood of purchase. However, crafting a successful product description is tedious and highly time-consuming. Due to its importance, automating the product description generation has attracted considerable interest from both research and industrial communities. Existing methods mainly use templates or statistical methods, and their performance could be rather limited. In this paper, we explore a new way to generate personalized product descriptions by combining the power of neural networks and knowledge base. Specifically, we propose a KnOwledge Based pErsonalized (or KOBE) product description generation model in the context of E-commerce.
Junyang Lin, Yichang Zhang, Hongxia Yang, Jingren Zhou 0001, Jie Tang 0001
KDD3
2015 Mining activation force defined dependency patterns for relation extraction
Chunyun Zhang, Yichang Zhang, Weiran Xu, Zhanyu Ma, Yan Leng, Jun Guo 0002
Knowl. Based Syst.2