VLDB 2026 Research / reviewers in the wild / expert
Yichang Zhang
dblp:165/9507
· DBLP profile ↗
14ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationabstractXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoyuan Li 0001, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang 0007, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu |
ACL (1) | 7 |
| 2026 | Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward ModelsabstractBinghai Wang, Yantao Liu, Yuxuan Liu, Tianyi Tang, Shenzhi Wang, Chang Gao, Chujie Zheng, Yichang Zhang, Le Yu, Shixuan Liu, Tao Gui, Qi Zhang, Xuanjing Huang, Bowen Yu, Fei Huang, Junyang Lin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Binghai Wang, Yantao Liu, Shenzhi Wang, Chujie Zheng, Yichang Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Bowen Yu 0002, Fei Huang 0002, Junyang Lin |
ACL (1) | 8 |
| 2025 | Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMsabstractLarge Language Models (LLMs) can correct their self-generated responses, but a decline in accuracy after self-correction is also witnessed.To have a deeper understanding of selfcorrection, we endeavor to decompose, evaluate, and analyze the self-correction behaviors of LLMs.By enumerating and analyzing answer correctness before and after self-correction, we decompose the self-correction capability into confidence (being confident to correct answers) and critique (turning wrong answers to correct) capabilities, and propose two metrics from a probabilistic perspective to measure these 2 capabilities, along with another metric for overall self-correction capability evaluation.Based on our decomposition and evaluation metrics, we conduct extensive experiments and draw some empirical conclusions.For example, we find different models can exhibit distinct behaviors: some models are confident while others are more critical.We also find the trade-off between the two capabilities (i.e.improving one can lead to a decline in the other) when manipulating model self-correction behavior by prompts or in-context learning.Further, we find a simple yet efficient strategy to improve self-correction capability by transforming Supervision Fine-Tuning (SFT) data format, and our strategy outperforms vanilla SFT in both capabilities and achieves much higher accuracy after self-correction.Our code is publicly available on GitHub. Zhe Yang 0013, Yichang Zhang, Yudong Wang 0005, Ziyao Xu 0001, Junyang Lin, Zhifang Sui |
ACL (1) | 2 |
| 2025 | CodeArena: Evaluating and Aligning CodeLLMs on Human PreferenceabstractJian Yang, Jiaxi Yang, Wei Zhang, Jin Ke, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li, Binyuan Hui, Junyang Lin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jian Yang 0003, Jiaxi Yang 0004, Wei Zhang 0021, Yibo Miao, Lei Zhang 0201, Liqun Yang, Zeyu Cui, Yichang Zhang, Zhoujun Li 0001, Binyuan Hui, Junyang Lin |
EMNLP | 9 |
| 2025 | A Probabilistic Inference Scaling Theory for LLM Self-CorrectionabstractLarge Language Models (LLMs) have demonstrated the capability to refine their generated answers through self-correction, enabling continuous performance improvement over multiple rounds. However, the mechanisms underlying how and why accuracy evolves during this iterative process remain unexplored. To fill this gap, we propose a probabilistic theory to model the dynamics of accuracy change and explain the performance improvements observed in multi-round self-correction. Through mathematical derivation, we establish that the accuracy after the t^{th} round of self-correction is given by: Acc_t = Upp - \alpha^t(Upp - Acc_0),where Acc_0 denotes the initial accuracy, Upp represents the upper bound of accuracy convergence, and \alpha determines the rate of convergence. Based on our theory, these parameters can be calculated and the predicted accuracy curve then can be obtained through only a single round of self-correction. Extensive experiments across diverse models and datasets demonstrate that our theoretical predictions align closely with empirical accuracy curves, validating the effectiveness of the theory. Our work provides a theoretical foundation for understanding LLM self-correction, thus paving the way for further explorations. Yichang Zhang, Junyang Lin, Zhifang Sui |
EMNLP | 2 |
| 2025 | Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsabstractRecent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities.
However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning. Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang |
ICLR | 17 |
| 2025 | PolyMath: Evaluating Mathematical Reasoning in Multilingual ContextsabstractIn this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs.We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level.From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning:(1) Reasoning performance varies widely across languages for current LLMs;(2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance;(3) The thinking length differs significantly by language for current LLMs.Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs. Yiming Wang 0011, Pei Zhang 0011, Jialong Tang, Baosong Yang, Rui Wang 0015, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang 0002, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
NeurIPS | 12 |
| 2025 | An explainable prediction of shale gas ultimate recovery based on Tree-Based ensemble machine learning and Shapley additive explanations
Min Pang, Zhaoming Zhou, Yichang Zhang |
Appl. Intell. | 4 |
| 2024 | Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?abstractLarge language models (LLMs) have demonstrated impressive capabilities, but still suffer from inconsistency issues (e.g.LLMs can react differently to disturbances like rephrasing or inconsequential order change).In addition to these inconsistencies, we also observe that LLMs, while capable of solving hard problems, can paradoxically fail at easier ones.To evaluate this hard-to-easy inconsistency, we develop the ConsisEval benchmark, where each entry comprises a pair of questions with a strict order of difficulty.Furthermore, we introduce the concept of consistency score to quantitatively measure this inconsistency and analyze the potential for improvement in consistency by relative consistency score.Based on comprehensive experiments across a variety of existing models, we find: (1) GPT-4 achieves the highest consistency score of 92.2% but is still inconsistent to specific questions due to distraction by redundant information, misinterpretation of questions, etc.; (2) models with stronger capabilities typically exhibit higher consistency, but exceptions also exist; (3) hard data enhances consistency for both fine-tuning and in-context learning.Our data and code will be publicly available on GitHub. 1 Zhe Yang 0013, Yichang Zhang, Tianyu Liu 0001, Jian Yang 0003, Junyang Lin, Chang Zhou 0005, Zhifang Sui |
EMNLP | 2 |
| 2021 | Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsabstractHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang, Yaliang Li, Jun Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haojie Pan, Chengyu Wang 0001, Minghui Qiu, Yichang Zhang, Yaliang Li, Jun Huang 0007 |
ACL/IJCNLP (1) | 4 |
| 2021 | M6: Multi-Modality-to-Multi-Modality Multitask Mega-transformer for Unified PretrainingabstractMultimodal pretraining has demonstrated success in the downstream tasks of cross-modal representation learning. However, it is limited to the English data, and there is still a lack of large-scale dataset for multimodal pretraining in Chinese. In this work, we propose the largest dataset for pretraining in Chinese, which consists of over 1.9TB images and 292GB texts. The dataset has large coverage over domains, including encyclopedia, question answering, forum discussion, etc. Besides, we propose a method called M6, referring to Multi-Modality-to-Multi-Modality Multitask Mega-transformer, for unified pretraining on the data of single modality and multiple modalities. The model is pretrained with our proposed tasks, including text-to-text transfer, image-to-text transfer, as well as multi-modality-to-text transfer. The tasks endow the model with strong capability of understanding and generation. We scale the model to 10 billion parameters, and build the largest pretrained model in Chinese. Experimental results show that our proposed M6 outperforms the baseline in a number of downstream tasks concerning both single modality and multiple modalities, and the 10B-parameter pretrained model demonstrates strong potential in the setting of zero-shot learning. Junyang Lin, Rui Men, An Yang, Chang Zhou 0005, Yichang Zhang, Peng Wang 0028, Jingren Zhou 0001, Jie Tang 0001, Hongxia Yang |
KDD | 5 |
| 2019 | Towards Knowledge-Based Recommender Dialog SystemabstractQibin Chen, Junyang Lin, Yichang Zhang, Ming Ding, Yukuo Cen, Hongxia Yang, Jie Tang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Junyang Lin, Yichang Zhang, Ming Ding 0004, Yukuo Cen, Hongxia Yang, Jie Tang 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Towards Knowledge-Based Personalized Product Description Generation in E-commerceabstractQuality product descriptions are critical for providing competitive customer experience in an E-commerce platform. An accurate and attractive description not only helps customers make an informed decision but also improves the likelihood of purchase. However, crafting a successful product description is tedious and highly time-consuming. Due to its importance, automating the product description generation has attracted considerable interest from both research and industrial communities. Existing methods mainly use templates or statistical methods, and their performance could be rather limited. In this paper, we explore a new way to generate personalized product descriptions by combining the power of neural networks and knowledge base. Specifically, we propose a KnOwledge Based pErsonalized (or KOBE) product description generation model in the context of E-commerce. Junyang Lin, Yichang Zhang, Hongxia Yang, Jingren Zhou 0001, Jie Tang 0001 |
KDD | 3 |
| 2015 | Mining activation force defined dependency patterns for relation extraction
Chunyun Zhang, Yichang Zhang, Weiran Xu, Zhanyu Ma, Yan Leng, Jun Guo 0002 |
Knowl. Based Syst. | 2 |