Zhihan Zhang 0001

dblp:245/8608-1 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0003-0922-9146ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
abstract
Tengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su, Yuanyang Liu, Zhihan Zhang, Meng Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tengchao Yang, Sichen Guo, Mengzhao Jia, Jiaming Su 0008, Yuanyang Liu, Zhihan Zhang 0001, Meng Jiang 0001
ACL (1)6
2025 Enhancing Mathematical Reasoning in LLMs by Stepwise Correction
abstract
Best-of-N decoding methods instruct large language models (LLMs) to generate multiple solutions, score each using a scoring function, and select the highest scored as the final answer to mathematical reasoning problems.However, this repeated independent process often leads to the same mistakes, making the selected solution still incorrect.We propose a novel prompting method named Stepwise Correction (STEPCO) that helps LLMs identify and revise incorrect steps in their generated reasoning paths.It iterates verification and revision phases that employ a process-supervised verifier.The verifythen-revise process not only improves answer correctness but also reduces token consumption with fewer paths needed to generate.With STEPCO, a series of LLMs demonstrate exceptional performance.Notably, using GPT-4o as the backend LLM, STEPCO achieves an average accuracy of 94.1 across eight datasets, significantly outperforming the state-of-the-art Best-of-N method by +2.4, while reducing token consumption by 77.8%.Our implementation is made publicly available at https: //wzy6642.github.io/stepco.github.io.
Zhenyu Wu 0004, Qingkai Zeng 0001, Zhihan Zhang 0001, Zhaoxuan Tan, Meng Jiang 0001
ACL (1)3
2025 Aligning Large Language Models with Implicit Preferences from User-Generated Content
abstract
Learning from preference feedback is essential for aligning large language models (LLMs) with human values and improving the quality of generated responses. However, existing preference learning methods rely heavily on curated data from humans or advanced LLMs, which is costly and difficult to scale. In this work, we present PUGC, a novel framework that leverages implicit human Preferences in unlabeled User-Generated Content (UGC) to generate preference data. Although UGC is not explicitly created to guide LLMs in generating human-preferred responses, it often reflects valuable insights and implicit preferences from its creators that has the potential to address readers’ questions. PUGC transforms UGC into user queries and generates responses from the policy model. The UGC is then leveraged as a reference text for response scoring, aligning the model with these implicit preferences. This approach improves the quality of preference data while enabling scalable, domain-specific alignment. Experimental results on Alpaca Eval 2 show that models trained with DPO and PUGC achieve a 9.37% performance improvement over traditional methods, setting a 35.93% state-of-the-art length-controlled win rate using Mistral-7B-Instruct. Further studies highlight gains in reward quality, domain-specific alignment effectiveness, robustness against UGC quality, and theory of mind capabilities. Our code and dataset are available at https://zhaoxuan.info/PUGC.github.io/.
Zhaoxuan Tan, Zheng Li 0018, Hyokun Yun, Ming Zeng 0001, Zhihan Zhang 0001, Yifan Gao 0001, Ruijie Wang 0004, Priyanka Nigam, Meng Jiang 0001
ACL (1)8
2025 RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
abstract
Large Language Models (LLMs) excel in code generation yet struggle with modern AI software engineering tasks. Unlike traditional function-level or file-level coding tasks, AI software engineering requires not only basic coding proficiency but also advanced skills in managing and interacting with code repositories. However, existing methods often overlook the need for repository-level code understanding, which is crucial for accurately grasping the broader context and developing effective solutions. On this basis, we present RepoGraph, a plug-in module that manages a repository-level structure for modern AI software engineering solutions. RepoGraph offers the desired guidance and serves as a repository-wide navigation for AI software engineers. We evaluate RepoGraph on the SWE-bench by plugging it into four different methods of two lines of approaches, where RepoGraph substantially boosts the performance of all systems, leading to a new state-of-the-art among open-source frameworks. Our analyses also demonstrate the extensibility and flexibility of RepoGraph by testing on another repo-level coding benchmark, CrossCodeEval. Our code is available at https://github.com/ozyyshr/RepoGraph.
Siru Ouyang, Wenhao Yu 0002, Kaixin Ma, Zilin Xiao, Zhihan Zhang 0001, Mengzhao Jia, Jiawei Han 0001, Hongming Zhang 0009, Dong Yu 0001
ICLR5
2025 BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
abstract
Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing **diverse function calls as tools** to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately understanding **complex instructions**. Fulfilling both of these characteristics can pose a great challenge for LLMs. To assess how well LLMs can solve challenging and practical tasks via programs, we introduce BigCodeBench, a benchmark that challenges LLMs to invoke multiple function calls as tools from 139 libraries and 7 domains for 1,140 fine-grained tasks. To evaluate LLMs rigorously, each task encompasses 5.6 test cases with an average branch coverage of 99%. In addition, we propose a natural-language-oriented variant of BigCodeBench, BigCodeBench-Instruct, that automatically transforms the original docstrings into short instructions containing only essential information. Our extensive evaluation of 60 LLMs shows that **LLMs are not yet capable of following complex instructions to use function calls precisely, with scores up to 60%, significantly lower than the human performance of 97%**. The results underscore the need for further advancements in this area.
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 0011, Wenhao Yu 0002, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong 0005, James Hoang, Armel Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Zhihan Zhang 0001, Prateek Yadav
ICLR19
2025 ALERT: An LLM-powered Benchmark for Automatic Evaluation of Recommendation Explanations
abstract
Yichuan Li, Xinyang Zhang, Chenwei Zhang, Mao Li, Tianyi Liu, Pei Chen, Yifan Gao, Kyumin Lee, Kaize Ding, Zhengyang Wang, Zhihan Zhang, Jingbo Shang, Xian Li, Trishul Chilimbi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yichuan Li 0001, Xinyang Zhang 0002, Yifan Gao 0001, Kyumin Lee, Kaize Ding, Zhihan Zhang 0001, Jingbo Shang, Trishul Chilimbi
NAACL (Long Papers)11
2025 IHEval: Evaluating Language Models on Following the Instruction Hierarchy
abstract
Zhihan Zhang, Shiyang Li, Zixuan Zhang, Xin Liu, Haoming Jiang, Xianfeng Tang, Yifan Gao, Zheng Li, Haodong Wang, Zhaoxuan Tan, Yichuan Li, Qingyu Yin, Bing Yin, Meng Jiang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhihan Zhang 0001, Xin Liu 0039, Haoming Jiang, Xianfeng Tang, Yifan Gao 0001, Zheng Li 0018, Zhaoxuan Tan, Yichuan Li 0001, Qingyu Yin, Meng Jiang 0001
NAACL (Long Papers)1
2025 MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
abstract
Zifeng Zhu, Mengzhao Jia, Zhihan Zhang, Lang Li, Meng Jiang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zifeng Zhu, Mengzhao Jia, Zhihan Zhang 0001, Meng Jiang 0001
NAACL (Long Papers)3
2024 PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning
abstract
Zhihan Zhang, Dong-Ho Lee, Yuwei Fang, Wenhao Yu, Mengzhao Jia, Meng Jiang, Francesco Barbieri. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhihan Zhang 0001, Yuwei Fang, Wenhao Yu 0002, Mengzhao Jia, Meng Jiang 0001, Francesco Barbieri
ACL (1)1
2024 Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited Examples
abstract
Automatic taxonomy induction is crucial for web search, recommendation systems, and question answering. Manual curation of taxonomies is expensive in terms of human effort, making automatic taxonomy construction highly desirable. In this work, we introduce Chain-of-Layer which is an in-context learning framework designed to induct taxonomies from a given set of entities. Chain-of-Layer breaks down the task into selecting relevant candidate entities in each layer and gradually building the taxonomy from top to bottom. To minimize errors, we introduce the Ensemble-based Ranking Filter to reduce the hallucinated content generated at each iteration. Through extensive experiments, we demonstrate that Chain-of-Layer achieves state-of-the-art performance on four real-world benchmarks. Source code available at: https://github.com/qingkaizeng/chain-of-layer.
Qingkai Zeng 0001, Yuyang Bai, Zhaoxuan Tan, Shangbin Feng, Zhenwen Liang, Zhihan Zhang 0001, Meng Jiang 0001
CIKM6
2024 Large Language Models Can Self-Correct with Key Condition Verification
abstract
Intrinsic self-correct was a method that instructed large language models (LLMs) to verify and correct their responses without external feedback.Unfortunately, the study concluded that the LLMs could not self-correct reasoning yet.We find that a simple yet effective prompting method enhances LLM performance in identifying and correcting inaccurate answers without external feedback.That is to mask a key condition in the question, add the current response to construct a verification question, and predict the condition to verify the response.The condition can be an entity in an open-domain question or a numerical value in an arithmetic question, which requires minimal effort (via prompting) to identify.We propose an iterative verify-then-correct framework to progressively identify and correct (probably) false responses, named PROCO.We conduct experiments on three reasoning tasks.On average, PROCO, with GPT-3.5-Turbo-1106 as the backend LLM, yields +6.8 exact match on four open-domain question answering datasets, +14.1 accuracy on three arithmetic reasoning datasets, and +9.6 accuracy on a commonsense reasoning dataset, compared to Self-Correct.Our implementation is made publicly avail
Zhenyu Wu 0004, Qingkai Zeng 0001, Zhihan Zhang 0001, Zhaoxuan Tan, Meng Jiang 0001
EMNLP3
2024 Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
abstract
Supervised fine-tuning enhances the problemsolving abilities of language models across various mathematical reasoning tasks.To maximize such benefits, existing research focuses on broadening the training set with various data augmentation techniques, which is effective for standard single-round question-answering settings.Our work introduces a novel technique aimed at cultivating a deeper understanding of the training problems at hand, enhancing performance not only in standard settings but also in more complex scenarios that require reflective thinking.Specifically, we propose reflective augmentation, a method that embeds problem reflection into each training instance.It trains the model to consider alternative perspectives and engage with abstractions and analogies, thereby fostering a thorough comprehension through reflective reasoning.Extensive experiments validate the achievement of our aim, underscoring the unique advantages of our method and its complementary nature relative to existing augmentation techniques. 1 Question Answer Question
Zhihan Zhang 0001, Tao Ge 0001, Zhenwen Liang, Wenhao Yu 0002, Dian Yu 0001, Mengzhao Jia, Dong Yu 0001, Meng Jiang 0001
EMNLP1
2023 A Survey of Multi-task Learning in Natural Language Processing: Regarding Task Relatedness and Training Methods
abstract
Multi-task learning (MTL) has become increasingly popular in natural language processing (NLP) because it improves the performance of related tasks by exploiting their commonalities and differences.Nevertheless, it is still not understood very well how multi-task learning can be implemented based on the relatedness of training tasks.In this survey, we review recent advances of multi-task learning methods in NLP, with the aim of summarizing them into two general multi-task training methods based on their task relatedness: (i) joint training and (ii) multi-step training.We present examples in various NLP downstream applications, summarize the task relationships and discuss future directions of this promising topic.
Zhihan Zhang 0001, Wenhao Yu 0002, Mengxia Yu, Zhichun Guo, Meng Jiang 0001
EACL1
2023 Pre-training Language Models for Comparative Reasoning
abstract
Comparative reasoning is a process of comparing objects, concepts, or entities to draw conclusions, which constitutes a fundamental cognitive ability.In this paper, we propose a novel framework to pre-train language models for enhancing their abilities of comparative reasoning over texts.While there have been approaches for NLP tasks that require comparative reasoning, they suffer from costly manual data labeling and limited generalizability to different tasks.Our approach introduces a novel method of collecting scalable data for text-based entity comparison, which leverages both structured and unstructured data.Moreover, we present a framework of pre-training language models via three novel objectives on comparative reasoning.Evaluation on downstream tasks including comparative question answering, question generation, and summarization shows that our pre-training framework significantly improves the comparative reasoning abilities of language models, especially under low-resource conditions.This work also releases the first integrated benchmark for comparative reasoning.
Mengxia Yu, Zhihan Zhang 0001, Wenhao Yu 0002, Meng Jiang 0001
EMNLP2
2023 Exploring Contrast Consistency of Open-Domain Question Answering Systems on Minimally Edited Questions
abstract
Abstract Contrast consistency, the ability of a model to make consistently correct predictions in the presence of perturbations, is an essential aspect in NLP. While studied in tasks such as sentiment analysis and reading comprehension, it remains unexplored in open-domain question answering (OpenQA) due to the difficulty of collecting perturbed questions that satisfy factuality requirements. In this work, we collect minimally edited questions as challenging contrast sets to evaluate OpenQA models. Our collection approach combines both human annotation and large language model generation. We find that the widely used dense passage retriever (DPR) performs poorly on our contrast sets, despite fitting the training set well and performing competitively on standard test sets. To address this issue, we introduce a simple and effective query-side contrastive loss with the aid of data augmentation to improve DPR training. Our experiments on the contrast sets demonstrate that DPR’s contrast consistency is improved without sacrificing its accuracy on the standard test sets.1
Zhihan Zhang 0001, Wenhao Yu 0002, Zheng Ning, Mingxuan Ju, Meng Jiang 0001
Trans. Assoc. Comput. Linguistics1
2023 Modeling Co-Evolution of Attributed and Structural Information in Graph Sequence
abstract
Most graph neural network models learn embeddings of nodes in static attributed graphs for predictive analysis. Recent attempts have been made to learn temporal proximity of the nodes. We find that real dynamic attributed graphs exhibit complex phenomenon of co-evolution between node attributes and graph structure. Learning node embeddings for forecasting change of node attributes and evolution of graph structure over time remains an open problem. In this work, we present a novel framework called CoEvoGNN for modeling dynamic attributed graph sequence. It preserves the impact of earlier graphs on the current graph by embedding generation through the sequence of attributed graphs. It has a temporal self-attention architecture to model long-range dependencies in the evolution. Moreover, CoEvoGNN optimizes model parameters jointly on two dynamic tasks, attribute inference and link prediction over time. So the model can capture the co-evolutionary patterns of attribute change and link formation. This framework can adapt to any graph neural algorithms so we implemented and investigated three methods based on it: CoEvoGCN, CoEvoGAT, and CoEvoSAGE. Experiments demonstrate the framework (and its methods) outperforms strong baseline methods on predicting an entire unseen graph snapshot of personal attributes and interpersonal links in dynamic social graphs and financial graphs.
Daheng Wang, Zhihan Zhang 0001, Yihong Ma, Tong Zhao 0003, Tianwen Jiang, Nitesh V. Chawla, Meng Jiang 0001
IEEE Trans. Knowl. Data Eng.2
2022 Retrieval Augmentation for Commonsense Reasoning: A Unified Approach
abstract
A common thread of retrieval-augmented methods in the existing literature focuses on retrieving encyclopedic knowledge, such as Wikipedia, which facilitates well-defined entity and relation spaces that can be modeled.However, applying such methods to commonsense reasoning tasks faces two unique challenges, i.e., the lack of a general large-scale corpus for retrieval and a corresponding effective commonsense retriever.In this paper, we systematically investigate how to leverage commonsense knowledge retrieval to improve commonsense reasoning tasks.We proposed a unified framework of Retrieval-Augmented Commonsense reasoning (called RACO), including a newly constructed commonsense corpus with over 20 million documents and novel strategies for training a commonsense retriever.We conducted experiments on four different commonsense reasoning tasks.Extensive evaluation results showed that our proposed RACO can significantly outperform other knowledgeenhanced method counterparts, achieving new SoTA performance on the CommonGen 1 and CREAK 2 leaderboards.Our code is available at https://github.com/wyu97/RACo.
Wenhao Yu 0002, Chenguang Zhu 0001, Zhihan Zhang 0001, Shuohang Wang, Zhuosheng Zhang 0001, Yuwei Fang, Meng Jiang 0001
EMNLP3
2022 A Unified Encoder-Decoder Framework with Entity Memory
abstract
Entities, as important carriers of real-world knowledge, play a key role in many NLP tasks.We focus on incorporating entity knowledge into an encoder-decoder framework for informative text generation.Existing approaches tried to index, retrieve, and read external documents as evidence, but they suffered from a large computational overhead.In this work, we propose an Encoder-Decoder framework with an entity Memory, namely EDMem.The entity knowledge is stored in the memory as latent representations, and the memory is pre-trained on Wikipedia along with encoder-decoder parameters.To precisely generate entity names, we design three decoding methods to constrain entity generation by linking entities in the memory.EDMem is a unified framework that can be used on various entity-intensive question answering and generation tasks.Extensive experimental results show that EDMem outperforms both memory-based auto-encoder models and non-memory encoder-decoder models.
Zhihan Zhang 0001, Wenhao Yu 0002, Chenguang Zhu 0001, Meng Jiang 0001
EMNLP1
2021 Knowledge-Aware Procedural Text Understanding with Multi-Stage Training
abstract
Procedural text describes dynamic state changes during a step-by-step natural process (e.g., photosynthesis). In this work, we focus on the task of procedural text understanding, which aims to comprehend such documents and track entities’ states and locations during a process. Although recent approaches have achieved substantial progress, their results are far behind human performance. Two challenges, the difficulty of commonsense reasoning and data insufficiency, still remain unsolved, which require the incorporation of external knowledge bases. Previous works on external knowledge injection usually rely on noisy web mining tools and heuristic rules with limited applicable scenarios. In this paper, we propose a novel KnOwledge-Aware proceduraL text understAnding (KoaLa) model, which effectively leverages multiple forms of external knowledge in this task. Specifically, we retrieve informative knowledge triples from ConceptNet and perform knowledge-aware reasoning while tracking the entities. Besides, we employ a multi-stage training schema which fine-tunes the BERT model over unlabeled data collected from Wikipedia before further fine-tuning it on the final model. Experimental results on two procedural text datasets, ProPara and Recipes, verify the effectiveness of the proposed methods, in which our model achieves state-of-the-art performance in comparison to various baselines.1
Zhihan Zhang 0001, Xiubo Geng, Tao Qin 0001, Yunfang Wu, Daxin Jiang
WWW1
2021 Domain-specific meta-embedding with latent semantic structures
Qian Liu 0012, Jie Lu 0001, Guangquan Zhang 0001, Tao Shen 0001, Zhihan Zhang 0001, Heyan Huang
Inf. Sci.5
2020 DCA: Diversified Co-attention Towards Informative Live Video Commenting
Zhihan Zhang 0001, Zhiyi Yin, Shuhuai Ren
NLPCC (2)1
2019 Cross-Modal Commentator: Automatic Machine Commenting Based on Cross-Modal Information
abstract
Automatic commenting of online articles can provide additional opinions and facts to the reader, which improves user experience and engagement on social media platforms.Previous work focuses on automatic commenting based solely on textual content.However, in real-scenarios, online articles usually contain multiple modal contents.For instance, graphic news contains plenty of images in addition to text.Contents other than text are also vital because they are not only more attractive to the reader but also may provide critical information.To remedy this, we propose a new task: cross-model automatic commenting (CMAC), which aims to make comments by integrating multiple modal contents.We construct a largescale dataset for this task and explore several representative methods.Going a step further, an effective co-attention model is presented to capture the dependency between textual and visual information.Evaluation results show that our proposed model can achieve better performance than competitive baselines.1
Zhihan Zhang 0001, Fuli Luo, Lei Li 0039, Chengyang Huang, Xu Sun 0001
ACL (1)2
2019 CTGA: Graph-based Biomedical Literature Search
abstract
Scientists are relying heavily on biomedical literature search (BLS) engines (e.g., PubMed) to acquire knowledge. Existing BLS systems adopt a “C-A” paradigm that is to design query-document similarity measurement based on words/phrases in the unstructured Content and to develop search Algorithms. In this work, we argue that structures should be extracted and utilized to bridge the gap between text content and knowledge-based search. And graph is one of the most effective structured forms of knowledge and more informative than words or phrases. So we carry out a paradigm shift from “C-A” to “CTGA”. Here “T” is for factual tuple of concepts and relations, and “G” is for knowledge graph. Our proposed graph-based BLS system has three parts: (1) it uses neural information extraction models to turn text into tuples; (2) it represents the tuples of a query or a document as a knowledge graph of linked concept nodes; (3) it has an efficient graph-based matching algorithm to return related documents at the level of structured knowledge. Experiments show that in both objective and subjective evaluation, our CTGA performs significantly better than traditional CA-based PubMed.
Tianwen Jiang, Zhihan Zhang 0001, Tong Zhao 0003, Bing Qin 0001, Ting Liu 0001, Nitesh V. Chawla, Meng Jiang 0001
BIBM2