EDBT 2026 Demo / reviewers in the wild / expert
Zhijiang Guo
dblp:43/6147
· DBLP profile ↗
44ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0002-6232-5957ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 6 first-author · 34 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Too Long, Do Re-weighting for Efficient LLM Reasoning CompressionabstractZhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhongzhi Li, Lei Ji 0001, Xing W, Haizhen Huang, Yeyun Gong, Zhijiang Guo, Xiao Liu 0029, Cheng-Lin Liu 0001 |
ACL (1) | 11 |
| 2026 | Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the EdgeabstractFederated fine-tuning enables privacypreserving LLM adaptation but faces a critical bottleneck: the disparity between LLMs' high memory demands and edge devices' limited capacity.To break the memory barrier, we propose Chain Federated Fine-Tuning (CHAINFED), an innovative paradigm that forgoes end-to-end updates in favor of a sequential, layer-by-layer manner.It first trains the initial adapter to convergence, freezes its weights, and then proceeds to the next.This iterative train-and-freeze process forms an optimization chain, gradually enhancing the model's task-specific proficiency.CHAINFED further integrates three core techniques: 1) Dynamic Layer Co-Tuning to bridge semantic gaps between sequentially tuned layers and facilitate information flow; 2) Globally Perceptive Optimization to endow each adapter with foresight beyond its local objective; 3) Function-Oriented Adaptive Tuning to automatically identify the optimal fine-tuning starting point.Extensive experiments on multiple benchmarks demonstrate the superiority of CHAINFED over existing methods, boosting average accuracy by up to 46.46%. Yebo Wu, Jingguang Li, Chunlin Tian, Kahou Tam, Zhijiang Guo, Li Li 0064 |
ACL (1) | 5 |
| 2026 | From System 1 to System 2: A Survey of Reasoning Large Language ModelsabstractAchieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical reasoning for more accurate judgments and reduced biases. Foundational Large Language Models (LLMs) excel at fast decision-making but lack the depth for complex reasoning, as they have not yet fully embraced the step-by-step analysis characteristic of true System 2 thinking. Recently, reasoning LLMs like OpenAI's o1/o3 and DeepSeek's R1 have demonstrated expert-level performance in fields such as mathematics and coding, closely mimicking the deliberate reasoning of System 2 and showcasing human-like cognitive abilities. This survey begins with a brief overview of the progress in foundational LLMs and the early development of System 2 technologies, exploring how their combination has paved the way for reasoning LLMs. Next, we discuss how to construct reasoning LLMs, trace the evolution of various reasoning models, and examine the core methods that enable advanced reasoning behind them. Additionally, we provide an overview of reasoning benchmarks, offering an in-depth comparison of the performance of representative reasoning LLMs. Finally, we explore promising directions for advancing reasoning LLMs and maintain a real-time GitHub Repository to track the latest developments. We hope this survey will serve as a valuable resource to inspire innovation and drive progress in this rapidly evolving field. Duzhen Zhang, Zhongzhi Li, Jiaxin Zhang 0024, Zengyan Liu, Junhao Zheng, Xiuyi Chen, Jiahua Dong 0001, Zhijiang Guo, Cheng-Lin Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 13 |
| 2025 | TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer ReviewabstractYuan Chang, Ziyue Li, Hengyuan Zhang, Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yuanbo Kong, Yanru Wu, Hayden Kwok-Hay So, Zhijiang Guo, Liya Zhu, Ngai Wong 0001 |
EMNLP | 7 |
| 2025 | ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific ChartsabstractScientific fact-checking has largely focused on textual and tabular sources, neglecting scientific charts-a primary medium for conveying quantitative evidence and supporting statistical reasoning in research communication.We introduce CLIMATEVIZ, the first largescale benchmark for scientific fact-checking grounded in real-world, expert-curated scientific charts.CLIMATEVIZ comprises 49,862 claims paired with 2,896 visualizations, each labeled as support, refute, or not enough information.To enable interpretable verification, each instance includes structured knowledge graph explanations that capture statistical patterns, temporal trends, spatial comparisons, and causal relations.We conduct a comprehensive evaluation of state-of-the-art multimodal large language models, including proprietary and open-source systems, under zero-shot and few-shot settings.Our results show that current models struggle to perform fact-checking when statistical reasoning over charts is required: even the best-performing systems, such as Gemini 2.5 and InternVL 2.5, achieve only 76.2-77.8%accuracy in label-only output settings, which is far below human performance (89.3% and 92.7%).While few-shot prompting yields limited improvements, explanationaugmented outputs significantly enhance performance in some closed-source models, notably o3 and Gemini 2.5.We released our dataset and code alongside the paper. 1(c) Subgraph of Relevant Facts Caption: Cumulative mass loss of the Greenland Ice Sheet from 1972 to 2022, showing accelerating ice loss and corresponding sea level rise. Ruiran Su, Jiasheng Si, Zhijiang Guo, Janet B. Pierrehumbert |
EMNLP | 3 |
| 2025 | UNComp: Can Matrix Entropy Uncover Sparsity? - A Compressor Design from an Uncertainty-Aware PerspectiveabstractJing Xiong, Jianghan Shen, Fanghua Ye, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Xun Wu, Chuanyang Zheng, Zhijiang Guo, Min Yang, Lingpeng Kong, Ngai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jianghan Shen, Fanghua Ye 0001, Chaofan Tao, Zhongwei Wan, Jianqiao Lu, Chuanyang Zheng, Zhijiang Guo, Min Yang 0007, Lingpeng Kong, Ngai Wong 0001 |
EMNLP | 9 |
| 2025 | FormalAlign: Automated Alignment Evaluation for AutoformalizationabstractAutoformalization aims to convert informal mathematical proofs into machine-verifiable formats, bridging the gap between natural and formal languages. However, ensuring semantic alignment between the informal and formalized statements remains challenging. Existing approaches heavily rely on manual verification, hindering scalability. To address this, we introduce FormalAlign, a framework for automatically evaluating the alignment between natural and formal languages in autoformalization. FormalAlign trains on both the autoformalization sequence generation task and the representational alignment between input and output, employing a dual loss that combines a pair of mutually enhancing autoformalization and alignment tasks. Evaluated across four benchmarks augmented by our proposed misalignment strategies, FormalAlign demonstrates superior performance. In our experiments, FormalAlign outperforms GPT-4, achieving an Alignment-Selection Score 11.58\% higher on \forml-Basic (99.21\% vs. 88.91\%) and 3.19\% higher on MiniF2F-Valid (66.39\% vs. 64.34\%). This effective alignment evaluation significantly reduces the need for manual verification. Jianqiao Lu, Yingjia Wan, Yinya Huang, Zhengying Liu, Zhijiang Guo |
ICLR | 6 |
| 2025 | OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization ModelingabstractLarge language models (LLMs) have exhibited their problem-solving abilities in mathematical reasoning. Solving realistic optimization (OPT) problems in application scenarios requires advanced and applied mathematics ability. However, current OPT benchmarks that merely solve linear programming are far from complex realistic situations. In this work, we propose **OptiBench**, a benchmark for End-to-end optimization problem-solving with human-readable inputs and outputs. **OptiBench** contains rich optimization problems, including linear and nonlinear programming with or without tabular data, which can comprehensively evaluate LLMs' solving ability. In our benchmark, LLMs are required to call a code solver to provide precise numerical answers.
Furthermore, to alleviate the data scarcity for optimization problems, and to bridge the gap between open-source LLMs on a small scale (e.g., Llama-3-8b) and closed-source LLMs (e.g., GPT-4), we further propose a data synthesis method namely ***ReSocratic***. Unlike general data synthesis methods that proceed from questions to answers, \ReSocratic first incrementally synthesizes formatted optimization demonstration with mathematical formulations step by step and then back-translates the generated demonstrations into questions. Based on this, we synthesize the ***ReSocratic-29k*** dataset. We further conduct supervised fine-tuning with ***ReSocratic-29k*** on multiple open-source models. Experimental results show that ***ReSocratic-29k*** significantly improves the performance of open-source models. Yiwei Wang 0001, Yinya Huang, Zhijiang Guo, Xiongwei Han, Liang Feng 0001, Linqi Song, Xiaodan Liang, Jing Tang 0004 |
ICLR | 4 |
| 2025 | Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model EnsemblingabstractLarge language models (LLMs) exhibit varying strengths and weaknesses across different tasks, prompting recent studies to explore the benefits of ensembling models to leverage their complementary advantages. However, existing LLM ensembling methods often overlook model compatibility and struggle with inefficient alignment of probabilities across the entire vocabulary. In this study, we empirically investigate the factors influencing ensemble performance, identifying model performance, vocabulary size, and response style as key determinants, revealing that compatibility among models is essential for effective ensembling. This analysis leads to the development of a simple yet effective model selection strategy that identifies compatible models. Additionally, we introduce the \textsc{Uni}on \textsc{T}op-$k$ \textsc{E}nsembling (\textsc{UniTE}), a novel approach that efficiently combines models by focusing on the union of the top-k tokens from each model, thereby avoiding the need for full vocabulary alignment and reducing computational overhead. Extensive evaluations across multiple benchmarks demonstrate that \textsc{UniTE} significantly enhances performance compared to existing methods, offering a more efficient framework for LLM ensembling. Han Wu 0004, Sichun Luo, Xiongwei Han, Jie Liu 0022, Zhijiang Guo, Linqi Song |
ICLR | 7 |
| 2025 | EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuningabstractAs large language models (LLMs) play an increasingly important role in code generation, enhancing both correctness and efficiency has become crucial. Current methods primarily focus on correctness, often overlooking efficiency. To address this gap, we introduce SWIFTCODE to improve both aspects by fine-tuning LLMs on a high-quality dataset comprising correct and efficient code samples. Our methodology involves leveraging multiple LLMs to generate diverse candidate code solutions for various tasks across different programming languages. We then evaluate these solutions by directly measuring their execution time and memory usage through local execution. The code solution with the lowest execution time and memory consumption is selected as the final output for each task. Experimental results demonstrate significant improvements when fine-tuning with SWIFTCODE. For instance, Qwen2.5-Coder-7B-Instruct’s pass@1 score increases from 44.8% to 57.7%, while the average execution time for correct tasks decreases by 48.4%. SWIFTCODE offers a scalable and effective solution for advancing AI-driven code generation, benefiting both software development and computational problem-solving. Dong Huang 0005, Guangtao Zeng, Jianbo Dai, Meng Luo 0010, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050 |
ICML | 8 |
| 2025 | Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language ModelsabstractLarge Language Models (LLMs) are expected to be predictable and trustworthy to support reliable decision-making systems. Yet current LLMs often show inconsistencies in their judgments. In this work, we examine \textit{logical preference consistency} as a foundational requirement for building more dependable LLM systems, ensuring stable and coherent decision-making while minimizing erratic or contradictory outputs.
To quantify the logical preference consistency, we propose a universal evaluation framework based on three fundamental properties: *transitivity*, *commutativity* and *negation invariance*.
Through extensive experimentation across diverse LLMs, we demonstrate that these properties serve as strong indicators of judgment robustness.
Furthermore, we introduce a data refinement and augmentation technique, REPAIR, that enhances logical consistency while maintaining alignment with human preferences. Finally, we show that improving consistency leads to better performance in LLM-driven logic-based algorithms, reinforcing stability and coherence in decision-making systems. Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi, Ivan Vulic, Nigel Collier |
ICML | 2 |
| 2025 | AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the WebabstractTextual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation.Existing datasets for automated verification of image-text claims remain limited, as they often consist of synthetic claims and lack evidence annotations to capture the reasoning behind the verdict.In this work, we introduce AVerImaTeC, a dataset consisting of 1,297 real-world image-text claims. Each claim is annotated with question-answer (QA) pairs containing evidence from the web, reflecting a decomposed reasoning regarding the verdict.We mitigate common challenges in fact-checking datasets such as contextual dependence, temporal leakage, and evidence insufficiency, via claim normalization, temporally constrained evidence annotation, and a two-stage sufficiency check. We assess the consistency of the annotation in AVerImaTeC via inter-annotator studies, achieving a $\kappa=0.742$ on verdicts and $74.7\%$ consistency on QA pairs. We also propose a novel evaluation method for evidence retrieval and conduct extensive experiments to establish baselines for verifying image-text claims using open-web evidence. Zifeng Ding, Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
NeurIPS | 3 |
| 2025 | EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated CodeabstractExisting code generation benchmarks primarily evaluate functional correctness, with limited attention to code efficiency, and they are often restricted to a single language such as Python. To address this gap, we introduce EffiBench‑X, the first large‑scale multi‑language benchmark specifically designed for robust efficiency evaluation of LLM‑generated code. EffiBench‑X supports Python, C++, Java, JavaScript, Ruby, and Go, and comprises competitive programming tasks paired with human‑expert solutions as efficiency baselines. Evaluating state‑of‑the‑art LLMs on EffiBench‑X reveals that while models frequently generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM‑generated solutions (e.g., Qwen3‑32B) achieve only around 62% of human efficiency on average, with significant language‑specific variation: models tend to perform better in Python, Ruby, and JavaScript than in Java, C++, and Go (e.g., DeepSeek‑R1’s Python code is markedly more efficient than its Java code). These findings highlight the need for research into optimization‑oriented methods to improve the efficiency of LLM‑generated code across diverse languages. The dataset and evaluation infrastructure are publicly available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x. Yuhao Qing, Boyu Zhu, Mingzhe Du, Zhijiang Guo, Terry Yue Zhuo, Qianru Zhang, Jie Zhang 0050, Heming Cui, Siu-Ming Yiu, Dong Huang 0005, See-Kiong Ng, Anh Tuan Luu |
NeurIPS | 4 |
| 2025 | Atom of Thoughts for Markov LLM Test-Time ScalingabstractLarge Language Models (LLMs) achieve superior performance through training-time scaling, and test-time scaling further enhances their capabilities by conducting effective reasoning during inference.
However, as the scale of reasoning increases, existing test-time scaling methods suffer from accumulated historical information, which not only wastes computational resources but also interferes with effective reasoning.
To address this issue, we observe that complex reasoning can be achieved by solving a series of independent and self-contained subquestions. These subquestions are essentially \textit{atomic questions}, exhibiting the memoryless property similar to Markov processes.
Based on this observation, we propose Atom of Thoughts (\our), where each state transition consists of decomposing the current question into a dependency-based directed acyclic graph and contracting its subquestions, forming a simplified question that maintains answer equivalence with the original problem. This answer preservation enables the iterative \textit{decomposition-contraction} process to naturally form a meaningful Markov reasoning process.
Furthermore, these atomic states can be seamlessly integrated into existing test-time scaling methods, enabling \our to serve as a plug-in enhancement for improving reasoning capabilities. Fengwei Teng, Zhaoyang Yu 0004, Jiayi Zhang 0017, Yuyu Luo, Chenglin Wu 0001, Zhijiang Guo |
NeurIPS | 7 |
| 2025 | TimE: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World ScenariosabstractTemporal reasoning is pivotal for Large Language Models (LLMs) to comprehend the real world. However, existing works neglect the real-world challenges for temporal reasoning: (1) intensive temporal information, (2) fast-changing event dynamics, and (3) complex temporal dependencies in social interactions. To bridge this gap, we propose a multi-level benchmark TimE, designed for temporal reasoning in real-world scenarios. TimE consists of 38,522 QA pairs, covering 3 levels with 11 fine-grained sub-tasks. This benchmark encompasses 3 sub-datasets reflecting different real-world challenges: TimE-Wiki, TimE-News, and TimE-Dial. We conduct extensive experiments on reasoning models and non-reasoning models. And we conducted an in-depth analysis of temporal reasoning performance across diverse real-world scenarios and tasks, and summarized the impact of test-time scaling on temporal reasoning capabilities. Additionally, we release TimE-Lite, a human-annotated subset to foster future research and standardized evaluation in temporal reasoning. Shaohang Wei, Wei Li 0101, Feifan Song 0001, Wen Luo 0001, Tianyi Zhuang, Haochen Tan, Zhijiang Guo, Houfeng Wang |
NeurIPS | 7 |
| 2025 | Activation-Guided Consensus Merging for Large Language ModelsabstractRecent research has increasingly focused on reconciling the reasoning capabilities of System 2 with the efficiency of System 1. While existing training-based and prompt-based approaches face significant challenges in terms of efficiency and stability, model merging emerges as a promising strategy to integrate the diverse capabilities of different Large Language Models (LLMs) into a unified model. However, conventional model merging methods often assume uniform importance across layers, overlooking the functional heterogeneity inherent in neural components. To address this limitation, we propose \textbf{A}ctivation-Guided \textbf{C}onsensus \textbf{M}erging (\textbf{ACM}), a plug-and-play merging framework that determines layer-specific merging coefficients based on mutual information between activations of pre-trained and fine-tuned models. ACM effectively preserves task-specific capabilities without requiring gradient computations or additional training. Extensive experiments on Long-to-Short (L2S) and general merging tasks demonstrate that ACM consistently outperforms all baseline methods. For instance, in the case of Qwen-7B models, TIES-Merging equipped with ACM achieves a \textbf{55.3\%} reduction in response length while simultaneously improving reasoning accuracy by \textbf{1.3} points. We submit the code with the paper for reproducibility, and it will be publicly available. Shuqi Liu 0001, Zehua Liu, Qintong Li, Xiongwei Han, Zhijiang Guo, Han Wu 0004, Linqi Song |
NeurIPS | 7 |
| 2025 | Analysis of longitudinal social media for monitoring symptoms during a pandemic
Shixu Lin, Lucas Garay, Yining Hua, Zhijiang Guo, Wanxin Li, Jie Yang 0039 |
J. Biomed. Informatics | 4 |
| 2024 | ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language ModelsabstractHaochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Haochen Tan, Zhijiang Guo, Zhan Shi 0001, Zhili Liu, Yunlong Feng, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Linqi Song |
ACL (1) | 2 |
| 2024 | Towards Human-aligned Evaluation for Linear Programming Word ProblemsabstractMath Word Problem (MWP) is a crucial NLP task aimed at providing solutions for given mathematical descriptions. A notable sub-category of MWP is the Linear Programming Word Problem (LPWP), which holds significant relevance in real-world decision-making and operations research. While the recent rise of generative large language models (LLMs) has brought more advanced solutions to LPWPs, existing evaluation methodologies for this task still diverge from human judgment and face challenges in recognizing mathematically equivalent answers. In this paper, we introduce a novel evaluation metric rooted in graph edit distance, featuring benefits such as permutation invariance and more accurate program equivalence identification. Human evaluations empirically validate the superior efficacy of our proposed metric when particularly assessing LLM-based solutions for LPWP. Linzi Xing, Xinglu Wang, Yuxi Feng, Zhenan Fan, Zhijiang Guo, Xiaojin Fu, Rindranirina Ramamonjison, Mahdi Mostajabdaveh, Xiongwei Han, Zirui Zhou, Yong Zhang 0004 |
LREC/COLING | 6 |
| 2024 | DVD: Dynamic Contrastive Decoding for Knowledge Amplification in Multi-Document Question AnsweringabstractLarge language models (LLMs) are widely used in question-answering (QA) systems but often generate information with hallucinations.Retrieval-augmented generation (RAG) offers a potential remedy, yet the uneven retrieval quality and irrelevant contents may distract LLMs.In this work, we address these issues at the generation phase by treating RAG as a multi-document QA task.We propose a novel decoding strategy, Dynamic Contrastive Decoding (DVD), which dynamically amplifies knowledge from selected documents during the generation phase.DVD involves constructing inputs batchwise, designing new selection criteria to identify documents worth amplifying, and applying contrastive decoding with a specialized weight calculation to adjust the final logits used for sampling answer tokens.Zero-shot experimental results on ALCE-ASQA, NQ, TQA and PopQA benchmarks show that our method outperforms other decoding strategies.Additionally, we conduct experiments to validate the effectiveness of our selection criteria, weight calculation, and general multi-document scenarios.Our method requires no training and can be integrated with other methods to improve the RAG performance. Houfeng Wang, Zhijiang Guo |
EMNLP | 5 |
| 2024 | Knowledge Conflicts for LLMs: A SurveyabstractThis survey provides an in-depth analysis of knowledge conflicts for large language models (LLMs), highlighting the complex challenges they encounter when blending contextual and parametric knowledge.Our focus is on three categories of knowledge conflicts: contextmemory, inter-context, and intra-memory conflict.These conflicts can significantly impact the trustworthiness and performance of LLMs, especially in real-world applications where noise and misinformation are common.By categorizing these conflicts, exploring the causes, examining the behaviors of LLMs under such conflicts, and reviewing available solutions, this survey aims to shed light on strategies for improving the robustness of LLMs, thereby serving as a valuable resource for advancing research in this evolving area. Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang 0003, Yue Zhang 0004, Wei Xu 0039 |
EMNLP | 3 |
| 2024 | Do We Need Language-Specific Fact-Checking Models? The Case of ChineseabstractThis paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese using CHEF dataset.To better reflect real-world fact-checking, we first develop a novel Chinese document-level evidence retriever, achieving state-of-the-art performance.We then demonstrate the limitations of translation-based methods and multilingual language models, highlighting the need for language-specific systems.To better analyze token-level biases in different systems, we construct an adversarial dataset based on the CHEF dataset, where each instance has a large word overlap with the original one but holds the opposite veracity label.Experimental results on the CHEF dataset and our adversarial dataset show that our proposed method outperforms translation-based methods and multilingual language models and is more robust toward biases, emphasizing the importance of language-specific fact-checking systems. 1 Verifiers Retrievers Caiqi Zhang, Zhijiang Guo, Andreas Vlachos 0001 |
EMNLP | 2 |
| 2024 | Towards Understanding Factual Knowledge of Large Language ModelsabstractLarge language models (LLMs) have recently driven striking performance improvements across a range of natural language processing tasks. The factual knowledge acquired during pretraining and instruction tuning can be useful in various downstream tasks, such as question answering, and language generation. Unlike conventional Knowledge Bases (KBs) that explicitly store factual knowledge, LLMs implicitly store facts in their parameters. Content generated by the LLMs can often exhibit inaccuracies or deviations from the truth, due to facts that can be incorrectly induced or become obsolete over time. To this end, we aim to explore the extent and scope of factual knowledge within LLMs by designing the benchmark Pinocchio. Pinocchio contains 20K diverse factual questions that span different sources, timelines, domains, regions, and languages. Furthermore, we investigate whether LLMs can compose multiple facts, update factual knowledge temporally, reason over multiple pieces of facts, identify subtle factual differences, and resist adversarial examples. Extensive experiments on different sizes and types of LLMs show that existing LLMs still lack factual knowledge and suffer from various spurious correlations. We believe this is a critical bottleneck for realizing trustworthy artificial intelligence. The dataset Pinocchio and our codes are publicly available at: https://github.com/THU-BPM/Pinocchio. Xuming Hu, Junzhe Chen 0001, Xiaochuan Li 0003, Yufei Guo 0001, Lijie Wen 0001, Philip S. Yu, Zhijiang Guo |
ICLR | 7 |
| 2024 | DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context LearningabstractRecent advances in natural language processing, primarily propelled by Large Language Models (LLMs), have showcased their remarkable capabilities grounded in in-context learning. A promising avenue for guiding LLMs in intricate reasoning tasks involves the utilization of intermediate reasoning steps within the Chain-of-Thought (CoT) paradigm. Nevertheless, the central challenge lies in the effective selection of exemplars for facilitating in-context learning. In this study, we introduce a framework that leverages Dual Queries and Low-rank approximation Re-ranking (DQ-LoRe) to automatically select exemplars for in-context learning. Dual Queries first query LLM to obtain LLM-generated knowledge such as CoT, then query the retriever to obtain the final exemplars via both question and the knowledge. Moreover, for the second query, LoRe employs dimensionality reduction techniques to refine exemplar selection, ensuring close alignment with the input question's knowledge. Through extensive experiments, we demonstrate that DQ-LoRe significantly outperforms prior state-of-the-art methods in the automatic selection of exemplars for GPT-4, enhancing performance from 92.5\% to 94.2\%. Our comprehensive analysis further reveals that DQ-LoRe consistently outperforms retrieval-based approaches in terms of both performance and adaptability, especially in scenarios characterized by distribution shifts. DQ-LoRe pushes the boundaries of in-context learning and opens up new avenues for addressing complex reasoning challenges. Chuanyang Zheng, Zhijiang Guo, Yichun Yin, Enze Xie, Qingxing Cao, Xiongwei Han, Jing Tang 0004, Xiaodan Liang |
ICLR | 4 |
| 2024 | EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationabstractLarge language models (LLMs) have shown remarkable progress in code generation, but their generated code often suffers from inefficiency, resulting in longer execution times and higher memory consumption. To address this issue, we propose EffiLearner, a self-optimization framework that utilizes execution overhead profiles to improve the efficiency of LLM-generated code. EffiLearner first generates code using an LLM, then executes it locally to capture execution time and memory usage profiles. These profiles are fed back to the LLM, which then revises the code to reduce overhead. To evaluate the effectiveness of EffiLearner, we conduct extensive experiments on EffiBench and two commonly used code generation benchmarks with 16 open-source and 6 closed-source models. Our evaluation results demonstrate that through iterative self-optimization, EffiLearner significantly enhances the efficiency of LLM-generated code. For example, the execution time (ET) of StarCoder2-15B for the EffiBench decreases from 0.93 (s) to 0.12 (s) which reduces 87.1\% execution time requirement compared with the initial code. The total memory usage (TMU) of StarCoder2-15B also decreases from 22.02 (Mb*s) to 2.03 (Mb*s), which decreases 90.8\% total memory consumption during the execution process. Dong Huang 0005, Jianbo Dai, Han Weng, Puzhen Wu, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050 |
NeurIPS | 7 |
| 2024 | AutoPSV: Automated Process-Supervised VerifierabstractIn this work, we propose a novel method named \textbf{Auto}mated \textbf{P}rocess-\textbf{S}upervised \textbf{V}erifier (\textbf{\textsc{AutoPSV}}) to enhance the reasoning capabilities of large language models (LLMs) by automatically annotating the reasoning steps.
\textsc{AutoPSV} begins by training a verification model on the correctness of final answers, enabling it to generate automatic process annotations.
This verification model assigns a confidence score to each reasoning step, indicating the probability of arriving at the correct final answer from that point onward.
We detect relative changes in the verification's confidence scores across reasoning steps to automatically annotate the reasoning process, enabling error detection even in scenarios where ground truth answers are unavailable.
This alleviates the need for numerous manual annotations or the high computational costs associated with model-induced annotation approaches.
We experimentally validate that the step-level confidence changes learned by the verification model trained on the final answer correctness can effectively identify errors in the reasoning steps.
We demonstrate that the verification model, when trained on process annotations generated by \textsc{AutoPSV}, exhibits improved performance in selecting correct answers from multiple LLM-generated outputs.
Notably, we achieve substantial improvements across five datasets in mathematics and commonsense reasoning. The source code of \textsc{AutoPSV} is available at \url{https://github.com/rookie-joe/AutoPSV}. Jianqiao Lu, Zhiyang Dou, Hongru Wang 0003, Zeyu Cao, Jianbo Dai, Yunlong Feng, Zhijiang Guo |
NeurIPS | 7 |
| 2024 | HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-TuningabstractAdapting Large Language Models (LLMs) to new tasks through fine-tuning has been made more efficient by the introduction of Parameter-Efficient Fine-Tuning (PEFT) techniques, such as LoRA. However, these methods often underperform compared to full fine-tuning, particularly in scenarios involving complex datasets. This issue becomes even more pronounced in complex domains, highlighting the need for improved PEFT approaches that can achieve better performance. Through a series of experiments, we have uncovered two critical insights that shed light on the training and parameter inefficiency of LoRA. Building on these insights, we have developed HydraLoRA, a LoRA framework with an asymmetric structure that eliminates the need for domain expertise. Our experiments demonstrate that HydraLoRA outperforms other PEFT approaches, even those that rely on domain knowledge during the training and inference phases. Our anonymous codes are submitted with the paper and will be publicly available. Code is available: https://github.com/Clin0212/HydraLoRA. Chunlin Tian, Zhijiang Guo, Li Li 0064, Cheng-Zhong Xu 0001 |
NeurIPS | 3 |
| 2024 | MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsabstractLarge language models (LLMs) have shown increasing capability in problem-solving and decision-making, largely based on the step-by-step chain-of-thought reasoning processes. However, evaluating these reasoning abilities has become increasingly challenging. Existing outcome-based benchmarks are beginning to saturate, becoming less effective in tracking meaningful progress. To address this, we present a process-based benchmark MR-Ben that demands a meta-reasoning skill, where LMs are asked to locate and analyse potential errors in automatically generated reasoning steps. Our meta-reasoning paradigm is especially suited for system-2 slow thinking, mirroring the human cognitive process of carefully examining assumptions, conditions, calculations, and logic to identify mistakes. MR-Ben comprises 5,975 questions curated by human experts across a wide range of subjects, including physics, chemistry, logic, coding, and more. Through our designed metrics for assessing meta-reasoning on this benchmark, we identify interesting limitations and weaknesses of current LLMs (open-source and closed-source models). For example, with models like the o1 series from OpenAI demonstrating strong performance by effectively scrutinizing the solution space, many other state-of-the-art models fall significantly behind on MR-Ben, exposing potential shortcomings in their training strategies and inference methodologies. Zhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li 0001, Pengguang Chen, Jianbo Dai, Rongwu Xu, Zehan Qi, Wanru Zhao, Linling Shen, Jianqiao Lu, Haochen Tan, Yukang Chen, Bailin Wang, Zhijiang Guo, Jiaya Jia |
NeurIPS | 18 |
| 2023 | TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language ModelsabstractJing Xiong, Jianhao Shen, Ye Yuan, Haiming Wang, Yichun Yin, Zhengying Liu, Lin Li, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang, Qun Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jianhao Shen, Ye Yuan 0016, Yichun Yin, Zhengying Liu, Zhijiang Guo, Qingxing Cao, Yinya Huang, Chuanyang Zheng, Xiaodan Liang, Ming Zhang 0004, Qun Liu 0001 |
EMNLP | 8 |
| 2023 | AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the WebabstractExisting datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the claim. In this paper we introduce AVeriTeC, a new dataset of 4,568 real-world claims covering fact-checks by 50 different organizations. Each claim is annotated with question-answer pairs supported by evidence available online, as well as textual justifications explaining how the evidence combines to produce a verdict. Through a multi-round annotation process, we avoid common pitfalls including context dependence, evidence insufficiency, and temporal leakage, and reach a substantial inter-annotator agreement of $\kappa=0.619$ on verdicts. We develop a baseline as well as an evaluation scheme for verifying claims through question-answering against the open web. Michael Sejr Schlichtkrull, Zhijiang Guo, Andreas Vlachos 0001 |
NeurIPS | 2 |
| 2023 | MR2: A Benchmark for Multimodal Retrieval-Augmented Rumor Detection in Social MediaabstractAs social media platforms are evolving from text-based forums into multi-modal environments, the nature of misinformation in social media is also transforming accordingly. Misinformation spreaders have recently targeted contextual connections between the modalities e.g., text and image. However, existing datasets for rumor detection mainly focus on a single modality i.e., text. To bridge this gap, we construct MR2, a multimodal multilingual retrieval-augmented dataset for rumor detection. The dataset covers rumors with images and texts, and provides evidence from both modalities that are retrieved from the Internet. Further, we develop established baselines and conduct a detailed analysis of the systems evaluated on the dataset. Extensive experiments show that MR2 will provide a challenging testbed for developing rumor detection systems designed to retrieve and reason over social media posts. Source code and data are available at: https://github.com/THU-BPM/MR2. Xuming Hu, Zhijiang Guo, Junzhe Chen 0001, Lijie Wen 0001, Philip S. Yu |
SIGIR | 2 |
| 2023 | Read it Twice: Towards Faithfully Interpretable Fact Verification by Revisiting EvidenceabstractReal-world fact verification task aims to verify the factuality of a claim by retrieving evidence from the source document. The quality of the retrieved evidence plays an important role in claim verification. Ideally, the retrieved evidence should be faithful (reflecting the model's decision-making process in claim verification) and plausible (convincing to humans), and can improve the accuracy of verification task. Although existing approaches leverage the similarity measure of semantic or surface form between claims and documents to retrieve evidence, they all rely on certain heuristics that prevent them from satisfying all three requirements. In light of this, we propose a fact verification model named ReRead to retrieve evidence and verify claim that: (1) Train the evidence retriever to obtain interpretable evidence (i.e., faithfulness and plausibility criteria); (2) Train the claim verifier to revisit the evidence retrieved by the optimized evidence retriever to improve the accuracy. The proposed system is able to achieve significant improvements upon best-reported models under different settings. Xuming Hu, Zhaochen Hong, Zhijiang Guo, Lijie Wen 0001, Philip S. Yu |
SIGIR | 3 |
| 2022 | Scene Graph Modification as Incremental Structure ExpandingabstractA scene graph is a semantic representation that expresses the objects, attributes, and relationships between objects in a scene. Scene graphs play an important role in many cross modality tasks, as they are able to capture the interactions between images and texts. In this paper, we focus on scene graph modification (SGM), where the system is required to learn how to update an existing scene graph based on a natural language query. Unlike previous approaches that rebuilt the entire scene graph, we frame SGM as a graph expansion task by introducing the incremental structure expanding (ISE). ISE constructs the target graph by incrementally expanding the source graph without changing the unmodified structure. Based on ISE, we further propose a model that iterates between nodes prediction and edges prediction, inferring more accurate and harmonious expansion decisions progressively. In addition, we construct a challenging dataset that contains more complicated queries and larger scene graphs than existing datasets. Experiments on four benchmarks demonstrate the effectiveness of our approach, which surpasses the previous state-of-the-art model by large margins. Xuming Hu, Zhijiang Guo, Lijie Wen 0001, Philip S. Yu |
COLING | 2 |
| 2022 | CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-CheckingabstractXuming Hu, Zhijiang Guo, GuanYu Wu, Aiwei Liu, Lijie Wen, Philip Yu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xuming Hu, Zhijiang Guo, Guanyu Wu, Aiwei Liu, Lijie Wen 0001, Philip S. Yu |
NAACL-HLT | 2 |
| 2022 | METS-CoV: A Dataset of Medical Entity and Targeted Sentiment on COVID-19 Related TweetsabstractThe COVID-19 pandemic continues to bring up various topics discussed or debated on social media. In order to explore the impact of pandemics on people's lives, it is crucial to understand the public's concerns and attitudes towards pandemic-related entities (e.g., drugs, vaccines) on social media. However, models trained on existing named entity recognition (NER) or targeted sentiment analysis (TSA) datasets have limited ability to understand COVID-19-related social media texts because these datasets are not designed or annotated from a medical perspective. In this paper, we release METS-CoV, a dataset containing medical entities and targeted sentiments from COVID-19 related tweets. METS-CoV contains 10,000 tweets with 7 types of entities, including 4 medical entity types (Disease, Drug, Symptom, and Vaccine) and 3 general entity types (Person, Location, and Organization). To further investigate tweet users' attitudes toward specific entities, 4 types of entities (Person, Organization, Drug, and Vaccine) are selected and annotated with user sentiments, resulting in a targeted sentiment dataset with 9,101 entities (in 5,278 tweets). To the best of our knowledge, METS-CoV is the first dataset to collect medical entities and corresponding sentiments of COVID-19 related tweets. We benchmark the performance of classical machine learning models and state-of-the-art deep learning models on NER and TSA tasks with extensive experiments. Results show that this dataset has vast room for improvement for both NER and TSA tasks. With rich annotations and comprehensive benchmark results, we believe METS-CoV is a fundamental resource for building better medical social media understanding tools and facilitating computational social science research, especially on epidemiological topics. Our data, annotation guidelines, benchmark models, and source code are publicly available (\url{https://github.com/YLab-Open/METS-CoV}) to ensure reproducibility. Peilin Zhou, Zeqiang Wang, Dading Chong, Zhijiang Guo, Yining Hua, Zichang Su, Zhiyang Teng, Jiageng Wu, Jie Yang 0039 |
NeurIPS | 4 |
| 2022 | A Survey on Automated Fact-CheckingabstractAbstract Fact-checking has become increasingly important due to the speed with which both information and misinformation can spread in the modern media ecosystem. Therefore, researchers have been exploring how fact-checking can be automated, using techniques based on natural language processing, machine learning, knowledge representation, and databases to automatically predict the veracity of claims. In this paper, we survey automated fact-checking stemming from natural language processing, and discuss its connections to related tasks and disciplines. In this process, we present an overview of existing datasets and models, aiming to unify the various definitions given and identify common concepts. Finally, we highlight challenges for future research. Zhijiang Guo, Michael Sejr Schlichtkrull, Andreas Vlachos 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Uncovering Main Causalities for Long-tailed Information ExtractionabstractInformation Extraction (IE) aims to extract structural information from unstructured texts.In practice, long-tailed distributions caused by the selection bias of a dataset, may lead to incorrect correlations, also known as spurious correlations, between entities and labels in the conventional likelihood models.This motivates us to propose counterfactual IE (CFIE), a novel framework that aims to uncover the main causalities behind data in the view of causal inference.Specifically, 1) we first introduce a unified structural causal model (SCM) for various IE tasks, describing the relationships among variables; 2) with our SCM, we then generate counterfactuals based on an explicit language structure to better calculate the direct causal effect during the inference stage; 3) we further propose a novel debiasing approach to yield more robust predictions.Experiments on three IE tasks across five public datasets show the effectiveness of our CFIE model in mitigating the spurious correlation issues. Guoshun Nan, Jiaqi Zeng, Rui Qiao 0006, Zhijiang Guo, Wei Lu 0011 |
EMNLP (1) | 4 |
| 2020 | Reasoning with Latent Structure Refinement for Document-Level Relation ExtractionabstractDocument-level relation extraction requires integrating information within and across multiple sentences of a document and capturing complex interactions between inter-sentence entities.However, effective aggregation of relevant information in the document remains a challenging research question.Existing approaches construct static document-level graphs based on syntactic trees, co-references or heuristics from the unstructured text to model the dependencies.Unlike previous methods that may not be able to capture rich non-local interactions for inference, we propose a novel model that empowers the relational reasoning across sentences by automatically inducing the latent document-level graph.We further develop a refinement strategy, which enables the model to incrementally aggregate relevant information for multi-hop reasoning.Specifically, our model achieves an F 1 score of 59.05 on a large-scale documentlevel dataset (DocRED), significantly improving over the previous results, and also yields new state-of-the-art results on the CDR and GDA dataset.Furthermore, extensive analyses show that the model is able to discover more accurate inter-sentence relations. Guoshun Nan, Zhijiang Guo, Ivan Sekulic, Wei Lu 0011 |
ACL | 2 |
| 2020 | Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text GenerationabstractAMR-to-text generation is used to transduce Abstract Meaning Representation structures (AMR) into text.A key challenge in this task is to efficiently learn effective graph representations.Previously, Graph Convolution Networks (GCNs) were used to encode input AMRs, however, vanilla GCNs are not able to capture non-local information and additionally, they follow a local (first-order) information aggregation scheme.To account for these issues, larger and deeper GCN models are required to capture more complex interactions.In this paper, we introduce a dynamic fusion mechanism, proposing Lightweight Dynamic Graph Convolutional Networks (LDGCNs) that capture richer non-local interactions by synthesizing higher order information from the input graphs.We further develop two novel parameter saving strategies based on the group graph convolutions and weight tied convolutions to reduce memory usage and model complexity.With the help of these strategies, we are able to train a model with fewer parameters while maintaining the model capacity.Experiments demonstrate that LDGCNs outperform stateof-the-art models on two benchmark datasets for AMR-to-text generation with significantly fewer parameters. Yan Zhang 0004, Zhijiang Guo, Zhiyang Teng, Wei Lu 0011, Shay B. Cohen, Zuozhu Liu, Lidong Bing |
EMNLP (1) | 2 |
| 2020 | Learning Latent Forests for Medical Relation ExtractionabstractThe goal of medical relation extraction is to detect relations among entities, such as genes, mutations and drugs in medical texts. Dependency tree structures have been proven useful for this task. Existing approaches to such relation extraction leverage off-the-shelf dependency parsers to obtain a syntactic tree or forest for the text. However, for the medical domain, low parsing accuracy may lead to error propagation downstream the relation extraction pipeline. In this work, we propose a novel model which treats the dependency structure as a latent variable and induces it from the unstructured text in an end-to-end fashion. Our model can be understood as composing task-specific dependency forests that capture non-local interactions for better relation extraction. Extensive results on four datasets show that our model is able to significantly outperform state-of-the-art systems without relying on any direct tree supervision or pre-training. Zhijiang Guo, Guoshun Nan, Wei Lu 0011, Shay B. Cohen |
IJCAI | 1 |
| 2019 | Attention Guided Graph Convolutional Networks for Relation ExtractionabstractDependency trees convey rich structural information that is proven useful for extracting relations among entities in text.However, how to effectively make use of relevant information while ignoring irrelevant information from the dependency trees remains a challenging research question.Existing approaches employing rule based hard-pruning strategies for selecting relevant partial dependency structures may not always yield optimal results.In this work, we propose Attention Guided Graph Convolutional Networks (AGGCNs), a novel model which directly takes full dependency trees as inputs.Our model can be understood as a soft-pruning approach that automatically learns how to selectively attend to the relevant sub-structures useful for the relation extraction task.Extensive results on various tasks including cross-sentence n-ary relation extraction and large-scale sentence-level relation extraction show that our model is able to better leverage the structural information of the full dependency trees, giving significantly better results than previous approaches. Zhijiang Guo, Yan Zhang 0004, Wei Lu 0011 |
ACL (1) | 1 |
| 2019 | Densely Connected Graph Convolutional Networks for Graph-to-Sequence LearningabstractWe focus on graph-to-sequence learning, which can be framed as transducing graph structures to sequences for text generation. To capture structural information associated with graphs, we investigate the problem of encoding graphs using graph convolutional networks (GCNs). Unlike various existing approaches where shallow architectures were used for capturing local structural information only, we introduce a dense connection strategy, proposing a novel Densely Connected Graph Convolutional Network (DCGCN). Such a deep architecture is able to integrate both local and non-local features to learn a better structural representation of a graph. Our model outperforms the state-of-the-art neural models significantly on AMR-to-text generation and syntax-based neural machine translation. Zhijiang Guo, Yan Zhang 0004, Zhiyang Teng, Wei Lu 0011 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | Better Transition-Based AMR Parsing with Refined Search SpaceabstractThis paper introduces a simple yet effective transition-based system for Abstract Meaning Representation (AMR) parsing.We argue that a well-defined search space for a transition system is crucial for building an effective parser.We propose to conduct the search in a refined search space based on a new compact AMR graph and an improved oracle.Our end-to-end parser achieves the state-of-the-art performance on various datasets with minimal additional information.1 Zhijiang Guo, Wei Lu 0011 |
EMNLP | 1 |
| 2006 | Realization of Micromanipulating Gough-Stewart Platforms with Desired DynamicsabstractA new design method for optimizing Gough-Stewart platform geometries is proposed in this paper. The prior studies are extended through optimizing both the kinematics and dynamics. Firstly, a new class of weighted orthogonal Gough-Stewart platforms (w-OGSPs) is proposed. Then the concept of simultaneously diagonal OGSPs (sd-OGSPs) is defined. The sd-OGSPs allow the Cartesian second order impedance to be specified, thus obtaining desired local system poles. Design algorithms and simulation results are given Zhijiang Guo, John E. McInroy, Farhad Jafari |
ICRA | 1 |