Liangming Pan

dblp:186/9707 · DBLP profile ↗
← Back
51ranked-venue papers
11as first author
41since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 9 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures
abstract
While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and safe deployment.Existing surveys in explainable AI largely focus on post-hoc explanation methods that interpret trained models through external approximations.In contrast, intrinsic interpretability, which builds transparency directly into model architectures and computations, has recently emerged as a promising alternative.This paper presents a systematic review of the recent advances in intrinsic interpretability for LLMs, categorizing existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction.We further discuss open challenges and outline future research directions in this emerging field.The paper list is available at: Survey-Intrinsic-Interpretability-of-LLMs
Liangming Pan
ACL (1)4
2026 Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures
abstract
Yi Hu, Jiaqi Gu, Ruxin Wang, Zijun Yao, Hao Peng, Xiaobao Wu, Jianhui Chen, Muhan Zhang, Liangming Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zijun Yao 0002, Hao Peng 0015, Xiaobao Wu, Muhan Zhang, Liangming Pan
ACL (1)9
2026 LEDOM: Reverse Language Model
abstract
Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu 0001, Li Lin 0014, Xinyi Wang 0003, Liangming Pan, William Yang Wang, Xiaojun Wan 0001
ACL (1)7
2025 Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
abstract
Recent advancements in multimodal large language models (MLLMs) have shown unprecedented capabilities in advancing various vision-language tasks. However, MLLMs face significant challenges with hallucinations, and misleading outputs that do not align with the input data. While existing efforts are paid to combat MLLM hallucinations, several pivotal challenges are still unsolved. First, while current approaches aggressively focus on addressing errors at the perception level, another important type at the cognition level requiring factual commonsense can be overlooked. In addition, existing methods might fall short in finding a more effective way to represent visual input, which is yet a key bottleneck that triggers visual hallucinations. Moreover, MLLMs can frequently be misled by faulty textual inputs and cause hallucinations, while unfortunately, this type of issue has long been overlooked by existing studies. Inspired by human intuition in handling hallucinations, this paper introduces a novel bottom-up reasoning framework. Our framework systematically addresses potential issues in both visual and textual inputs by verifying and integrating perception-level information with cognition-level commonsense knowledge, ensuring more reliable outputs. Extensive experiments demonstrate significant improvements in multiple hallucination benchmarks after integrating MLLMs with the proposed framework. In-depth analyses reveal the great potential of our methods in addressing perception- and cognition-level hallucinations.
Shengqiong Wu, Hao Fei 0001, Liangming Pan, William Yang Wang, Shuicheng Yan, Tat-Seng Chua
AAAI3
2025 SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation
abstract
Zijun Yao, Weijian Qi, Liangming Pan, Shulin Cao, Linmei Hu, Liu Weichuan, Lei Hou, Juanzi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zijun Yao 0002, Weijian Qi, Liangming Pan, Shulin Cao, Linmei Hu, Weichuan Liu, Lei Hou 0001, Juan-Zi Li
ACL (1)3
2025 InductionBench: LLMs Fail in the Simplest Complexity Class
abstract
Large language models (LLMs) have shown remarkable improvements in reasoning and many existing benchmarks have been addressed by models such as o1 and o3 either fully or partially.However, a majority of these benchmarks emphasize deductive reasoning, including mathematical and coding tasks in which rules such as mathematical axioms or programming syntax are clearly defined, based on which LLMs can plan and apply these rules to arrive at a solution.In contrast, inductive reasoning, where one infers the underlying rules from observed data, remains less explored.Such inductive processes lie at the heart of scientific discovery, as they enable researchers to extract general principles from empirical observations.To assess whether LLMs possess this capacity, we introduce InductionBench, a new benchmark designed to evaluate the inductive reasoning ability of LLMs.Our experimental findings reveal that even the most advanced modelw available struggle to master the simplest complexity classes within the subregular hierarchy of functions, highlighting a notable deficiency in current LLMs' inductive reasoning capabilities.Coda and data are available https://github.com/wenyueh/ inductive_reasoning_benchmark.
Wenyue Hua, Tyler Wong, Fei Sun 0001, Liangming Pan, Adam Jardine, William Yang Wang
ACL (1)4
2025 AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
abstract
Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao, Yubo Ma, Mingzhe Du, Rui Mao, Anh Tuan Luu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao 0007, Yubo Ma, Mingzhe Du, Rui Mao 0010, Anh Tuan Luu, William Yang Wang
ACL (1)2
2025 Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework
abstract
Jundong Xu, Hao Fei, Meng Luo, Qian Liu, Liangming Pan, William Yang Wang, Preslav Nakov, Mong-Li Lee, Wynne Hsu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jundong Xu, Hao Fei 0001, Meng Luo 0010, Qian Liu 0012, Liangming Pan, William Yang Wang, Preslav Nakov, Mong-Li Lee, Wynne Hsu
ACL (1)5
2025 Gödel Agent: A Self-Referential Agent Framework for Recursively Self-Improvement
abstract
The rapid advancement of large language models (LLMs) has significantly enhanced the capabilities of agents across various tasks.However, existing agentic systems, whether based on fixed pipeline algorithms or pre-defined meta-learning frameworks, cannot search the whole agent design space due to the restriction of human-designed components, and thus might miss the more optimal agent design.In this paper, we introduce Gödel Agent, a selfevolving framework inspired by the Gödel machine, enabling agents to recursively improve themselves without relying on predefined routines or fixed optimization algorithms.Gödel Agent leverages LLMs to dynamically modify its own logic and behavior, guided solely by high-level objectives through prompting.Experimental results on multiple domains demonstrate that implementation of Gödel Agent can achieve continuous self-improvement, surpassing manually crafted agents in performance, efficiency, and generalizability.
Xunjian Yin, Xinyi Wang 0003, Liangming Pan, Li Lin 0014, Xiaojun Wan 0001, William Yang Wang
ACL (1)3
2025 RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
abstract
Ruiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng, Xiaobao Wu, En Yu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruiwen Zhou, Wenyue Hua, Liangming Pan, Sitao Cheng, Xiaobao Wu, En Yu, William Yang Wang
ACL (1)3
2025 How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
abstract
We introduce Grade School Math with Distracting Context (GSM-DC 1 ), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC).GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation.Our experiments demonstrate that LLMs are significantly sensitive to IC, affecting both reasoning path selection and arithmetic accuracy.Additionally, training models with strong distractors improves performance in both in-distribution and out-of-distribution scenarios.We further propose a stepwise tree search guided by a process reward model, which notably enhances robustness in out-of-distribution conditions.
Minglai Yang 0002, Ethan Huang, Mihai Surdeanu, William Yang Wang, Liangming Pan
EMNLP6
2025 Towards Temporal-Aware Multi-Modal Retrieval Augemented Generation in Finance
abstract
Finance decision-making often relies on in-depth data analysis across various data sources, including financial tables, news articles, stock prices, etc. In this work, we introduce FINTMMBench, the first comprehensive benchmark for evaluating temporal-aware multi-modal Retrieval-Augmented Generation (RAG) systems in finance. Built from heterologous data of NASDAQ 100 companies, FINTMMBench offers three significant advantages. 1) Multi-modal Corpus: It encompasses a hybrid of financial tables, news articles, daily stock prices, and visual technical charts as the corpus. 2) Temporal-aware Questions: Each question requires the retrieval and interpretation of its relevant data over a specific time period, including daily, weekly, monthly, quarterly, and annual periods. 3) Diverse Financial Analysis Tasks: The questions involve 10 different financial analysis tasks designed by domain experts, including information extraction, trend analysis, sentiment analysis and event detection, etc. We further propose a novel TMMHybridRAG method, which first leverages a multi-modal LLM to convert data from other modalities (e.g., tabular, visual and time-series data) into textual format and then incorporates temporal information in each node when constructing graphs and dense indexes. Its effectiveness has been validated in extensive experiments, but notable gaps remain, highlighting the challenges presented by our FINTMMBench. The benchmark and source code will be made publicly available.
Fengbin Zhu, Liangming Pan, Wenjie Wang 0007, Fuli Feng, Chao Wang 0049, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia3
2025 CausalEval: Towards Better Causal Reasoning in Language Models
abstract
Longxuan Yu, Delin Chen, Siheng Xiong, Qingyang Wu, Dawei Li, Zhikai Chen, Xiaoze Liu, Liangming Pan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Longxuan Yu, Delin Chen, Siheng Xiong, Qingyang Wu, Xiaoze Liu, Liangming Pan
NAACL (Long Papers)8
2025 MuSLR: Multimodal Symbolic Logical Reasoning
abstract
Multimodal symbolic logical reasoning, which aims to deduce new facts from multimodal input via formal logic, is critical in high-stakes applications such as autonomous driving and medical diagnosis, as its rigorous, deterministic reasoning helps prevent serious consequences. To evaluate such capabilities of current state-of-the-art vision language models (VLMs), we introduce the first benchmark MuSLR for multimodal symbolic logical reasoning grounded in formal logical rules. MuSLR comprises 1,093 instances across 7 domains, including 35 atomic symbolic logic and 976 logical combinations, with reasoning depths ranging from 2 to 9. We evaluate 7 state-of-the-art VLMs on MuSLR and find that they all struggle with multimodal symbolic reasoning, with the best model, GPT-4.1, achieving only 46.8%. Thus, we propose LogiCAM, a modular framework that applies formal logical rules to multimodal inputs, boosting GPT-4.1’s Chain-of-Thought performance by 14.13%, and delivering even larger gains on complex logics such as first-order logic. We also conduct a comprehensive error analysis, showing that around 70% of failures stem from logical misalignment between modalities, offering key insights to guide future improvements.
Jundong Xu, Hao Fei 0001, Liangming Pan, Qijun Huang, Qian Liu 0012, Preslav Nakov, Min-Yen Kan, William Yang Wang, Mong-Li Lee, Wynne Hsu
NeurIPS4
2025 How do Transformers Learn Implicit Reasoning?
abstract
Recent work suggests that large language models (LLMs) can perform multi-hop reasoning implicitly---producing correct answers without explicitly verbalizing intermediate steps---but the underlying mechanisms remain poorly understood. In this paper, we study how such implicit reasoning emerges by training transformers from scratch in a controlled symbolic environment. Our analysis reveals a three-stage developmental trajectory: early memorization, followed by in-distribution generalization, and eventually cross-distribution generalization. We find that training with atomic triples is not necessary but accelerates learning, and that second-hop generalization relies on query-level exposure to specific compositional structures. To interpret these behaviors, we introduce two diagnostic tools: cross-query semantic patching, which identifies semantically reusable intermediate representations, and a cosine-based representational lens, which reveals that successful reasoning correlates with the cosine-base clustering in hidden space. This clustering phenomenon in turn provides a coherent explanation for the behavioral dynamics observed across training, linking representational structure to reasoning capability. These findings provide new insights into the interpretability of implicit multi-hop reasoning in LLMs, helping to clarify how complex reasoning processes unfold internally and offering pathways to enhance the transparency of such models.
Jiaran Ye, Zijun Yao 0002, Zhidian Huang, Liangming Pan, Jinxin Liu 0002, Yushi Bai, Amy Xin, Weichuan Liu, Xiaoyin Che, Lei Hou 0001, Juan-Zi Li
NeurIPS4
2025 Long Context vs. RAG: Strategies for Processing Long Documents in LLMs
abstract
Large Language Models (LLMs) excel at zero- and few-shot learning but are restricted by the length of context windows when processing long documents. Two strategies have emerged to overcome this limitation: (1) Long Context (LC) methods, which extend or compress transformer architectures to input more text; and (2) Retrieval-Augmented Generation (RAG), which integrates external knowledge sources via embedding- or index-based retrieval. This half-day tutorial offers a unified, beginner-friendly introduction to both approaches. We first review transformer fundamentals-positional encoding, attention complexity, and common LC techniques. Next, we explain the classic RAG pipeline and recent RAG strategies, alongside evaluation metrics and benchmarks. We also analyze recent empirical studies to highlight strengths, limitations, and trade-offs of LC vs. RAG in terms of scalability, computational cost, and retrieval effectiveness. We conclude with best practices for real-world deployments, emerging hybrid architectures, and open research directions, equipping IR researchers and practitioners with actionable guidelines for processing long documents in LLMs.
Xinze Li 0001, Yushi Bai, Bowen Jin, Fengbin Zhu, Liangming Pan, Yixin Cao 0002
SIGIR5
2024 Faithful Logical Reasoning via Symbolic Chain-of-Thought
abstract
While the recent Chain-of-Thought (CoT) technique enhances the reasoning ability of large language models (LLMs) with the theory of mind, it might still struggle in handling logical reasoning that relies much on symbolic expressions and rigid deducing rules.To strengthen the logical reasoning capability of LLMs, we propose a novel Symbolic Chain-of-Thought, namely SymbCoT, a fully LLM-based framework that integrates symbolic expressions and logic rules with CoT prompting.Technically, building upon an LLM, SymbCoT 1) first translates the natural language context into the symbolic format, and then 2) derives a step-by-step plan to solve the problem with symbolic logical rules, 3) followed by a verifier to check the translation and reasoning chain.Via thorough evaluations on 5 standard datasets with both First-Order Logic and Constraint Optimization symbolic expressions, SymbCoT shows striking improvements over the CoT method consistently, meanwhile refreshing the current stateof-the-art performances.We further demonstrate that our system advances in more faithful, flexible, and explainable logical reasoning.To our knowledge, this is the first to combine symbolic expressions and rules into CoT for logical reasoning with LLMs.Code is open at https://github.com/Aiden0526/SymbCoT.
Jundong Xu, Hao Fei 0001, Liangming Pan, Qian Liu 0012, Mong-Li Lee, Wynne Hsu
ACL (1)3
2024 Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
abstract
Recent studies show that large language models (LLMs) improve their performance through self-feedback on certain tasks while degrade on others.We discovered that such a contrary is due to LLM's bias in evaluating their own output.In this paper, we formally define LLM's self-bias -the tendency to favor its own generation -using two statistics.We analyze six LLMs (GPT-4, GPT-3.5, Gemini, LLaMA2, Mixtral and DeepSeek) on translation, constrained text generation, and mathematical reasoning tasks.We find that self-bias is prevalent in all examined LLMs across multiple languages and tasks.Our analysis reveals that while the self-refine pipeline improves the fluency and understandability of model outputs, it further amplifies self-bias.To mitigate such biases, we discover that larger model size and external feedback with accurate assessment can significantly reduce bias in the self-refine pipeline, leading to actual performance improvement in downstream tasks.The code and data are released at https://github. com/xu1998hz/llm_self_bias.
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li 0005, William Yang Wang
ACL (1)4
2024 SciAgent: Tool-augmented Language Models for Scientific Reasoning
abstract
Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang, Yixin Cao, Aixin Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang 0001, Yixin Cao 0002, Aixin Sun
EMNLP6
2024 AKEW: Assessing Knowledge Editing in the Wild
abstract
Knowledge editing injects knowledge updates into language models to keep them correct and up-to-date.However, its current evaluations deviate significantly from practice: their knowledge updates solely consist of structured facts derived from meticulously crafted datasets, instead of practical sources-unstructured texts like news articles, and they often overlook practical real-world knowledge updates.To address these issues, in this paper we propose AKEW (Assessing Knowledge Editing in the Wild), a new practical benchmark for knowledge editing.AKEW fully covers three editing settings of knowledge updates: structured facts, unstructured texts as facts, and extracted triplets.It further introduces new datasets featuring both counterfactual and real-world knowledge updates.Through extensive experiments, we demonstrate the considerable gap between state-of-the-art knowledge-editing methods and practical scenarios.Our analyses further highlight key insights to motivate future research for practical knowledge editing 1 .
Xiaobao Wu, Liangming Pan, William Yang Wang, Anh Tuan Luu
EMNLP2
2024 Understanding Reasoning Ability of Language Models From the Perspective of Reasoning Paths Aggregation
abstract
Pre-trained language models (LMs) are able to perform complex reasoning without explicit fine-tuning. To understand how pre-training with a next-token prediction objective contributes to the emergence of such reasoning capability, we propose that we can view an LM as deriving new conclusions by aggregating indirect reasoning paths seen at pre-training time. We found this perspective effective in two important cases of reasoning: logic reasoning with knowledge graphs (KGs) and chain-of-thought (CoT) reasoning. More specifically, we formalize the reasoning paths as random walk paths on the knowledge/reasoning graphs. Analyses of learned LM distributions suggest that a weighted sum of relevant random walk path probabilities is a reasonable way to explain how LMs reason. Experiments and analysis on multiple KG and CoT datasets reveal the effect of training on random walk paths and suggest that augmenting unlabeled random walk reasoning paths can improve real-world multi-step reasoning performance.
Xinyi Wang 0003, Alfonso Amayuelas, Kexun Zhang, Liangming Pan, Wenhu Chen, William Yang Wang
ICML4
2024 Position: AI/ML Influencers Have a Place in the Academic Process
abstract
As the number of accepted papers at AI and ML conferences reaches into the thousands, it has become unclear how researchers access and read research publications. In this paper, we investigate the role of social media influencers in enhancing the visibility of machine learning research, particularly the citation counts of papers they share. We have compiled a comprehensive dataset of over 8,000 papers, spanning tweets from December 2018 to October 2023, alongside controls precisely matched by 9 key covariates. Our statistical and causal inference analysis reveals a significant increase in citations for papers endorsed by these influencers, with median citation counts 2-3 times higher than those of the control group. Additionally, the study delves into the geographic, gender, and institutional diversity of highlighted authors. Given these findings, we advocate for a responsible approach to curation, encouraging influencers to uphold the journalistic standard that includes showcasing diverse research topics, authors, and institutions.
Iain Weissburg, Mehir Arora, Xinyi Wang 0003, Liangming Pan, William Yang Wang
ICML4
2024 MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations
abstract
Understanding documents with rich layouts and multi-modal components is a long-standing and practical task. Recent Large Vision-Language Models (LVLMs) have made remarkable strides in various tasks, particularly in single-page document understanding (DU). However, their abilities on long-context DU remain an open problem. This work presents MMLONGBENCH-DOC, a long-context, multi- modal benchmark comprising 1,082 expert-annotated questions. Distinct from previous datasets, it is constructed upon 135 lengthy PDF-formatted documents with an average of 47.5 pages and 21,214 textual tokens. Towards comprehensive evaluation, answers to these questions rely on pieces of evidence from (1) different sources (text, image, chart, table, and layout structure) and (2) various locations (i.e., page number). Moreover, 33.7\% of the questions are cross-page questions requiring evidence across multiple pages. 20.6\% of the questions are designed to be unanswerable for detecting potential hallucinations. Experiments on 14 LVLMs demonstrate that long-context DU greatly challenges current models. Notably, the best-performing model, GPT-4o, achieves an F1 score of only 44.9\%, while the second-best, GPT-4V, scores 30.5\%. Furthermore, 12 LVLMs (all except GPT-4o and GPT-4V) even present worse performance than their LLM counterparts which are fed with lossy-parsed OCR documents. These results validate the necessity of future research toward more capable long-context LVLMs.
Yubo Ma, Yuhang Zang, Liangyu Chen 0005, Meiqi Chen 0001, Yizhu Jiao, Xinze Li 0001, Xinyuan Lu, Xiaoyi Dong, Pan Zhang 0001, Liangming Pan, Yu-Gang Jiang 0001, Jiaqi Wang 0003, Yixin Cao 0002, Aixin Sun
NeurIPS12
2024 Automatically Correcting Large Language Models: Surveying the Landscape of Diverse Automated Correction Strategies
abstract
Abstract While large language models (LLMs) have shown remarkable effectiveness in various NLP tasks, they are still prone to issues such as hallucination, unfaithful reasoning, and toxicity. A promising approach to rectify these flaws is correcting LLMs with feedback, where the LLM itself is prompted or guided with feedback to fix problems in its own output. Techniques leveraging automated feedback—either produced by the LLM itself (self-correction) or some external system—are of particular interest as they make LLM-based solutions more practical and deployable with minimal human intervention. This paper provides an exhaustive review of the recent advances in correcting LLMs with automated feedback, categorizing them into training-time, generation-time, and post-hoc approaches. We also identify potential challenges and future directions in this emerging field.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang 0003, William Yang Wang
Trans. Assoc. Comput. Linguistics1
2023 InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
abstract
Cross-lingual topic models have been prevalent for cross-lingual text analysis by revealing aligned latent topics. However, most existing methods suffer from producing repetitive topics that hinder further analysis and performance decline caused by low-coverage dictionaries. In this paper, we propose the Cross-lingual Topic Modeling with Mutual Information (InfoCTM). Instead of the direct alignment in previous work, we propose a topic alignment with mutual information method. This works as a regularization to properly align topics and prevent degenerate topic representations of words, which mitigates the repetitive topic issue. To address the low-coverage dictionary issue, we further propose a cross-lingual vocabulary linking method that finds more linked cross-lingual words for topic alignment beyond the translations of a given dictionary. Extensive experiments on English, Chinese, and Japanese datasets demonstrate that our method outperforms state-of-the-art baselines, producing more coherent, diverse, and well-aligned topics and showing better transferability for cross-lingual classification tasks.
Xiaobao Wu, Xinshuai Dong, Thong Nguyen 0003, Chaoqun Liu, Liangming Pan, Anh Tuan Luu
AAAI5
2023 Modeling What-to-ask and How-to-ask for Answer-unaware Conversational Question Generation
abstract
Xuan Long Do, Bowei Zou, Shafiq Joty, Tran Tai, Liangming Pan, Nancy Chen, Ai Ti Aw. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Do Xuan Long, Bowei Zou, Shafiq R. Joty, Anh Tran Tai, Liangming Pan, Nancy F. Chen, AiTi Aw
ACL (1)5
2023 Fact-Checking Complex Claims with Program-Guided Reasoning
abstract
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, Preslav Nakov
ACL (1)1
2023 Doolittle: Benchmarks and Corpora for Academic Writing Formalization
abstract
Shizhe Diao, Yongyu Lei, Liangming Pan, Tianqing Fang, Wangchunshu Zhou, Sedrick Keh, Min-Yen Kan, Tong Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Shizhe Diao, Yongyu Lei, Liangming Pan, Tianqing Fang, Wangchunshu Zhou, Sedrick Keh, Min-Yen Kan, Tong Zhang 0001
EMNLP3
2023 SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables
abstract
Current scientific fact-checking benchmarks exhibit several shortcomings, such as biases arising from crowd-sourced claims and an overreliance on text-based evidence.We present SCITAB, a challenging evaluation dataset consisting of 1.2K expert-verified scientific claims that 1) originate from authentic scientific publications and 2) require compositional reasoning for verification.The claims are paired with evidence-containing scientific tables annotated with labels.Through extensive evaluations, we demonstrate that SCITAB poses a significant challenge to state-of-the-art models, including table-based pretraining models and large language models.All models except GPT-4 achieved performance barely above random guessing.Popular prompting techniques, such as Chain-of-Thought, do not achieve much performance gains on SCITAB.Our analysis uncovers several unique challenges posed by SCITAB, including table grounding, claim ambiguity, and compositional reasoning.
Xinyuan Lu, Liangming Pan, Qian Liu 0033, Preslav Nakov, Min-Yen Kan
EMNLP2
2023 MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language Models
abstract
Language Models (LMs) have shown impressive performance in various natural language tasks.However, when it comes to natural language reasoning, LMs still face challenges such as hallucination, generating incorrect intermediate reasoning steps, and making mathematical errors.Recent research has focused on enhancing LMs through self-improvement using feedback.Nevertheless, existing approaches relying on a single generic feedback source fail to address the diverse error types found in LMgenerated reasoning chains.In this work, we propose Multi-Aspect Feedback, an iterative refinement framework that integrates multiple feedback modules, including frozen LMs and external tools, each focusing on a specific error category.Our experimental results demonstrate the efficacy of our approach to addressing several errors in the LM-generated reasoning chain and thus improving the overall performance of an LM in several reasoning tasks.We see a relative improvement of up to 20% in Mathematical Reasoning and up to 18% in Logical Entailment.We release our source code, prompts, and data 1 to accelerate future research.
Deepak Nathani, Liangming Pan, William Yang Wang
EMNLP3
2023 INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback
abstract
Automatically evaluating the quality of language generation is critical.Although recent learned metrics show high correlation with human judgement, these metrics do not provide explicit explanation of their verdict, nor associate the scores with defects in the generated text.To address this limitation, we present IN-STRUCTSCORE, a fine-grained explainable evaluation metric for text generation.By harnessing both explicit human instruction and the implicit knowledge of GPT-4, we fine-tune a text evaluation metric based on LLaMA, producing both a score for generated text and a human readable diagnostic report.We evaluate INSTRUCTSCORE on a variety of generation tasks, including translation, captioning, data-to-text, and commonsense generation.Experiments show that our 7B model surpasses all other unsupervised metrics, including those based on 175B GPT-3 and GPT-4.Surprisingly, our INSTRUCTSCORE, even without direct supervision from human-rated data, achieves performance levels on par with state-of-the-art metrics like COMET22, which were fine-tuned on human ratings.Prompt: You are evaluating a model output based on a reference.Reference: Normally the administration office downstairs would call me when there's a delivery.Output: Usually when there is takeaway, the management office downstairs will call.
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, Lei Li 0005
EMNLP3
2023 FollowupQG: Towards information-seeking follow-up question generation
abstract
Yan Meng, Liangming Pan, Yixin Cao, Min-Yen Kan. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liangming Pan, Yixin Cao 0002, Min-Yen Kan
IJCNLP (1)2
2023 Attacking Open-domain Question Answering by Injecting Misinformation
abstract
Liangming Pan, Wenhu Chen, Min-Yen Kan, William Yang Wang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liangming Pan, Wenhu Chen, Min-Yen Kan, William Yang Wang
IJCNLP (1)1
2023 Investigating Zero- and Few-shot Generalization in Fact Verification
abstract
Liangming Pan, Yunxiang Zhang, Min-Yen Kan. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liangming Pan, Yunxiang Zhang 0002, Min-Yen Kan
IJCNLP (1)1
2023 Hashtag-Guided Low-Resource Tweet Classification
abstract
Social media classification tasks (e.g., tweet sentiment analysis, tweet stance detection) are challenging because social media posts are typically short, informal, and ambiguous. Thus, training on tweets is challenging and demands large-scale human-annotated labels, which are time-consuming and costly to obtain. In this paper, we find that providing hashtags to social media tweets can help alleviate this issue because hashtags can enrich short and ambiguous tweets in terms of various information, such as topic, sentiment, and stance. This motivates us to propose a novel Hashtag-guided Tweet Classification model (HashTation), which automatically generates meaningful hashtags for the input tweet to provide useful auxiliary signals for tweet classification. To generate high-quality and insightful hashtags, our hashtag generation model retrieves and encodes the post-level and entity-level information across the whole corpus. Experiments show that HashTation achieves significant improvements on seven low-resource tweet classification tasks, in which only a limited amount of training data is provided, showing that automatically enriching tweets with model-generated hashtags could significantly reduce the demand for large-scale human-labeled data. Further analysis demonstrates that HashTation is able to generate high-quality hashtags that are consistent with the tweets and their labels. The code is available at https://github.com/shizhediao/HashTation.
Shizhe Diao, Sedrick Keh, Liangming Pan, Zhiliang Tian, Yan Song 0003, Tong Zhang 0001
WWW3
2022 KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base
abstract
Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou, Juanzi Li, Bin He, Hanwang Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shulin Cao, Jiaxin Shi, Liangming Pan, Lunyiu Nie, Yutong Xiang, Lei Hou 0001, Juan-Zi Li, Hanwang Zhang
ACL (1)3
2022 CoHS-CQG: Context and History Selection for Conversational Question Generation
abstract
Conversational question generation (CQG) serves as a vital task for machines to assist humans, such as interactive reading comprehension, through conversations. Compared to traditional single-turn question generation (SQG), CQG is more challenging in the sense that the generated question is required not only to be meaningful, but also to align with the provided conversation. Previous studies mainly focus on how to model the flow and alignment of the conversation, but do not thoroughly study which parts of the context and history are necessary for the model. We believe that shortening the context and history is crucial as it can help the model to optimise more on the conversational alignment property. To this end, we propose CoHS-CQG, a two-stage CQG framework, which adopts a novel CoHS module to shorten the context and history of the input. In particular, it selects the top-p sentences and history turns by calculating the relevance scores of them. Our model achieves state-of-the-art performances on CoQA in both the answer-aware and answer-unaware settings.
Do Xuan Long, Bowei Zou, Liangming Pan, Nancy F. Chen, Shafiq R. Joty, AiTi Aw
COLING3
2022 KHANQ: A Dataset for Generating Deep Questions in Education
abstract
Designing in-depth educational questions is a time-consuming and cognitively demanding task. Therefore, it is intriguing to study how to build Question Generation (QG) models to automate the question creation process. However, existing QG datasets are not suitable for educational question generation because the questions are not real questions asked by humans during learning and can be solved by simply searching for information. To bridge this gap, we present KHANQ, a challenging dataset for educational question generation, containing 1,034 high-quality learner-generated questions seeking an in-depth understanding of the taught online courses in Khan Academy. Each data sample is carefully paraphrased and annotated as a triple of 1) Context: an independent paragraph on which the question is based; 2) Prompt: a text prompt for the question (e.g., the learner’s background knowledge); 3) Question: a deep question based on Context and coherent with Prompt. By conducting a human evaluation on the aspects of appropriateness, coverage, coherence, and complexity, we show that state-of-the-art QG models which perform well on shallow question generation datasets have difficulty in generating useful educational questions. This makes KHANQ a challenging testbed for educational question generation.
Huanli Gong, Liangming Pan, Hengchang Hu
COLING2
2022 Ingredient-enriched Recipe Generation from Cooking Videos
abstract
Cooking video captioning aims to generate the text instructions that describes the cooking procedures presented in the video. Current approaches tend to use large neural models or use more robust feature extractors to increase the expressive ability of features, ignoring the strong correlation between consecutive cooking steps in the video. However, it is intuitive that previous cooking steps can provide clues for the next cooking step. Specially, consecutive cooking steps tend to share the same ingredients. Therefore, accurate ingredients recognition can help to introduce more fine-grained information in captioning. To improve the performance of video procedural caption in cooking video, this paper proposes a framework that introduces ingredient recognition module which uses the copy mechanism to fuse the predicted ingredient information into the generated sentence. Moreover, we integrate the visual information of the previous step into the generation of the current step, and the visual information of the two steps together assist in the generation process. Extensive experiments verify the effectiveness of our propose framework and it achieves the promising performances on both YouCookII and Cooking-COIN datasets.
Jianlong Wu, Liangming Pan, Jingjing Chen 0001, Yu-Gang Jiang 0001
ICMR2
2021 Unsupervised Multi-hop Question Answering by Question Generation
abstract
Liangming Pan, Wenhu Chen, Wenhan Xiong, Min-Yen Kan, William Yang Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Liangming Pan, Wenhu Chen, Wenhan Xiong, Min-Yen Kan, William Yang Wang
NAACL-HLT1
2021 A Hybrid Approach for Detecting Prerequisite Relations in Multi-Modal Food Recipes
abstract
Modeling the structure of culinary recipes is the core of recipe representation learning. Current approaches mostly focus on extracting the workflow graph from recipes based on text descriptions. Process images, which constitute an important part of cooking recipes, has rarely been investigated in recipe structure modeling. We study this recipe structure problem from a multi-modal learning perspective, by proposing aprerequisite treeto represent recipes with cooking images at a step-level granularity. We propose a simple-yet-effective two-stage framework to automatically construct the prerequisite tree for a recipe by (1) utilizing a trained classifier to detect pairwise prerequisite relations that fuses multi-modal features as input; then (2) applying different strategies (greedy method, maximum weight, and beam search) to build the tree structure. Experiments on the MM-ReS dataset demonstrates the advantages of introducing process images for recipe structure modeling. Also, compared with neural methods which require large numbers of training data, we show that our two-stage pipeline can achieve promising results using only 400 labeled prerequisite trees as training data.
Liangming Pan, Jingjing Chen 0001, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Tat-Seng Chua
IEEE Trans. Multim.1
2020 Zero-Shot Ingredient Recognition by Multi-Relational Graph Convolutional Network
abstract
Recognizing ingredients for a given dish image is at the core of automatic dietary assessment, attracting increasing attention from both industry and academia. Nevertheless, the task is challenging due to the difficulty of collecting and labeling sufficient training data. On one hand, there are hundred thousands of food ingredients in the world, ranging from the common to rare. Collecting training samples for all of the ingredient categories is difficult. On the other hand, as the ingredient appearances exhibit huge visual variance during the food preparation, it requires to collect the training samples under different cooking and cutting methods for robust recognition. Since obtaining sufficient fully annotated training data is not easy, a more practical way of scaling up the recognition is to develop models that are capable of recognizing unseen ingredients. Therefore, in this paper, we target the problem of ingredient recognition with zero training samples. More specifically, we introduce multi-relational GCN (graph convolutional network) that integrates ingredient hierarchy, attribute as well as co-occurrence for zero-shot ingredient recognition. Extensive experiments on both Chinese and Japanese food datasets are performed to demonstrate the superior performance of multi-relational GCN and shed light on zero-shot ingredients recognition.
Jingjing Chen 0001, Liangming Pan, Zhipeng Wei 0001, Xiang Wang 0010, Chong-Wah Ngo, Tat-Seng Chua
AAAI2
2020 Expertise Style Transfer: A New Task Towards Better Communication between Experts and Laymen
abstract
The curse of knowledge can impede communication between experts and laymen.We propose a new task of expertise style transfer and contribute a manually annotated dataset with the goal of alleviating such cognitive biases.Solving this task not only simplifies the professional language, but also improves the accuracy and expertise level of laymen descriptions using simple words.This is a challenging task, unaddressed in previous work, as it requires the models to have expert intelligence in order to modify text with a deep understanding of domain knowledge and structures.We establish the benchmark performance of five stateof-the-art models for style transfer and text simplification.The results demonstrate a significant gap between machine and human performance.We also discuss the challenges of automatic evaluation, to provide insights into future research directions.
Yixin Cao 0002, Ruihao Shui, Liangming Pan, Min-Yen Kan, Zhiyuan Liu 0001, Tat-Seng Chua
ACL3
2020 Semantic Graphs for Generating Deep Questions
abstract
This paper proposes the problem of Deep Question Generation (DQG), which aims to generate complex questions that require reasoning over multiple pieces of information of the input passage.In order to capture the global structure of the document and facilitate reasoning, we propose a novel framework which first constructs a semantic-level graph for the input document and then encodes the semantic graph by introducing an attention-based GGNN (Att-GGNN).Afterwards, we fuse the document-level and graphlevel representations to perform joint training of content selection and question decoding.On the HotpotQA deep-question centric dataset, our model greatly improves performance over questions requiring reasoning over multiple facts, leading to state-of-theart performance.The code is publicly available at https://github.com/WING-NUS/ SG-Deep-Question-Generation.
Liangming Pan, Yuxi Xie, Yansong Feng 0002, Tat-Seng Chua, Min-Yen Kan
ACL1
2020 Exploring Question-Specific Rewards for Generating Deep Questions
abstract
Recent question generation (QG) approaches often utilize the sequence-to-sequence framework (Seq2Seq) to optimize the log-likelihood of ground-truth questions using teacher forcing.However, this training objective is inconsistent with actual question quality, which is often reflected by certain global properties such as whether the question can be answered by the document.As such, we directly optimize for QG-specific objectives via reinforcement learning to improve question quality.We design three different rewards that target to improve the fluency, relevance, and answerability of generated questions.We conduct both automatic and human evaluations in addition to a thorough analysis to explore the effect of each QG-specific reward.We find that optimizing question-specific rewards generally leads to better performance in automatic evaluation metrics.However, only the rewards that correlate well with human judgement (e.g., relevance) lead to real improvement in question quality.Optimizing for the others, especially answerability, introduces incorrect bias to the model, resulting in poor question quality.
Yuxi Xie, Liangming Pan, Dongzhe Wang, Min-Yen Kan, Yansong Feng 0002
COLING2
2020 Hyperbolic Visual Embedding Learning for Zero-Shot Recognition
abstract
This paper proposes a Hyperbolic Visual Embedding Learning Network for zero-shot recognition. The network learns image embeddings in hyperbolic space, which is capable of preserving the hierarchical structure of semantic classes in low dimensions. Comparing with existing zeroshot learning approaches, the network is more robust because the embedding feature in hyperbolic space better represents class hierarchy and thereby avoid misleading resulted from unrelated siblings. Our network outperforms exiting baselines under hierarchical evaluation with an extremely challenging setting, i.e., learning only from 1,000 categories to recognize 20,841 unseen categories. While under flat evaluation, it has competitive performance as state-of-the-art methods but with five times lower embedding dimensions. Our code is publicly available.
Shaoteng Liu, Jingjing Chen 0001, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, Yu-Gang Jiang 0001
CVPR3
2020 Exploring and Evaluating Attributes, Values, and Structures for Entity Alignment
abstract
Entity alignment (EA) aims at building a unified Knowledge Graph (KG) of rich content by linking the equivalent entities from various KGs.GNN-based EA methods present promising performance by modeling the KG structure defined by relation triples.However, attribute triples can also provide crucial alignment signal but have not been well explored yet.In this paper, we propose to utilize an attributed value encoder and partition the KG into subgraphs to model the various types of attribute triples efficiently.Besides, the performances of current EA methods are overestimated because of the name-bias of existing EA datasets.To make an objective evaluation, we propose a hard experimental setting where we select equivalent entity pairs with very different names as the test set.Under both the regular and hard settings, our method achieves significant improvements (5.10% on average Hits@1 in DBP15k) over 12 baselines in crosslingual and monolingual datasets.Ablation studies on different subgraphs and a case study about attribute types further demonstrate the effectiveness of our method.Source code and data can be found at https://github.com/ thunlp/explore-and-evaluate.
Zhiyuan Liu 0001, Yixin Cao 0002, Liangming Pan, Juan-Zi Li, Tat-Seng Chua
EMNLP (1)3
2020 Multi-modal Cooking Workflow Construction for Food Recipes
abstract
Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial task that involves common-sense reasoning. However, existing efforts rely on hand-crafted features to extract the workflow graph from recipes due to the lack of large-scale labeled datasets. Moreover, they fail to utilize the cooking images, which constitute an important part of food recipes. In this paper, we build MM-ReS, the first large-scale dataset for cooking workflow construction, consisting of 9,850 recipes with human-labeled workflow graphs. Cooking steps are multi-modal, featuring both text instructions and cooking images. We then propose a neural encoder-decoder model that utilizes both visual and textual information to construct the cooking workflow, which achieved over 20% performance gain over existing hand-crafted baselines.
Liangming Pan, Jingjing Chen 0001, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yu-Gang Jiang 0001, Tat-Seng Chua
ACM Multimedia1
2017 Prerequisite Relation Learning for Concepts in MOOCs
abstract
What prerequisite knowledge should students achieve a level of mastery before moving forward to learn subsequent coursewares?We study the extent to which the prerequisite relation between knowledge concepts in Massive Open Online Courses (MOOCs) can be inferred automatically.In particular, what kinds of information can be leveraged to uncover the potential prerequisite relation between knowledge concepts.We first propose a representation learning-based method for learning latent representations of course concepts, and then investigate how different features capture the prerequisite relations between concepts.Our experiments on three datasets form Coursera show that the proposed method achieves significant improvements (+5.9-48.0%by F1-score) comparing with existing methods.
Liangming Pan, Chengjiang Li, Juan-Zi Li, Jie Tang 0001
ACL (1)1
2017 Course Concept Extraction in MOOCs via Embedding-Based Graph Propagation
abstract
Massive Open Online Courses (MOOCs), offering a new way to study online, are revolutionizing education. One challenging issue in MOOCs is how to design effective and fine-grained course concepts such that students with different backgrounds can grasp the essence of the course. In this paper, we conduct a systematic investigation of the problem of course concept extraction for MOOCs. We propose to learn latent representations for candidate concepts via an embedding-based method. Moreover, we develop a graph-based propagation algorithm to rank the candidate concepts based on the learned representations. We evaluate the proposed method using different courses from XuetangX and Coursera. Experimental results show that our method significantly outperforms all the alternative methods (+0.013-0.318 in terms of R-precision; p<<0.01, t-test).
Liangming Pan, Xiaochen Wang 0002, Chengjiang Li, Juan-Zi Li, Jie Tang 0001
IJCNLP(1)1
2016 Domain Specific Cross-Lingual Knowledge Linking Based on Similarity Flooding
Liangming Pan, Juan-Zi Li, Jie Tang 0001
KSEM1