Moxin Li

dblp:266/2836 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
abstract
Xiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoyuan Li 0001, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang 0007, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu
ACL (1)4
2026 TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards
abstract
Xiqiao Xiong, Ouxiang Li, Zhuo Liu, Moxin Li, Wentao Shi, Fengbin Zhu, Qifan Wang, Fuli Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiqiao Xiong, Ouxiang Li, Moxin Li, Wentao Shi 0002, Fengbin Zhu, Qifan Wang 0001, Fuli Feng
ACL (1)4
2025 Knowledge Boundary of Large Language Models: A Survey
abstract
Moxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li, Wenya Xie, See-Kiong Ng, Tat-Seng Chua, Yang Deng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Moxin Li, Wenxuan Zhang 0001, Shuaiyi Li, Wenya Xie, See-Kiong Ng, Tat-Seng Chua, Yang Deng 0002
ACL (1)1
2025 Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs
abstract
Large Language Models (LLMs) excel in various natural language processing tasks but struggle with hallucination issues. Existing solutions have considered utilizing LLMs’ inherent reasoning abilities to alleviate hallucination, such as self-correction and diverse sampling methods. However, these methods often overtrust LLMs’ initial answers due to inherent biases. The key to alleviating this issue lies in overriding LLMs’ inherent biases for answer inspection. To this end, we propose a CounterFactual Multi-Agent Debate (CFMAD) framework. CFMAD presets the stances of LLMs to override their inherent biases by compelling LLMs to generate justifications for a predetermined answer’s correctness. The LLMs with different predetermined stances are engaged with a skeptical critic for counterfactual debate on the rationality of generated justifications. Finally, the debate process is evaluated by a third-party judge to determine the final answer. Extensive experiments on four datasets of three tasks demonstrate the superiority of CFMAD over existing methods.
Yi Fang 0010, Moxin Li, Wenjie Wang 0007, Fuli Feng
COLING2
2025 Large Language Models with Multi-faceted Relation Alignment for User Novel Interest Discovery
Shuxian Bi, Wenjie Wang 0007, Moxin Li, Chongming Gao, Fuli Feng
PAKDD (7)3
2025 Unveiling Knowledge Boundary of Large Language Models for Trustworthy Information Access
abstract
Large Language Models (LLMs) have emerged as powerful tools for generating content and facilitating information seeking across diverse domains. While their integration into conversational systems opens new avenues for interactive information-seeking experiences, their effectiveness is constrained by their knowledge boundaries-the limits of what they know and their ability to provide reliable, truthful, and contextually appropriate information. Understanding these boundaries is essential for maximizing the utility of LLMs for real-time information seeking while ensuring their reliability and trustworthiness. In this tutorial, we will explore the taxonomy of knowledge boundary in LLMs, addressing their handling of uncertainty, response calibration, and mitigation of unintended behaviors that can arise during interaction with users. We will also present advanced techniques for optimizing LLM behavior in generative information-seeking tasks, ensuring that models align with user expectations of accuracy and transparency. Attendees will gain insights into research trends and practical methods for enhancing the reliability and utility of LLMs for trustworthy information access.
Yang Deng 0002, Moxin Li, Liang Pang 0001, Wenxuan Zhang 0001, Wai Lam
SIGIR2
2024 Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
abstract
Table-text document (e.g., financial reports) understanding has attracted increasing attention in recent two years. TAT-DQA is a realistic setting for the understanding of visually-rich table-text documents, which involves answering associated questions requiring discrete reasoning. Most existing work relies on token-level semantics, falling short in the reasoning across document elements such as quantities and dates. To address this limitation, we propose a novel Doc2SoarGraph model that exploits element-level semantics and employs Semantic-oriented hierarchical Graph structures to capture the differences and correlations among different elements within the given document and question. Extensive experiments on the TAT-DQA dataset reveal that our model surpasses the state-of-the-art conventional method (i.e., MHST) and large language model (i.e., ChatGPT) by 17.73 and 6.49 points respectively in terms of Exact Match (EM) metric, demonstrating exceptional effectiveness.
Fengbin Zhu, Chao Wang 0049, Fuli Feng, Zifeng Ren, Moxin Li, Tat-Seng Chua
LREC/COLING5
2024 Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
abstract
Despite the remarkable abilities of Large Language Models (LLMs) to answer questions, they often display a considerable level of overconfidence even when the question does not have a definitive answer.To avoid providing hallucinated answers to these unknown questions, existing studies typically investigate approaches to refusing to answer these questions.In this work, we propose a novel and scalable self-alignment method to utilize the LLM itself to enhance its response-ability to different types of unknown questions, being capable of not just refusing to answer but further proactively providing explanations to the unanswerability of unknown questions.Specifically, the Self-Align method first employ a two-stage classaware self-augmentation approach to generate a large amount of unknown question-response data.Then we conduct disparity-driven selfcuration to select qualified data for fine-tuning the LLM itself for aligning the responses to unknown questions as desired.Experimental results on two datasets across four types of unknown questions validate the superiority of the Self-Aligned method over existing baselines in terms of three types of task formulation. 1 * Equal contribution.Q: What animal can be found at the top of the men's Wimbledon trophy? Direct AnswerA: The animal that can be found at the top of the men's Wimbledon trophy is a falcon.
Yang Deng 0002, Moxin Li, See-Kiong Ng, Tat-Seng Chua
EMNLP3
2023 Robust Prompt Optimization for Large Language Models Against Distribution Shifts
abstract
Large Language Model (LLM) has demonstrated significant ability in various Natural Language Processing tasks.However, their effectiveness is highly dependent on the phrasing of the task prompt, leading to research on automatic prompt optimization using labeled task data.We reveal that these prompt optimization techniques are vulnerable to distribution shifts such as subpopulation shifts, which are common for LLMs in real-world scenarios such as customer reviews analysis.In this light, we propose a new problem of robust prompt optimization for LLMs against distribution shifts, which requires the prompt optimized over the labeled source group can simultaneously generalize to an unlabeled target group.To solve this problem, we propose Generalized Prompt Optimization framework , which incorporates the unlabeled data from the target group into prompt optimization.Extensive experimental results demonstrate the effectiveness of the proposed framework with significant performance improvement on the target group and comparable performance on the source group.
Moxin Li, Wenjie Wang 0007, Fuli Feng, Yixin Cao 0002, Jizhi Zhang, Tat-Seng Chua
EMNLP1
2022 Learning to Imagine: Integrating Counterfactual Thinking in Neural Discrete Reasoning
abstract
Neural discrete reasoning (NDR) has shown remarkable progress in combining deep models with discrete reasoning.However, we find that existing NDR solution suffers from large performance drop on hypothetical questions, e.g., "what the annualized rate of return would be if the revenue in 2020 was doubled".The key to hypothetical question answering (HQA) is counterfactual thinking, which is a natural ability of human reasoning but difficult for deep models.In this work, we devise a Learning to Imagine (L2I) module, which can be seamlessly incorporated into NDR models to perform the imagination of unseen counterfactual.In particular, we formulate counterfactual thinking into two steps: 1) identifying the fact to intervene, and 2) deriving the counterfactual from the fact and assumption, which are designed as neural networks.Based on TAT-QA, we construct a very challenging HQA dataset with 8,283 hypothetical questions.We apply the proposed L2I to TAGOP, the state-of-theart solution on TAT-QA, validating the rationality and effectiveness of our approach.
Moxin Li, Fuli Feng, Hanwang Zhang, Xiangnan He 0001, Fengbin Zhu, Tat-Seng Chua
ACL (1)1
2021 Hybrid Learning to Rank for Financial Event Ranking
abstract
The financial markets are moved by events such as the issuance of administrative orders. The participants in financial markets (e.g., traders) thus pay constant attention to financial news relevant to the financial asset (e.g., oil) of interest. Due to the large scale of news stream, it is time and labor intensive to manually identify influential events that can move the price of the financial asset, pushing the financial participants to embrace automatic financial event ranking, which has received relatively little scrutiny to date. In this work, we formulate the financial event ranking task, which aims to score financial news (document) according to its influence to the given asset (query). To solve this task, we propose a Hybrid News Ranking framework that, from the asset perspective, evaluates the influence of news articles by comparing their contents; and from the event perspective, accesses the influence over all query assets. Moreover, we resolve the dilemma between the essential requirement of sufficient labels for training the framework and the unaffordable cost of hiring domain experts for labeling the news. In particular, we design a cost-friendly system for news labeling that leverages the knowledge within published financial analyst reports. In this way, we construct three financial event ranking datasets. Extensive experiments on the datasets validate the effectiveness of the proposed framework and the rationality of solving financial event ranking through learning to rank.
Fuli Feng, Moxin Li, Cheng Luo 0001, Ritchie Ng, Tat-Seng Chua
SIGIR2
2020 Learning Sense Representation from Word Representation for Unsupervised Word Sense Disambiguation (Student Abstract)
abstract
Unsupervised WSD methods do not rely on annotated training datasets and can use WordNet. Since each ambiguous word in the WSD task exists in WordNet and each sense of the word has a gloss, we propose SGM and MGM to learn sense representations for words in WordNet using the glosses. In the WSD task, we calculate the similarity between each sense of the ambiguous word and its context to select the sense with the highest similarity. We evaluate our method on several benchmark WSD datasets and achieve better performance than the state-of-the-art unsupervised WSD systems.
Zhenxin Fu, Moxin Li, Haisong Zhang, Dongyan Zhao 0001, Rui Yan 0001
AAAI3