Mayi Xu

dblp:292/2996 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
abstract
Large Language Models (LLMs) are increasingly employed in applications that require processing information from heterogeneous formats, including texts, tables, infoboxes, and knowledge graphs. However, systematic biases toward particular formats may undermine LLMs' ability to integrate heterogeneous data impartially, potentially resulting in reasoning errors and increased risks in downstream tasks. Yet it remains unclear whether such biases are systematic, which data-level factors drive them, and what internal mechanisms underlie their emergence. In this paper, we present the first comprehensive study of format bias in LLMs through a three-stage empirical analysis. The first stage explores the presence and direction of bias across a diverse range of LLMs. The second stage examines how key data-level factors influence these biases. The third stage analyzes how format bias emerges within LLMs' attention patterns and evaluates a lightweight intervention to test its effectiveness. Our results show that format bias is consistent across model families, driven by information richness, structure quality, and representation type, and is closely associated with attention imbalance within the LLMs. Based on these investigations, we identify three future research directions to reduce format bias: enhancing data pre-processing through format repair and normalization, introducing inference-time interventions such as attention re-weighting, and developing format-balanced training corpora. These directions will support the design of more robust and fair heterogeneous data processing systems.
Mayi Xu, Qiankun Pi, Ming Zhong 0002, Yuanyuan Zhu 0001, Mengchi Liu, Tieyun Qian
AAAI2
2026 Privacy-protected Retrieval-Augmented Generation for Knowledge Graph Question Answering
abstract
Large Language Models (LLMs) often suffer from hallucinations and outdated or incomplete knowledge. Retrieval-Augmented Generation (RAG) is proposed to address these issues by integrating external knowledge like that in knowledge graphs (KGs) into LLMs. However, leveraging private KGs in RAG systems poses significant privacy risks due to the black-box nature of LLMs and potential insecure data transmission. In this paper, we investigate the privacy-protected RAG scenario for the first time, where entities in KGs are anonymous for LLMs, thus preventing them from accessing entity semantics. Due to the loss of semantics of entities, previous RAG systems cannot retrieve question-relevant knowledge from KGs by matching questions with the meaningless identifiers of anonymous entities. To realize an effective RAG system in this scenario, two key challenges must be addressed: (1) How can anonymous entities be converted into retrievable information? (2) How to retrieve question-relevant anonymous entities? To address these challenges, we propose a novel Abstraction Reasoning on Graph (ARoG) framework including relation-centric abstraction and structure-oriented abstraction strategies. For challenge (1), the first strategy abstracts entities into high-level concepts by dynamically capturing the semantics of their adjacent relations. Hence, it supplements meaningful semantics which can further support the retrieval process. For challenge (2), the second strategy transforms unstructured natural language questions into structured abstract concept paths. These paths can be more effectively aligned with the abstracted concepts in KGs, thereby improving retrieval performance. In addition to guiding LLMs to effectively retrieve knowledge from KGs, these abstraction strategies also strictly protect privacy from being exposed to LLMs. Experiments on three datasets demonstrate that ARoG achieves strong performance and privacy-robustness, establishing a new practical direction for privacy-protected RAG systems.
Yunfeng Ning, Mayi Xu, Jintao Wen, Qiankun Pi, Yuanyuan Zhu 0001, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian
AAAI2
2026 Debiasing LLMs in Knowledge-Intensive Tasks via Information-Gain Guided Front-Door Adjustment
Yongqi Li 0002, Hankun Kang, Mayi Xu, Jintao Wen, Yuanyuan Zhu 0001, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian
DASFAA (3)4
2026 Balanced Heterogeneous Multi-teacher Distillation Framework for Biomedical Relation Extraction
Yingchang Liu, Mayi Xu, Wanli Li 0002, Zaiwen Feng
KSEM (2)4
2026 Recursive Short-to-Long Generalization for Multi-hop Reasoning
Mayi Xu, Ke Sun 0010, Jianhao Chen 0003, Qiankun Pi, Guixin Su, Yunfeng Ning, Yongqi Li 0002, Yuanyuan Zhu 0001, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian
SIGIR1
2026 ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
abstract
Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop evasive perturbations to disguise toxic content and evade detectors. Traditional detectors or methods are static over time and are inadequate in addressing these evolving evasion tactics. Thus, continual learning emerges as a logical approach to dynamically update detection ability against evolving perturbations. Nevertheless, disparities across perturbations hinder the detector's continual learning on perturbed text. More importantly, perturbation-induced noises distort semantics to degrade comprehension and also impair critical feature learning to render detection sensitive to perturbations. These amplify the challenge of continual learning against evolving perturbations.
Hankun Kang, Jianhao Chen 0003, Jintao Wen, Mayi Xu, Weiyu Zhang 0001, Wenpeng Lu, Tieyun Qian
WWW5
2026 Developing continuous toxicity detection against increasing types of perturbed toxic text
Hankun Kang, Jianhao Chen 0003, Yongqi Li 0002, Mayi Xu, Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian
Inf. Process. Manag.5
2026 Reasoning based on symbolic and parametric knowledge bases: A survey
Mayi Xu, Yunfeng Ning, Yongqi Li 0002, Jianhao Chen 0003, Jintao Wen, Birong Pan, Zepeng Bao, Hankun Kang, Ke Sun 0010, Tieyun Qian
Inf. Process. Manag.1
2026 How Robust are Large Language Models Against Word-Level Spurious Correlations? A Causal Discovery Approach
Yongqi Li 0002, Hankun Kang, Mayi Xu, Jintao Wen, Yuyang Ren, Tieyun Qian
Mach. Learn.4
2026 Local and Global Exploration for Next New POI Recommendation
abstract
The next Point-of-Interest (POI) recommendation is a hotspot for both industry and academia, which helps users better experience the physical world. However, existing methods suffer from a severe bias towards recommending repeat POIs that have been visited by the target user before, and perform inefficiently when recommending new POIs that have not been visited by the target user yet. To overcome this issue, we delve into the next new POI recommendation and uncover the coexistence of local and global exploration patterns in users’ visits to new POIs, showing their willingness to explore not only nearby new POIs but also those distant ones. Subsequently, we develop a novel Local and Global Exploration ( LGE ) framework for the next new POI recommendation. In particular, LGE involves three key modules: (1) a Zone-Aware Local Exploration (ZLE) module, which encourages users to explore POIs in the local area by learning zone-aware POI representations and regularizing POI prediction with zone information; (2) an Intention-Aware Global Exploration (IGE) module, which recommends POIs that meet user intentions without distance constraints by extracting static and dynamic intentions from category information; (3) a fusion module, which contains a Mean Pooling (MP) strategy and a Weighted Pooling (WP) strategy to aggregate the outputs of local and global exploration modules for the final recommendation. Experiments carried out on real-world datasets have shown the effectiveness of LGE in recommending new POIs.
Ke Sun 0010, Liyu Zhou, Mayi Xu, Tieyun Qian
ACM Trans. Knowl. Discov. Data3
2025 Strong Empowered and Aligned Weak Mastered Annotation for Weak-to-Strong Generalization
abstract
The super-alignment problem of how humans can effectively supervise super-human AI has garnered increasing attention. Recent research has focused on investigating the weak-to-strong generalization (W2SG) scenario as an analogy for super-alignment. This scenario examines how a pre-trained strong model, supervised by an aligned weak model, can outperform its weak supervisor. Despite good progress, current W2SG methods face two main issues: 1) The annotation quality is limited by the knowledge scope of the weak model; 2) It is risky to position the strong model as the final corrector. To tackle these issues, we propose a ``Strong Empowered and Aligned Weak Mastered'' (SEAM) framework for weak annotations in W2SG. This framework can leverage the vast intrinsic knowledge of the pre-trained strong model to empower the annotation and position the aligned weak model as the annotation master. Specifically, the pre-trained strong model first generates principle fast-and-frugal trees for samples to be annotated, encapsulating rich sample-related knowledge. Then, the aligned weak model picks informative nodes based on the tree's information distribution for final annotations. Experiments on six datasets for preference tasks in W2SG scenarios validate the effectiveness of our proposed method.
Yongqi Li 0002, Mayi Xu, Tieyun Qian
AAAI3
2025 Enhancing Relation Extraction via Supervised Rationale Verification and Feedback
abstract
Despite the rapid progress that existing automated feedback methods have made in correcting the output of large language models (LLMs), these methods cannot be well applied to the relation extraction (RE) task due to their designated feedback objectives and correction manner. To address this problem, we propose a novel automated feedback framework for RE, which presents a rationale supervisor to verify the rationale and provides re-selected demonstrations as feedback to correct the initial prediction. Specifically, we first design a causal intervention and observation method to collect biased/unbiased rationales for contrastive training the rationale supervisor. Then, we present a verification-feedback-correction procedure to iteratively enhance LLMs' capability of handling the RE task. Extensive experiments prove that our proposed framework significantly outperforms existing methods.
Yongqi Li 0002, Mayi Xu, Yuyang Ren, Tieyun Qian
AAAI4
2025 Aligning VLM Assistants with Personalized Situated Cognition
abstract
Yongqi Li, Shen Zhou, Xiaohu Li, Xin Miao, Jintao Wen, Mayi Xu, Jianhao Chen, Birong Pan, Hankun Kang, Yuanyuan Zhu, Ming Zhong, Tieyun Qian. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yongqi Li 0002, Xiaohu Li, Jintao Wen, Mayi Xu, Jianhao Chen 0003, Birong Pan, Hankun Kang, Yuanyuan Zhu 0001, Ming Zhong 0002, Tieyun Qian
ACL (1)6
2024 Prompting Large Language Models for Counterfactual Generation: An Empirical Study
abstract
Large language models (LLMs) have made remarkable progress in a wide range of natural language understanding and generation tasks. However, their ability to generate counterfactuals has not been examined systematically. To bridge this gap, we present a comprehensive evaluation framework on various types of NLU tasks, which covers all key factors in determining LLMs’ capability of generating counterfactuals. Based on this framework, we 1) investigate the strengths and weaknesses of LLMs as the counterfactual generator, and 2) disclose the factors that affect LLMs when generating counterfactuals, including both the intrinsic properties of LLMs and prompt designing. The results show that, though LLMs are promising in most cases, they face challenges in complex tasks like RE since they are bounded by task-specific performance, entity constraints, and inherent selection bias. We also find that alignment techniques, e.g., instruction-tuning and reinforcement learning from human feedback, may potentially enhance the counterfactual generation ability of LLMs. On the contrary, simply increasing the parameter size does not yield the desired improvements. Besides, from the perspective of prompt designing, task guidelines unsurprisingly play an important role. However, the chain-of-thought approach does not always help due to inconsistency issues.
Yongqi Li 0002, Mayi Xu, Tieyun Qian
LREC/COLING2
2024 Adaption-of-Thought: Learning Question Difficulty Improves Large Language Models for Reasoning
abstract
Large language models (LLMs) have shown excellent capability for solving reasoning problems.Existing approaches do not differentiate the question difficulty when designing prompting methods for them.Clearly, a simple method cannot elicit sufficient knowledge from LLMs to answer a hard question.Meanwhile, a sophisticated one will force the LLMs to generate redundant or even inaccurate intermediate steps for a simple question.Consequently, the performance of existing methods fluctuates among various questions.In this work, we propose Adaption-of-Thought (ADOT), an adaptive method, to improve LLMs for the reasoning problem, which first measures the question difficulty and then tailors demonstration set construction and difficulty-adapted retrieval strategies for the adaptive demonstration construction.Experimental results on three reasoning tasks prove the superiority of our proposed method, showing an absolute improvement of up to 5.5% on arithmetic reasoning, 7.4% on symbolic reasoning, and 2.3% on commonsense reasoning.
Mayi Xu, Yongqi Li 0002, Ke Sun 0010, Tieyun Qian
EMNLP1
2023 Cold-Start Multi-hop Reasoning by Hierarchical Guidance and Self-verification
Mayi Xu, Ke Sun 0010, Yongqi Li 0002, Tieyun Qian
ECML/PKDD (2)1
2023 RFM: response-aware feedback mechanism for background based conversation
Jiatao Chen, Zhibin Du, Huimin Deng, Mayi Xu, Zibang Gan, Meirong Ding
Appl. Intell.5
2023 Improving Span-Based Aspect Sentiment Triplet Extraction with Abundant Syntax Knowledge
Lingcong Feng, Lewei He, Mayi Xu, Huimin Deng, Zipeng Huang, Weihua Du
Neural Process. Lett.4
2022 Combining dynamic local context focus and dependency cluster attention for aspect-level sentiment classification
Mayi Xu, Heng Yang 0008, Junlong Chi, Jiatao Chen, Hongye Liu
Neurocomputing1
2022 Learning for target-dependent sentiment based on local context-aware embedding
Heng Yang 0008, Shuai Liu 0004, Mayi Xu
J. Supercomput.4