EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyuan Yi
dblp:179/2248
· DBLP profile ↗
41ranked-venue papers
6as first author
33since 2021 · last 2026
0000-0003-2710-1613ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 6 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective OptimizationabstractTrained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values) by prompting, benefiting applications like personalized LLMs and social simulations. However, existing methods suffer from the superficial elicitation problem: LLMs can only be steered to mimic shallow and unstable stylistic patterns, failing to embody the desired traits precisely and consistently across diverse tasks like humans. To address this challenge, we propose IROTE, a novel in-context method for stable and transferable trait elicitation. Drawing on psychological theories suggesting that traits are formed through identity-related reflection, our method automatically generates and optimizes a textual self-reflection within prompts, which comprises self-perceived experience, to stimulate LLMs' trait-driven behavior. The optimization is performed by iteratively maximizing an information-theoretic objective that enhances the connections between LLMs' behavior and the target trait, while reducing noisy redundancy in reflection without any fine-tuning, leading to evocative and compact trait reflection. Extensive experiments across three human trait systems manifest that one single IROTE-generated self-reflection can induce LLMs' stable impersonation of the target trait across diverse downstream tasks beyond simple questionnaire answering, consistently outperforming existing strong baselines. Yuzhuo Bai, Shitong Duan, Muhua Huang, Jing Yao 0003, Zhenghao Liu 0001, Peng Zhang 0060, Tun Lu, Xiaoyuan Yi, Maosong Sun 0001, Xing Xie 0001 |
AAAI | 8 |
| 2026 | MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual QuestionsabstractRecently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despite substantial work investigating the trustworthiness of language models, MMLMs' capability to act honestly, especially when faced with visually unanswerable questions, remains largely underexplored. This work presents the first systematic assessment of honesty behaviors across various MLLMs. We ground honesty in models' response behaviors to unanswerable visual questions, define four representative types of such questions, and construct MoHoBench, a large-scale MMLM honest benchmark, consisting of 12k+ visual question samples, whose quality is guaranteed by multi-stage filtering and human verification. Using MoHoBench, we benchmarked the honesty of 28 popular MMLMs and conducted a comprehensive analysis. Our findings show that: (1) most models fail to appropriately refuse to answer when necessary, and (2) MMLMs' honesty is not solely a language modeling issue, but is deeply influenced by visual information, necessitating the development of dedicated methods for multimodal honesty alignment. Therefore, we implemented initial alignment methods using supervised and preference learning to improve honesty behavior, providing a foundation for future work on trustworthy MLLMs. Yanxu Zhu, Shitong Duan, Xiangxu Zhang, Peng Zhang 0060, Tun Lu, Xiao Zhou 0005, Jing Yao 0003, Xiaoyuan Yi, Xing Xie 0001 |
AAAI | 9 |
| 2026 | Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference OptimizationabstractShaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu, Xiaoyuan Yi, Yukun Yan, Shuo Wang, Yu Gu, Ge Yu, Maosong Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaohua Duan, Pengcheng Huang 0004, Zhenghao Liu 0001, Xiaoyuan Yi, Yukun Yan, Shuo Wang 0013, Yu Gu 0002, Ge Yu 0001, Maosong Sun 0001 |
ACL (1) | 5 |
| 2026 | Influence-based Online Experience Selection for Effective RLHFabstractReinforcement Learning from Human Feedback (RLHF) has emerged as a crucial technique for aligning large language models (LLMs) with human preferences.However, existing RLHF methods face key challenges, including poor sample efficiency, high computational overhead, and slow convergence.Recent studies highlight the importance of data selection in RL, but how to effectively select the most beneficial experiences for RL training remains an open problem.Existing data selection methods for RL rely on heuristic metrics, failing to establish an interpretable connection between data and optimization objectives.To address this problem, we propose InfOES (Influence-based Online Experience Selection), a novel data selection method for RLHF that dynamically estimates the influence of individual training samples on policy optimization.By incorporating data attribution into the policy gradient, InfOES can identify and filter out detrimental samples on the fly, ensuring effective convergence toward alignment objectives.Our approach is compatible with various RL algorithms (e.g., PPO, GRPO, RE-INFORCE++).Extensive experiments demonstrate that InfOES significantly enhances training effectiveness, achieving superior alignment performance with fewer optimization steps. Yifan Gong 0001, Jing Yao 0003, Xiting Wang, Xunlong Wang, Xiaoyuan Yi, Xing Xie 0001 |
ACL (1) | 5 |
| 2026 | Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural AlignmentabstractBryan Chen Zhengyu Tan, Zhengyuan Liu, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Roy Ka-Wei Lee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bryan Chen Zhengyu Tan, Zhengyuan Liu, Xiaoyuan Yi, Jing Yao 0003, Xing Xie 0001, Nancy F. Chen, Roy Ka-Wei Lee |
ACL (1) | 3 |
| 2026 | MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationabstractWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengyuan Liu, Tanmoy Chakraborty 0002, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu 0071, Xing Xie 0001, Xiaoyuan Yi, Jing Yao 0003, Chaojun Wang, Rui Liu 0019, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Lingyu Ye, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen |
ACL (1) | 11 |
| 2025 | Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsabstractThe application scope of Large Language Models (LLMs) continues to expand, leading to increasing interest in personalized LLMs that align with human values.However, aligning these models with individual values raises significant safety concerns, as certain values may correlate with harmful information.In this paper, we identify specific safety risks associated with value-aligned LLMs and investigate the psychological principles behind these challenges.Our findings reveal two key insights.(1) Value-aligned LLMs are more prone to harmful behavior compared to non-fine-tuned models and exhibit slightly higher risks in traditional safety evaluations than other fine-tuned models.(2) These safety issues arise because value-aligned LLMs genuinely generate text according to the aligned values, which can amplify harmful outcomes.Using a dataset with detailed safety categories, we find significant correlations between value alignment and safety risks, supported by psychological hypotheses.This study offers insights into the "black box" of value alignment and proposes in-context alignment methods to enhance the safety of value-aligned LLMs. 1 Sooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao 0003, Xing Xie 0001, JinYeong Bak |
ACL (1) | 3 |
| 2025 | Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and EnhancementabstractAs Large Language Models (LLMs) advance, aligning them with human values is critical for their responsible development.Value principles serve as the foundation for clarifying alignment goals.Multiple sets of value principles have been proposed, such as HHH (helpful, honest, harmless) and instructions for data synthesis in reinforcement learning from AI feedback (RLAIF).However, most of them are heuristically crafted, without consideration of three primary challenges in practical LLM alignment: 1) Comprehensiveness to deal with diverse and even unforeseen scenarios in which LLMs could be applied; 2) Precision to provide LLMs with clear and actionable guidance in specific scenarios; and 3) Compatability to avoid internal contracts between principles.In this paper, we formalize quantitative metrics to evaluate value principles along the three desirable properties.Building on these metrics, we propose the Hierarchical Value Principle framework (HiVaP) 1 , which constructs a hierarchical principle set and retrieves principles tailored to each scenario in a cascading way, addressing above challenges.Experimental results validate that the three metrics capture the effectiveness of value principles for LLM alignment, and our HiVaP framework that enhances these metrics leads to superior alignment. Bingbing Xu 0009, Jing Yao 0003, Xiaoyuan Yi, Aishan Maoliniyazi, Xing Xie 0001, Xiaofeng Meng 0001 |
ACL (1) | 3 |
| 2025 | LegalDuet: Learning Fine-Grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
Buqiang Xu, Zhenghao Liu 0001, Huiyuan Xie, Xiaoyuan Yi, Shuo Wang 0013, Yukun Yan, Liner Yang, Yu Gu 0002, Ge Yu 0001 |
ADMA (1) | 5 |
| 2025 | MoVa: Towards Generalizable Classification of Human Morals and ValuesabstractZiyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen, Jing Yao, Xiaoyuan Yi, Xing Xie, Chenhao Tan, Lexing Xie. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Junfei Sun, Tuan Dung Nguyen, Jing Yao 0003, Xiaoyuan Yi, Xing Xie 0001, Chenhao Tan, Lexing Xie |
EMNLP | 6 |
| 2025 | Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testingabstract*Warning: Contains harmful model outputs.*
Despite significant advancements, the propensity of Large Language Models (LLMs) to generate harmful and unethical content poses critical challenges.
Measuring value alignment of LLMs becomes crucial for their regulation and responsible deployment. Although numerous benchmarks have been constructed to assess social bias, toxicity, and ethical issues in LLMs, those static benchmarks suffer from *evaluation chronoeffect*, in which, as models rapidly evolve, existing benchmarks may leak into training data or become saturated, *overestimating* ever-developing LLMs. To tackle this problem, we propose GETA, a novel *generative evolving testing* approach based on adaptive testing methods in measurement theory. Unlike traditional adaptive testing methods that rely on a static test item pool, GETA probes the underlying moral boundaries of LLMs by dynamically generating test items tailored to model capability. GETA co-evolves with LLMs by learning a joint distribution of item difficulty and model value conformity, thus effectively addressing evaluation chronoeffect.
We evaluated various popular LLMs with GETA and demonstrated that 1) GETA can dynamically create difficulty-tailored test items and 2) GETA's evaluation results are more consistent with models' performance on unseen OOD and i.i.d. items, laying the groundwork for future evaluation paradigms. Han Jiang 0007, Xiaoyuan Yi, Zhihua Wei 0001, Ziang Xiao, Xing Xie 0001 |
ICML | 2 |
| 2025 | Benchmarking Retrieval-Augmented Generation in Multi-Modal ContextsabstractWith the rapid advancement of Multi-modal Large Language Models (MLLMs), their capability in understanding both images and text has greatly improved. However, their potential for leveraging multi-modal contextual information in Retrieval-Augmented Generation (RAG) remains largely underexplored. To address this gap, this paper introduces Multi-Modal Retrieval-Augmented Generation (M2RAG), a benchmark designed to evaluate the effectiveness of Multi-modal Large Language Models in leveraging knowledge from multi-modal retrieval documents. The benchmark comprises four tasks: image captioning, multi-modal question answering, multi-modal fact verification, and image reranking. All tasks are set in an open-domain setting, requiring RAG models to retrieve query-relevant information from a multi-modal document collection and use it as contextual input for RAG modeling. To enhance the context utilization capabilities of MLLMs, we also introduce Multi-Modal Retrieval-Augmented Instruction Tuning (MM-RAIT), an instruction tuning method that optimizes MLLMs within multi-modal contexts. Our experiments demonstrate the effectiveness of MM-RAIT by significantly improving the quality of responses generated by different RAG models, outperforming MiniCPM-V 2.6 and Qwen2-VL with 34% and 33% gains, respectively. All data and code are available at https://github.com/NEUIR/M2RAG. Zhenghao Liu 0001, Xingsheng Zhu, Tianshuo Zhou, Xiaoyuan Yi, Yukun Yan, Ge Yu 0001, Maosong Sun 0001 |
ACM Multimedia | 5 |
| 2025 | Specify Privacy Yourself: Assessing Inference-Time Personalized Privacy Preservation Ability of Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable capabilities but raise significant privacy concerns due to their abilities to infer sensitive personal information from images with high precision. While current LVLMs are relatively well aligned to protect universal privacy, e.g., credit card data, we argue that privacy is inherently personalized and context-dependent. This work pivots towards a novel task: can LVLMs achieve Inference-Time Personalized Privacy Protection (ITP3), allowing users to dynamically specify privacy boundaries through language specifications? To this end, we present SPY-Bench, the first systematic assessment of ITP3 ability, which comprises (1) 32,700 unique samples with image-question pairs and personalized privacy instructions across 67 categories and 24 real-world scenarios, and (2) novel metrics grounded in user specifications and context awareness. Benchmarking the ITP3 ability of 21 SOTA LVLMs, we reveal that: (i) most models, even the top-performing o4-mini, perform poorly, with only ~24% compliance accuracy; (ii) they show quite limited contextual privacy understanding capability. Therefore, we implemented initial ITP3 alignment methods, including a novel Noise Contrastive Alignment variant which achieves 96.88% accuracy while maintaining reasonable general performance. These results mark an initial step towards the ethical deployment of more controllable LVLMs. Code and data are at https://github.com/achernarwang/specify-privacy-yourself. Xingqi Wang 0003, Xiaoyuan Yi, Xing Xie 0001, Jia Jia 0001 |
ACM Multimedia | 2 |
| 2025 | Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language ModelsabstractAs large language models (LLMs) become increasingly integrated into applications serving users across diverse cultures, communities, and demographics, it is critical to align LLMs with pluralistic human values beyond average principles (e.g., HHH).
In psychological and social value theories such as Schwartz’s Value Theory, pluralistic values are represented by multiple value dimensions paired with various priorities. However, existing methods encounter two challenges when aligning with such fine-grained value objectives: 1) they often treat multiple values as independent and equally important, ignoring their interdependence and relative priorities (value complexity); 2) they struggle to precisely control nuanced value priorities, especially those underrepresented ones (value steerability). To handle these challenges, we propose COUPLE, a COUnterfactual reasoning framework for PLuralistic valuE alignment. It introduces a structural causal model (SCM) to feature complex interdependency and prioritization among features, as well as the causal relationship between high-level value dimensions and behaviors. Moreover, it applies counterfactual reasoning to generate outputs aligned with any desired value objectives. Benefitting from explicit causal modeling, COUPLE also provides better interpretability. We evaluate COUPLE on two datasets with different value systems and demonstrate that COUPLE advances other baselines across diverse types of value objectives. Our code is available at https://github.com/microsoft/COUPLE. Hanze Guo, Jing Yao 0003, Xiao Zhou 0005, Xiaoyuan Yi, Xing Xie 0001 |
NeurIPS | 4 |
| 2025 | ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented GenerationabstractLarge language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external evidence. However, they remain susceptible to unfaithful generation, where outputs contradict retrieved context despite its relevance and accuracy. Existing approaches aiming to improve faithfulness primarily focus on enhancing the utilization of external context, but often overlook the persistent influence of internal parametric knowledge during generation. In this work, we investigate the internal mechanisms behind unfaithful generation and identify a subset of mid-to-deep feed-forward networks (FFNs) that are disproportionately activated in such cases. Building on this insight, we propose Parametric Knowledge Muting through FFN Suppression (ParamMute), a framework that improves contextual faithfulness by suppressing the activation of unfaithfulness-associated FFNs and calibrating the model toward retrieved knowledge. To evaluate our approach, we introduce CoFaithfulQA, a benchmark specifically designed to evaluate faithfulness in scenarios where internal knowledge conflicts with accurate external evidence. Experimental results show that ParamMute significantly enhances faithfulness across both CoFaithfulQA and the established ConFiQA benchmark, achieving substantial reductions in reliance on parametric memory. These findings underscore the importance of mitigating internal knowledge dominance and provide a new direction for improving LLM trustworthiness in RAG. All codes are available at https://github.com/OpenBMB/ParamMute. Pengcheng Huang 0004, Zhenghao Liu 0001, Yukun Yan, Xiaoyuan Yi, Zhiyuan Liu 0001, Maosong Sun 0001, Tong Xiao 0001, Ge Yu 0001, Chenyan Xiong |
NeurIPS | 5 |
| 2025 | Multi-Evidence Based Fact Verification via A Confidential Graph Neural NetworkabstractFact verification tasks aim to identify the integrity of textual contents according to the truthful corpus. Existing fact verification models usually build a fully connected reasoning graph, which regards claim-evidence pairs as nodes and connects them with edges. They employ the graph to propagate the semantics of the nodes. Nevertheless, the noisy nodes usually propagate their semantics via the edges of the reasoning graph, which misleads the semantic representations of other nodes and amplifies the noise signals. To mitigate the propagation of noisy semantic information, we introduce a Confidential Graph Attention Network (CO-GAT), which proposes a node masking mechanism for modeling the nodes. Specifically, CO-GAT calculates the node confidence score by estimating the relevance between the claim and evidence pieces. Then, the node masking mechanism uses the node confidence scores to control the noise information flow from the vanilla node to the other graph nodes. CO-GAT achieves a 73.59% FEVER score on the FEVER dataset and shows the generalization ability by broadening the effectiveness to the science-specific domain. Yuqing Lan, Zhenghao Liu 0001, Yu Gu 0002, Xiaoyuan Yi, Xiaohua Li 0004, Liner Yang, Ge Yu 0001 |
IEEE Trans. Big Data | 4 |
| 2025 | Neural Recommendation Reasoning with Logic RulesabstractExplainability is critical for recommender systems to ensure good user experience and facilitate designers to debug. However, generating explanations in recommender systems usually requires large efforts due to the dependency on additional data and case-by-case model design. One possible solution to these challenges is reasoning with logic rules, whose validity or confidence can automatically indicate high-quality explanations and formats are general. However, pioneer methods can be hardly applied in recommendation due to the high sparsity of interaction data, which raises the difficulty in accurately computing the rule validity, and the specific ranking-oriented task. To bridge this gap, we propose a general framework for Reco mmendation with lo gic r ule reasoning ( Recolor ) that satisfies three desirable properties. First, we explicitly estimate the rule validity to ensure well-grounded decisions, where a fuzzy logic validity module is designed for accurate estimation on highly sparse recommendation data. Second, we ensure the generality for both the types of input data and model architectures by designing a neural logic generation module, which decouples the user–item representation learning from the rule construction. Third, we integrate the two above-mentioned modules with a ranking-oriented BPR loss and achieve a unified optimization of explainability and accuracy. For any given neural recommendation model, our proposed logic rule reasoning framework can upgrade it to a self-explainable version. Numerical experiments and user studies on four public recommendation datasets with different levels of sparsity demonstrate that our framework shows high-validity rule explanations, generality in architecture and data, and high recommendation accuracy. Jing Yao 0003, Xiting Wang, Jianxun Lian, Xiaoyuan Yi, Xing Xie 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2024 | Denevil: towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction LearningabstractLarge Language Models (LLMs) have made unprecedented breakthroughs, yet their increasing integration into everyday life might raise societal risks due to generated unethical content. Despite extensive study on specific issues like bias, the intrinsic values of LLMs remain largely unexplored from a moral philosophy perspective. This work delves into ethical values utilizing Moral Foundation Theory. Moving beyond conventional discriminative evaluations with poor reliability, we propose DeNEVIL, a novel prompt generation algorithm tailored to dynamically exploit LLMs’ value vulnerabilities and elicit the violation of ethics in a generative manner, revealing their underlying value inclinations. On such a basis, we construct MoralPrompt, a high-quality dataset comprising 2,397 prompts covering 500+ value principles, and then benchmark the intrinsic values across a spectrum of LLMs. We discovered that most models are essentially misaligned, necessitating further ethical value alignment. In response, we develop VILMO, an in-context alignment method that substantially enhances the value compliance of LLM outputs by learning to generate appropriate value instructions, outperforming existing competitors. Our methods are suitable for black-box and open-source models, offering a promising initial step in studying the ethical values of LLMs. Shitong Duan, Xiaoyuan Yi, Peng Zhang 0060, Tun Lu, Xing Xie 0001, Ning Gu 0001 |
ICLR | 2 |
| 2024 | On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
Xinpeng Wang 0001, Shitong Duan, Xiaoyuan Yi, Jing Yao 0003, Shanlin Zhou, Zhihua Wei 0001, Peng Zhang 0060, Dongkuan Xu, Maosong Sun 0001, Xing Xie 0001 |
IJCAI | 3 |
| 2024 | Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value OptimizationabstractRecent advancements in diffusion models trained on large-scale data have enabled the generation of indistinguishable human-level images, yet they often produce harmful content misaligned with human values, e.g., social bias, and offensive content. Despite extensive research on Large Language Models (LLMs), the challenge of Text-to-Image (T2I) model alignment remains largely unexplored. Addressing this problem, we propose LiVO (Lightweight Value Optimization), a novel lightweight method for aligning T2I models with human values. LiVO only optimizes a plug-and-play value encoder to integrate a specified value principle with the input prompt, allowing the control of generated images over both semantics and values. Specifically, we design a diffusion model-tailored preference optimization loss, which theoretically approximates the Bradley-Terry model used in LLM alignment but provides a more flexible trade-off between image quality and value conformity. To optimize the value encoder, we also develop a framework to automatically construct a text-image preference dataset of 86k (prompt, aligned image, violating image, value principle) samples. Without updating most model parameters and through adaptive value selection from the input prompt, LiVO significantly reduces harmful outputs and achieves faster convergence, surpassing several strong baselines and taking an initial step towards ethically aligned T2I models. Warning: This paper involves descriptions and images depicting discriminatory, pornographic, bloody, and horrific scenes. Xingqi Wang 0003, Xiaoyuan Yi, Xing Xie 0001, Jia Jia 0001 |
ACM Multimedia | 2 |
| 2024 | Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human ValueabstractJing Yao, Xiaoyuan Yi, Yifan Gong, Xiting Wang, Xing Xie. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jing Yao 0003, Xiaoyuan Yi, Yifan Gong 0001, Xiting Wang, Xing Xie 0001 |
NAACL-HLT | 2 |
| 2024 | CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated ResponsesabstractThe rapid progress in Large Language Models (LLMs) poses potential risks such as generating unethical content. Assessing the values embedded in LLMs' generated responses can help expose their misalignment, but this relies on reference-free value evaluators, e.g. fine-tuned LLMs or closed-source models like GPT-4. Nevertheless, two key challenges emerge in open-ended value evaluation: the evaluator should adapt to changing human value definitions with minimal annotation, against their own bias (adaptability); and remain robust across varying value expressions and scenarios (generalizability). To handle these challenges, we introduce CLAVE, a novel framework that integrates two complementary LLMs: a large model to extract high-level value concepts from diverse responses, leveraging its extensive knowledge and generalizability, and a small model fine-tuned on these concepts to adapt to human value annotations. This dual-model framework enables adaptation to any value system using <100 human-labeled samples per value type. We also present ValEval, a comprehensive dataset comprising 13k+ (text,value,label) tuples across diverse domains, covering three major value systems. We benchmark the performance of 15+ popular LLM evaluators and fully analyze their strengths and weaknesses. Our findings reveal that CLAVE combining a large prompt-based model and a small fine-tuned one serves as an optimal balance in value evaluation. Jing Yao 0003, Xiaoyuan Yi, Xing Xie 0001 |
NeurIPS | 2 |
| 2024 | A Survey on Evaluation of Large Language ModelsabstractLarge language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey Yupeng Chang, Jindong Wang 0001, Yuan Wu 0002, Linyi Yang, Kaijie Zhu, Hao Chen 0102, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang 0003, Wei Ye 0004, Yue Zhang 0004, Yi Chang 0001, Philip S. Yu, Qiang Yang 0001, Xing Xie 0001 |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2023 | DuNST: Dual Noisy Self Training for Semi-Supervised Controllable Text GenerationabstractSelf-training (ST) has prospered again in language understanding by augmenting the finetuning of big pre-trained models when labeled data is insufficient.However, it remains challenging to incorporate ST into attributecontrollable language generation.Augmented only by self-generated pseudo text, generation models over-exploit the previously learned text space and fail to explore a larger one, suffering from a restricted generalization boundary and limited controllability.In this work, we propose DuNST, a novel ST framework to tackle these problems.DuNST jointly models text generation and classification as a dual process and further perturbs and escapes from the collapsed space by adding two kinds of flexible noise.In this way, our model could construct and utilize both pseudo text generated from given labels and pseudo labels predicted from available unlabeled text, which are gradually refined during the ST phase.Theoretically, we show that DuNST can be viewed as enhancing the exploration of the potentially larger real text space while maintaining exploitation, guaranteeing improved performance.Experiments on three controllable generation tasks show that DuNST significantly boosts control accuracy with comparable generation fluency and diversity against several strong baselines. Yuxi Feng, Xiaoyuan Yi, Xiting Wang, Laks V. S. Lakshmanan, Xing Xie 0001 |
ACL (1) | 2 |
| 2023 | ToViLaG: Your Visual-Language Generative Model is Also An EvildoerabstractWarning: this paper includes model outputs showing offensive content.Recent large-scale Visual-Language Generative Models (VLGMs) have achieved unprecedented improvement in multimodal image/text generation.However, these models might also generate toxic content, e.g., offensive text and pornography images, raising significant ethical risks.Despite exhaustive studies on toxic degeneration of language models, this problem remains largely unexplored within the context of visual-language generation.This work delves into the propensity for toxicity generation and susceptibility to toxic data across various VLGMs.For this purpose, we built ToViLaG, a dataset comprising 32K co-toxic/mono-toxic text-image pairs and 1K innocuous but evocative text that tends to stimulate toxicity.Furthermore, we propose WInToRe, a novel toxicity metric tailored to visual-language generation, which theoretically reflects different aspects of toxicity considering both input and output.On such a basis, we benchmarked the toxicity of a diverse spectrum of VLGMs and discovered that some models do more evil than expected while some are more vulnerable to infection, underscoring the necessity of VLGMs detoxification.Therefore, we develop an innovative information bottleneckbased detoxification method.Our method reduces toxicity while maintaining acceptable generation quality, providing a promising initial solution to this line of research.* Work done as an intern at MSRA mentored by X. Yi. Xinpeng Wang 0001, Xiaoyuan Yi, Han Jiang 0007, Shanlin Zhou, Zhihua Wei 0001, Xing Xie 0001 |
EMNLP | 2 |
| 2023 | Unified Detoxifying and Debiasing in Language Generation via Inference-time Adaptive Optimization
Zonghan Yang, Xiaoyuan Yi, Peng Li 0030, Yang Liu 0005, Xing Xie 0001 |
ICLR | 2 |
| 2023 | KEST: Kernel Distance Based Efficient Self-Training for Improving Controllable Text GenerationabstractSelf-training (ST) has come to fruition in language understanding tasks by producing pseudo labels, which reduces the labeling bottleneck of language model fine-tuning. Nevertheless, in facilitating semi-supervised controllable language generation, ST faces two key challenges. First, augmented by self-generated pseudo text, generation models tend to over-exploit the previously learned text distribution, suffering from mode collapse and poor generation diversity. Second, generating pseudo text in each iteration is time-consuming, severely decelerating the training process. In this work, we propose KEST, a novel and efficient self-training framework to handle these problems. KEST utilizes a kernel-based loss, rather than standard cross entropy, to learn from the soft pseudo text produced by a shared non-autoregressive generator. We demonstrate both theoretically and empirically that KEST can benefit from more diverse pseudo text in an efficient manner, which allows not only refining and exploiting the previously fitted distribution but also enhanced exploration towards a larger potential text space, providing a guarantee of improved performance. Experiments on three controllable generation tasks demonstrate that KEST significantly improves control accuracy while maintaining comparable text fluency and generation diversity against several strong baselines. Yuxi Feng, Xiaoyuan Yi, Laks V. S. Lakshmanan, Xing Xie 0001 |
IJCAI | 2 |
| 2022 | Evade the Trap of Mediocrity: Promoting Diversity and Novelty in Text Generation via Concentrating AttentionabstractRecently, powerful Transformer architectures have proven superior in generating high-quality sentences.Nevertheless, these models tend to produce dull high-frequency phrases, severely hurting the diversity and novelty of generated text.In this work, we dig into the intrinsic mechanism of this problem and found that sparser attention values in Transformer could improve diversity.To understand such a phenomenon, we first conduct both empirical and theoretical analysis and then attribute it to representation degeneration caused by the attentive mixture of the hidden states during training.We term this process the Trap of Mediocrity.To escape from such a trap, we introduce a novel attention regularization loss to control the sharpness of the attention distribution, which is transparent to model structures and can be easily implemented within 20 lines of python code.We prove that this method could be mathematically regarded as learning a Bayesian approximation of posterior attention.Experiments show that our method improved the diversity and novelty of the generated text while maintaining comparable quality on a variety of conditional and unconditional generation tasks. Wenhao Li 0003, Xiaoyuan Yi, Jinyi Hu, Maosong Sun 0001, Xing Xie 0001 |
EMNLP | 2 |
| 2022 | Clickbait Detection via Contrastive Variational Modelling of Text and LabelabstractClickbait refers to deliberately created sensational or deceptive text for tricking readers into clicking, which severely hurts the web ecosystem. With a growing number of clickbaits on social media, developing automatic detection methods becomes essential. Nonetheless, the performance of existing neural classifiers is limited due to the underutilization of small labelled datasets. Inspired by related pedagogy theories that learning to write can promote comprehension ability, we propose a novel Contrastive Variational Modelling (CVM) framework to exploit the labelled data better. CVM models the conditional distributions of text and clickbait labels by predicting labels from text and generating text from labels simultaneously with Variational AutoEncoder and further differentiates the learned spaces under each label by a mixed contrastive learning loss. In this way, CVM can capture more underlying textual properties and hence utilize label information to its full potential, boosting detection performance. We theoretically demonstrate CVM as learning a joint distribution of text, clickbait label, and latent variable. Experiments on three clickbait detection datasets show our method's robustness to inadequate and biased labels, outperforming several recent strong baselines. Xiaoyuan Yi, Jiarui Zhang 0002, Wenhao Li 0003, Xiting Wang, Xing Xie 0001 |
IJCAI | 1 |
| 2022 | Personalized Chit-Chat Generation for Recommendation Using External Chat CorporaabstractChit-chat has been shown effective in engaging users in human-computer interaction. We find with a user study that generating appropriate chit-chat for news articles can help expand user interest and increase the probability that a user reads a recommended news article. Based on this observation, we propose a method to generate personalized chit-chat for news recommendation. Different from existing methods for personalized text generation, our method only requires an external chat corpus obtained from an online forum, which can be disconnected from the recommendation dataset from both the user and item (news) perspectives. This is achieved by designing a weak supervision method for estimating users' personalized interest in a chit-chat post by transferring knowledge learned by a news recommendation model. Based on the method for estimating user interest, a reinforcement learning framework is proposed to generate personalized chit-chat. Extensive experiments, including the automatic offline evaluation and user studies, demonstrate the effectiveness of our method. Changyu Chen, Xiting Wang, Xiaoyuan Yi, Fangzhao Wu, Xing Xie 0001, Rui Yan 0001 |
KDD | 3 |
| 2022 | Fuse It More Deeply! A Variational Transformer with Layer-Wise Latent Variable Inference for Text GenerationabstractJinyi Hu, Xiaoyuan Yi, Wenhao Li, Maosong Sun, Xing Xie. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jinyi Hu, Xiaoyuan Yi, Wenhao Li 0003, Maosong Sun 0001, Xing Xie 0001 |
NAACL-HLT | 2 |
| 2022 | Self-explaining deep models with logic rule reasoningabstractWe present SELOR, a framework for integrating self-explaining capabilities into a given deep model to achieve both high prediction performance and human precision. By “human precision”, we refer to the degree to which humans agree with the reasons models provide for their predictions. Human precision affects user trust and allows users to collaborate closely with the model. We demonstrate that logic rule explanations naturally satisfy them with the expressive power required for good predictive performance. We then illustrate how to enable a deep model to predict and explain with logic rules. Our method does not require predefined logic rule sets or human annotations and can be learned efficiently and easily with widely-used deep learning modules in a differentiable way. Extensive experiments show that our method gives explanations closer to human decision logic than other methods while maintaining the performance of the deep learning model. SeungEon Lee 0001, Xiting Wang, Sungwon Han 0001, Xiaoyuan Yi, Xing Xie 0001, Meeyoung Cha |
NeurIPS | 4 |
| 2021 | Neural Quality Estimation with Multiple Hypotheses for Grammatical Error CorrectionabstractZhenghao Liu, Xiaoyuan Yi, Maosong Sun, Liner Yang, Tat-Seng Chua. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhenghao Liu 0001, Xiaoyuan Yi, Maosong Sun 0001, Liner Yang, Tat-Seng Chua |
NAACL-HLT | 2 |
| 2020 | MixPoet: Diverse Poetry Generation via Learning Controllable Mixed Latent SpaceabstractAs an essential step towards computer creativity, automatic poetry generation has gained increasing attention these years. Though recent neural models make prominent progress in some criteria of poetry quality, generated poems still suffer from the problem of poor diversity. Related literature researches show that different factors, such as life experience, historical background, etc., would influence composition styles of poets, which considerably contributes to the high diversity of human-authored poetry. Inspired by this, we propose MixPoet, a novel model that absorbs multiple factors to create various styles and promote diversity. Based on a semi-supervised variational autoencoder, our model disentangles the latent space into some subspaces, with each conditioned on one influence factor by adversarial training. In this way, the model learns a controllable latent variable to capture and mix generalized factor-related properties. Different factor mixtures lead to diverse styles and hence further differentiate generated poems from each other. Experiment results on Chinese poetry demonstrate that MixPoet improves both diversity and quality against three state-of-the-art models. Xiaoyuan Yi, Cheng Yang 0002, Wenhao Li 0003, Maosong Sun 0001 |
AAAI | 1 |
| 2020 | Text Style Transfer via Learning Style Instance Supported Latent SpaceabstractText style transfer pursues altering the style of a sentence while remaining its main content unchanged. Due to the lack of parallel corpora, most recent work focuses on unsupervised methods and has achieved noticeable progress. Nonetheless, the intractability of completely disentangling content from style for text leads to a contradiction of content preservation and style transfer accuracy. To address this problem, we propose a style instance supported method, StyIns. Instead of representing styles with embeddings or latent variables learned from single sentences, our model leverages the generative flow technique to extract underlying stylistic properties from multiple instances of each style, which form a more discriminative and expressive latent style space. By combining such a space with the attention-based structure, our model can better maintain the content and simultaneously achieve high transfer accuracy. Furthermore, the proposed method can be flexibly extended to semi-supervised learning so as to utilize available limited paired data. Experiments on three transfer tasks, sentiment modification, formality rephrasing, and poeticness generation, show that StyIns obtains a better balance between content and style, outperforming several recent baselines. Xiaoyuan Yi, Zhenghao Liu 0001, Wenhao Li 0003, Maosong Sun 0001 |
IJCAI | 1 |
| 2019 | Sentiment-Controllable Chinese Poetry GenerationabstractExpressing diverse sentiments is one of the main purposes of human poetry creation. Existing Chinese poetry generation models have made great progress in poetry quality, but they all neglected to endow generated poems with specific sentiments. Such defect leads to strong sentiment collapse or bias and thus hurts the diversity and semantics of generated poems. Meanwhile, there are few sentimental Chinese poetry resources for studying. To address this problem, we first collect a manually-labelled sentimental poetry corpus with fine-grained sentiment labels. Then we propose a novel semi-supervised conditional Variational Auto-Encoder model for sentiment-controllable poetry generation. Besides, since poetry is discourse-level text where the polarity and intensity of sentiment could transfer among lines, we incorporate a temporal module to capture sentiment transition patterns among different lines. Experimental results show our model can control the sentiment of not only a whole poem but also each line, and improve the poetry diversity against the state-of-the-art models without losing quality. Xiaoyuan Yi, Maosong Sun 0001, Wenhao Li 0003, Cheng Yang 0002, Zhipeng Guo 0001 |
IJCAI | 2 |
| 2018 | Chinese Poetry Generation with a Salient-Clue MechanismabstractAs a precious part of the human cultural heritage, Chinese poetry has influenced people for generations.Automatic poetry composition is a challenge for AI.In recent years, significant progress has been made in this area benefiting from the development of neural networks.However, the coherence in meaning, theme or even artistic conception for a generated poem as a whole still remains a big problem.In this paper, we propose a novel Salient-Clue mechanism for Chinese poetry generation.Different from previous work which tried to exploit all the context information, our model selects the most salient characters automatically from each so-far generated line to gradually form a salient clue, which is utilized to guide successive poem generation process so as to eliminate interruptions and improve coherence.Besides, our model can be flexibly extended to control the generated poem in different aspects, for example, poetry style, which further enhances the coherence.Experimental results show that our model is very effective, outperforming three strong baselines. Xiaoyuan Yi, Maosong Sun 0001 |
CoNLL | 1 |
| 2018 | Stylistic Chinese Poetry Generation via Unsupervised Style DisentanglementabstractThe ability to write diverse poems in different styles under the same poetic imagery is an important characteristic of human poetry writing.Most previous works on automatic Chinese poetry generation focused on improving the coherency among lines.Some work explored style transfer but suffered from expensive expert labeling of poem styles.In this paper, we target on stylistic poetry generation in a fully unsupervised manner for the first time.We propose a novel model which requires no supervised style labeling by incorporating mutual information, a concept in information theory, into modeling.Experimental results show that our model is able to generate stylistic poems without losing fluency and coherency. Cheng Yang 0002, Maosong Sun 0001, Xiaoyuan Yi, Wenhao Li 0003 |
EMNLP | 3 |
| 2018 | Automatic Poetry Generation with Mutual Reinforcement LearningabstractPoetry is one of the most beautiful forms of human language art.As a crucial step towards computer creativity, automatic poetry generation has drawn researchers' attention for decades.In recent years, some neural models have made remarkable progress in this task.However, they are all based on maximum likelihood estimation, which only learns common patterns of the corpus and results in lossevaluation mismatch.Human experts evaluate poetry in terms of some specific criteria, instead of word-level likelihood.To handle this problem, we directly model the criteria and use them as explicit rewards to guide gradient update by reinforcement learning, so as to motivate the model to pursue higher scores.Besides, inspired by writing theories, we propose a novel mutual reinforcement learning schema.We simultaneously train two learners (generators) which learn not only from the teacher (rewarder) but also from each other to further improve performance.We experiment on Chinese poetry.Based on a strong basic model, our method achieves better results and outperforms the current state-of-theart method. Xiaoyuan Yi, Maosong Sun 0001, Wenhao Li 0003 |
EMNLP | 1 |
| 2018 | Chinese Poetry Generation with a Working Memory ModelabstractAs an exquisite and concise literary form, poetry is a gem of human culture. Automatic poetry generation is an essential step towards computer creativity. In recent years, several neural models have been designed for this task. However, among lines of a whole poem, the coherence in meaning and topics still remains a big challenge. In this paper, inspired by the theoretical concept in cognitive psychology, we propose a novel Working Memory model for poetry generation. Different from previous methods, our model explicitly maintains topics and informative limited history in a neural memory. During the generation process, our model reads the most relevant parts from memory slots to generate the current line. After each line is generated, it writes the most salient parts of the previous line into memory slots. By dynamic manipulation of the memory, our model keeps a coherent information flow and learns to express each topic flexibly and naturally. We experiment on three different genres of Chinese poetry: quatrain, iambic and chinoiserie lyric. Both automatic and human evaluation results show that our model outperforms current state-of-the-art methods. Xiaoyuan Yi, Maosong Sun 0001, Zonghan Yang |
IJCAI | 1 |
| 2016 | Inferring users' emotions for human-mobile voice dialogue applicationsabstractIn this paper, we tackle the problem of inferring users' emotions in real-world Voice Dialogue Applications (VDAs, Siri1, Cortana2, etc.). We first conduct an investigation, indicating that besides the text information of users' queries, the acoustic information and query attributes are very important in inferring emotions in VDAs. To integrate the information above, we propose a Hybrid Emotion Inference Model (HEIM), which involves a Latent Dirichlet Allocation (LDA) to extract text features and a Long Short-Term Memory (LSTM) to model the acoustic features. To further improve accuracy, a Recurrent Autoencoder Guided by Query Attributes (RAGQA) which incorporates other emotion-related query attributes is proposed in HEIM to pre-train LSTM. The accuracy of HEIM on a data set collected from Sogou Voice Assistant3(Chinese Siri) containing 93,000 utterances achieves 75.2%, which outperforms state-of-the-art methods for 33.5–38.5%. Specifically, we discover that on average, the acoustic information enhances the performance for 46.6%, while query attributes further enhance the performance for 6.5%. Boya Wu, Jia Jia 0001, Tao He 0016, Xiaoyuan Yi, Yishuang Ning |
ICME | 5 |