Yekun Chai

dblp:252/0188 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Curiosity-Driven Reinforcement Learning from Human Feedback
abstract
Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity and alignment quality remains a significant challenge. Drawing inspiration from curiosity-driven exploration in reinforcement learning, we introduce curiosity-driven RLHF (CD-RLHF), a framework that incorporates intrinsic rewards for novel states, alongside traditional sparse extrinsic rewards, to optimize both output diversity and alignment quality. We demonstrate the effectiveness of CD-RLHF through extensive experiments on a range of tasks, including text summarization and instruction following. Our approach achieves significant gains in diversity on multiple diversity-oriented metrics while maintaining alignment with human preferences comparable to standard RLHF. We will make our code publicly available.
Yekun Chai, Shuohuan Wang, Hua Wu 0003, Haifeng Wang 0001
ACL (1)2
2025 Understanding Subword Compositionality of Large Language Models
abstract
Large language models (LLMs) take sequences of subwords as input, requiring them to effective compose subword representations into meaningful word-level representations.In this paper, we present a comprehensive set of experiments to probe how LLMs compose subword information, focusing on three key aspects: structural similarity, semantic decomposability, and form retention.Our analysis of the experiments suggests that five LLM families can be classified into three distinct groups, likely reflecting difference in their underlying composition strategies.Specifically, we observe (i) three distinct patterns in the evolution of structural similarity between subword compositions and whole-word representations across layers; (ii) great performance when probing layer by layer their sensitivity to semantic decompositionality; and (iii) three distinct patterns when probing sensitivity to formal features, e.g., character sequence length.These findings provide valuable insights into the compositional dynamics of LLMs and highlight different compositional pattens in how LLMs encode and integrate subword information.
Qiwei Peng 0003, Yekun Chai, Anders Søgaard
EMNLP2
2025 Debiasing Multilingual LLMs in Cross-lingual Latent Space
abstract
Debiasing techniques such as SentDebias aim to reduce bias in large language models (LLMs).Previous studies have evaluated their cross-lingual transferability by directly applying these methods to LLM representations, revealing their limited effectiveness across languages.In this work, we therefore propose to perform debiasing in a joint latent space rather than directly on LLM representations.We construct a well-aligned cross-lingual latent space using an autoencoder trained on parallel TED talk scripts.Our experiments with Aya-expanse and two debiasing techniques across four languages (English, French, German, Dutch) demonstrate that a) autoencoders effectively construct a well-aligned cross-lingual latent space, and b) applying debiasing techniques in the learned cross-lingual latent space significantly improves both the overall debiasing performance and cross-lingual transferability.
Qiwei Peng 0003, Guimin Hu, Yekun Chai, Anders Søgaard
EMNLP3
2025 CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages
abstract
Code-mixing, the practice of switching between languages within a conversation, poses unique challenges for traditional NLP.Existing benchmarks are limited by their narrow language pairs and tasks, failing to adequately assess large language models' (LLMs) codemixing abilities.Despite the recognized importance of code-mixing for multilingual users, research on LLMs in this context remains sparse.Additionally, current techniques for synthesizing code-mixed data are underdeveloped to generate code-mixing.In response, we introduce CodeMixBench, a comprehensive benchmark covering eight tasks, including three specific to LLMs and five traditional NLP tasks, and 18 languages across seven language families.We also propose a new method for generating largescale synthetic code-mixed texts by combining word substitution with GPT-4 prompting.Our evaluation reveals consistent underperformance of LLMs on code-mixed datasets involving different language families.Enhancements in training data size, model scale, and fewshot learning could improve their performance.The code and dataset are available at https:// github.com/Jeromeyluck/CodeMixBench.
Yilun Yang, Yekun Chai
EMNLP2
2025 MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
abstract
Reinforcement learning from human feedback (RLHF) has demonstrated effectiveness in aligning large language models (LLMs) with human preferences. However, token-level RLHF suffers from the credit assignment problem over long sequences, where delayed rewards make it challenging for the model to discern which actions contributed to preferred outcomes. This hinders learning efficiency and slows convergence.In this paper, we propose MA-RLHF, a simple yet effective RLHF framework that incorporates macro actions --- sequences of tokens or higher-level language constructs --- into the learning process. By operating at higher level of abstraction, our approach reduces the temporal distance between actions and rewards, facilitating faster and more accurate credit assignment. This results in more stable policy gradient estimates and enhances learning efficiency within each episode, all without increasing computational complexity during training or inference. We validate our approach through extensive experiments across various model sizes and tasks, including text summarization, dialogue generation, question answering, and program synthesis. Our method achieves substantial performance improvements over standard RLHF, with performance gains of up to 30\% in text summarization and code generation, 18\% in dialogue, and 8\% in question answering tasks. Notably, our approach reaches parity with vanilla RLHF $1.7 \sim 2$ times faster in terms of training time and continues to outperform it with further training. We make our code and data publicly available at \url{https://github.com/ernie-research/MA-RLHF}.
Yekun Chai, Huang Fang, Shuohuan Wang, Hua Wu 0003
ICLR1
2024 HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
abstract
Large language models (LLMs) have made significant progress in generating codes from textual prompts. However, existing benchmarks have mainly concentrated on translating English prompts to multilingual codes or have been constrained to very limited natural languages (NLs). These benchmarks have overlooked the vast landscape of massively multilingual NL to multilingual code, leaving a critical gap in the evaluation of multilingual LLMs. In response, we introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency. HumanEval-XL establishes connections between 23 NLs and 12 programming languages (PLs), and comprises of a collection of 22,080 prompts with an average of 8.33 test cases. By ensuring parallel data across multiple NLs and PLs, HumanEval-XL offers a comprehensive evaluation platform for multilingual LLMs, allowing the assessment of the understanding of different NLs. Our work serves as a pioneering step towards filling the void in evaluating NL generalization in the area of multilingual code generation. We make our evaluation code and data publicly available at https://github.com/FloatAI/HumanEval-XL.
Qiwei Peng 0002, Yekun Chai, Xuhong Li 0002
LREC/COLING2
2024 On Training Data Influence of GPT Models
abstract
Amidst the rapid advancements in generative language models, the investigation of how training data shapes the performance of GPT models is still emerging.This paper presents GPTfluence, a novel approach that leverages a featurized simulation to assess the impact of training examples on the training dynamics of GPT models.Our approach not only traces the influence of individual training instances on performance trajectories, such as loss and other key metrics, on targeted test points but also enables a comprehensive comparison with existing methods across various training scenarios in GPT models, ranging from 14 million to 2.8 billion parameters, across a range of downstream tasks.Contrary to earlier methods that struggle with generalization to new data, GPTfluence introduces a parameterized simulation of training dynamics, demonstrating robust generalization capabilities to unseen training data.This adaptability is evident across both fine-tuning and instruction-tuning scenarios, spanning tasks in natural language understanding and generation.We make our
Yekun Chai, Qingyi Liu, Shuohuan Wang, Yu Sun 0004, Qiwei Peng 0002, Hua Wu 0003
EMNLP1
2024 Autoregressive Pre-Training on Pixels and Texts
abstract
The integration of visual and textual information represents a promising direction in the advancement of language models.In this paper, we explore the dual modality of language-both visual and textual-within an autoregressive framework, pre-trained on both document images and texts.Our method employs a multimodal training strategy, utilizing visual data through next patch prediction with a regression head and/or textual data through next token prediction with a classification head.We focus on understanding the interaction between these two modalities and their combined impact on model performance.Our extensive evaluation across a wide range of benchmarks shows that incorporating both visual and textual data significantly improves the performance of pixel-based language models.Remarkably, we find that a unidirectional pixelbased model trained solely on visual data can achieve comparable results to state-of-the-art bidirectional models on several language understanding tasks.This work uncovers the untapped potential of integrating visual and textual modalities for more effective language modeling.We release our code, data, and model checkpoints at
Yekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang, Hua Wu 0003
EMNLP1
2024 Tool-Augmented Reward Modeling
abstract
Reward modeling (*a.k.a.*, preference modeling) is instrumental for aligning large language models with human preferences, particularly within the context of reinforcement learning from human feedback (RLHF). While conventional reward models (RMs) have exhibited remarkable scalability, they oft struggle with fundamental functionality such as arithmetic computation, code execution, and factual lookup. In this paper, we propose a tool-augmented preference modeling approach, named Themis, to address these limitations by empowering RMs with access to external environments, including calculators and search engines. This approach not only fosters synergy between tool utilization and reward grading but also enhances interpretive capacity and scoring reliability. Our study delves into the integration of external tools into RMs, enabling them to interact with diverse external sources and construct task-specific tool engagement and reasoning traces in an autoregressive manner. We validate our approach across a wide range of domains, incorporating seven distinct external tools. Our experimental results demonstrate a noteworthy overall improvement of 17.7% across eight tasks in preference ranking. Furthermore, our approach outperforms Gopher 280B by 7.3% on TruthfulQA task in zero-shot evaluation. In human evaluations, RLHF trained with Themis attains an average win rate of 32% when compared to baselines across four distinct tasks. Additionally, we provide a comprehensive collection of tool-related RM datasets, incorporating data from seven distinct tool APIs, totaling 15,000 instances. We have made the code, data, and model checkpoints publicly available to facilitate and inspire further research advancements (https://github.com/ernie-research/Tool-Augmented-Reward-Model).
Lei Li 0040, Yekun Chai, Shuohuan Wang, Yu Sun 0004, Hao Tian 0005, Ningyu Zhang 0001, Hua Wu 0003
ICLR2
2024 GiLOT: Interpreting Generative Language Models via Optimal Transport
abstract
While large language models (LLMs) surge with the rise of generative AI, algorithms to explain LLMs highly desire. Existing feature attribution methods adequate for discriminative language models like BERT often fail to deliver faithful explanations for LLMs, primarily due to two issues: (1) For every specific prediction, the LLM outputs a probability distribution over the vocabulary–a large number of tokens with unequal semantic distance; (2) As an autoregressive language model, the LLM handles input tokens while generating a sequence of probability distributions of various tokens. To address above two challenges, this work proposes GiLOT that leverages Optimal Transport to measure the distributional change of all possible generated sequences upon the absence of every input token, while taking into account the tokens’ similarity, so as to faithfully estimate feature attribution for LLMs. We have carried out extensive experiments on top of Llama families and their fine-tuned derivatives across various scales to validate the effectiveness of GiLOT for estimating the input attributions. The results show that GiLOT outperforms existing solutions on a number of faithfulness metrics under fair comparison settings. Source code is publicly available at https://github.com/holyseven/GiLOT.
Xuhong Li 0002, Yekun Chai, Haoyi Xiong
ICML3
2023 Improved Training Of Mixture-Of-Experts Language GANs
abstract
Despite the dramatic success in image generation, Generative Adversarial Networks (GANs) still face great challenges in text generation. The difficulty in generator training arises from the limited representation capacity and uninformative learning signals obtained from the discriminator. In this work, we (1) first empirically show that the multi-generator approach is able to enhance the representation capacity of the generator for sequence GANs and (2) harness the Feature Statistics Alignment (FSA) paradigm to render fine-grained learning signals to advance the generator training. Specifically, FSA forces the mean statistics of the distribution of fake data to approach that of real samples as close as possible in the finite-dimensional feature space. Empirical study on synthetic and real benchmarks shows the superior performance in quantitative evaluation and demonstrates the effectiveness of our approach to adversarial text generation.
Yekun Chai, Qiyue Yin, Junge Zhang
ICASSP1
2023 Neural Text Classification by Jointly Learning to Cluster and Align
abstract
Distributional text clustering delivers semantically informative representations and captures the relevance between each word and semantic clustering centroids. We extend the neural text clustering approach to text classification tasks by inducing cluster centers via a variational autoencoder and interacting with distributional word embeddings, to enrich the text representation and measure the relatedness between tokens and each learnable cluster centroid. The proposed method jointly learns word clustering centroids and cluster-token alignments, achieving competitive results on multiple benchmark datasets and proving that the proposed cluster-token alignment mechanism is favorable to text classification. Notably, the learned text representations are well-clustered, which matches the ground-truth categories. Experimental results show that our model can also improve the classification performance on top of BERT representations. To the best of our knowledge, we are the first adopting the variational autoencoder to update clustering centroids for text classification.
Yekun Chai, Qiyue Yin, Junge Zhang
IJCNN1
2023 M4: A Unified XAI Benchmark for Faithfulness Evaluation of Feature Attribution Methods across Metrics, Modalities and Models
Xuhong Li 0002, Mengnan Du, Yekun Chai, Himabindu Lakkaraju, Haoyi Xiong
NeurIPS4
2022 Predicate-Argument Based Bi-Encoder for Paraphrase Identification
abstract
Paraphrase identification involves identifying whether a pair of sentences express the same or similar meanings.While cross-encoders have achieved high performances across several benchmarks, bi-encoders such as SBERT have been widely applied to sentence pair tasks.They exhibit substantially lower computation complexity and are better suited to symmetric tasks.In this work, we adopt a biencoder approach to the paraphrase identification task, and investigate the impact of explicitly incorporating predicate-argument information into SBERT through weighted aggregation.Experiments on six paraphrase identification datasets demonstrate that, with a minimal increase in parameters, the proposed model is able to outperform SBERT/SRoBERTa significantly.Further, ablation studies reveal that the predicate-argument based component plays a significant role in the performance gain.
Qiwei Peng 0002, David J. Weir, Julie Weeds, Yekun Chai
ACL (1)4
2020 Highway Transformer: Self-Gating Enhanced Self-Attentive Networks
abstract
Self-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, standing on the multiheaded dot product attention by attending to all the global contexts at different locations.Through a pseudo information highway, we introduce a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations.The subsidiary content-based SDU gates allow for the information flow of modulated latent embeddings through skipped connections, leading to a clear margin of convergence speed with gradient descent algorithms.We may unveil the role of gating mechanism to aid in the contextbased Transformer modules, with hypothesizing that SDU gates, especially on shallow layers, could push it faster to step towards suboptimal points during the optimization process.
Yekun Chai, Jin Shuo, Xinwen Hou
ACL1
2019 Exponential Moving Averaged Q-Network for DDPG
Xiangxiang Shen, Chuanhuan Yin, Yekun Chai, Xinwen Hou
PRCV (1)3