Chenglei Si

dblp:251/8778 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 12 since 2021
YearPublicationVenuePosition
2025 Contextual Experience Replay for Self-Improvement of Language Agents
abstract
Large language model (LLM) agents have been applied to sequential decision-making tasks such as web navigation, but without any environment-specific experiences, they often fail in these complex tasks.Moreover, current LLM agents are not designed to continually learn from past experiences during inference time, which could be crucial for them to gain these environment-specific experiences.To address this, we propose Contextual Experience Replay (CER), a training-free framework to enable efficient self-improvement for language agents in their context window.Specifically, CER accumulates and synthesizes past experiences into a dynamic memory buffer.These experiences encompass environment dynamics and common decision-making patterns, allowing the agents to retrieve and augment themselves with relevant knowledge in new tasks, enhancing their adaptability in complex environments.We evaluate CER on the challenging WEBARENA and VISUALWEBARENA benchmarks.On VISUALWEBARENA, CER achieves competitive performance of 31.9%.On WEBARENA, CER also gets a competitive average success rate of 36.7%, relatively improving the success rate of the GPT-4o agent baseline by 51.0%.We also conduct a comprehensive analysis on it to prove its efficiency, validity and understand it better.
Yitao Liu, Chenglei Si, Karthik Narasimhan, Shunyu Yao 0006
ACL (1)2
2025 Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
abstract
Recent advancements in large language models (LLMs) have sparked optimism about their potential to accelerate scientific discovery, with a growing number of works proposing research agents that autonomously generate and validate new ideas. Despite this, no evaluations have shown that LLM systems can take the very first step of producing novel, expert-level ideas, let alone perform the entire research process. We address this by establishing an experimental design that evaluates research idea generation while controlling for confounders and performs the first comparison between expert NLP researchers and an LLM ideation agent. By recruiting over 100 NLP researchers to write novel ideas and blind reviews of both LLM and human ideas, we obtain the first statistically significant conclusion on current LLM capabilities for research ideation: we find LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility. Studying our agent baselines closely, we identify open problems in building and evaluating research agents, including failures of LLM self-evaluation and their lack of diversity in generation.
Chenglei Si, Diyi Yang, Tatsunori B. Hashimoto
ICLR1
2025 Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
abstract
Chenglei Si, Yanzhe Zhang, Ryan Li, Zhengyuan Yang, Ruibo Liu, Diyi Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Chenglei Si, Zhengyuan Yang, Ruibo Liu, Diyi Yang
NAACL (Long Papers)1
2025 Position: Towards Bidirectional Human-AI Alignment
abstract
Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the research community should explicitly define and critically reflect on "alignment" to account for the bidirectional and dynamic relationship between humans and AI. Through a systematic review of over 400 papers spanning HCI, NLP, ML, and more, we examine how alignment is currently defined and operationalized. Building on this analysis, we introduce the Bidirectional Human-AI Alignment framework, which not only incorporates traditional efforts to align AI with human values but also introduces the critical, underexplored dimension of aligning humans with AI – supporting cognitive, behavioral, and societal adaptation to rapidly advancing AI technologies. Our findings reveal significant gaps in current literature, especially in long-term interaction design, human value modeling, and mutual understanding. We conclude with three central challenges and actionable recommendations to guide future research toward more nuanced, reciprocal, and human-AI alignment approaches.
Hua Shen 0005, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Savvas Petridis, Yi-Hao Peng, Li Qiwei, Chenglei Si, Yutong Xie 0007, Jeffrey P. Bigham, Frank Bentley, Joyce Y. Chai, Zachary C. Lipton, Qiaozhu Mei, Michael Terry, Diyi Yang, Meredith Ringel Morris, Paul Resnick, David Jurgens
NeurIPS10
2025 Predicting Empirical AI Research Outcomes with Language Models
abstract
Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks. We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs \textit{published after our base model's cut-off date} for testing, and 6,000 pairs for training. We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with. In the NLP domain, our system beats human experts by a large margin (64.4\% v.s. 48.9\%). On the full test set, our system achieves 77\% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation. We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests. Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent. Our system achieves 63.6\% accuracy, demonstrating its potential as a reward model for improving idea generation models. Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research.
Jiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 0001, Shi Feng 0005
NeurIPS2
2024 Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly Wrong
abstract
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao 0013, Shi Feng 0005, Hal Daumé III, Jordan L. Boyd-Graber
NAACL-HLT1
2023 Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations
abstract
In-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood.We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of underspecified demonstrations in which two features are equally predictive of the labels.First, we characterize the feature biases of GPT-3 models by constructing underspecified demonstrations from a range of NLP datasets and feature combinations.We find that LLMs exhibit clear feature biases-for example, demonstrating a strong bias to predict labels according to sentiment rather than shallow lexical features, like punctuation.Second, we evaluate the effect of different interventions that are designed to impose an inductive bias in favor of a particular feature, such as adding a natural language instruction or using semantically relevant label words.We find that, while many interventions can influence the learner to prefer a particular feature, it can be difficult to overcome strong prior biases.Overall, our results provide a broader picture of the types of features that ICL may be more likely to exploit and how to impose inductive biases that are better aligned with the intended task. 1
Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng 0005, Danqi Chen 0001, He He 0001
ACL (1)1
2023 READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input Noises
abstract
For many real-world applications, the usergenerated inputs usually contain various noises due to speech recognition errors caused by linguistic variations 1 or typographical errors (typos).Thus, it is crucial to test model performance on data with realistic input noises to ensure robustness and fairness.However, little study has been done to construct such benchmarks for Chinese, where various languagespecific input noises happen in the real world.In order to fill this important gap, we construct READIN: a Chinese multi-task benchmark with REalistic And Diverse Input Noises.READIN contains four diverse tasks and requests annotators to re-enter the original test data with two commonly used Chinese input methods: Pinyin input and speech input.We designed our annotation pipeline to maximize diversity, for example by instructing the annotators to use diverse input method editors (IMEs) for keyboard noises and recruiting speakers from diverse dialectical groups for speech noises.We experiment with a series of strong pretrained language models as well as robust training methods, we find that these models often suffer significant performance drops on READIN even with robustness methods like data augmentation.As the first large-scale attempt in creating a benchmark with noises geared towards user-generated inputs, we believe that READIN serves as an important complement to existing Chinese NLP benchmarks.The source code and dataset can be obtained from https://github.com/ thunlp/READIN.
Chenglei Si, Zhengyan Zhang, Yingfa Chen, Xiaozhi Wang, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)1
2023 Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition
abstract
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Kost, Christopher Carnahan, Jordan Boyd-Graber. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, Jordan L. Boyd-Graber
EMNLP5
2023 Prompting GPT-3 To Be Reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jordan L. Boyd-Graber
ICLR1
2023 Sub-Character Tokenization for Chinese Pretrained Language Models
abstract
Abstract Tokenization is fundamental to pretrained language models (PLMs). Existing tokenization methods for Chinese PLMs typically treat each character as an indivisible token. However, they ignore the unique feature of the Chinese writing system where additional linguistic information exists below the character level, i.e., at the sub-character level. To utilize such information, we propose sub-character (SubChar for short) tokenization. Specifically, we first encode the input text by converting each Chinese character into a short sequence based on its glyph or pronunciation, and then construct the vocabulary based on the encoded text with sub-word segmentation. Experimental results show that SubChar tokenizers have two main advantages over existing tokenizers: 1) They can tokenize inputs into much shorter sequences, thus improving the computational efficiency. 2) Pronunciation-based SubChar tokenizers can encode Chinese homophones into the same transliteration sequences and produce the same tokenization output, hence being robust to homophone typos. At the same time, models trained with SubChar tokenizers perform competitively on downstream tasks. We release our code and models at https://github.com/thunlp/SubCharTokenization to facilitate future work.
Chenglei Si, Zhengyan Zhang, Yingfa Chen, Fanchao Qi, Xiaozhi Wang, Zhiyuan Liu 0001, Yasheng Wang, Qun Liu 0001, Maosong Sun 0001
Trans. Assoc. Comput. Linguistics1
2021 What's in a Name? Answer Equivalence For Open-Domain Question Answering
abstract
A flaw in QA evaluation is that annotations often only provide one gold answer.Thus, model predictions semantically equivalent to the answer but superficially different are considered incorrect.This work explores mining alias entities from knowledge bases and using them as additional gold answers (i.e., equivalent answers).We incorporate answers for two settings: evaluation with additional answers and model training with equivalent answers.We analyse three QA benchmarks: Natural Questions, TriviaQA and SQuAD.Answer expansion increases the exact match score on all datasets for evaluation, while incorporating it helps model training over real-world datasets.We ensure the additional answers are valid through a human post hoc evaluation. 1
Chenglei Si, Chen Zhao 0013, Jordan L. Boyd-Graber
EMNLP (1)1
2020 CharBERT: Character-aware Pre-trained Language Model
abstract
Most pre-trained language models (PLMs) construct word representations at subword level with Byte-Pair Encoding (BPE) or its variations, by which OOV (out-of-vocab) words are almost avoidable.However, those methods split a word into subword units and make the representation incomplete and fragile.In this paper, we propose a character-aware pre-trained language model named CharBERT improving on the previous methods (such as BERT, RoBERTa) to tackle these problems.We first construct the contextual word embedding for each token from the sequential character representations, then fuse the representations of characters and the subword representations by a novel heterogeneous interaction module.We also propose a new pre-training task named NLM (Noisy LM) for unsupervised character representation learning.We evaluate our method on question answering, sequence labeling, and text classification tasks, both on the original datasets and adversarial misspelling test sets.The experimental results show that our method can significantly improve the performance and robustness of PLMs simultaneously.Pretrained models, evaluation sets, and code are available at https
Yiming Cui 0001, Chenglei Si, Ting Liu 0001, Shijin Wang 0001
COLING3