VLDB 2026 Research / reviewers in the wild / expert
Letian Peng
dblp:303/0630
· DBLP profile ↗
13ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0001-5039-3496ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deriving Character Logic from Storyline as Codified Decision TreesabstractRole-playing (RP) agents rely on behavioral profiles to act consistently across diverse narrative contexts, yet existing profiles are largely unstructured, non-executable, and weakly validated, leading to brittle agent behavior.We propose Codified Decision Trees (CDT), a datadriven framework that induces an executable and interpretable decision structure from largescale narrative data.CDT represents behavioral profiles as a tree of conditional rules, where internal nodes correspond to validated scene conditions and leaves encode grounded behavioral statements, enabling deterministic retrieval of context-appropriate rules at execution time.The tree is learned by iteratively inducing candidate scene-action rules, validating them against data, and refining them through hierarchical specialization, yielding profiles that support transparent inspection and principled updates.Across multiple benchmarks, CDT substantially outperforms human-written profiles and prior profile induction methods on 85 characters across 16 artifacts, indicating that codified and validated behavioral representations lead to more reliable agent grounding.1 Letian Peng, Kun Zhou 0002, Longfei Yun, Yupeng Hou, Jingbo Shang |
ACL (1) | 1 |
| 2026 | C-World: A Computer Use Agent Environment CreatorabstractZiqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, Kun Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziqiao Xi, Letian Peng, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, Kun Zhou 0002 |
ACL (1) | 5 |
| 2026 | Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge InjectionabstractYuwei Zhang, Wenhao Yu, Shangbin Feng, Yifan Zhu, Letian Peng, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuwei Zhang 0001, Wenhao Yu 0002, Shangbin Feng, Letian Peng, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang |
ACL (1) | 5 |
| 2025 | Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's NestabstractMassive high-quality data, both pre-training raw texts and post-training annotations, have been carefully prepared to incubate advanced large language models (LLMs).In contrast, for information extraction (IE), pre-training data, such as BIO-tagged sequences, are hard to scale up.We show that IE models can act as free riders on LLM resources by reframing next-token prediction into extraction for tokens already present in the context.Specifically, our proposed next tokens extraction (NTE) paradigm learns a versatile IE model, Cuckoo 1 , with 102.6M extractive data converted from LLM's pre-training and post-training data.Under the few-shot setting, Cuckoo adapts effectively to traditional and complex instruction-following IE with better performance than existing pretrained IE models.As a free rider, Cuckoo can naturally evolve with the ongoing advancements in LLM data preparation, benefiting from improvements in LLM training pipelines without additional manual effort.2 Letian Peng, Zilong Wang 0002, Jingbo Shang |
ACL (1) | 1 |
| 2025 | Correlation and Navigation in the Vocabulary Key Representation Space of Language ModelsabstractLanguage model (LM) decoding is based on the next-token prediction (NTP) probability distribution. For neural LMs (e.g., Transformer-based), NTP distribution is
essentially a softmax-regularized dot product between an encoded input context
(query) and fixed vocabulary representations (keys). In this paper, we study the
effect of the key distribution on the NTP distribution, with a focus on whether
the similarity between keys will trigger spurious correlations in NTP. Through
knowledge-probing tasks, we show that in the NTP distribution, the few top-ranked
tokens are typically accurate. However, the middle-ranked prediction is highly biased
towards the tokens that are distributionally (not necessarily semantically) similar to
these top ones. For instance, if “P” is predicted as the top-1 token, “A”-“Z” will all
be ranked high in NTP, no matter whether they can lead to correct decoding results.
This hurts the sampling diversity and makes the sampling of correct, long-tail
results hopeless and noisy. We attempt to alleviate this issue via a novel in-context
method that iteratively pushes the query representation away from explored regions.
Specifically, we include the explored decoding results in the context and prompt
the LM to generate something else, which encourages the LM to produce a query
representation that has small dot products with explored keys. Experiments on
knowledge-probing tasks show that our method leads to efficient navigation away
from explored keys to correct new keys. We further extend our method to open-ended and chain-of-thought (for reasoning) generation. Experiment results show
that ICN contributes to better generation diversity and improved self-consistency
voting performance. Finally, we discuss potential training issues caused by the
fixed key space together with the challenges and possible ways to address them in
future research. Letian Peng, Chenyang An, Jingbo Shang |
ICLR | 1 |
| 2025 | Codifying Character Logic in Role-PlayingabstractThis paper introduces Codified Profiles for role-playing, a novel approach that represents character logic as structured, executable functions for behavioral decision-making.
Converted by large language model (LLM) from textual profiles, each codified profile defines a set of functions parse_by_scene(scene) that output multiple logic-grounded assertions according to scene, using both explicit control structures (e.g., if-then-else) and flexible check_condition(scene, question) functions where each question is a semantically meaningful prompt about the scene (e.g., "Is the character in danger?") discriminated by the role-playing LLM as true, false, or unknown.
This explicit representation offers three key advantages over traditional prompt-based textual profiles, which append character descriptions directly into text prompts:
(1) Persistence, by enforcing complete and consistent execution of character logic, rather than relying on the model's implicit reasoning;
(2) Updatability, through systematic inspection and revision of behavioral logic, which is difficult to track or debug in prompt-only approaches;
(3) Controllable Randomness, by supporting stochastic behavior directly within the logic, enabling fine-grained variability that prompting alone struggles to achieve.
To validate these advantages, we introduce a new benchmark constructed from 83 characters and 5,141 scenes curated from Fandom, using natural language inference (NLI)-based scoring to compare character responses against ground-truths.
Our experiments demonstrate the significant benefits of codified profiles in improving persistence, updatability, and behavioral diversity.
Notably, by offloading a significant portion of reasoning to preprocessing, codified profiles enable even 1B-parameter models to perform high-quality role-playing, providing an efficient, lightweight foundation for local deployment of role-play agents. Letian Peng, Jingbo Shang |
NeurIPS | 1 |
| 2024 | Learn from Failure: Fine-tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic ProvingabstractChenyang An, Zhibo Chen, Qihao Ye, Emily First, Letian Peng, Jiayun Zhang, Zihan Wang, Sorin Lerner, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenyang An, Zhibo Chen 0009, Qihao Ye, Emily First, Letian Peng, Jiayun Zhang, Zihan Wang 0001, Sorin Lerner, Jingbo Shang |
ACL (1) | 5 |
| 2024 | Answer is All You Need: Instruction-following Text Embedding via Answering the QuestionabstractLetian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa, Gaowen Liu, Zihan Wang, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Letian Peng, Yuwei Zhang 0001, Zilong Wang 0002, Jayanth Srinivasa, Gaowen Liu, Zihan Wang 0001, Jingbo Shang |
ACL (1) | 1 |
| 2024 | Incubating Text Classifiers Following User Instruction with Nothing but LLMabstractIn this paper, we aim to generate text classification data given arbitrary class definitions (i.e., user instruction), so one can train a text classifier without any human annotation or raw corpus.Recent advances in large language models (LLMs) lead to pioneer attempts to individually generate texts for each class via prompting.In this paper, we propose Incubator, the first framework that can handle complicated and even mutually dependent classes (e.g., "TED Talk given by Educator" and "Other").Specifically, our Incubator is a fine-tuned LLM that takes the instruction of all class definitions as input, and in each inference, it can jointly generate one sample for every class.First, we tune Incubator on the instruction-to-data mappings that we obtained from classification datasets and descriptions on Hugging Face together with in-context augmentation by GPT-4.To emphasize the uniformity and diversity in generations, we refine Incubator by fine-tuning with the cluster centers of semantic textual embeddings of the generated samples.We compare Incubator on various classification tasks with strong baselines such as direct LLM-based inference and training data generation by prompt engineering.Experiments show Incubator is able to (1) outperform previous methods on traditional benchmarks, (2) take label interdependency and user preference into consideration, and (3) enable logical text mining by incubating multiple classifiers. Letian Peng, Zilong Wang 0002, Jingbo Shang |
EMNLP | 1 |
| 2024 | Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text ClassificationabstractFor extremely weak-supervised text classification, pioneer research generates pseudo labels by mining texts similar to the class names from the raw corpus, which may end up with very limited or even no samples for the minority classes.Recent works have started to generate the relevant texts by prompting LLMs using the class names or definitions; however, there is a high risk that LLMs cannot generate indistribution (i.e., similar to the corpus where the text classifier will be applied) data, leading to ungeneralizable classifiers.In this paper, we combine the advantages of these two approaches and propose to bridge the gap via a novel framework, text grafting, which aims to obtain clean and near-distribution weak supervision for minority classes.Specifically, we first use LLM-based logits to mine masked templates from the raw corpus, which have a high potential for data synthesis into the target minority class.Then, the templates are filled by state-of-the-art LLMs to synthesize neardistribution texts falling into minority classes.Text grafting shows significant improvement over direct mining or synthesis on minority classes.We also use analysis and case studies to comprehend the property of text grafting. Letian Peng, Yi Gu 0002, Chengyu Dong, Zihan Wang 0001, Jingbo Shang |
EMNLP | 1 |
| 2024 | Semantics-Preserved Distortion for Personal Privacy Protection in Information Management
Jiajia Li 0005, Lu Yang 0008, Letian Peng, Shitou Zhang, Ping Wang 0028, Zuchao Li, Hai Zhao 0001 |
ICANN (5) | 3 |
| 2024 | Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playingabstractPersona-driven role-playing (PRP) aims to build AI characters that can respond to user queries by faithfully sticking with \emph{all} (factual) statements in persona documents.
Unfortunately, existing faithfulness criteria for PRP are limited to coarse-grained LLM-based scoring without a clear definition or formulation.
This paper presents a pioneering exploration to quantify PRP faithfulness evaluation as a fine-grained and explainable criterion, which also serves as a reliable reference for faithfulness optimization.
Our criterion first discriminates persona statements into \emph{active} and \emph{passive} constraints by identifying the query-statement relevance.
Then, we incorporate all constraints following the principle that the AI character's response should be (a) entailed by active constraints and (b) not contradicted by passive constraints.
We translate this principle mathematically into a novel Active-Passive-Constraint (APC) score, a constraint-wise sum of statement-to-response natural language inference (NLI) scores weighted by constraint-query relevance scores.
In practice, we build the APC scoring system by symbolically distilling small NLI and relevance discriminators (300M parameters) from GPT-4 for efficiency, and both show high consistency with GPT-4's discrimination.
We validate the quality of the APC score against human evaluation based on example personas with tens of statements, and the results show a high correlation.
As the APC score could faithfully reflect the PRP quality, we further leverage it as a reward system in direct preference optimization (DPO) for better AI characters.
Our experiments offer a fine-grained and explainable comparison between existing PRP techniques, revealing their advantages and limitations.
We further find APC-based DPO to be one of the most competitive techniques for sticking with all constraints and can be well incorporated with other techniques.
We then extend the scale of the experiments to real persons with hundreds of statements and reach a consistent conclusion.
Finally, we provide comprehensive analyses and case studies to support the effectiveness of APC and APC-based DPO. Letian Peng, Jingbo Shang |
NeurIPS | 1 |
| 2023 | iRe2f: Rethinking Effective Refinement in Language Structure Prediction via Efficient Iterative Retrospecting and ReasoningabstractRefinement plays a critical role in language structure prediction, a process that deals with complex situations such as structural edge interdependencies. Since language structure prediction usually modeled as graph parsing, typical refinement methods involve taking an initial parsing graph as input and refining it using language input and other relevant information. Intuitively, a refinement component, i.e., refiner, should be lightweight and efficient, as it is only responsible for correcting faults in the initial graph. However, current refiners add a significant burden to the parsing process due to their reliance on time-consuming encoding-decoding procedure on the language input and graph. To make the refiner more practical for real-world applications, this paper proposes a lightweight but effective iterative refinement framework, iRe^2f, based on iterative retrospecting and reasoning without involving the re-encoding process on the graph. iRe^2f iteratively refine the parsing graph based on interaction between graph and sequence and efficiently learns the shortcut to update the sequence and graph representations in each iteration. The shortcut is calculated based on the graph representation in the latest iteration. iRe^2f reduces the number of refinement parameters by 90% compared to the previous smallest refiner. Experiments on a variety of language structure prediction tasks show that iRe^2f performs comparably or better than current state-of-the-art refiners, with a significant increase in efficiency. Zuchao Li, Xingyi Guo, Letian Peng, Lefei Zhang, Hai Zhao 0001 |
IJCAI | 3 |