VLDB 2026 Research / reviewers in the wild / expert
Yu Bai 0018
dblp:03/6325-18
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0001-2036-0789ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying and Analyzing Performance-Critical Tokens in Large Language ModelsabstractIn-context learning (ICL) has emerged as an effective solution for few-shot learning with large language models (LLMs). However, how LLMs leverage demonstrations to specify a task and learn a corresponding computational function through ICL is underexplored. Drawing from the way humans learn from content-label mappings in demonstrations, we categorize the tokens in an ICL prompt into content, stopword, and template tokens. Our goal is to identify the types of tokens whose representations directly influence LLM's performance, a property we refer to as being performance-critical. By ablating representations from the attention of the test example, we find that the representations of informative content tokens have less influence on performance compared to template and stopword tokens, which contrasts with the human attention to informative words. We give evidence that the representations of performance-critical tokens aggregate information from the content tokens. Moreover, we demonstrate experimentally that lexical meaning, repetition, and structural cues are the main distinguishing characteristics of these tokens. Our work sheds light on how LLMs learn to perform tasks from demonstrations and deepens our understanding of the roles different types of tokens play in LLMs. Yu Bai 0018, Heyan Huang, Cesare Spinoso Di Piano, Sanxing Chen, Marc-Antoine Rondeau, Yang Gao 0016, Jackie Chi Kit Cheung |
AAAI | 1 |
| 2026 | EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational ScenariosabstractBin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yu Bai 0018, Huashan Sun, Yiguan Lin, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Yang Gao 0016, Heyan Huang |
ACL (1) | 2 |
| 2026 | Fusing generation task data to enhance generalization of LLMs in zero-shot relational triplet extraction
Xingpeng Si, Yuhao Ye, Yuming Shang, Mengyuan Cao, Yu Bai 0018, Jiawei Li 0020, Yang Gao 0016 |
Expert Syst. Appl. | 5 |
| 2026 | Dynamic token halting for efficient abstractive summarization with importance-aware regularization
Heyan Huang, Yu Bai 0018, Yang Gao 0016, Minpeng Liao |
Knowl. Based Syst. | 2 |
| 2024 | Fundamental Capabilities of Large Language Models and their Applications in Domain Scenarios: A SurveyabstractJiawei Li, Yizhe Yang, Yu Bai, Xiaofeng Zhou, Yinghao Li, Huashan Sun, Yuhang Liu, Xingpeng Si, Yuhao Ye, Yixiao Wu, Yiguan Lin, Bin Xu, Bowen Ren, Chong Feng, Yang Gao, Heyan Huang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jiawei Li 0020, Yizhe Yang, Yu Bai 0018, Xiaofeng Zhou 0004, Huashan Sun, Xingpeng Si, Yuhao Ye, Yixiao Wu, Yiguan Lin, Ren Bowen, Chong Feng 0001, Yang Gao 0016, Heyan Huang |
ACL (1) | 3 |
| 2024 | CItruS: Chunked Instruction-aware State Eviction for Long Sequence ModelingabstractLong sequence modeling has gained broad interest as large language models (LLMs) continue to advance.Recent research has identified that a large portion of hidden states within the key-value caches of Transformer models can be discarded (also termed evicted) without affecting the perplexity performance in generating long sequences.However, we show that these methods, despite preserving perplexity performance, often drop information that is important for solving downstream tasks, a problem which we call information neglect.To address this issue, we introduce Chunked Instruction-aware State Eviction (CItruS), a novel modeling technique that integrates the attention preferences useful for a downstream task into the eviction process of hidden states.In addition, we design a method for chunked sequence processing to further improve efficiency.Our training-free method exhibits superior performance on long sequence comprehension and retrieval tasks over several strong baselines under the same memory budget, while preserving language modeling perplexity.The code and data have been released at https: //github.com/ybai-nlp/CItruS. Yu Bai 0018, Xiyuan Zou, Heyan Huang, Sanxing Chen, Marc-Antoine Rondeau, Yang Gao 0016, Jackie Chi Kit Cheung |
EMNLP | 1 |
| 2024 | Can Pretrained English Language Models Benefit Non-English NLP Systems in Low-Resource Scenarios?abstractPretrained language models have achieved great success in a wide range of natural language processing (NLP) problems, because they learn language representations from large-scale text corpora and can adapt to downstream tasks by finetuning them on annotated task data. However, such success relies on both large-scale text and annotated data, so the lack of training data is a major practical problem for many languages, especially low-resource languages. In this paper, we explore whether a pretrained English language model can benefit non-English NLP systems in low-resource scenarios, i.e., with limited text corpora or annotated data. To achieve this, we first propose cross-lingual knowledge transfer methods and then validate our methods in low-resource scenarios. Specifically, our cross-lingual knowledge transfer methods are applied in the training stages of language model pretraining or downstream finetuning. At the two stages, the methods are designed for the transfer of upstream general knowledge or downstream task-specific knowledge, respectively. In the experiments, we perform pretraining and finetuning with limited non-English data to simulate the low-resource scenarios. We evaluate our methods on ten downstream tasks over a wide range of languages, and present systematic comparisons among various knowledge transfer methods. Experimental results show that our methods successfully leverage a pretrained English language model to improve task performance in other languages. Besides, we demonstrate the multilinguality of the English language model in various application scenarios. Our findings imply the possibility to improve low-resource-language NLP systems with large-scale English language models. Zewen Chi, Heyan Huang, Yu Bai 0018, Xiaoyan Gao 0001, Xianling Mao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | PSP: Pre-trained Soft Prompts for Few-Shot Abstractive SummarizationabstractFew-shot abstractive summarization has become a challenging task in natural language generation. To support it, we developed a novel soft prompts architecture coupled with a prompt pre-training plus prompt fine-tuning paradigm, which is effective and tunes only extremely light parameters. To meet the structure of the generation models, the soft prompts comprise continuous input embeddings across an encoder and a decoder. Importantly, a new inner-prompt placed in the text is introduced to capture document-level information. The aim is to devote attention to understanding the document that better prompts the model to generate document-related content. In the training process, the prompt pre-training with self-supervised pseudo-data firstly teaches the model basic summarizing capability. Then, with few-shot examples, only the designed lightweight soft prompts are fine-tuned. Experimental results on the CNN/DailyMail and XSum datasets show that our method, with only 0.1% of the parameters, outperforms full-model tuning where all model parameters are tuned. It also surpasses Prompt Tuning by a large margin and delivers competitive results against Prefix-Tuning with 3% of the parameters. Yang Gao 0016, Yu Bai 0018, Jiawei Li 0020, Yinan Hu, Heyan Huang, Boxing Chen |
COLING | 3 |
| 2022 | Stage-wise Stylistic Headline Generation: Style Generation and Summarized Content InsertionabstractA quality headline with a high click-rate should not only summarize the content of an article, but also reflect a style that attracts users. Such demand has drawn rising attention to the task of stylistic headline generation (SHG). An intuitive method is to first generate plain headlines leveraged by document-headline parallel data then transfer them to a target style. However, this inevitably suffers from error propagation. Therefore, to unify the two sub-tasks and explicitly decompose style-relevant attributes and summarize content, we propose an end-to-end stage-wise SHG model containing the style generation component and the content insertion component, where the former generates stylistic-relevant intermediate outputs and the latter receives these outputs then inserts the summarized content. The intermediate outputs are observable, making the style generation easy to control. Our system is comprehensively evaluated by both quantitative and qualitative metrics, and it achieves state-of-the-art results in SHG over three different stylistic datasets. Jiaao Zhan, Yang Gao 0016, Yu Bai 0018, Qianhui Liu |
IJCAI | 3 |
| 2022 | Unifying Cross-lingual Summarization and Machine Translation with Compression RateabstractCross-Lingual Summarization (CLS) is a task that extracts important information from a source document and summarizes it into a summary in another language. It is a challenging task that requires a system to understand, summarize, and translate at the same time, making it highly related to Monolingual Summarization (MS) and Machine Translation (MT). In practice, the training resources for Machine Translation are far more than that for cross-lingual and monolingual summarization. Thus incorporating the Machine Translation corpus into CLS would be beneficial for its performance. However, the present work only leverages a simple multi-task framework to bring Machine Translation in, lacking deeper exploration. Yu Bai 0018, Heyan Huang, Kai Fan 0002, Yang Gao 0016, Jiaao Zhan, Zewen Chi, Boxing Chen |
SIGIR | 1 |
| 2021 | Exploring Explainable Selection to Control Abstractive SummarizationabstractLike humans, document summarization models can interpret a document’s contents in a number of ways. Unfortunately, the neural models of today are largely black boxes that provide little explanation of how or why they generated a summary in the way they did. Therefore, to begin prying open the black box and to inject a level of control into the substance of the final summary, we developed a novel select-and-generate framework that focuses on explainability. By revealing the latent centrality and interactions between sentences, along with scores for novelty and relevance, users are given a window into the choices a model is making and an opportunity to guide those choices in a more desirable direction. A novel pair-wise matrix captures the sentence interactions, centrality and attribute scores, and a mask with tunable attribute thresholds allows the user to control which sentences are likely to be included in the extraction. A sentence-deployed attention mechanism in the abstractor ensures the final summary emphasizes the desired content. Additionally, the encoder is adaptable, supporting both Transformer- and BERT-based configurations. In a series of experiments assessed with ROUGE metrics and two human evaluations, ESCA outperformed eight state-of-the-art models on the CNN/DailyMail and NYT50 benchmark datasets. Yang Gao 0016, Yu Bai 0018, Mirella Lapata, Heyan Huang |
AAAI | 3 |
| 2021 | Cross-Lingual Abstractive Summarization with Limited Parallel ResourcesabstractYu Bai, Yang Gao, Heyan Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yu Bai 0018, Yang Gao 0016, Heyan Huang |
ACL/IJCNLP (1) | 1 |
| 2019 | Multiple Perspective Answer Reranking for Multi-passage Reading Comprehension
Mucheng Ren, Heyan Huang, Hongyu Liu 0001, Yu Bai 0018, Yang Gao 0016 |
NLPCC (2) | 5 |