Yi-An Lai

dblp:125/6980 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 45% Information extraction and text analysis · 18% Graph learning · 10%
Software engineering, system software, and programming languages
2 papers
Software maintenance and evolution · 68% Empirical software engineering · 32%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
0.912025
DeAL: Decoding-time Alignment for Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › alignment
inference-time alignment
0.912025
DeAL: Decoding-time Alignment for Large Language Models · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.612022
Measuring and Reducing Model Update Regression in Structured Prediction for NLP · NeurIPS 2022
Natural language and speech › Language models and text generation › pre-trained language model › conversational language models
dialogue pre-training
0.612022
Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System · ACL (1) 2022
Machine learning › Learning paradigms › multi-task learning › multi-task transfer learning
multi-task pretraining
0.612022
Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System · ACL (1) 2022
Natural language and speech › Language models and text generation
pre-trained language model
0.612022
Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System · ACL (1) 2022
Natural language and speech › Information extraction and text analysis
semantic parsing
0.612022
Measuring and Reducing Model Update Regression in Structured Prediction for NLP · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.612022
Measuring and Reducing Model Update Regression in Structured Prediction for NLP · NeurIPS 2022
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.612022
Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System · ACL (1) 2022
Software maintenance and evolution › software compatibility
backward compatibility
0.612022
Measuring and Reducing Model Update Regression in Structured Prediction for NLP · NeurIPS 2022
Software maintenance and evolution › software defects
regression bugs
0.512021
Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates · ACL/IJCNLP (1) 2021
Machine learning › Graph learning
network embedding
0.312017
PRUNE: Preserving Proximity and Global Ranking for Network Embedding · NIPS 2017
Machine learning › Graph learning › network embedding
proximity-preserving network embedding
0.312017
PRUNE: Preserving Proximity and Global Ranking for Network Embedding · NIPS 2017
Information retrieval › ranking
graph-based ranking
0.312017
Unsupervised Ranking using Graph Structures and Node Attributes · WSDM 2017
Information retrieval
ranking
0.312017
Unsupervised Ranking using Graph Structures and Node Attributes · WSDM 2017
Machine learning › Graph learning
link prediction
0.112017
PRUNE: Preserving Proximity and Global Ranking for Network Embedding · NIPS 2017

Methods — techniques the papers use, named apart from their topics

model ensemble · 1.1knowledge distillation · 1.1backward-congruent re-ranking · 1.1large language model · 0.9decoding-time intervention · 0.9multi-task learning · 0.9pre-training · 0.6regression analysis · 0.5siamese neural network · 0.3pagerank · 0.3markov chain · 0.3
YearPublicationVenuePosition
2025 DeAL: Decoding-time Alignment for Large Language Models
abstract
James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas, Saab Mansour, Katrin Kirchhoff, Dan Roth. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
James Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai, Arshit Gupta, Nikolaos Pappas 0004, Saab Mansour, Katrin Kirchhoff, Dan Roth 0001
ACL (1)4
2024 Backward Compatibility During Data Updates by Weight Interpolation
abstract
Raphael Schumann, Elman Mansimov, Yi-An Lai, Nikolaos Pappas, Xibin Gao, Yi Zhang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Raphael Schumann, Elman Mansimov, Yi-An Lai, Nikolaos Pappas 0004, Xibin Gao, Yi Zhang 0001
EACL (1)3
2022 Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System
abstract
Pre-trained language models have been recently shown to benefit task-oriented dialogue (TOD) systems.Despite their success, existing methods often formulate this task as a cascaded generation problem which can lead to error accumulation across different sub-tasks and greater data annotation overhead.In this study, we present PPTOD, a unified plug-andplay model for task-oriented dialogue.In addition, we introduce a new dialogue multi-task pre-training strategy that allows the model to learn the primary TOD task completion skills from heterogeneous dialog corpora.We extensively test our model on three benchmark TOD tasks, including end-to-end dialogue modelling, dialogue state tracking, and intent classification.Experimental results show that PPTOD achieves new state of the art on all evaluated tasks in both high-resource and lowresource scenarios.Furthermore, comparisons against previous SOTA methods show that the responses generated by PPTOD are more factually correct and semantically coherent as judged by human annotators. 1
Yixuan Su, Lei Shu 0004, Elman Mansimov, Arshit Gupta, Deng Cai 0002, Yi-An Lai
ACL (1)6
2022 Measuring and Reducing Model Update Regression in Structured Prediction for NLP
abstract
Recent advance in deep learning has led to rapid adoption of machine learning based NLP models in a wide range of applications. Despite the continuous gain in accuracy, backward compatibility is also an important aspect for industrial applications, yet it received little research attention. Backward compatibility requires that the new model does not regress on cases that were correctly handled by its predecessor. This work studies model update regression in structured prediction tasks. We choose syntactic dependency parsing and conversational semantic parsing as representative examples of structured prediction tasks in NLP. First, we measure and analyze model update regression in different model update settings. Next, we explore and benchmark existing techniques for reducing model update regression including model ensemble and knowledge distillation. We further propose a simple and effective method, Backward-Congruent Re-ranking (BCR), by taking into account the characteristics of structured output. Experiments show that BCR can better mitigate model update regression than model ensemble and knowledge distillation approaches.
Deng Cai 0002, Elman Mansimov, Yi-An Lai, Yixuan Su, Lei Shu 0004
NeurIPS3
2021 Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates
abstract
Yuqing Xie, Yi-An Lai, Yuanjun Xiong, Yi Zhang, Stefano Soatto. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuqing Xie 0001, Yi-An Lai, Yuanjun Xiong, Stefano Soatto
ACL/IJCNLP (1)2
2020 Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text Collections
abstract
Summarizing data samples by quantitative measures has a long history, with descriptive statistics being a case in point. However, as natural language processing methods flourish, there are still insufficient characteristic metrics to describe a collection of texts in terms of the words, sentences, or paragraphs they comprise. In this work, we propose metrics of diversity, density, and homogeneity that quantitatively measure the dispersion, sparsity, and uniformity of a text collection. We conduct a series of simulations to verify that each metric holds desired properties and resonates with human intuitions. Experiments on real-world datasets demonstrate that the proposed characteristic metrics are highly correlated with text classification performance of a renowned model, BERT, which could inspire future applications.
Yi-An Lai, Xuan Zhu 0002, Mona T. Diab
LREC1
2019 Goal-Embedded Dual Hierarchical Model for Task-Oriented Dialogue Generation
abstract
Hierarchical neural networks are often used to model inherent structures within dialogues.For goal-oriented dialogues, these models miss a mechanism adhering to the goals and neglect the distinct conversational patterns between two interlocutors.In this work, we propose Goal-Embedded Dual Hierarchical Attentional Encoder-Decoder (G-DuHA) able to center around goals and capture interlocutorlevel disparity while modeling goal-oriented dialogues.Experiments on dialogue generation, response generation, and human evaluations demonstrate that the proposed model successfully generates higher-quality, more diverse and goal-centric dialogues.Moreover, we apply data augmentation via goal-oriented dialogue generation for task-oriented dialog systems with better performance achieved.
Yi-An Lai, Arshit Gupta
CoNLL1
2019 DeepRank: improving unsupervised node ranking via link discovery
Yi-An Lai, Chin-Chi Hsu, Mi-Yen Yeh, Shou-De Lin
Data Min. Knowl. Discov.1
2017 PRUNE: Preserving Proximity and Global Ranking for Network Embedding
abstract
We investigate an unsupervised generative approach for network embedding. A multi-task Siamese neural network structure is formulated to connect embedding vectors and our objective to preserve the global node ranking and local proximity of nodes. We provide deeper analysis to connect the proposed proximity objective to link prediction and community detection in the network. We show our model can satisfy the following design properties: scalability, asymmetry, unity and simplicity. Experiment results not only verify the above design properties but also demonstrate the superior performance in learning-to-rank, classification, regression, and link prediction tasks.
Yi-An Lai, Chin-Chi Hsu, Mi-Yen Yeh, Shou-De Lin
NIPS1
2017 Unsupervised Ranking using Graph Structures and Node Attributes
abstract
PageRank has been the signature unsupervised ranking model for ranking node importance in a graph. One potential drawback of PageRank is that its computation depends only on input graph structures, not considering external information such as the attributes of nodes. This work proposes AttriRank, an unsupervised ranking model that considers not only graph structure but also the attributes of nodes. AttriRank is unsupervised and domain-independent, which is different from most of the existing works requiring either ground-truth labels or specific domain knowledge. Combining two reasonable assumptions about PageRank and node attributes, AttriRank transfers extra node information into a Markov chain model to obtain the ranking. We further develop approximation for AttriRank and reduce its complexity to be linear to the number of nodes or links in the graph, which makes it feasible for large network data. The experiments show that AttriRank outperforms competing models in diverse graph ranking applications.
Chin-Chi Hsu, Yi-An Lai, Ming-Han Feng, Shou-De Lin
WSDM2