VLDB 2026 Research / reviewers in the wild / expert
Yang Janet Liu
dblp:302/4022
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Probing LLMs for Multilingual Discourse Generalization Through a Unified Label SetabstractDiscourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations.This work investigates whether large language models (LLMs) capture discourse knowledge that generalizes across languages and frameworks.We address this question along two dimensions: (1) developing a unified discourse relation label set to facilitate cross-lingual and cross-framework discourse analysis, and (2) probing LLMs to assess whether they encode generalizable discourse abstractions.Using multilingual discourse relation classification as a testbed, we examine a comprehensive set of 23 LLMs of varying sizes and multilingual capabilities.Our results show that LLMs, especially those with multilingual training corpora, can generalize discourse information across languages and frameworks.Further layer-wise analyses reveal that language generalization at the discourse level is most salient in the intermediate layers.Lastly, our error analysis provides an account of challenging relation classes. Florian Eichin, Yang Janet Liu, Barbara Plank, Michael A. Hedderich |
ACL (1) | 2 |
| 2025 | Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesabstractUnderstanding pragmatics—the use of language in context—is crucial for developing NLP systems capable of interpreting nuanced language use. Despite recent advances in language technologies, including large language models, evaluating their ability to handle pragmatic phenomena such as implicatures and references remains challenging. To advance pragmatic abilities in models, it is essential to understand current evaluation trends and identify existing limitations. In this survey, we provide a comprehensive review of resources designed for evaluating pragmatic capabilities in NLP, categorizing datasets by the pragmatic phenomena they address. We analyze task designs, data collection methods, evaluation approaches, and their relevance to real-world applications. By examining these resources in the context of modern language models, we highlight emerging trends, challenges, and gaps in existing benchmarks. Our survey aims to clarify the landscape of pragmatic evaluation and guide the development of more comprehensive and targeted benchmarks, ultimately contributing to more nuanced and context-aware NLP models. Bolei Ma, Wei Zhou 0067, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank |
ACL (1) | 5 |
| 2025 | Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label VariationabstractThe recent rise of reasoning-tuned Large Language Models (LLMs)-which generate chains of thought (CoTs) before giving the final answer-has attracted significant attention and offers new opportunities for gaining insights into human label variation, which refers to plausible differences in how multiple annotators label the same data instance.Prior work has shown that LLM-generated explanations can help align model predictions with human label distributions, but typically adopt a reverse paradigm: producing explanations based on given answers.In contrast, CoTs provide a forward reasoning path that may implicitly embed rationales for each answer option, before generating the answers.We thus propose a novel LLM-based pipeline enriched with linguistically-grounded discourse segmenters to extract supporting and opposing statements for each answer option from CoTs with improved accuracy.We also propose a rank-based HLV evaluation framework that prioritizes the ranking of answers over exact scores, which instead favor direct comparison of label distributions.Our method outperforms a direct generation method as well as baselines on three datasets, and shows better alignment of ranking methods with humans, highlighting the effectiveness of our approach. Beiduo Chen, Yang Janet Liu, Anna Korhonen, Barbara Plank |
EMNLP | 2 |
| 2025 | References Matter: Investigating the Impact of Reference Set Variation on Summarization EvaluationabstractHuman language production exhibits remarkable richness and variation, reflecting diverse communication styles and intents. However, this variation is often overlooked in summarization evaluation. While having multiple reference summaries is known to improve correlation with human judgments, the impact of the reference set on reference-based metrics has not been systematically investigated. This work examines the sensitivity of widely used reference-based metrics in relation to the choice of reference sets, analyzing three diverse multi-reference summarization datasets: SummEval, GUMSum, and DUC2004. We demonstrate that many popular metrics exhibit significant instability. This instability is particularly concerning for n-gram-based metrics like ROUGE, where model rankings vary depending on the reference sets, undermining the reliability of model comparisons. We also collect human judgments on LLM outputs for genre-diverse data and examine their correlation with metrics to supplement existing findings beyond newswire summaries, finding weak-to-no correlation. Taken together, we recommend incorporating reference set variation into summarization evaluation to enhance consistency alongside correlation with human judgments, especially when evaluating LLMs. Silvia Casola, Yang Janet Liu, Siyao Peng, Oliver Kraus, Albert Gatt, Barbara Plank |
INLG | 2 |
| 2025 | eRST: A Signaled Graph Theory of Discourse Relations and OrganizationabstractAbstract In this article we present Enhanced Rhetorical Structure Theory (eRST), a new theoretical framework for computational discourse analysis, based on an expansion of Rhetorical Structure Theory (RST). The framework encompasses discourse relation graphs with tree-breaking, non-projective and concurrent relations, as well as implicit and explicit signals which give explainable rationales to our analyses. We survey shortcomings of RST and other existing frameworks, such as Segmented Discourse Representation Theory, the Penn Discourse Treebank, and Discourse Dependencies, and address these using constructs in the proposed theory. We provide annotation, search, and visualization tools for data, and present and evaluate a freely available corpus of English annotated according to our framework, encompassing 12 spoken and written genres with over 200K tokens. Finally, we discuss automatic parsing, evaluation metrics, and applications for data in our framework. Amir Zeldes, Tatsuya Aoyama, Yang Janet Liu, Siyao Peng, Debopam Das, Luke Gessler |
Comput. Linguistics | 3 |
| 2024 | DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse ProcessingabstractThis paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about 4 millions tokens and around 250,000 discourse relation instances from 4 discourse frameworks: RST, SDRT, PDTB, and Discourse Dependencies. We present an overview of the data, its development across three NLP shared tasks on discourse processing carried out in the past five years, and the latest modifications and added extensions. We also carry out an evaluation of state-of-the-art multilingual systems trained on the data for each task, showing plateau performance on segmentation, but important room for improvement for connective identification and relation classification. The DISRPT benchmark employs a unified format that we make available on GitHub and HuggingFace in order to encourage future work on discourse processing across languages, domains, and frameworks. Chloé Braud, Amir Zeldes, Laura Rivière, Yang Janet Liu, Philippe Muller, Damien Sileo, Tatsuya Aoyama |
LREC/COLING | 4 |
| 2024 | GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and DomainsabstractYang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Elizabeth Levine, Jessica Lin, Devika Tiwari, Amir Zeldes. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu 0001, Shabnam Behzad, Lauren Levine, Jessica Lin 0004, Devika Tiwari, Amir Zeldes |
EMNLP | 1 |
| 2023 | Why Can't Discourse Parsing Generalize? A Thorough Investigation of the Impact of Data DiversityabstractRecent advances in discourse parsing performance create the impression that, as in other NLP tasks, performance for high-resource languages such as English is finally becoming reliable.In this paper we demonstrate that this is not the case, and thoroughly investigate the impact of data diversity on RST parsing stability.We show that state-of-the-art architectures trained on the standard English newswire benchmark do not generalize well, even within the news domain.Using the two largest RST corpora of English with text from multiple genres, we quantify the impact of genre diversity in training data for achieving generalization to text types unseen during training.Our results show that a heterogeneous training regime is critical for stable and generalizable models, across parser architectures.We also provide error analyses of model outputs and out-ofdomain performance.To our knowledge, this study is the first to fully evaluate cross-corpus RST parsing generalizability on complete trees, examine between-genre degradation within an RST corpus, and investigate the impact of genre diversity in training data composition. Yang Janet Liu, Amir Zeldes |
EACL | 1 |
| 2023 | Lightweight and Efficient Spoken Language Identification of Long-form Audio
Winstead Zhu, Md. Iftekhar Tanveer, Yang Janet Liu, Seye Ojumu, Rosie Jones |
INTERSPEECH | 3 |
| 2023 | What's Hard in English RST Parsing? Predictive Models for Error AnalysisabstractDespite recent advances in Natural Language Processing (NLP), hierarchical discourse parsing in the framework of Rhetorical Structure Theory remains challenging, and our understanding of the reasons for this are as yet limited.In this paper, we examine and model some of the factors associated with parsing difficulties in previous work: the existence of implicit discourse relations, challenges in identifying long-distance relations, out-of-vocabulary items, and more.In order to assess the relative importance of these variables, we also release two annotated English test-sets with explicit correct and distracting discourse markers associated with gold standard RST relations.Our results show that as in shallow discourse parsing, the explicit/implicit distinction plays a role, but that long-distance dependencies are the main challenge, while lack of lexical overlap is less of a problem, at least for in-domain parsing.Our final model is able to predict where errors will occur with an accuracy of 76.3% for the bottom-up parser and 76.6% for the top-down parser. Yang Janet Liu, Tatsuya Aoyama, Amir Zeldes |
SIGDIAL | 1 |