VLDB 2026 Research / reviewers in the wild / expert
Sean Papay
dblp:229/3086
· DBLP profile ↗
4ranked-venue papers
3as first author
2since 2021 · last 2025
0009-0006-2330-7349ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Information extraction and text analysis · 71% Trustworthy machine learning · 29% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Which Demographics do LLMs Default to During Annotation? · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › data annotation
text annotation |
0.9 | 1 | 2025 | Which Demographics do LLMs Default to During Annotation? · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.6 | 1 | 2022 | Constraining Linear-chain CRFs to Regular Languages · ICLR 2022 |
Natural language and speech › Information extraction and text analysis
span selection |
0.4 | 1 | 2020 | Dissecting Span Identification Tasks with Performance Prediction · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
chunking |
0.1 | 1 | 2020 | Dissecting Span Identification Tasks with Performance Prediction · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 1 | 2020 | Dissecting Span Identification Tasks with Performance Prediction · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
large language model annotation · 0.9demographic analysis · 0.9meta-learning · 0.4LSTM · 0.4CRF · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Which Demographics do LLMs Default to During Annotation?abstractJohannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li, Nadine Probol, Lynn Greschner, Sean Papay, Yarik Menchaca Resendiz, Aswathy Velutharambath, Amelie Wuehrl, Sabine Weber, Roman Klinger. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Johannes Schäfer, Aidan Combs, Christopher Bagdon, Nadine Probol, Lynn Greschner, Sean Papay, Yarik Menchaca Resendiz, Aswathy Velutharambath, Amelie Wührl, Sabine Weber, Roman Klinger |
ACL (1) | 7 |
| 2022 | Constraining Linear-chain CRFs to Regular Languages
Sean Papay, Roman Klinger, Sebastian Padó |
ICLR | 1 |
| 2020 | Dissecting Span Identification Tasks with Performance PredictionabstractSpan identification (in short, span ID) tasks such as chunking, NER, or code-switching detection, ask models to identify and classify relevant spans in a text.Despite being a staple of NLP, and sharing a common structure, there is little insight on how these tasks' properties influence their difficulty, and thus little guidance on what model families work well on span ID tasks, and why.We analyze span ID tasks via performance prediction, estimating how well neural architectures do on different tasks.Our contributions are: (a) we identify key properties of span ID tasks that can inform performance prediction; (b) we carry out a large-scale experiment on English data, building a model to predict performance for unseen span ID tasks that can support architecture choices; (c), we investigate the parameters of the meta model, yielding new insights on how model and task properties interact to affect span ID performance.We find, e.g., that span frequency is especially important for LSTMs, and that CRFs help when spans are infrequent and boundaries non-distinctive. Sean Papay, Roman Klinger, Sebastian Padó |
EMNLP (1) | 1 |
| 2020 | RiQuA: A Corpus of Rich Quotation Annotation for English Literary TextabstractWe introduce RiQuA (RIch QUotation Annotations), a corpus that provides quotations, including their interpersonal structure (speakers and addressees) for English literary text. The corpus comprises 11 works of 19th-century literature that were manually doubly annotated for direct and indirect quotations. For each quotation, its span, speaker, addressee, and cue are identified (if present). This provides a rich view of dialogue structures not available from other available corpora. We detail the process of creating this dataset, discuss the annotation guidelines, and analyze the resulting corpus in terms of inter-annotator agreement and its properties. RiQuA, along with its annotations guidelines and associated scripts, are publicly available for use, modification, and experimentation. Sean Papay, Sebastian Padó |
LREC | 1 |