VLDB 2026 Research / reviewers in the wild / expert
Nelleke Oostdijk
dblp:93/0
· DBLP profile ↗
30ranked-venue papers
10as first author
4since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 10 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 67% Language models and text generation · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
discourse analysis |
1.0 | 1 | 2026 | Discourse Realization of Generics in Human and LLM-generated Texts · ACL (1) 2026 |
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse parsing
rhetorical structure theory |
1.0 | 1 | 2026 | Discourse Realization of Generics in Human and LLM-generated Texts · ACL (1) 2026 |
Information retrieval › question answering
answer extraction |
0.1 | 1 | 2007 | Evaluating discourse-based answer extraction for why-question answering · SIGIR 2007 |
Information retrieval
question answering and dialogue systems |
0.1 | 1 | 2007 | Evaluating discourse-based answer extraction for why-question answering · SIGIR 2007 |
Information retrieval › question answering
why-question answering |
0.1 | 1 | 2007 | Evaluating discourse-based answer extraction for why-question answering · SIGIR 2007 |
Methods — techniques the papers use, named apart from their topics
genericity scoring · 1.0clause-level annotation · 1.0discourse analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discourse Realization of Generics in Human and LLM-generated TextsabstractLarge Language Models (LLMs) often produce texts that appear coherent and credible, even when their factual reliability is uncertain.This paper investigates whether such perceived credibility correlates with the pervasive use of generics-generalizations without explicit quantification.We introduce a text-level genericity score derived from clause-level annotations and apply it to argumentative essays produced by humans and LLMs.To analyze how generics are realized in discourse, we employ Rhetorical Structure Theory to examine coherence relations across varying levels of genericity.Results show that according to our genericity metric, human texts are less generic than LLM-produced texts.As regards discourse, higher genericity correlates with less structured, paratactic structures, while for some models coherence is maintained through ELABORA-TION relations.Our findings suggest that some LLMs maintain well-structured coherence even in highly generic texts, which might enable them to "camouflage" argumentative texts as informative, enhancing their perceived credibility and persuasiveness. Søren Fomsgaard, Martial Pastor, Gaël Dias, Nelleke Oostdijk |
ACL (1) | 4 |
| 2025 | Enhancing Discourse Parsing for Local Structures from Social Media with LLM-Generated DataabstractWe explore the use of discourse parsers for extracting a particular discourse structure in a real-world social media scenario. Specifically, we focus on enhancing parser performance through the integration of synthetic data generated by large language models (LLMs). We conduct experiments using a newly developed dataset of 1,170 local RST discourse structures, including 900 synthetic and 270 gold examples, covering three social media platforms: online news comments sections, a discussion forum (Reddit), and a social media messaging platform (Twitter). Our primary goal is to assess the impact of LLM-generated synthetic training data on parser performance in a raw text setting without pre-identified discourse units. While both top-down and bottom-up RST architectures greatly benefit from synthetic data, challenges remain in classifying evaluative discourse structures. Martial Pastor, Nelleke Oostdijk, Patricia Martín-Rodilla, Javier Parapar |
COLING | 2 |
| 2023 | RECESS: Resource for Extracting Cause, Effect, and Signal SpansabstractFiona Anting Tan, Hansi Hettiarachchi, Ali Hürriyetoğlu, Nelleke Oostdijk, Tommaso Caselli, Tadashi Nomoto, Onur Uca, Farhana Ferdousi Liza, See-Kiong Ng. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Fiona Anting Tan, Hansi Hettiarachchi, Ali Hurriyetoglu, Nelleke Oostdijk, Tommaso Caselli, Tadashi Nomoto, Onur Uca, Farhana Ferdousi Liza, See-Kiong Ng |
IJCNLP (1) | 4 |
| 2022 | The Causal News Corpus: Annotating Causal Relations in Event Sentences from NewsabstractDespite the importance of understanding causality, corpora addressing causal relations are limited. There is a discrepancy between existing annotation guidelines of event causality and conventional causality corpora that focus more on linguistics. Many guidelines restrict themselves to include only explicit relations or clause-based arguments. Therefore, we propose an annotation schema for event causality that addresses these concerns. We annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not. Our corpus is known as the Causal News Corpus (CNC). A neural network built upon a state-of-the-art pre-trained language model performed well with 81.20% F1 score on test set, and 83.46% in 5-folds cross-validation. CNC is transferable across two external corpora: CausalTimeBank (CTB) and Penn Discourse Treebank (PDTB). Leveraging each of these external datasets for training, we achieved up to approximately 64% F1 on the CNC test set without additional fine-tuning. CNC also served as an effective training and pre-training dataset for the two external corpora. Lastly, we demonstrate the difficulty of our task to the layman in a crowd-sourced annotation exercise. Our annotated corpus is publicly available, providing a valuable resource for causal text mining researchers. Fiona Anting Tan, Ali Hurriyetoglu, Tommaso Caselli, Nelleke Oostdijk, Tadashi Nomoto, Hansi Hettiarachchi, Iqra Ameer, Onur Uca, Farhana Ferdousi Liza, Tiancheng Hu |
LREC | 4 |
| 2020 | The CLARIN Knowledge Centre for Atypical Communication ExpertiseabstractThis paper introduces a new CLARIN Knowledge Center which is the K-Centre for Atypical Communication Expertise (ACE for short) which has been established at the Centre for Language and Speech Technology (CLST) at Radboud University. Atypical communication is an umbrella term used here to denote language use by second language learners, people with language disorders or those suffering from language disabilities, but also more broadly by bilinguals and users of sign languages. It involves multiple modalities (text, speech, sign, gesture) and encompasses different developmental stages. ACE closely collaborates with The Language Archive (TLA) at the Max Planck Institute for Psycholinguistics in order to safeguard GDPR-compliant data storage and access. We explain the mission of ACE and show its potential on a number of showcases and a use case. Henk van den Heuvel, Nelleke Oostdijk, Caroline F. Rowland, Paul Trilsbeek |
LREC | 2 |
| 2020 | The Connection between the Text and Images of News Articles: New Insights for Multimedia AnalysisabstractWe report on a case study of text and images that reveals the inadequacy of simplistic assumptions about their connection and interplay. The context of our work is a larger effort to create automatic systems that can extract event information from online news articles about flooding disasters. We carry out a manual analysis of 1000 articles containing a keyword related to flooding. The analysis reveals that the articles in our data set cluster into seven categories related to different topical aspects of flooding, and that the images accompanying the articles cluster into five categories related to the content they depict. The results demonstrate that flood-related news articles do not consistently report on a single, currently unfolding flooding event and we should also not assume that a flood-related image will directly relate to a flooding-event described in the corresponding article. In particular, spatiotemporal distance is important. We validate the manual analysis with an automatic classifier demonstrating the technical feasibility of multimedia analysis approaches that admit more realistic relationships between text and images. In sum, our case study confirms that closer attention to the connection between text and images has the potential to improve the collection of multimodal information from news articles. Nelleke Oostdijk, Hans van Halteren, Mustafa Erkan Basar, Martha A. Larson |
LREC | 1 |
| 2018 | Metadata Collection Records for Language Resources
Henk van den Heuvel, Erwin Komen, Nelleke Oostdijk |
LREC | 3 |
| 2017 | Supporting Experts to Handle Tweet Collections About Significant Events
Ali Hurriyetoglu, Nelleke Oostdijk, Mustafa Erkan Basar, Antal van den Bosch |
NLDB | 2 |
| 2016 | Falling silent, lost for words ... Tracing personal involvement in interviews with Dutch war veterans
Henk van den Heuvel, Nelleke Oostdijk |
LREC | 2 |
| 2014 | The evolving infrastructure for language resources and the role for data scientists
Nelleke Oostdijk, Henk van den Heuvel |
LREC | 1 |
| 2014 | Dealing with temporal variation in patent categorization
Eva D'hondt, Suzan Verberne, Nelleke Oostdijk, Jean Beney, Cornelis H. A. Koster, Lou Boves |
Inf. Retr. | 3 |
| 2013 | Shallow parsing for recognizing threats in Dutch tweetsabstractIn this paper, we investigate the recognition of threats in Dutch tweets. As tweets often display irregular grammatical form and deviant orthography, analysis by standard means is problematic. Therefore, we have implemented a new shallow parsing mechanism which is driven by handcrafted rules. Experimental results are encouraging, with an F-measure of about 40% on a random sample of Dutch tweets. Moreover, the error analysis shows some clear avenues for further improvement. Nelleke Oostdijk, Hans van Halteren |
ASONAM | 1 |
| 2013 | N-Gram-Based Recognition of Threatening Tweets
Nelleke Oostdijk, Hans van Halteren |
CICLing (2) | 1 |
| 2012 | Beyond SoNaR: towards the facilitation of large corpus building efforts
Martin Reynaert, Ineke Schuurman, Véronique Hoste, Nelleke Oostdijk, Maarten van Gompel |
LREC | 4 |
| 2012 | Collection of a corpus of Dutch SMS
Maaske Treurniet, Orphée De Clercq, Henk van den Heuvel, Nelleke Oostdijk |
LREC | 4 |
| 2010 | Constructing a Broad-coverage Lexicon for Text Mining in the Patent Domain
Nelleke Oostdijk, Suzan Verberne, Cornelis H. A. Koster |
LREC | 1 |
| 2010 | Balancing SoNaR: IPR versus Processing Issues in a 500-Million-Word Written Dutch Reference Corpus
Martin Reynaert, Nelleke Oostdijk, Orphée De Clercq, Henk van den Heuvel, Franciska de Jong |
LREC | 2 |
| 2010 | What Is Not in the Bag of Words for Why-QA?abstractWhile developing an approach to why-QA, we extended a passage retrieval system that uses off-the-shelf retrieval technology with a re-ranking step incorporating structural information. We get significantly higher scores in terms of MRR@150 (from 0.25 to 0.34) and success@10. The 23% improvement that we reach in terms of MRR is comparable to the improvement reached on different QA tasks by other researchers in the field, although our re-ranking approach is based on relatively lightweight overlap measures incorporating syntactic constituents, cue words, and document structure. Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen |
Comput. Linguistics | 3 |
| 2008 | Using Syntactic Information for Improving Why-Question Answering
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen |
COLING | 3 |
| 2008 | Evaluating Paragraph Retrieval for
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen |
ECIR | 3 |
| 2008 | From D-Coi to SoNaR: a reference corpus for Dutch
Nelleke Oostdijk, Martin Reynaert, Paola Monachesi, Gertjan van Noord, Roeland Ordelman, Ineke Schuurman, Vincent Vandeghinste |
LREC | 1 |
| 2007 | Evaluating discourse-based answer extraction for why-question answeringabstractNo abstract available. Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen |
SIGIR | 3 |
| 2006 | User requirements analysis for the design of a reference corpus of written Dutch
Nelleke Oostdijk, Lou Boves |
LREC | 1 |
| 2006 | Data for question answering: The case of why
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen |
LREC | 3 |
| 2005 | On temporal aspects of turn taking in conversational dialogues
Louis ten Bosch, Nelleke Oostdijk, Lou Boves |
Speech Commun. | 2 |
| 2004 | Linguistic profiling of texts for the purpose of language verification
Hans van Halteren, Nelleke Oostdijk |
COLING | 2 |
| 2004 | Using Large Multi-purpose Corpora for Specific Research Questions: Discourse Phenomena Related to Wh-questions in the Spoken Dutch Corpus
Nelleke Oostdijk, Lou Boves |
LREC | 1 |
| 2004 | Linguistic Annotation of the Spoken Dutch Corpus: If We Had To Do It All Over Again
Ineke Schuurman, Wim Goedertier, Heleen Hoekstra, Nelleke Oostdijk, Richard Piepenbrock, Machteld Schouppe |
LREC | 4 |
| 2002 | Experiences from the Spoken Dutch Corpus Project
Nelleke Oostdijk, Wim Goedertier, Frank Van Eynde, Lou Boves, Jean-Pierre Martens, Michael Moortgat, R. Harald Baayen |
LREC | 1 |
| 2000 | The Spoken Dutch Corpus. Overview and First Evaluation
Nelleke Oostdijk |
LREC | 1 |