Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nelleke Oostdijk

dblp:93/0 · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
4since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 10 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 67% Language models and text generation · 33%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
discourse analysis
1.012026
Discourse Realization of Generics in Human and LLM-generated Texts · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse parsing
rhetorical structure theory
1.012026
Discourse Realization of Generics in Human and LLM-generated Texts · ACL (1) 2026
Information retrieval › question answering
answer extraction
0.112007
Evaluating discourse-based answer extraction for why-question answering · SIGIR 2007
Information retrieval
question answering and dialogue systems
0.112007
Evaluating discourse-based answer extraction for why-question answering · SIGIR 2007
Information retrieval › question answering
why-question answering
0.112007
Evaluating discourse-based answer extraction for why-question answering · SIGIR 2007

Methods — techniques the papers use, named apart from their topics

genericity scoring · 1.0clause-level annotation · 1.0discourse analysis · 0.1
YearPublicationVenuePosition
2026 Discourse Realization of Generics in Human and LLM-generated Texts
abstract
Large Language Models (LLMs) often produce texts that appear coherent and credible, even when their factual reliability is uncertain.This paper investigates whether such perceived credibility correlates with the pervasive use of generics-generalizations without explicit quantification.We introduce a text-level genericity score derived from clause-level annotations and apply it to argumentative essays produced by humans and LLMs.To analyze how generics are realized in discourse, we employ Rhetorical Structure Theory to examine coherence relations across varying levels of genericity.Results show that according to our genericity metric, human texts are less generic than LLM-produced texts.As regards discourse, higher genericity correlates with less structured, paratactic structures, while for some models coherence is maintained through ELABORA-TION relations.Our findings suggest that some LLMs maintain well-structured coherence even in highly generic texts, which might enable them to "camouflage" argumentative texts as informative, enhancing their perceived credibility and persuasiveness.
Søren Fomsgaard, Martial Pastor, Gaël Dias, Nelleke Oostdijk
ACL (1)4
2025 Enhancing Discourse Parsing for Local Structures from Social Media with LLM-Generated Data
abstract
We explore the use of discourse parsers for extracting a particular discourse structure in a real-world social media scenario. Specifically, we focus on enhancing parser performance through the integration of synthetic data generated by large language models (LLMs). We conduct experiments using a newly developed dataset of 1,170 local RST discourse structures, including 900 synthetic and 270 gold examples, covering three social media platforms: online news comments sections, a discussion forum (Reddit), and a social media messaging platform (Twitter). Our primary goal is to assess the impact of LLM-generated synthetic training data on parser performance in a raw text setting without pre-identified discourse units. While both top-down and bottom-up RST architectures greatly benefit from synthetic data, challenges remain in classifying evaluative discourse structures.
Martial Pastor, Nelleke Oostdijk, Patricia Martín-Rodilla, Javier Parapar
COLING2
2023 RECESS: Resource for Extracting Cause, Effect, and Signal Spans
abstract
Fiona Anting Tan, Hansi Hettiarachchi, Ali Hürriyetoğlu, Nelleke Oostdijk, Tommaso Caselli, Tadashi Nomoto, Onur Uca, Farhana Ferdousi Liza, See-Kiong Ng. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Fiona Anting Tan, Hansi Hettiarachchi, Ali Hurriyetoglu, Nelleke Oostdijk, Tommaso Caselli, Tadashi Nomoto, Onur Uca, Farhana Ferdousi Liza, See-Kiong Ng
IJCNLP (1)4
2022 The Causal News Corpus: Annotating Causal Relations in Event Sentences from News
abstract
Despite the importance of understanding causality, corpora addressing causal relations are limited. There is a discrepancy between existing annotation guidelines of event causality and conventional causality corpora that focus more on linguistics. Many guidelines restrict themselves to include only explicit relations or clause-based arguments. Therefore, we propose an annotation schema for event causality that addresses these concerns. We annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not. Our corpus is known as the Causal News Corpus (CNC). A neural network built upon a state-of-the-art pre-trained language model performed well with 81.20% F1 score on test set, and 83.46% in 5-folds cross-validation. CNC is transferable across two external corpora: CausalTimeBank (CTB) and Penn Discourse Treebank (PDTB). Leveraging each of these external datasets for training, we achieved up to approximately 64% F1 on the CNC test set without additional fine-tuning. CNC also served as an effective training and pre-training dataset for the two external corpora. Lastly, we demonstrate the difficulty of our task to the layman in a crowd-sourced annotation exercise. Our annotated corpus is publicly available, providing a valuable resource for causal text mining researchers.
Fiona Anting Tan, Ali Hurriyetoglu, Tommaso Caselli, Nelleke Oostdijk, Tadashi Nomoto, Hansi Hettiarachchi, Iqra Ameer, Onur Uca, Farhana Ferdousi Liza, Tiancheng Hu
LREC4
2020 The CLARIN Knowledge Centre for Atypical Communication Expertise
abstract
This paper introduces a new CLARIN Knowledge Center which is the K-Centre for Atypical Communication Expertise (ACE for short) which has been established at the Centre for Language and Speech Technology (CLST) at Radboud University. Atypical communication is an umbrella term used here to denote language use by second language learners, people with language disorders or those suffering from language disabilities, but also more broadly by bilinguals and users of sign languages. It involves multiple modalities (text, speech, sign, gesture) and encompasses different developmental stages. ACE closely collaborates with The Language Archive (TLA) at the Max Planck Institute for Psycholinguistics in order to safeguard GDPR-compliant data storage and access. We explain the mission of ACE and show its potential on a number of showcases and a use case.
Henk van den Heuvel, Nelleke Oostdijk, Caroline F. Rowland, Paul Trilsbeek
LREC2
2020 The Connection between the Text and Images of News Articles: New Insights for Multimedia Analysis
abstract
We report on a case study of text and images that reveals the inadequacy of simplistic assumptions about their connection and interplay. The context of our work is a larger effort to create automatic systems that can extract event information from online news articles about flooding disasters. We carry out a manual analysis of 1000 articles containing a keyword related to flooding. The analysis reveals that the articles in our data set cluster into seven categories related to different topical aspects of flooding, and that the images accompanying the articles cluster into five categories related to the content they depict. The results demonstrate that flood-related news articles do not consistently report on a single, currently unfolding flooding event and we should also not assume that a flood-related image will directly relate to a flooding-event described in the corresponding article. In particular, spatiotemporal distance is important. We validate the manual analysis with an automatic classifier demonstrating the technical feasibility of multimedia analysis approaches that admit more realistic relationships between text and images. In sum, our case study confirms that closer attention to the connection between text and images has the potential to improve the collection of multimodal information from news articles.
Nelleke Oostdijk, Hans van Halteren, Mustafa Erkan Basar, Martha A. Larson
LREC1
2018 Metadata Collection Records for Language Resources
Henk van den Heuvel, Erwin Komen, Nelleke Oostdijk
LREC3
2017 Supporting Experts to Handle Tweet Collections About Significant Events
Ali Hurriyetoglu, Nelleke Oostdijk, Mustafa Erkan Basar, Antal van den Bosch
NLDB2
2016 Falling silent, lost for words ... Tracing personal involvement in interviews with Dutch war veterans
Henk van den Heuvel, Nelleke Oostdijk
LREC2
2014 The evolving infrastructure for language resources and the role for data scientists
Nelleke Oostdijk, Henk van den Heuvel
LREC1
2014 Dealing with temporal variation in patent categorization
Eva D'hondt, Suzan Verberne, Nelleke Oostdijk, Jean Beney, Cornelis H. A. Koster, Lou Boves
Inf. Retr.3
2013 Shallow parsing for recognizing threats in Dutch tweets
abstract
In this paper, we investigate the recognition of threats in Dutch tweets. As tweets often display irregular grammatical form and deviant orthography, analysis by standard means is problematic. Therefore, we have implemented a new shallow parsing mechanism which is driven by handcrafted rules. Experimental results are encouraging, with an F-measure of about 40% on a random sample of Dutch tweets. Moreover, the error analysis shows some clear avenues for further improvement.
Nelleke Oostdijk, Hans van Halteren
ASONAM1
2013 N-Gram-Based Recognition of Threatening Tweets
Nelleke Oostdijk, Hans van Halteren
CICLing (2)1
2012 Beyond SoNaR: towards the facilitation of large corpus building efforts
Martin Reynaert, Ineke Schuurman, Véronique Hoste, Nelleke Oostdijk, Maarten van Gompel
LREC4
2012 Collection of a corpus of Dutch SMS
Maaske Treurniet, Orphée De Clercq, Henk van den Heuvel, Nelleke Oostdijk
LREC4
2010 Constructing a Broad-coverage Lexicon for Text Mining in the Patent Domain
Nelleke Oostdijk, Suzan Verberne, Cornelis H. A. Koster
LREC1
2010 Balancing SoNaR: IPR versus Processing Issues in a 500-Million-Word Written Dutch Reference Corpus
Martin Reynaert, Nelleke Oostdijk, Orphée De Clercq, Henk van den Heuvel, Franciska de Jong
LREC2
2010 What Is Not in the Bag of Words for Why-QA?
abstract
While developing an approach to why-QA, we extended a passage retrieval system that uses off-the-shelf retrieval technology with a re-ranking step incorporating structural information. We get significantly higher scores in terms of MRR@150 (from 0.25 to 0.34) and success@10. The 23% improvement that we reach in terms of MRR is comparable to the improvement reached on different QA tasks by other researchers in the field, although our re-ranking approach is based on relatively lightweight overlap measures incorporating syntactic constituents, cue words, and document structure.
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
Comput. Linguistics3
2008 Using Syntactic Information for Improving Why-Question Answering
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
COLING3
2008 Evaluating Paragraph Retrieval for
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
ECIR3
2008 From D-Coi to SoNaR: a reference corpus for Dutch
Nelleke Oostdijk, Martin Reynaert, Paola Monachesi, Gertjan van Noord, Roeland Ordelman, Ineke Schuurman, Vincent Vandeghinste
LREC1
2007 Evaluating discourse-based answer extraction for why-question answering
abstract
No abstract available.
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
SIGIR3
2006 User requirements analysis for the design of a reference corpus of written Dutch
Nelleke Oostdijk, Lou Boves
LREC1
2006 Data for question answering: The case of why
Suzan Verberne, Lou Boves, Nelleke Oostdijk, Peter-Arno Coppen
LREC3
2005 On temporal aspects of turn taking in conversational dialogues
Louis ten Bosch, Nelleke Oostdijk, Lou Boves
Speech Commun.2
2004 Linguistic profiling of texts for the purpose of language verification
Hans van Halteren, Nelleke Oostdijk
COLING2
2004 Using Large Multi-purpose Corpora for Specific Research Questions: Discourse Phenomena Related to Wh-questions in the Spoken Dutch Corpus
Nelleke Oostdijk, Lou Boves
LREC1
2004 Linguistic Annotation of the Spoken Dutch Corpus: If We Had To Do It All Over Again
Ineke Schuurman, Wim Goedertier, Heleen Hoekstra, Nelleke Oostdijk, Richard Piepenbrock, Machteld Schouppe
LREC4
2002 Experiences from the Spoken Dutch Corpus Project
Nelleke Oostdijk, Wim Goedertier, Frank Van Eynde, Lou Boves, Jean-Pierre Martens, Michael Moortgat, R. Harald Baayen
LREC1
2000 The Spoken Dutch Corpus. Overview and First Evaluation
Nelleke Oostdijk
LREC1