Jelke Bloem

dblp:151/8437 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0003-2221-0554ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Examining Large Language Models' form-meaning mappings of information structure constructions in Mandarin Chinese
abstract
Construction Grammar (CxG) knowledge in language models has been extensively studied for English, but remains underexplored in other languages.In Mandarin Chinese, the ba (把, disposal) and bei (被, passive) constructions are widely used for managing information structure.They foreground topical elements (information structure) and encode systematic form-meaning mappings (CxG), particularly with respect to the semantic role of the object.We probe language models' linguistic competence with these constructions using minimal pairs, constructing a new minimalpair dataset comprising seven paradigms that target both syntactic constraints and verbconstruction compatibility.Our results show that it remains a challenge for many models to capture the form-meaning mappings underlying the ba construction, although they achieve high accuracy on paradigms driven by surface syntactic cues.
Xiaojuan Tan, Jelke Bloem
CoNLL3
2026 Towards Dynamic Metaphor Identification: Evaluating GPT O-Series Models on Five Metaphoricity Cues in U.S. Trade Corpora
abstract
Although recent advances have focused on detecting metaphors, existing models generally treat them as static entities. There has been little research into identifying dynamic metaphors in discourse. This article addresses this gap by focusing on metaphoricity cues: Linguistic signals that may indicate the activation of metaphoric meaning in different discourse contexts. This study examines the ability of OpenAI’s O-series models (O4-mini, O4-mini-high and O3) in detecting five metaphoricity cues in the U.S. trade discourse, including cues of explicit mapping, emphasis, marking, repetition and novelisation. Research results show that the models performed best on repetition and emphasis, while novelisation was the most difficult cue to detect.
Berkay Bas, Jelke Bloem, Xiaojuan Tan
LREC2
2026 Multi-SimLex for Dutch: Benchmarking Embedding- and Prompt-Based Model Performance on Semantic Similarity
abstract
We introduce Dutch Multi-SimLex, a 1,888–pair extension of the Multi-SimLex benchmark for evaluating lexical semantic similarity in Dutch. The dataset was rated by 100 native speakers on a 0–6 scale and shows high reliability (overall ICC(2,k)=0.82) as well as strong alignment with English (ρ=0.73). Using this resource, we evaluate eighteen models across four architectural families: static embeddings, encoder-only transformers, encoder–decoders, and decoder-only LLMs. We evaluate models using two complementary approaches: embedding-based cosine similarity and prompted similarity judgments in Dutch. In embedding-based evaluation, FastText (ρ=0.485) and the monolingual Dutch encoder BERTje (ρ=0.468) achieve the strongest alignment with human ratings, while multilingual encoders such as mBERT (ρ=0.208) and XLM-R (ρ=0.186) perform weaker. Prompt-based evaluation yields substantially higher correlations, with GPT-4 (ρ=0.761) performing best, followed by DeepSeek-V3 (ρ=0.753) and Gemini 1.5 Pro (ρ=0.722). Together, the results show that model performance depends strongly on how meaning is tested. Dutch Multi-SimLex provides a reliable foundation for evaluating meaning across architectures and advancing Dutch semantic evaluation.
Lizzy Brans, Jelke Bloem
LREC2
2026 Graph-TempCZ: A Graph Representation of Software Mentions for Predicting Software Usage in Scientific Publications
abstract
Predicting how software is used, shared, and evolves across publications is essential to studying scientific progress. Existing methods for representing software usage in publications rely mainly on tabular or textual formats, which limit their structural expressiveness and consequently their ability to predict software usage. We address these gaps by representing software mentions and citations as a graph and formulating software usage prediction as a link prediction task. To support this study, we construct the first large-scale graph dataset of publication and software mentions, Graph-TempCZ, covering 1959-2022 with over six million mention relationships. Experiments using both traditional machine learning and Graph Neural Network (GNN) show that graph-based models substantially outperform feature-based baselines, achieving a 5.98% improvement in test accuracy. Temporal experiments further reveal that models trained on one year generalize effectively to nearby years but show gradual performance decay as the temporal gap increases. This work provides the first comprehensive foundation for analyzing software usage through a temporal graph representation.
Congfeng Cao, Jelke Bloem
LREC3
2026 Investigating How LLMs Propagate Female Stereotypes: Comparing What Models Say via Prompts with What They Represent in Their Embeddings
abstract
As Large Language Models (LLMs) are increasingly deployed in sensitive domains, concerns about their encoding and reproduction of social bias have intensified. We examine how gender stereotypes are represented in embeddings and expressed in outputs across three models: BERT, base LLaMA-2-7b, and instruction-tuned LLaMA-2-7b-Chat. Focusing on seven female-oriented stereotype categories, we compare embedding-level bias using Directional Embedding Probing with output-level behavior measured via masked token prediction (BERT) and narrative prompt completions (LLaMA models). LLaMA-2-Chat showed the strongest representational–behavioral alignment, with female-aligned scores ranging from 60% to 100% and a significant point-biserial correlation (r = 0.55, p = 0.0008). BERT exhibited weaker alignment (0%–60%; r = 0.39, p = 0.054), while base LLaMA-2 showed intermediate but inconsistent patterns. These findings suggest that instruction tuning is associated with clearer alignment between internal representations and generated outputs, while prompt design plays a critical role in surfacing latent bias. The study contributes to fairness research by emphasizing the need to assess both internal representations and their behavioral expression in LLMs.
Andrea Valderrey Nuñez, Jelke Bloem
LREC2
2026 Prompting Instruction-tuned LLMs for Semantic Similarity Values
abstract
The impressive few-shot performance of generative decoder transformer language models at novel tasks has raised interest in using them to estimate lexical-semantic properties of words, word pairs or multi-word expressions. We explore the task of eliciting semantic similarity scores between word pairs through prompting, comparing these scores to human benchmarks. We investigate different prompting approaches, different model architectures and different languages using the Dutch, English and Mandarin Chinese SimLex-999 benchmarks. The results show that prompting each word pair individually yields better correlations, and that models struggle with the distinction between similarity and relatedness, just as static and contextual word embedding models did. The new, open-weight gpt-oss-20b model yields the highest correlation with human ratings out of the models we evaluated.
Xander Akiko Snelder, Yunchong Huang, Jelke Bloem
LREC3
2025 Mapping semantic networks to Dutch word embeddings as a diagnostic tool for cognitive decline
abstract
We explore the possibility of semantic networks as a diagnostic tool for cognitive decline by using Dutch verbal fluency data to investigate the relationship between semantic networks and cognitive health.In psychology, semantic networks serve as abstract representations of the semantic memory system.Semantic verbal fluency data can be used to estimate said networks.Traditionally, this is done by counting the number of raw items produced by participants in a verbal fluency task.We used static and contextual word embedding models to connect the elicited words through semantic similarity scores, and extracted three network distance metrics.We then tested how well these metrics predict participants' cognitive health scores on the Mini-Mental State Examination (MMSE).While the significant predictors differed per model, the traditional number-of-words measure was not significant in any case.These findings suggest that semantic network metrics may provide a more sensitive measure of cognitive health than traditional scoring.
Maithe van Noort, Michal Korenar, Jelke Bloem
EMNLP3
2024 SimLex-999 for Dutch
abstract
Word embeddings revolutionised natural language processing by effectively representing words as dense vectors. Although many datasets exist to evaluate English embeddings, few cater to Dutch. We developed a Dutch variant of the SimLex-999 word similarity dataset by gathering similarity judgements from 235 native Dutch speakers. Subsequently, we evaluated two popular Dutch language models, Bertje and RobBERT, finding that Bertje showed superior alignment with human semantic similarity judgments compared to RobBERT. This study provides the first intrinsic Dutch word embedding evaluation dataset, which enables accurate assessment of these embeddings and fosters the development of effective Dutch language models.
Lizzy Brans, Jelke Bloem
LREC/COLING2
2024 Impact of Task Adapting on Transformer Models for Targeted Sentiment Analysis in Croatian Headlines
abstract
Transformer models, such as BERT, are often taken off-the-shelf and then fine-tuned on a downstream task. Although this is sufficient for many tasks, low-resource settings require special attention. We demonstrate an approach of performing an extra stage of self-supervised task-adaptive pre-training to a number of Croatian-supporting Transformer models. In particular, we focus on approaches to language, domain, and task adaptation. The task in question is targeted sentiment analysis for Croatian news headlines. We produce new state-of-the-art results (F1 = 0.781), but the highest performing model still struggles with irony and implicature. Overall, we find that task-adaptive pre-training benefits massively multilingual models but not Croatian-dominant models.
Sofia Lee, Jelke Bloem
LREC/COLING2
2024 Automatic Animacy Classification for Romanian Nouns
abstract
We introduce the first Romanian animacy classifier, specifically a type-based binary classifier of Romanian nouns into the classes human/non-human, using pre-trained word embeddings and animacy information derived from Romanian WordNet. By obtaining a seed set of labeled nouns and their embeddings, we are able to train classifiers that generalize to unseen nouns. We compare three different architectures and observe good performance on classifying word types. In addition, we manually annotate a small corpus for animacy to perform a token-based evaluation of Romanian animacy classification in a naturalistic setting, which reveals limitations of the type-based classification approach.
Maria Tepei, Jelke Bloem
LREC/COLING2
2021 Challenging distributional models with a conceptual network of philosophical terms
abstract
Yvette Oortwijn, Jelke Bloem, Pia Sommerauer, Francois Meyer, Wei Zhou, Antske Fokkens. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yvette Oortwijn, Jelke Bloem, Pia Sommerauer, Francois Meyer, Wei Zhou 0067, Antske Fokkens
NAACL-HLT2
2020 Expert Concept-Modeling Ground Truth Construction for Word Embeddings Evaluation in Concept-Focused Domains
abstract
We present a novel, domain expert-controlled, replicable procedure for the construction of concept-modeling ground truths with the aim of evaluating the application of word embeddings.In particular, our method is designed to evaluate the application of word and paragraph embeddings in concept-focused textual domains, where a generic ontology does not provide enough information.We illustrate the procedure, and validate it by describing the construction of an expert ground truth, QuiNE-GT.QuiNE-GT is built to answer research questions concerning the concept of naturalized epistemology in QUINE, a 2-million-token, single-author, 20th-century English philosophy corpus of outstanding quality, cleaned up and enriched for the purpose.To the best of our ken, expert conceptmodeling ground truths are extremely rare in current literature, nor has the theoretical methodology behind their construction ever been explicitly conceptualised and properly systematised.Expert-controlled concept-modeling ground truths are however essential to allow proper evaluation of word embeddings techniques, and increase their trustworthiness in specialised domains in which the detection of concepts through their expression in texts is important.We highlight challenges, requirements, and prospects for future work.
Arianna Betti, Martin Reynaert, Thijs Ossenkoppele, Yvette Oortwijn, Andrew Salway, Jelke Bloem
COLING6
2014 Applying automatically parsed corpora to the study of language variation
Jelke Bloem, Arjen Versloot, Fred Weerman
COLING1