Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marcel Bollmann

dblp:148/8804 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-2598-8150ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 80% Deep learning architectures and training · 9% Learning paradigms · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 77% Computing education · 23%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
multilingual NLP
1.012026
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP · ACL (1) 2026
Natural language and speech › Information extraction and text analysis › text normalization
historical text normalization
0.722019
Historical Text Normalization with Delayed Rewards · ACL (1) 2019
Learning attention for historical text normalization by learning to pronounce · ACL (1) 2017
Computational social science and digital humanities › scientometrics
bibliometric analysis
0.412020
On Forgetting to Cite Older Papers: An Analysis of the ACL Anthology · ACL 2020
Machine learning › Deep learning architectures and training › attention mechanism
attention learning
0.312017
Learning attention for historical text normalization by learning to pronounce · ACL (1) 2017
Machine learning › Learning paradigms
multi-task learning
0.312017
Learning attention for historical text normalization by learning to pronounce · ACL (1) 2017
Computing education
research community analysis
0.112020
On Forgetting to Cite Older Papers: An Analysis of the ACL Anthology · ACL 2020
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model
0.112019
Historical Text Normalization with Delayed Rewards · ACL (1) 2019

Methods — techniques the papers use, named apart from their topics

auditing · 1.0bibliographic analysis · 0.4reinforcement learning · 0.4policy gradient · 0.4delayed rewards · 0.4multi-task learning · 0.3encoder-decoder · 0.3attention mechanism · 0.3
YearPublicationVenuePosition
2026 How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
abstract
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather C. Lent, Miryam de Lhoneux
ACL (1)5
2026 A Dataset for Probing Translationese Preferences in English-to-Swedish Translation
Jenny Kunz, Anja Jarochenko, Marcel Bollmann
LREC3
2024 CreoleVal: Multilingual Multitask Benchmarks for Creoles
abstract
Abstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe.
Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva
Trans. Assoc. Comput. Linguistics15
2021 Error Analysis and the Role of Morphology
abstract
We evaluate two common conjectures in error analysis of NLP models: (i) Morphology is predictive of errors; and (ii) the importance of morphology increases with the morphological complexity of a language. We show across four different tasks and up to 57 languages that of these conjectures, somewhat surprisingly, only (i) is true. Using morphological features does improve error prediction across tasks; however, this effect is less pronounced with morphologically complex languages. We speculate this is because morphology is more discriminative in morphologically simple languages. Across all four tasks, case and gender are the morphological features most predictive of error.
Marcel Bollmann, Anders Søgaard
EACL1
2020 On Forgetting to Cite Older Papers: An Analysis of the ACL Anthology
abstract
The field of natural language processing is experiencing a period of unprecedented growth, and with it a surge of published papers.This represents an opportunity for us to take stock of how we cite the work of other researchers, and whether this growth comes at the expense of "forgetting" about older literature.In this paper, we address this question through bibliographic analysis.We analyze the age of outgoing citations in papers published at selected ACL venues between 2010 and 2019, finding that there is indeed a tendency for recent papers to cite more recent work, but the rate at which papers older than 15 years are cited has remained relatively stable.
Marcel Bollmann, Desmond Elliott
ACL1
2019 Historical Text Normalization with Delayed Rewards
abstract
Training neural sequence-to-sequence models with simple token-level log-likelihood is now a standard approach to historical text normalization, albeit often outperformed by phrasebased models.Policy gradient training enables direct optimization for exact matches, and while the small datasets in historical text normalization are prohibitive of from-scratch reinforcement learning, we show that policy gradient fine-tuning leads to significant improvements across the board.Policy gradient training, in particular, leads to more accurate normalizations for long or unseen words.
Simon Flachs, Marcel Bollmann, Anders Søgaard
ACL (1)2
2017 Learning attention for historical text normalization by learning to pronounce
abstract
Automated processing of historical texts often relies on pre-normalization to modern word forms.Training encoder-decoder architectures to solve such problems typically requires a lot of training data, which is not available for the named task.We address this problem by using several novel encoder-decoder architectures, including a multi-task learning (MTL) architecture using a grapheme-to-phoneme dictionary as auxiliary data, pushing the state-of-theart by an absolute 2% increase in performance.We analyze the induced models across 44 different texts from Early New High German.Interestingly, we observe that, as previously conjectured, multi-task learning can learn to focus attention during decoding, in ways remarkably similar to recently proposed attention mechanisms.This, we believe, is an important step toward understanding how MTL works.
Marcel Bollmann, Joachim Bingel, Anders Søgaard
ACL (1)1
2016 Improving historical spelling normalization with bi-directional LSTMs and multi-task learning
abstract
Natural-language processing of historical documents is complicated by the abundance of variant spellings and lack of annotated data. A common approach is to normalize the spelling of historical words to modern forms. We explore the suitability of a deep neural network architecture for this task, particularly a deep bi-LSTM network applied on a character level. Our model compares well to previously established normalization algorithms when evaluated on a diverse set of texts from Early New High German. We show that multi-task learning with additional normalization data can improve our model’s performance further.
Marcel Bollmann, Anders Søgaard
COLING1