Gertjan van Noord

dblp:09/4143 · DBLP profile ↗
← Back
44ranked-venue papers
13as first author
6since 2021 · last 2024
0000-0001-5564-6341ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 13 first-author · 6 since 2021Theory of computation · 2Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Information extraction and text analysis · 52% Transfer learning and domain adaptation · 36% Efficient and distributed learning · 7%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 72% Programming languages and type systems · 28%
Theoretical computer science
3 papers
Automata and formal languages · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.822020
UDapter: Language Adaptation for Truly Universal Dependency Parsing · EMNLP (1) 2020
Modeling Input Uncertainty in Neural Network Dependency Parsing · EMNLP 2018
Natural language and speech › Information extraction and text analysis › text normalization
lexical normalization
0.312018
Modeling Input Uncertainty in Neural Network Dependency Parsing · EMNLP 2018
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.212022
Hyper-X: A Unified Hypernetwork for Multi-Task Multilingual Transfer · EMNLP 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation › domain adaptation for NLP
multilingual model adaptation
0.112020
UDapter: Language Adaptation for Truly Universal Dependency Parsing · EMNLP (1) 2020
Compilers and program optimization
parsing
0.112011
Effective Measures of Domain Similarity for Parsing · ACL 2011
Natural language and speech › Information extraction and text analysis › computational morphology
unknown word handling
0.112010
Using Unknown Word Techniques to Learn Known Words · EMNLP 2010
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.112018
Modeling Input Uncertainty in Neural Network Dependency Parsing · EMNLP 2018
Automata and formal languages
finite automata
0.122004
Error Mining for Wide-Coverage Grammar Engineering · ACL 2004
The Intersection of Finite State Automata and Definite Clause Grammars · ACL 1995
Programming languages and type systems
grammar engineering
0.012004
Error Mining for Wide-Coverage Grammar Engineering · ACL 2004
Automata and formal languages
parsing
0.021995
The Intersection of Finite State Automata and Definite Clause Grammars · ACL 1995
Head Corner Parsing for Discontinuous Constituency · ACL 1991
Automata and formal languages › grammar formalisms › logic grammars
definite clause grammars
0.011995
The Intersection of Finite State Automata and Definite Clause Grammars · ACL 1995
Automata and formal languages › parsing
discontinuous constituency parsing
0.011991
Head Corner Parsing for Discontinuous Constituency · ACL 1991
Natural language and speech › Question answering and dialogue systems › knowledge base question answering
logical form generation
0.011989
A Semantic-Head-Driven Generation Algorithm for Unification-Based Formalisms · ACL 1989
Natural language and speech › Language models and text generation
text generation
0.011989
A Semantic-Head-Driven Generation Algorithm for Unification-Based Formalisms · ACL 1989
Knowledge, reasoning and agents › Knowledge representation and reasoning › constraint-based grammar
unification-based grammar
0.011989
A Semantic-Head-Driven Generation Algorithm for Unification-Based Formalisms · ACL 1989

Methods — techniques the papers use, named apart from their topics

adapter modules · 1.0hypernetwork · 0.6typological features · 0.4contextual parameter generation · 0.4word embeddings · 0.3normalization candidate ranking · 0.3character-level modeling · 0.3domain similarity metrics · 0.2suffix array · 0.1recursive constraints · 0.0delayed evaluation · 0.0
YearPublicationVenuePosition
2024 Neural-agent Language Learning and Communication: Emergence of Dependency Length Minimization
Yuqing Zhang 0003, Tessa Verhoef, Gertjan van Noord, Arianna Bisazza
CogSci3
2024 Endowing Neural Language Learners with Human-like Biases: A Case Study on Dependency Length Minimization
abstract
Natural languages show a tendency to minimize the linear distance between heads and their dependents in a sentence, known as dependency length minimization (DLM). Such a preference, however, has not been consistently replicated with neural agent simulations. Comparing the behavior of models with that of human learners can reveal which aspects affect the emergence of this phenomenon. In this work, we investigate the minimal conditions that may lead neural learners to develop a DLM preference. We add three factors to the standard neural-agent language learning and communication framework to make the simulation more realistic, namely: (i) the presence of noise during listening, (ii) context-sensitivity of word use through non-uniform conditional word distributions, and (iii) incremental sentence processing, or the extent to which an utterance’s meaning can be guessed before hearing it entirely. While no preference appears in production, we show that the proposed factors can contribute to a small but significant learning advantage of DLM for listeners of verb-initial languages.
Yuqing Zhang 0003, Tessa Verhoef, Gertjan van Noord, Arianna Bisazza
LREC/COLING3
2024 Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
abstract
Abstract Pretrained character-level and byte-level language models have been shown to be competitive with popular subword models across a range of Natural Language Processing tasks. However, there has been little research on their effectiveness for neural machine translation (NMT), particularly within the popular pretrain-then-finetune paradigm. This work performs an extensive comparison across multiple languages and experimental conditions of character- and subword-level pretrained models (ByT5 and mT5, respectively) on NMT. We show the effectiveness of character-level modeling in translation, particularly in cases where fine-tuning data is limited. In our analysis, we show how character models’ gains in translation quality are reflected in better translations of orthographically similar words and rare words. While evaluating the importance of source texts in driving model predictions, we highlight word-level patterns within ByT5, suggesting an ability to modulate word-level and character-level information during generation. We conclude by assessing the efficiency tradeoff of byte models, suggesting their usage in non-time-critical scenarios to boost translation quality.
Lukas Edman, Gabriele Sarti, Antonio Toral, Gertjan van Noord, Arianna Bisazza
Trans. Assoc. Comput. Linguistics4
2022 Hyper-X: A Unified Hypernetwork for Multi-Task Multilingual Transfer
abstract
Massively multilingual models are promising for transfer learning across tasks and languages.However, existing methods are unable to fully leverage training data when it is available in different task-language combinations.To exploit such heterogeneous supervision, we propose Hyper-X, a single hypernetwork that unifies multi-task and multilingual learning with efficient adaptation.This model generates weights for adapter modules conditioned on both tasks and language embeddings.By learning to combine task and language-specific knowledge, our model enables zero-shot transfer for unseen languages and task-language combinations.Our experiments on a diverse set of languages demonstrate that Hyper-X achieves the best or competitive gain when a mixture of multiple resources is available, while being on par with strong baselines in the standard scenario.Hyper-X is also considerably more efficient in terms of parameters and resources compared to methods that train separate adapters.Finally, Hyper-X consistently produces strong results in few-shot scenarios for new languages, showing the versatility of our approach beyond zero-shot transfer.1 NER en Pre-trained Model Pre-trained Model Fine-tuned Model ar tr Fine-tuned Model ar tr POS en Single-Task Pre-trained Model Fine-tuned Model ar tr ar tr NER POS en en Multi-Task Pre-trained Model Fine-tuned Model tr NER POS ar tr NER ar POS en en Mixed-Language Multi-Task Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš.2020a.From Zero to Hero: On the Limitations of Zero-Shot
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord, Sebastian Ruder
EMNLP4
2022 Evaluating Pre-training Objectives for Low-Resource Translation into Morphologically Rich Languages
abstract
The scarcity of parallel data is a major limitation for Neural Machine Translation (NMT) systems, in particular for translation into morphologically rich languages (MRLs). An important way to overcome the lack of parallel data is to leverage target monolingual data, which is typically more abundant and easier to collect. We evaluate a number of techniques to achieve this, ranging from back-translation to random token masking, on the challenging task of translating English into four typologically diverse MRLs, under low-resource settings. Additionally, we introduce Inflection Pre-Training (or PT-Inflect), a novel pre-training objective whereby the NMT system is pre-trained on the task of re-inflecting lemmatized target sentences before being trained on standard source-to-target language translation. We conduct our evaluation on four typologically diverse target MRLs, and find that PT-Inflect surpasses NMT systems trained only on parallel data. While PT-Inflect is outperformed by back-translation overall, combining the two techniques leads to gains in some of the evaluated language pairs.
Prajit Dhar, Arianna Bisazza, Gertjan van Noord
LREC3
2022 UDapter: Typology-based Language Adapters for Multilingual Dependency Parsing and Sequence Labeling
abstract
Abstract Recent advances in multilingual language modeling have brought the idea of a truly universal parser closer to reality. However, such models are still not immune to the “curse of multilinguality”: Cross-language interference and restrained model capacity remain major obstacles. To address this, we propose a novel language adaptation approach by introducing contextual language adapters to a multilingual parser. Contextual language adapters make it possible to learn adapters via language embeddings while sharing model parameters across languages based on contextual parameter generation. Moreover, our method allows for an easy but effective integration of existing linguistic typology features into the parsing model. Because not all typological features are available for every language, we further combine typological feature prediction with parsing in a multi-task model that achieves very competitive parsing performance without the need for an external prediction system for missing features. The resulting parser, UDapter, can be used for dependency parsing as well as sequence labeling tasks such as POS tagging, morphological tagging, and NER. In dependency parsing, it outperforms strong monolingual and multilingual baselines on the majority of both high-resource and low-resource (zero-shot) languages, showing the success of the proposed adaptation approach. In sequence labeling tasks, our parser surpasses the baseline on high resource languages, and performs very competitively in a zero-shot setting. Our in-depth analyses show that adapter generation via typological features of languages is key to this success.1
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord
Comput. Linguistics4
2020 Low-Resource Unsupervised NMT: Diagnosing the Problem and Providing a Linguistically Motivated Solution
abstract
Unsupervised Machine Translation has been advancing our ability to translate without parallel data, but state-of-the-art methods assume an abundance of monolingual data. This paper investigates the scenario where monolingual data is limited as well, finding that current unsupervised methods suffer in performance under this stricter setting. We find that the performance loss originates from the poor quality of the pretrained monolingual embeddings, and we offer a potential solution: dependency-based word embeddings. These embeddings result in a complementary word representation which offers a boost in performance of around 1.5 BLEU points compared to standard word2vec when monolingual data is limited to 1 million sentences per language. We also find that the inclusion of sub-word information is crucial to improving the quality of the embeddings.
Lukas Edman, Antonio Toral, Gertjan van Noord
EAMT3
2020 UDapter: Language Adaptation for Truly Universal Dependency Parsing
abstract
Recent advances in multilingual dependency parsing have brought the idea of a truly universal parser closer to reality.However, crosslanguage interference and restrained model capacity remain major obstacles.To address this, we propose a novel multilingual task adaptation approach based on contextual parameter generation and adapter modules.This approach enables to learn adapters via language embeddings while sharing model parameters across languages.It also allows for an easy but effective integration of existing linguistic typology features into the parsing network.The resulting parser, UDapter, outperforms strong monolingual and multilingual baselines on the majority of both high-resource and lowresource (zero-shot) languages, showing the success of the proposed adaptation approach.Our in-depth analyses show that soft parameter sharing via typological features is key to this success.1
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, Gertjan van Noord
EMNLP (1)4
2020 A Shared Task of a New, Collaborative Type to Foster Reproducibility: A First Exercise in the Area of Language Science and Technology with REPROLANG2020
abstract
n this paper, we introduce a new type of shared task — which is collaborative rather than competitive — designed to support and fosterthe reproduction of research results. We also describe the first event running such a novel challenge, present the results obtained, discussthe lessons learned and ponder on future undertakings.
António Branco, Nicoletta Calzolari, Piek Vossen, Gertjan van Noord, Dieter Van Uytvanck, João Silva 0004, Luís Gomes 0002, Willem Elbers
LREC4
2018 Modeling Input Uncertainty in Neural Network Dependency Parsing
abstract
Recently introduced neural network parsers allow for new approaches to circumvent data sparsity issues by modeling character level information and by exploiting raw data in a semi-supervised setting.Data sparsity is especially prevailing when transferring to nonstandard domains.In this setting, lexical normalization has often been used in the past to circumvent data sparsity.In this paper, we investigate whether these new neural approaches provide similar functionality as lexical normalization, or whether they are complementary.We provide experimental results which show that a separate normalization component improves performance of a neural network parser even if it has access to character level information as well as external word embeddings.Further improvements are obtained by a straightforward but novel approach in which the top-N best candidates provided by the normalization component are available to the parser.
Rob van der Goot, Gertjan van Noord
EMNLP2
2018 A Taxonomy for In-depth Evaluation of Normalization for User Generated Content
Rob van der Goot, Rik van Noord, Gertjan van Noord
LREC3
2018 Simple Embedding-Based Word Sense Disambiguation
abstract
We present a simple knowledge-based WSD method that uses word and sense embeddings to compute the similarity between the gloss of a sense and the context of the word.Our method is inspired by the Lesk algorithm as it exploits both the context of the words and the definitions of the senses.It only requires large unlabeled corpora and a sense inventory such as WordNet, and therefore does not rely on annotated data.We explore whether additional extensions to Lesk are compatible with our method.The results of our experiments show that by lexically extending the amount of words in the gloss and context, although it works well for other implementations of Lesk, harms our method.Using a lexical selection method on the context words, on the other hand, improves it.The combination of our method with lexical selection enables our method to outperform state-of the art knowledgebased systems.
Dieke Oele, Gertjan van Noord
GWC2
2018 Reproducibility in Computational Linguistics: Are We Willing to Share?
abstract
This study focuses on an essential precondition for reproducibility in computational linguistics: the willingness of authors to share relevant source code and data. Ten years after Ted Pedersen’s influential “Last Words” contribution in Computational Linguistics, we investigate to what extent researchers in computational linguistics are willing and able to share their data and code. We surveyed all 395 full papers presented at the 2011 and 2016 ACL Annual Meetings, and identified whether links to data and code were provided. If working links were not provided, authors were requested to provide this information. Although data were often available, code was shared less often. When working links to code or data were not provided in the paper, authors provided the code in about one third of cases. For a selection of ten papers, we attempted to reproduce the results using the provided data and code. We were able to reproduce the results approximately for six papers. For only a single paper did we obtain the exact same results. Our findings show that even though the situation appears to have improved comparing 2016 to 2011, empiricism in computational linguistics still largely remains a matter of faith. Nevertheless, we are somewhat optimistic about the future. Ensuring reproducibility is not only important for the field as a whole, but also seems worthwhile for individual researchers: The median citation count for studies with working links to the source code is higher.
Martijn Wieling 0001, Josine Rawee, Gertjan van Noord
Comput. Linguistics3
2016 Bilingual Learning of Multi-sense Embeddings with Discrete Autoencoders
abstract
We present an approach to learning multi-sense word embeddings relying both on monolingual and bilingual information.Our model consists of an encoder, which uses monolingual and bilingual context (i.e. a parallel sentence) to choose a sense for a given word, and a decoder which predicts context words based on the chosen sense.The two components are estimated jointly.We observe that the word representations induced from bilingual data outperform the monolingual counterparts across a range of evaluation tasks, even though crosslingual information is not available at test time.
Simon Suster, Ivan Titov 0001, Gertjan van Noord
HLT-NAACL3
2016 In Memoriam: Susan Armstrong
abstract
She had a fundamental role in the founding and successful development of SIGDAT.Susan arrived in Switzerland from the United States in 1978 to work at the University of Lausanne in the German Department teaching literature and linguistics, as well as pursuing her interest in computers and language as translator and consultant to Logitech.She then moved to Geneva to work at the ISSCO research institute where she participated in a number of European research projects related to natural language processing (NLP) and machine translation, also working with corpora in translation
Pierrette Bouillon, Paola Merlo, Gertjan van Noord, Mike Rosner
Comput. Linguistics3
2014 From neighborhood to parenthood: the advantages of dependency representation over bigrams in Brown clustering
Simon Suster, Gertjan van Noord
COLING2
2014 Treelet Probabilities for HPSG Parsing and Error Correction
Angelina Ivanova, Gertjan van Noord
LREC2
2011 Effective Measures of Domain Similarity for Parsing
Barbara Plank, Gertjan van Noord
ACL2
2011 An Empirical Comparison of Unknown Word Prediction Methods
Kostadin Cholakov, Gertjan van Noord, Valia Kordoni, Yi Zhang 0003
IJCNLP2
2010 Using Unknown Word Techniques to Learn Known Words
Kostadin Cholakov, Gertjan van Noord
EMNLP2
2010 POS Multi-tagging Based on Combined Models
Gertjan van Noord
LREC2
2009 Learning Efficient Parsing
Gertjan van Noord
EACL1
2008 From D-Coi to SoNaR: a reference corpus for Dutch
Nelleke Oostdijk, Martin Reynaert, Paola Monachesi, Gertjan van Noord, Roeland Ordelman, Ineke Schuurman, Vincent Vandeghinste
LREC4
2006 Syntactic Annotation of Large Corpora in STEVIN
Gertjan van Noord, Ineke Schuurman, Vincent Vandeghinste
LREC1
2004 Error Mining for Wide-Coverage Grammar Engineering
abstract
Parsing systems which rely on hand-coded linguistic descriptions can only perform adequately in as far as these descriptions are correct and complete.The paper describes an error mining technique to discover problems in hand-coded linguistic descriptions for parsing such as grammars and lexicons. By analysing parse results for very large unannotated corpora, the technique discovers missing, incorrect or incomplete linguistic descriptions.The technique uses the frequency of n-grams of words for arbitrary values of n. It is shown how a new combination of suffix arrays and perfect hash finite automata allows an efficient implementation.
Gertjan van Noord
ACL1
2004 Finite automata for compact representation of tuple dictionaries
Jan Daciuk, Gertjan van Noord
Theor. Comput. Sci.2
2003 Finite state methods in natural language processing
abstract
Finite state methods have been in common use in various areas of natural language processing (NLP) for many years. A series of specialized workshops in this area illustrates this. In 1996, András Kornai organized a very successful workshop entitled Extended Finite State Models of Language. One of the results of that workshop was a special issue of Natural Language Engineering (Volume 2, Number 4). In 1998, Kemal Oflazer organized a workshop called Finite State Methods in Natural Language Processing. A selection of submissions for this workshop were later included in a special issue of Computational Linguistics (Volume 26, Number 1). Inspired by these events, Lauri Karttunen, Kimmo Koskenniemi and Gertjan van Noord took the initiative for a workshop on finite state methods in NLP in Helsinki, as part of the European Summer School in Language, Logic and Information. As a related special event, the 20th anniversary of two-level morphology was celebrated. The appreciation of these events led us to believe that once again it should be possible, with some additional submissions, to compose an interesting special issue of this journal.
Lauri Karttunen, Kimmo Koskenniemi, Gertjan van Noord
Nat. Lang. Eng.3
2001 Finite Automata for Compact Representation of Language Models in NLP
Jan Daciuk, Gertjan van Noord
CIAA2
2000 Treatment of Epsilon Moves in Subset Construction
abstract
The paper discusses the problem of determinizing finite-state automata containing large numbers of ε-moves. Experiments with finite-state approximations of natural language grammars often give rise to very large automata with a very large number of ε-moves. The paper identifies and compares a number of subset construction algorithms that treat ε-moves. Experiments have been performed which indicate that the algorithms differ considerably in practice, both with respect to the size of the resulting deterministic automaton, and with respect to practical efficiency. Furthermore, the experiments suggest that the average number of ε-moves per state can be used to predict which algorithm is likely to be the fastest for a given input automaton.
Gertjan van Noord
Comput. Linguistics1
1999 Transducers from Rewrite Rules with Backreferences
Dale Gerdemann, Gertjan van Noord
EACL2
1999 Robust grammatical analysis for spoken dialogue systems
abstract
We argue that grammatical analysis is a viable alternative to concept spotting for processing spoken input in a practical spoken dialogue system. We discuss the structure of the grammar, and a model for robust parsing which combines linguistic sources of information and statistical sources of information. We discuss test results suggesting that grammatical processing allows fast and accurate processing of spoken input.
Gertjan van Noord, Gosse Bouma, Rob Koeling, Mark-Jan Nederhof
Nat. Lang. Eng.1
1997 An Efficient Implementation of the Head-Corner Parser
Gertjan van Noord
Comput. Linguistics1
1995 The Intersection of Finite State Automata and Definite Clause Grammars
abstract
Bernard Lang defines parsing as the calculation of the intersection of a FSA (the input) and a CFG. Viewing the input for parsing as a FSA rather than as a string combines well with some approaches in speech understanding systems, in which parsing takes a word lattice as input (rather than a word string). Furthermore, certain techniques for robust parsing can be modelled as finite state transducers.In this paper we investigate how we can generalize this approach for unification grammars. In particular we will concentrate on how we might the calculation of the intersection of a FSA and a DCG. It is shown that existing parsing algorithms can be easily extended for FSA inputs. However, we also show that the termination properties change drastically: we show that it is undecidable whether the intersection of a FSA and a DCG is empty (even if the DCG is off-line parsable).Furthermore we discuss approaches to cope with the problem.
Gertjan van Noord
ACL1
1994 Constraint-Based Categorical Grammar
abstract
We propose a generalization of Categorial Grammar in which lexical categories are defined by means of recursive constraints. In particular, the introduction of relational constraints allows one to capture the effects of (recursive) lexical rules in a computationally attractive manner. We illustrate the linguistic merits of the new approach by showing how it accounts for the syntax of Dutch cross-serial dependencies and the position and scope of adjuncts in such constructions. Delayed evaluation is used to process grammars containing recursive constraints.
Gosse Bouma, Gertjan van Noord
ACL2
1994 Adjuncts and the Processing of Lexical Rules
Gertjan van Noord, Gosse Bouma
COLING1
1994 Head-Corner Parsing for TAG
abstract
This paper describes a bidirectional head‐corner parser for (unification‐based versions of) lexicalized tree‐adjoining grammars.
Gertjan van Noord
Comput. Intell.1
1993 Head-driven Parsing for Lexicalist Grammars: Experimental Results
Gosse Bouma, Gertjan van Noord
EACL2
1992 Self-Monitoring with Reversible Grammars
Günter Neumann, Gertjan van Noord
COLING2
1991 Head Corner Parsing for Discontinuous Constituency
abstract
I describe a head-driven parser for a class of grammars that handle discontinuous constituency by a richer notion of string combination than ordinary concatenation. The parser is a generalization of the left-corner parser (Matsumoto et al., 1983) and can be used for grammars written in powerful formalisms such as non-concatenative versions of HPSG (Pollard, 1984; Reape, 1989).
Gertjan van Noord
ACL1
1991 An overview of MiMo2
Gertjan van Noord, Joke Dorrepaal, Pim van der Eijk, Maria Florenza, Herbert Ruessink, Louis des Tombe
Mach. Transl.1
1990 Reversible Unification Based Machine Translation
Gertjan van Noord
COLING1
1990 Semantic-Head-Driven Generation
Stuart M. Shieber, Gertjan van Noord, Fernando Pereira 0003, Robert C. Moore
Comput. Linguistics2
1989 A Semantic-Head-Driven Generation Algorithm for Unification-Based Formalisms
abstract
We present an algorithm for generating strings from logical form encodings that improves upon previous algorithms in that it places fewer restrictions on the class of grammars to which it is applicable. In particular, unlike an Earley deduction generator (Shieber, 1988), it allows use of semantically nonmonotonic grammars, yet unlike topdown methods, it also permits left-recursion. The enabling design feature of the algorithm is its implicit traversal of the analysis tree for the string being generated in a semantic-head-driven fashion.
Stuart M. Shieber, Gertjan van Noord, Robert C. Moore, Fernando Pereira 0003
ACL2
1989 An Approach To Sentence-Level Anaphora In Machine Translation
Gertjan van Noord, Joke Dorrepaal, Doug Arnold, Steven Krauwer, Louisa Sadler, Louis des Tombe
EACL1