Francis Bond

dblp:61/1808 · DBLP profile ↗
← Back
30ranked-venue papers in the field
9as first author
9since 2021 · last 2023
0000-0003-4973-8068ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 30 (9 first)
YearPublicationVenuePosition
2023 Documenting the Open Multilingual Wordnet
abstract
In this project note we describe our work to make better documentation for the Open Multilingual Wordnet (OMW), a platform integrating many open wordnets.This includes the documentation of the OMW website itself as well as of semantic relations used by the component wordnets.Some of this documentation work was done with the support of the Google Season of Docs.The OMW project page, which links both to the actual OMW server and the documentation has been moved to a new location: https://omwn.org.
Francis Bond, Michael Wayne Goodman, Ewa Rudnicka, Luís Morgado da Costa, Alexandre Rademaker, John P. McCrae
GWC1
2023 The Japanese Wordnet 2.0
abstract
This paper describes a new release of the Japanese wordnet.It uses the new global wordnet formats (McCrae et al., 2021) to incorporate a range of new information: orthographic variants (including hiragana, katakana and Latin representations) first described in Kuroda et al. (2011), classifiers, pronouns and exclamatives (Morgado da Costa and Bond, 2016) and many new senses, motivated both from corpus annotation and linking to the TUFs basic vocabulary (Bond et al., 2020).The wordnet has been moved to github and is available at https://bond-lab.github. io/wnja/.
Francis Bond, Takayuki Kuribayashi
GWC1
2023 Linking SIL Semantic Domains to Wordnet and Expanding the Abui Wordnet through Rapid Word Collection Methodology
abstract
In this paper we describe a new methodology to expand the Abui Wordnet through data collected using the Rapid Word Collection (RWC) method – based on SIL’s Semantic Domains. Using a multilingual sense-intersection algorithm, we created a ranked list of concept suggestions for each domain, and then used the ranked list as a filter to link the Abui RWC data to wordnet. This used translations from both SIL’s Semantic Domain’s structure and example words, both available through SIL’s Fieldworks software and the RWC project. We release both the new mapping of the SIL Semantic Domains to wordnet and an expansion of the Abui Wordnet.
Luís Morgado da Costa, Frantisek Kratochvil, George Saad, Benidiktus Delpada, Daniel Simon Lanma, Francis Bond, Natálie Wolfová, A. l. Blake
GWC6
2021 Taboo Wordnet
abstract
This paper describes the development of an online lexical resource to help detection systems regulate and curb the use of offensive words online.With the growing prevalence of social media platforms, many conversations are now conducted online.The increase of online conversations for leisure, work and socializing has led to an increase in harassment.In particular, we create a specialized sense-based vocabulary of Japanese offensive words for the Open Multilingual Wordnet.This vocabulary expands on an existing list of Japanese offensive words and provides categorization and proper linking to synsets within the multilingual wordnet.This paper then discusses the evaluation of the vocabulary as a resource for representing and classifying offensive words and as a possible resource for offensive word use detection in social media.Content Warning: this paper deals with obscene words and contains many examples of them.
Francis Bond, Merrick Yeu Herng Choo
GWC1
2021 Teaching Through Tagging - Interactive Lexical Semantics
abstract
In this paper we discuss an ongoing effort to enrich students' learning by involving them in sense tagging.The main goal is to lead students to discover how we can represent meaning and where the limits of our current theories lie.A subsidiary goal is to create sense tagged corpora and an accompanying linked lexicon (in our case wordnets).We present the results of tagging several texts and suggest some ways in which the tagging process could be improved.Two authors of this paper present their own experience as students.Overall, students reported that they found the tagging an enriching experience.The annotated corpora and changes to the wordnet are made available through the NTU multilingual corpus and associated wordnets (NTU-MC).
Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, Luís Morgado da Costa
GWC1
2021 Intrinsically Interlingual: The Wn Python Library for Wordnets
abstract
This paper introduces Wn, a new Python library for working with wordnets.Unlike previous libraries, Wn is built from the beginning to accommodate multiple wordnets-for multiple languages or multiple versions of the same wordnet-while retaining the ability to query and traverse them independently.It is also able to download and incorporate wordnets published online.These features are made possible through Wn's adoption of standard formats and methods for interoperability, namely the WN-LMF schema (Vossen et al., 2013;Bond et al., 2020) and the Collaborative Interlingual Index (Bond et al., 2016).Wn is open-source, easily available, 1 and well-documented. 2
Michael Wayne Goodman, Francis Bond
GWC2
2021 Testing agreement between lexicographers: A case of homonymy and polysemy
abstract
In this paper we compare Oxford Lexico and Merriam Webster dictionaries with Princeton WordNet with respect to the description of semantic (dis)similarity between polysemous and homonymous senses that could be inferred from them.WordNet lacks any explicit description of polysemy or homonymy, but as a network of linked senses it may be used to compute semantic distances between word senses.To compare WordNet with the dictionaries, we transformed sample entry microstructures of the latter into graphs and crosslinked them with the equivalent senses of the former.We found that dictionaries are in high agreement with each other, if one considers polysemy and homonymy altogether, and in moderate concordance, if one focuses merely on polysemy descriptions.Measuring the shortest path lengths on WordNet gave results comparable to those on the dictionaries in predicting semantic dissimilarity between polysemous senses, but was less felicitous while recognising homonymy.
Marek Maziarz, Francis Bond, Ewa Rudnicka
GWC2
2021 The GlobalWordNet Formats: Updates for 2020
abstract
The Global Wordnet Formats have been introduced to enable wordnets to have a common representation that can be integrated through the Global WordNet Grid.As a result of their adoption, a number of shortcomings of the format were identified, and in this paper we describe the extensions to the formats that address these issues.These include: ordering of senses, dependencies between wordnets, pronunciation, syntactic modelling, relations, sense keys, metadata and RDF support.Furthermore, we provide some perspectives on how these changes help in the integration of wordnets.
John P. McCrae, Michael Wayne Goodman, Francis Bond, Alexandre Rademaker, Ewa Rudnicka, Luís Morgado da Costa
GWC3
2021 OdeNet: Compiling a GermanWordNet from other Resources
abstract
The Princeton WordNet for the English language has been used worldwide in NLP projects for many years.With the OMW initiative, wordnets for different languages of the world are being linked via identifiers.The parallel development and linking allows new multilingual application perspectives.The development of a wordnet for the German language is also in this context.To save development time, existing resources were combined and recompiled.The result was then evaluated and improved.In a relatively short time a resource was created that can be used in projects and continuously improved and extended.
Melanie Siegel, Francis Bond
GWC2
2019 GeoNames Wordnet (geown): extracting wordnets from GeoNames
abstract
This paper introduces a new multilingual lexicon of geographical place names.The names are based on (and linked to) the GeoNames collection.Each location is treated as a new synset, which is linked by instance_hypernym to a small set of supertypes.These supertypes are linked to the collaborative interlingual index, based on mappings from GeoDomainWordnet.If a location is already in the interlingual index, then it is also linked to the entry, using mappings from the Geo-Wordnet.Finally, if GeoNames places the location in a larger location, this is linked using the mero_location link.Wordnets can be built for any language in GeoNames, we give results for those wordnets in the Open Multilingual Wordnet.We discuss how it is mapped and the characteristics of the extracted wordnets.
Francis Bond, Arthur Bond
GWC1
2019 Testing Zipf's meaning-frequency law with wordnets as sense inventories
abstract
According to George K. Zipf, more frequent words have more senses.We have tested this law using corpora and wordnets of English, Spanish, Portuguese, French, Polish, Japanese, Indonesian and Chinese.We have proved that the law works pretty well for all of these languages if we takeas Zipf did -mean values of meaning count and averaged ranks.On the other hand, the law disastrously fails in predicting the number of senses for a single lemma.We have also provided the evidence that slope coefficients of Zipfian log-log linear model may vary from language to language.
Francis Bond, Arkadiusz Janz, Marek Maziarz, Ewa Rudnicka
GWC1
2019 A Comparison of Sense-level Sentiment Scores
abstract
In this paper, we compare a variety of sense-tagged sentiment resources, including SentiWordNet, ML-Senticon, plWord-Net emo and the NTU Multilingual Corpus.The goal is to investigate the quality of the resources and see how well the sentiment polarity annotation maps across languages.
Francis Bond, Arkadiusz Janz, Maciej Piasecki
GWC1
2019 New Polysemy Structures in Wordnets Induced by Vertical Polysemy
abstract
This paper aims to study auto-hyponymy and auto-troponymy relations (or vertical polysemy) in 11 wordnets uploaded into the new Open Multilingual Wordnet (OMW) webpage.We investigate how vertical polysemy forms polysemy structures (or sense clusters) in semantic hierarchies of the wordnets.Our main results and discoveries are new polysemy structures that have not previously been associated with vertical polysemy, along with some inconsistencies of semantic relations analysis in the studied wordnets, which should not be there.In the case study, we turn attention to polysemy structures in the Estonian Wordnet (version 2.2.0), analyzing them and giving the lexicographers comments.In addition, we describe the detection algorithm of polysemy structures and an overview of the state of polysemy structures in 11 wordnets.
Ahti Lohk, Heili Orav, Kadri Vare, Francis Bond, Rasmus Vaik
GWC4
2019 English WordNet 2019 - An Open-Source WordNet for English
abstract
We describe the release of a new wordnet for English based on the Princeton WordNet, but now developed under an open-source model.In particular, this version of WordNet, which we call English WordNet 2019, which has been developed by multiple people around the world through GitHub, fixes many errors in previous wordnets for English.We give some details of the changes that have been made in this version and give some perspectives about likely future changes that will be made as this project continues to evolve.
John P. McCrae, Alexandre Rademaker, Francis Bond, Ewa Rudnicka, Christiane Fellbaum
GWC3
2018 The Company They Keep: Extracting Japanese Neologisms Using Language Patterns
abstract
We describe an investigation into the identification and extraction of unrecorded potential lexical items in Japanese text by detecting text passages containing selected language patterns typically associated with such items.We identified a set of suitable patterns, then tested them with two large collections of text drawn from the WWW and Twitter.Samples of the extracted items were evaluated, and it was demonstrated that the approach has considerable potential for identifying terms for later lexicographic analysis.
James Breen, Timothy Baldwin, Francis Bond
GWC3
2018 Toward Constructing the National Cancer Institute Thesaurus Derived WordNet (ncitWN)
abstract
We describe preliminary work in the creation of the first specialized vocabulary to be integrated into the Open Multilingual Wordnet (OMW).The NCIt Derived WordNet (ncitWN) is based on the National Cancer Institute Thesaurus (NCIt), a controlled biomedical terminology that includes formal class restrictions and English definitions developed by groups of clinicians and terminologists.The ncitWN is created by converting the NCIt to the WordNet Lexical Markup Framework and adding semantic types.We report the development of a prototype ncitWN and first steps towards integrating it into the OMW.
Amanda Hicks, Selja Seppälä, Francis Bond
GWC3
2018 Automatic Identification of Basic-Level Categories
abstract
Basic-level categories have been shown to be both psychologically significant and useful in a wide range of practical applications.We build a rule-based system to identify basic-level categories in WordNet, achieving 77% accuracy on a test set derived from prior psychological experiments.With additional annotations we found our system also has low precision, in part due to the existence of many categories that do not fit into the three classes (superordinate, basic-level, and subordinate) relied on in basiclevel category research.
Chad Mills, Francis Bond, Gina-Anne Levow
GWC2
2018 Putting Figures on Influences on Moroccan Darija from Arabic, French and Spanish using the WordNet
abstract
Moroccan Darija is a variant of Arabic with many influences.Using the Open Multilingual WordNet (OMW), we compare the lemmas in the Moroccan Darija Wordnet (MDW) with the standard Arabic, French and Spanish ones.We then compared the lemmas in each synset with their translation equivalents.Transliteration is used to bridge alphabet differences and match lemmas in the closest phonological way.The results put figures on the similarity Moroccan Darija has with Arabic, French and Spanish: respectively 42.0%, 2.8% and 2.2%.
Khalil Mrini, Francis Bond
GWC2
2018 Lexical Perspective on Wordnet to Wordnet Mapping
abstract
The paper presents a feature-based model of equivalence targeted at (manual) sense linking between Princeton WordNet and plWordNet.The model incorporates insights from lexicographic and translation theories on bilingual equivalence and draws on the results of earlier synsetlevel mapping of nouns between Princeton WordNet and plWordNet.It takes into account all basic aspects of language such as form, meaning and function and supplements them with (parallel) corpus frequency and translatability.Three types of equivalence are distinguished, namely strong, regular and weak depending on the conformity with the proposed features.The presented solutions are languageneutral and they can be easily applied to language pairs other than Polish and English.Sense-level mapping is a more finegrained mapping than the existing synset mappings and is thus of great potential to human and machine translation.
Ewa Rudnicka, Francis Bond, Lukasz Grabowski, Maciej Piasecki, Tadeusz Piotrowski
GWC2
2018 Enchancing the Collaborative Interlingual Index for Digital Humanities: Cross-linguistic Analysis in the Domain of Theology
abstract
We aim to support digital humanities work related to the study of sacred texts. To do this, we propose to build a cross-lingual wordnet within the do-main of theology. We target the Collaborative Interlingual Index (CILI) directly instead of each individual wordnet. The paper presents background for this proposal: (1) an overview of concepts relevant to theology and (2) a summary of the domain-associated issues observed in the Princeton WordNet (PWN). We have found that definitions for concepts in this domain can be too restrictive, inconsistent, and unclear. Necessary synsets are missing, with the PWN being skewed towards Christianity. We argue that tackling problems in a single domain is a better method for improving CILI. By focusing on a single topic rather than a single language, this will result in the proper construction of definitions, romanization/translation of lemmas, and also improvements in use of/creation of a cross-lingual domain hierarchy.
Laura A. Slaughter, Luís Morgado da Costa, Francis Bond
GWC4
2018 Multilingual Wordnet sense Ranking using nearest context
abstract
In this paper, we combine methods to estimate sense rankings from raw text with recent work on word embeddings to provide sense ranking estimates for the entries in the Open Multilingual WordNet (OMW).The existing Word2Vec pre-trained models from Polygot2 are only built for single word entries, we, therefore, re-train them with multiword expressions from the wordnets, so that multiword expressions can also be ranked.Thus this trained model gives embeddings for both single words and multiwords.The resulting lexicon gives a WSD baseline for five languages.The results are evaluated for Semcor sense corpora for 5 languages using Word2Vec and Glove models.The Glove model achieves an average accuracy of 0.47 and Word2Vec achieves 0.31 for languages such as English, Italian, Indonesian, Chinese and Japanese.The experimentation on OMW sense ranking proves that the rank correlation is generally similar to the human ranking.Hence distributional semantics can aid in Wordnet Sense Ranking.
E. Umamaheswari Vasanthakumar, Francis Bond
GWC2
2016 Multilingual Sense Intersection in a Parallel Corpus with Diverse Language Families
abstract
Supervised methods for Word Sense Disambiguation (WSD) benefit from highquality sense-annotated resources, which are lacking for many languages less common than English.There are, however, several multilingual parallel corpora that can be inexpensively annotated with senses through cross-lingual methods.We test the effectiveness of such an approach by attempting to disambiguate English texts through their translations in Italian, Romanian and Japanese.Specifically, we try to find the appropriate word senses for the English words by comparison with all the word senses associated to their translations.The main advantage of this approach is in that it can be applied to any parallel corpus, as long as large, highquality inter-linked sense inventories exist for all the languages considered.
Giulia Bonansinga, Francis Bond
GWC2
2016 CILI: the Collaborative Interlingual Index
abstract
This paper introduces the motivation for and design of the Collaborative InterLingual Index (CILI).It is designed to make possible coordination between multiple loosely coupled wordnet projects.The structure of the CILI is based on the Interlingual index first proposed in the Eu-roWordNet project with several pragmatic extensions: an explicit open license, definitions in English and links to wordnets in the Global Wordnet Grid.
Francis Bond, Piek Vossen, John P. McCrae, Christiane Fellbaum
GWC1
2016 Mapping and Generating Classifiers using an Open Chinese Ontology
abstract
In languages such as Chinese, classifiers (CLs) play a central role in the quantification of noun-phrases.This can be a problem when generating text from input that does not specify the classifier, as in machine translation (MT) from English to Chinese.Many solutions to this problem rely on dictionaries of noun-CL pairs.However, there is no open large-scale machine-tractable dictionary of noun-CL associations.Many published resources exist, but they tend to focus on how a CL is used (e.g.what kinds of nouns can be used with it, or what features seem to be selected by each CL).In fact, since nouns are open class words, producing an exhaustive definite list of noun-CL associations is not possible, since it would quickly get out of date.Our work tries to address this problem by providing an algorithm for automatic building of a frequency based dictionary of noun-CL pairs, mapped to concepts in the Chinese Open Wordnet (Wang and Bond, 2013), an open machinetractable dictionary for Chinese.All results will released under an open license.
Luís Morgado da Costa, Francis Bond, Helena Gao
GWC2
2016 Identifying and Exploiting Definitions in Wordnet Bahasa
abstract
This paper describes our attempts to add Indonesian definitions to synsets in the Wordnet Bahasa (Nurril Hirfana Mohamed Noor et al., 2011;Bond et al., 2014), to extract semantic relations between lemmas and definitions for nouns and verbs, such as synonym, hyponym, hypernym and instance hypernym, and to generally improve Wordnet.The original, somewhat noisy, definitions for Indonesian came from the Asian Wordnet project (Riza et al., 2010).The basic method of extracting the relations is based on Bond et al. (2004).Before the relations can be extracted, the definitions were cleaned up and tokenized.We found that the definitions cannot be completely cleaned up because of many misspellings and bad translations.However, we could identify four semantic relations in 57.10% of noun and verb definitions.For the remaining 42.90%, we propose to add 149 new Indonesian lemmas and make some improvements to Wordnet Bahasa and Wordnet in general.
David Moeljadi, Francis Bond
GWC2
2016 Toward a truly multilingual GlobalWordnet Grid
abstract
In this paper, we describe a new and improved Global Wordnet Grid that takes advantage of the Collaborative InterLingual Index (CILI).Currently, the Open Multilingal Wordnet has made many wordnets accessible as a single linked wordnet, but as it used the Princeton Wordnet of English (PWN) as a pivot, it loses concepts that are not part of PWN.The technical solution to this, a central registry of concepts, as proposed in the EuroWordnet project through the InterLingual Index, has been known for many years.However, the practical issues of how to host this index and who decides what goes in remained unsolved.Inspired by current practice in the Semantic Web and the Linked Open Data community, we propose a way to solve this issue.In this paper we define the principles and protocols for contributing to the Grid.We tested them on two use cases, adding version 3.1 of the Princeton WordNet to a CILI based on 3.0 and adding the Open Dutch Wordnet, to validate the current set up.This paper aims to be a call for action that we hope will be further discussed and ultimately taken up by the whole wordnet community.
Piek Vossen, Francis Bond, John P. McCrae
GWC2
2014 Issues in building English-Chinese parallel corpora with WordNets
abstract
We discuss some of the issues in producing sense-tagged parallel corpora: including pre-processing, adding new entries and linking.We have preliminary results for three genres: stories, essays and tourism web pages, in both Chinese and English.
Francis Bond, Shan Wang 0002
GWC1
2014 A Survey of WordNet Annotated Corpora
abstract
This paper surveys the current state of wordnet sense annotated corpora.We look at corpora in any language, and describe them in terms of accessibility and usefulness.We finally discuss possibilities in increasing the interoperability of the corpora, especially across languages.
Tommaso Petrolito, Francis Bond
GWC2
2014 Bringing together over- and under- represented languages: Linking WordNet to the SIL Semantic Domains
abstract
We have created an open-source mapping between the SIL's semantic domains (used for rapid lexicon building and organization for under-resourced languages) and WordNet, the standard resource for lexical semantics in natural language processing.We show that the resources complement each other, and suggest ways in which the mapping can be improved even further.The semantic domains give more general domain and associative links, which wordnet still has few of, while wordnet gives explicit semantic relations between senses, which the domains lack.
Muhammad Zulhelmy Bin Mohd Rosman, Frantisek Kratochvil, Francis Bond
GWC3
2014 Parse Ranking with Semantic Dependencies and WordNet
abstract
In this paper, we investigate which features are useful for ranking semantic representations of text.We show that two methods of generalization improved results: extended grand-parenting and supertypes.The models are tested on a subset of SemCor that has been annotated with both Dependency Minimal Recursion Semantic representations and WordNet senses.Using both types of features gives a significant improvement in whole sentence parse selection accuracy over the baseline model.
Xiaocheng Yin, Jung-Jae Kim 0001, Zinaida Pozen, Francis Bond
GWC4