Francis Bond

dblp:61/1808 · DBLP profile ↗
← Back
85ranked-venue papers
27as first author
17since 2021 · last 2026
0000-0003-4973-8068ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 85 · 27 first-author · 17 since 2021Databases, data management, data science and information retrieval · 30 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 More Aligned, Less Diverse? Analyzing the Grammar and Lexicon of Two Generations of LLMs
abstract
Adrián Gude, Roi Santos-Rios, Francis Bond, Dan Flickinger, Carlos Gómez-Rodríguez, Olga Zamaraeva. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Adrián Gude, Roi Santos-Rios, Francis Bond, Dan Flickinger, Carlos Gómez-Rodríguez, Olga Zamaraeva
ACL (1)3
2026 The DELPH-IN Grammary: A Curated Repository of Grammars and Treebanks
Francis Bond, Dan Flickinger
LREC1
2026 Cygnet: Refactoring the Open Multilingual Wordnet
Rowan Hall Maudslay, Francis Bond
LREC2
2025 Comparing LLM-generated and human-authored news text using formal syntactic theory
abstract
This study provides the first comprehensive comparison of New York Times-style text generated by six large language models against real, human-authored NYT writing. The comparison is based on a formal syntactic theory. We use Head-driven Phrase Structure Grammar (HPSG) to analyze the grammatical structure of the texts. We then investigate and illustrate the differences in the distributions of HPSG grammar types, revealing systematic distinctions between human and LLM-generated writing. These findings contribute to a deeper understanding of the syntactic behavior of LLMs as well as humans, within the NYT genre.
Olga Zamaraeva, Dan Flickinger, Francis Bond, Carlos Gómez-Rodríguez
ACL (1)3
2024 ChainNet: Structured Metaphor and Metonymy in WordNet
abstract
The senses of a word exhibit rich internal structure. In a typical lexicon, this structure is overlooked: A word’s senses are encoded as a list, without inter-sense relations. We present ChainNet, a lexical resource which for the first time explicitly identifies these structures, by expressing how senses in the Open English Wordnet are derived from one another. In ChainNet, every nominal sense of a word is either connected to another sense by metaphor or metonymy, or is disconnected (in the case of homonymy). Because WordNet senses are linked to resources which capture information about their meaning, ChainNet represents the first dataset of grounded metaphor and metonymy.
Rowan Hall Maudslay, Simone Teufel, Francis Bond, James Pustejovsky
LREC/COLING3
2023 Documenting the Open Multilingual Wordnet
abstract
In this project note we describe our work to make better documentation for the Open Multilingual Wordnet (OMW), a platform integrating many open wordnets.This includes the documentation of the OMW website itself as well as of semantic relations used by the component wordnets.Some of this documentation work was done with the support of the Google Season of Docs.The OMW project page, which links both to the actual OMW server and the documentation has been moved to a new location: https://omwn.org.
Francis Bond, Michael Wayne Goodman, Ewa Rudnicka, Luís Morgado da Costa, Alexandre Rademaker, John P. McCrae
GWC1
2023 The Japanese Wordnet 2.0
abstract
This paper describes a new release of the Japanese wordnet.It uses the new global wordnet formats (McCrae et al., 2021) to incorporate a range of new information: orthographic variants (including hiragana, katakana and Latin representations) first described in Kuroda et al. (2011), classifiers, pronouns and exclamatives (Morgado da Costa and Bond, 2016) and many new senses, motivated both from corpus annotation and linking to the TUFs basic vocabulary (Bond et al., 2020).The wordnet has been moved to github and is available at https://bond-lab.github. io/wnja/.
Francis Bond, Takayuki Kuribayashi
GWC1
2023 Linking SIL Semantic Domains to Wordnet and Expanding the Abui Wordnet through Rapid Word Collection Methodology
abstract
In this paper we describe a new methodology to expand the Abui Wordnet through data collected using the Rapid Word Collection (RWC) method – based on SIL’s Semantic Domains. Using a multilingual sense-intersection algorithm, we created a ranked list of concept suggestions for each domain, and then used the ranked list as a filter to link the Abui RWC data to wordnet. This used translations from both SIL’s Semantic Domain’s structure and example words, both available through SIL’s Fieldworks software and the RWC project. We release both the new mapping of the SIL Semantic Domains to wordnet and an expansion of the Abui Wordnet.
Luís Morgado da Costa, Frantisek Kratochvil, George Saad, Benidiktus Delpada, Daniel Simon Lanma, Francis Bond, Natálie Wolfová, A. l. Blake
GWC6
2022 Sense and Sentiment
abstract
In this paper we examine existing sentiment lexicons and sense-based sentiment-tagged corpora to find out how sense and concept-based semantic relations effect sentiment scores (for polarity and valence). We show that some relations are good predictors of sentiment of related words: antonyms have similar valence and opposite polarity, synonyms similar valence and polarity, as do many derivational relations. We use this knowledge and existing resources to build a sentiment annotated wordnet of English, and show how it can be used to produce sentiment lexicons for other languages using the Open Multilingual Wordnet.
Francis Bond, Merrick Yeu Herng Choo
LREC1
2022 Singlish Where Got Rules One? Constructing a Computational Grammar for Singlish
abstract
Singlish is a variety of English spoken in Singapore. In this paper, we share some of its grammar features and how they are implemented in the construction of a computational grammar of Singlish as a branch of English grammar. New rules were created and existing ones from standard English grammar of the English Resource Grammar (ERG) were changed in this branch to cater to how Singlish works. In addition, Singlish lexicon was added into the grammar together with some new lexical types. We used Head-driven Phrase Structure Grammar (HPSG) as the framework for this project of a creating a working computational grammar. As part of building the language resource, we also collected and formatted some data from the internet as part of a test suite for Singlish. Finally, the computational grammar was tested against a set of gold standard trees and compared with the standard English grammar to find out how well the grammar fares in analysing Singlish.
Siew Yeng Chow, Francis Bond
LREC2
2022 The Tembusu Treebank: An English Learner Treebank
abstract
This paper reports on the creation and development of the Tembusu Learner Treebank — an open treebank created from the NTU Corpus of Learner English, unique for incorporating mal-rules in the annotation of ungrammatical sentences. It describes the motivation and development of the treebank, as well as its exploitation to build a new parse-ranking model for the English Resource Grammar, designed to help improve the parse selection of ungrammatical sentences and diagnose these sentences through mal-rules. The corpus contains 25,000 sentences, of which 4,900 are treebanked. The paper concludes with an evaluation experiment that shows the usefulness of this new treebank in the tasks of grammatical error detection and diagnosis.
Luís Morgado da Costa, Francis Bond, Roger Vivek Placidus Winder
LREC2
2021 Taboo Wordnet
abstract
This paper describes the development of an online lexical resource to help detection systems regulate and curb the use of offensive words online.With the growing prevalence of social media platforms, many conversations are now conducted online.The increase of online conversations for leisure, work and socializing has led to an increase in harassment.In particular, we create a specialized sense-based vocabulary of Japanese offensive words for the Open Multilingual Wordnet.This vocabulary expands on an existing list of Japanese offensive words and provides categorization and proper linking to synsets within the multilingual wordnet.This paper then discusses the evaluation of the vocabulary as a resource for representing and classifying offensive words and as a possible resource for offensive word use detection in social media.Content Warning: this paper deals with obscene words and contains many examples of them.
Francis Bond, Merrick Yeu Herng Choo
GWC1
2021 Teaching Through Tagging - Interactive Lexical Semantics
abstract
In this paper we discuss an ongoing effort to enrich students' learning by involving them in sense tagging.The main goal is to lead students to discover how we can represent meaning and where the limits of our current theories lie.A subsidiary goal is to create sense tagged corpora and an accompanying linked lexicon (in our case wordnets).We present the results of tagging several texts and suggest some ways in which the tagging process could be improved.Two authors of this paper present their own experience as students.Overall, students reported that they found the tagging an enriching experience.The annotated corpora and changes to the wordnet are made available through the NTU multilingual corpus and associated wordnets (NTU-MC).
Francis Bond, Andrew Devadason, Melissa Rui Lin Teo, Luís Morgado da Costa
GWC1
2021 Intrinsically Interlingual: The Wn Python Library for Wordnets
abstract
This paper introduces Wn, a new Python library for working with wordnets.Unlike previous libraries, Wn is built from the beginning to accommodate multiple wordnets-for multiple languages or multiple versions of the same wordnet-while retaining the ability to query and traverse them independently.It is also able to download and incorporate wordnets published online.These features are made possible through Wn's adoption of standard formats and methods for interoperability, namely the WN-LMF schema (Vossen et al., 2013;Bond et al., 2020) and the Collaborative Interlingual Index (Bond et al., 2016).Wn is open-source, easily available, 1 and well-documented. 2
Michael Wayne Goodman, Francis Bond
GWC2
2021 Testing agreement between lexicographers: A case of homonymy and polysemy
abstract
In this paper we compare Oxford Lexico and Merriam Webster dictionaries with Princeton WordNet with respect to the description of semantic (dis)similarity between polysemous and homonymous senses that could be inferred from them.WordNet lacks any explicit description of polysemy or homonymy, but as a network of linked senses it may be used to compute semantic distances between word senses.To compare WordNet with the dictionaries, we transformed sample entry microstructures of the latter into graphs and crosslinked them with the equivalent senses of the former.We found that dictionaries are in high agreement with each other, if one considers polysemy and homonymy altogether, and in moderate concordance, if one focuses merely on polysemy descriptions.Measuring the shortest path lengths on WordNet gave results comparable to those on the dictionaries in predicting semantic dissimilarity between polysemous senses, but was less felicitous while recognising homonymy.
Marek Maziarz, Francis Bond, Ewa Rudnicka
GWC2
2021 The GlobalWordNet Formats: Updates for 2020
abstract
The Global Wordnet Formats have been introduced to enable wordnets to have a common representation that can be integrated through the Global WordNet Grid.As a result of their adoption, a number of shortcomings of the format were identified, and in this paper we describe the extensions to the formats that address these issues.These include: ordering of senses, dependencies between wordnets, pronunciation, syntactic modelling, relations, sense keys, metadata and RDF support.Furthermore, we provide some perspectives on how these changes help in the integration of wordnets.
John P. McCrae, Michael Wayne Goodman, Francis Bond, Alexandre Rademaker, Ewa Rudnicka, Luís Morgado da Costa
GWC3
2021 OdeNet: Compiling a GermanWordNet from other Resources
abstract
The Princeton WordNet for the English language has been used worldwide in NLP projects for many years.With the OMW initiative, wordnets for different languages of the world are being linked via identifiers.The parallel development and linking allows new multilingual application perspectives.The development of a wordnet for the German language is also in this context.To save development time, existing resources were combined and recompiled.The result was then evaluated and improved.In a relatively short time a resource was created that can be used in projects and continuously improved and extended.
Melanie Siegel, Francis Bond
GWC2
2020 Some Issues with Building a Multilingual Wordnet
abstract
In this paper we discuss the experience of bringing together over 40 different wordnets. We introduce some extensions to the GWA wordnet LMF format proposed in Vossen et al. (2016) and look at how this new information can be displayed. Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalized synsets and lemmas, new parts of speech, and more. Many of these extensions already exist in multiple wordnets – the challenge was to find a compatible representation. To this end, we introduce a new version of the Open Multilingual Wordnet (Bond and Foster, 2013), that integrates a new set of tools that tests the extensions introduced by this new format, while also ensuring the integrity of the Collaborative Interlingual Index (CILI: Bond et al., 2016), avoiding the same new concept to be introduced through multiple projects.
Francis Bond, Luís Morgado da Costa, Michael Wayne Goodman, John P. McCrae, Ahti Lohk
LREC1
2020 Linking the TUFS Basic Vocabulary to the Open Multilingual Wordnet
abstract
We describe the linking of the TUFS Basic Vocabulary Modules, created for online language learning, with the Open Multilingual Wordnet. The TUFS modules have roughly 500 lexical entries in 30 languages, each with the lemma, a link across the languages, an example sentence, usage notes and sound files. The Open Multilingual Wordnet has 34 languages (11 shared with TUFS) organized into synsets linked by semantic relations, with examples and definitions for some languages. The links can be used to (i) evaluate existing wordnets, (ii) add data to these wordnets and (iii) create new open wordnets for Khmer, Korean, Lao, Mongolian, Russian, Tagalog, Urdua nd Vietnamese
Francis Bond, Hiroki Nomoto, Luís Morgado da Costa, Arthur Bond
LREC1
2020 Automated Writing Support Using Deep Linguistic Parsers
abstract
This paper introduces a new web system that integrates English Grammatical Error Detection (GED) and course-specific stylistic guidelines to automatically review and provide feedback on student assignments. The system is being developed as a pedagogical tool for English Scientific Writing. It uses both general NLP methods and high precision parsers to check student assignments before they are submitted for grading. Instead of generalized error detection, our system aims to identify, with high precision, specific classes of problems that are known to be common among engineering students. Rather than correct the errors, our system generates constructive feedback to help students identify and correct them on their own. A preliminary evaluation of the system’s in-class performance has shown measurable improvements in the quality of student assignments.
Luís Morgado da Costa, Roger Vivek Placidus Winder, Shu Yun Li, Benedict Christopher Tzer Liang Lin, Joseph MacKinnon, Francis Bond
LREC6
2019 GeoNames Wordnet (geown): extracting wordnets from GeoNames
abstract
This paper introduces a new multilingual lexicon of geographical place names.The names are based on (and linked to) the GeoNames collection.Each location is treated as a new synset, which is linked by instance_hypernym to a small set of supertypes.These supertypes are linked to the collaborative interlingual index, based on mappings from GeoDomainWordnet.If a location is already in the interlingual index, then it is also linked to the entry, using mappings from the Geo-Wordnet.Finally, if GeoNames places the location in a larger location, this is linked using the mero_location link.Wordnets can be built for any language in GeoNames, we give results for those wordnets in the Open Multilingual Wordnet.We discuss how it is mapped and the characteristics of the extracted wordnets.
Francis Bond, Arthur Bond
GWC1
2019 Testing Zipf's meaning-frequency law with wordnets as sense inventories
abstract
According to George K. Zipf, more frequent words have more senses.We have tested this law using corpora and wordnets of English, Spanish, Portuguese, French, Polish, Japanese, Indonesian and Chinese.We have proved that the law works pretty well for all of these languages if we takeas Zipf did -mean values of meaning count and averaged ranks.On the other hand, the law disastrously fails in predicting the number of senses for a single lemma.We have also provided the evidence that slope coefficients of Zipfian log-log linear model may vary from language to language.
Francis Bond, Arkadiusz Janz, Marek Maziarz, Ewa Rudnicka
GWC1
2019 A Comparison of Sense-level Sentiment Scores
abstract
In this paper, we compare a variety of sense-tagged sentiment resources, including SentiWordNet, ML-Senticon, plWord-Net emo and the NTU Multilingual Corpus.The goal is to investigate the quality of the resources and see how well the sentiment polarity annotation maps across languages.
Francis Bond, Arkadiusz Janz, Maciej Piasecki
GWC1
2019 New Polysemy Structures in Wordnets Induced by Vertical Polysemy
abstract
This paper aims to study auto-hyponymy and auto-troponymy relations (or vertical polysemy) in 11 wordnets uploaded into the new Open Multilingual Wordnet (OMW) webpage.We investigate how vertical polysemy forms polysemy structures (or sense clusters) in semantic hierarchies of the wordnets.Our main results and discoveries are new polysemy structures that have not previously been associated with vertical polysemy, along with some inconsistencies of semantic relations analysis in the studied wordnets, which should not be there.In the case study, we turn attention to polysemy structures in the Estonian Wordnet (version 2.2.0), analyzing them and giving the lexicographers comments.In addition, we describe the detection algorithm of polysemy structures and an overview of the state of polysemy structures in 11 wordnets.
Ahti Lohk, Heili Orav, Kadri Vare, Francis Bond, Rasmus Vaik
GWC4
2019 English WordNet 2019 - An Open-Source WordNet for English
abstract
We describe the release of a new wordnet for English based on the Princeton WordNet, but now developed under an open-source model.In particular, this version of WordNet, which we call English WordNet 2019, which has been developed by multiple people around the world through GitHub, fixes many errors in previous wordnets for English.We give some details of the changes that have been made in this version and give some perspectives about likely future changes that will be made as this project continues to evolve.
John P. McCrae, Alexandre Rademaker, Francis Bond, Ewa Rudnicka, Christiane Fellbaum
GWC3
2018 Toward An Epic Epigraph Graph
Francis Bond, Graham Matthews
LREC1
2018 The Company They Keep: Extracting Japanese Neologisms Using Language Patterns
abstract
We describe an investigation into the identification and extraction of unrecorded potential lexical items in Japanese text by detecting text passages containing selected language patterns typically associated with such items.We identified a set of suitable patterns, then tested them with two large collections of text drawn from the WWW and Twitter.Samples of the extracted items were evaluated, and it was demonstrated that the approach has considerable potential for identifying terms for later lexicographic analysis.
James Breen, Timothy Baldwin, Francis Bond
GWC3
2018 Toward Constructing the National Cancer Institute Thesaurus Derived WordNet (ncitWN)
abstract
We describe preliminary work in the creation of the first specialized vocabulary to be integrated into the Open Multilingual Wordnet (OMW).The NCIt Derived WordNet (ncitWN) is based on the National Cancer Institute Thesaurus (NCIt), a controlled biomedical terminology that includes formal class restrictions and English definitions developed by groups of clinicians and terminologists.The ncitWN is created by converting the NCIt to the WordNet Lexical Markup Framework and adding semantic types.We report the development of a prototype ncitWN and first steps towards integrating it into the OMW.
Amanda Hicks, Selja Seppälä, Francis Bond
GWC3
2018 Automatic Identification of Basic-Level Categories
abstract
Basic-level categories have been shown to be both psychologically significant and useful in a wide range of practical applications.We build a rule-based system to identify basic-level categories in WordNet, achieving 77% accuracy on a test set derived from prior psychological experiments.With additional annotations we found our system also has low precision, in part due to the existence of many categories that do not fit into the three classes (superordinate, basic-level, and subordinate) relied on in basiclevel category research.
Chad Mills, Francis Bond, Gina-Anne Levow
GWC2
2018 Putting Figures on Influences on Moroccan Darija from Arabic, French and Spanish using the WordNet
abstract
Moroccan Darija is a variant of Arabic with many influences.Using the Open Multilingual WordNet (OMW), we compare the lemmas in the Moroccan Darija Wordnet (MDW) with the standard Arabic, French and Spanish ones.We then compared the lemmas in each synset with their translation equivalents.Transliteration is used to bridge alphabet differences and match lemmas in the closest phonological way.The results put figures on the similarity Moroccan Darija has with Arabic, French and Spanish: respectively 42.0%, 2.8% and 2.2%.
Khalil Mrini, Francis Bond
GWC2
2018 Lexical Perspective on Wordnet to Wordnet Mapping
abstract
The paper presents a feature-based model of equivalence targeted at (manual) sense linking between Princeton WordNet and plWordNet.The model incorporates insights from lexicographic and translation theories on bilingual equivalence and draws on the results of earlier synsetlevel mapping of nouns between Princeton WordNet and plWordNet.It takes into account all basic aspects of language such as form, meaning and function and supplements them with (parallel) corpus frequency and translatability.Three types of equivalence are distinguished, namely strong, regular and weak depending on the conformity with the proposed features.The presented solutions are languageneutral and they can be easily applied to language pairs other than Polish and English.Sense-level mapping is a more finegrained mapping than the existing synset mappings and is thus of great potential to human and machine translation.
Ewa Rudnicka, Francis Bond, Lukasz Grabowski, Maciej Piasecki, Tadeusz Piotrowski
GWC2
2018 Enchancing the Collaborative Interlingual Index for Digital Humanities: Cross-linguistic Analysis in the Domain of Theology
abstract
We aim to support digital humanities work related to the study of sacred texts. To do this, we propose to build a cross-lingual wordnet within the do-main of theology. We target the Collaborative Interlingual Index (CILI) directly instead of each individual wordnet. The paper presents background for this proposal: (1) an overview of concepts relevant to theology and (2) a summary of the domain-associated issues observed in the Princeton WordNet (PWN). We have found that definitions for concepts in this domain can be too restrictive, inconsistent, and unclear. Necessary synsets are missing, with the PWN being skewed towards Christianity. We argue that tackling problems in a single domain is a better method for improving CILI. By focusing on a single topic rather than a single language, this will result in the proper construction of definitions, romanization/translation of lemmas, and also improvements in use of/creation of a cross-lingual domain hierarchy.
Laura A. Slaughter, Luís Morgado da Costa, Francis Bond
GWC4
2018 Multilingual Wordnet sense Ranking using nearest context
abstract
In this paper, we combine methods to estimate sense rankings from raw text with recent work on word embeddings to provide sense ranking estimates for the entries in the Open Multilingual WordNet (OMW).The existing Word2Vec pre-trained models from Polygot2 are only built for single word entries, we, therefore, re-train them with multiword expressions from the wordnets, so that multiword expressions can also be ranked.Thus this trained model gives embeddings for both single words and multiwords.The resulting lexicon gives a WSD baseline for five languages.The results are evaluated for Semcor sense corpora for 5 languages using Word2Vec and Glove models.The Glove model achieves an average accuracy of 0.47 and Word2Vec achieves 0.31 for languages such as English, Italian, Indonesian, Chinese and Japanese.The experimentation on OMW sense ranking proves that the rank correlation is generally similar to the human ranking.Hence distributional semantics can aid in Wordnet Sense Ranking.
E. Umamaheswari Vasanthakumar, Francis Bond
GWC2
2016 LexSemTm: A Semantic Dataset Based on All-words Unsupervised Sense Distribution Learning
Andrew Bennett, Timothy Baldwin, Jey Han Lau, Diana McCarthy, Francis Bond
ACL (1)5
2016 Wow! What a Useful Extension! Introducing Non-Referential Concepts to Wordnet
Luís Morgado da Costa, Francis Bond
LREC2
2016 The Open Linguistics Working Group: Developing the Linguistic Linked Open Data Cloud
John P. McCrae, Christian Chiarcos, Francis Bond, Philipp Cimiano, Thierry Declerck, Gerard de Melo, Jorge Gracia, Sebastian Hellmann 0001, Bettina Klimek, Steven Moran, Petya Osenova, Antonio Pareja-Lora, Jonathan Pool
LREC3
2016 Multilingual Sense Intersection in a Parallel Corpus with Diverse Language Families
abstract
Supervised methods for Word Sense Disambiguation (WSD) benefit from highquality sense-annotated resources, which are lacking for many languages less common than English.There are, however, several multilingual parallel corpora that can be inexpensively annotated with senses through cross-lingual methods.We test the effectiveness of such an approach by attempting to disambiguate English texts through their translations in Italian, Romanian and Japanese.Specifically, we try to find the appropriate word senses for the English words by comparison with all the word senses associated to their translations.The main advantage of this approach is in that it can be applied to any parallel corpus, as long as large, highquality inter-linked sense inventories exist for all the languages considered.
Giulia Bonansinga, Francis Bond
GWC2
2016 CILI: the Collaborative Interlingual Index
abstract
This paper introduces the motivation for and design of the Collaborative InterLingual Index (CILI).It is designed to make possible coordination between multiple loosely coupled wordnet projects.The structure of the CILI is based on the Interlingual index first proposed in the Eu-roWordNet project with several pragmatic extensions: an explicit open license, definitions in English and links to wordnets in the Global Wordnet Grid.
Francis Bond, Piek Vossen, John P. McCrae, Christiane Fellbaum
GWC1
2016 Mapping and Generating Classifiers using an Open Chinese Ontology
abstract
In languages such as Chinese, classifiers (CLs) play a central role in the quantification of noun-phrases.This can be a problem when generating text from input that does not specify the classifier, as in machine translation (MT) from English to Chinese.Many solutions to this problem rely on dictionaries of noun-CL pairs.However, there is no open large-scale machine-tractable dictionary of noun-CL associations.Many published resources exist, but they tend to focus on how a CL is used (e.g.what kinds of nouns can be used with it, or what features seem to be selected by each CL).In fact, since nouns are open class words, producing an exhaustive definite list of noun-CL associations is not possible, since it would quickly get out of date.Our work tries to address this problem by providing an algorithm for automatic building of a frequency based dictionary of noun-CL pairs, mapped to concepts in the Chinese Open Wordnet (Wang and Bond, 2013), an open machinetractable dictionary for Chinese.All results will released under an open license.
Luís Morgado da Costa, Francis Bond, Helena Gao
GWC2
2016 Identifying and Exploiting Definitions in Wordnet Bahasa
abstract
This paper describes our attempts to add Indonesian definitions to synsets in the Wordnet Bahasa (Nurril Hirfana Mohamed Noor et al., 2011;Bond et al., 2014), to extract semantic relations between lemmas and definitions for nouns and verbs, such as synonym, hyponym, hypernym and instance hypernym, and to generally improve Wordnet.The original, somewhat noisy, definitions for Indonesian came from the Asian Wordnet project (Riza et al., 2010).The basic method of extracting the relations is based on Bond et al. (2004).Before the relations can be extracted, the definitions were cleaned up and tokenized.We found that the definitions cannot be completely cleaned up because of many misspellings and bad translations.However, we could identify four semantic relations in 57.10% of noun and verb definitions.For the remaining 42.90%, we propose to add 149 new Indonesian lemmas and make some improvements to Wordnet Bahasa and Wordnet in general.
David Moeljadi, Francis Bond
GWC2
2016 Toward a truly multilingual GlobalWordnet Grid
abstract
In this paper, we describe a new and improved Global Wordnet Grid that takes advantage of the Collaborative InterLingual Index (CILI).Currently, the Open Multilingal Wordnet has made many wordnets accessible as a single linked wordnet, but as it used the Princeton Wordnet of English (PWN) as a pivot, it loses concepts that are not part of PWN.The technical solution to this, a central registry of concepts, as proposed in the EuroWordnet project through the InterLingual Index, has been known for many years.However, the practical issues of how to host this index and who decides what goes in remained unsolved.Inspired by current practice in the Semantic Web and the Linked Open Data community, we propose a way to solve this issue.In this paper we define the principles and protocols for contributing to the Grid.We tested them on two use cases, adding version 3.1 of the Princeton WordNet to a CILI based on 3.0 and adding the Open Dutch Wordnet, to validate the current set up.This paper aims to be a call for action that we hope will be further discussed and ultimately taken up by the whole wordnet community.
Piek Vossen, Francis Bond, John P. McCrae
GWC2
2014 Identifying Idioms in Chinese Translations
Wan Yu Ho, Christine Kng, Shan Wang 0002, Francis Bond
LREC4
2014 Building The Sense-Tagged Multilingual Parallel Corpus
Shan Wang 0002, Francis Bond
LREC2
2014 Issues in building English-Chinese parallel corpora with WordNets
abstract
We discuss some of the issues in producing sense-tagged parallel corpora: including pre-processing, adding new entries and linking.We have preliminary results for three genres: stories, essays and tourism web pages, in both Chinese and English.
Francis Bond, Shan Wang 0002
GWC1
2014 A Survey of WordNet Annotated Corpora
abstract
This paper surveys the current state of wordnet sense annotated corpora.We look at corpora in any language, and describe them in terms of accessibility and usefulness.We finally discuss possibilities in increasing the interoperability of the corpora, especially across languages.
Tommaso Petrolito, Francis Bond
GWC2
2014 Bringing together over- and under- represented languages: Linking WordNet to the SIL Semantic Domains
abstract
We have created an open-source mapping between the SIL's semantic domains (used for rapid lexicon building and organization for under-resourced languages) and WordNet, the standard resource for lexical semantics in natural language processing.We show that the resources complement each other, and suggest ways in which the mapping can be improved even further.The semantic domains give more general domain and associative links, which wordnet still has few of, while wordnet gives explicit semantic relations between senses, which the domains lack.
Muhammad Zulhelmy Bin Mohd Rosman, Frantisek Kratochvil, Francis Bond
GWC3
2014 Parse Ranking with Semantic Dependencies and WordNet
abstract
In this paper, we investigate which features are useful for ranking semantic representations of text.We show that two methods of generalization improved results: extended grand-parenting and supertypes.The models are tested on a subset of SemCor that has been annotated with both Dependency Minimal Recursion Semantic representations and WordNet senses.Using both types of features gives a significant improvement in whole sentence parse selection accuracy over the baseline model.
Xiaocheng Yin, Jung-Jae Kim 0001, Zinaida Pozen, Francis Bond
GWC4
2013 Linking and Extending an Open Multilingual Wordnet
Francis Bond, Ryan Foster
ACL (1)1
2013 On the semantics of noun compounds
abstract
The noun compound – a sequence of nouns which functions as a single noun – is very common in English texts. No language processing system should ignore expressions like steel soup pot cover if it wants to be serious about such high-end applications of computational linguistics as question answering, information extraction, text summarization, machine translation – the list goes on. Processing noun compounds, however, is far from trouble-free. For one thing, they can be bracketed in various ways: is it steel soup, steel pot, or steel cover? Then there are relations inside a compound, annoyingly not signalled by any words: does potcontainsoup or is it for cookingsoup? These and many other research challenges are the subject of this special issue.
Stan Szpakowicz, Francis Bond, Preslav Nakov, Su Nam Kim
Nat. Lang. Eng.2
2012 Comparing Classifier use in Chinese and Japanese
Yue Hui Ting, Francis Bond
PACLIC2
2011 Creating the Open Wordnet Bahasa
Nurril Hirfana Bte Mohamed Noor, Suerya Sapuan, Francis Bond
PACLIC3
2011 Building and Annotating the Linguistically Diverse NTU-MC (NTU-Multilingual Corpus)
Liling Tan, Francis Bond
PACLIC2
2011 Language, Technology, and Society Richard Sproat (Oregon Health & Science University) Oxford: Oxford University Press, 2010, xiii+286 pp; hardbound, ISBN 978-0-19-954938-2, £25.00
Francis Bond
Comput. Linguistics1
2011 Deep open-source machine translation
Francis Bond, Stephan Oepen, Eric Nichols, Dan Flickinger, Erik Velldal, Petter Haugereid
Mach. Transl.1
2010 A Reexamination of MRD-Based Word Sense Disambiguation
abstract
This article reconsiders the task of MRD-based word sense disambiguation, in extending the basic Lesk algorithm to investigate the impact on WSD performance of different tokenization schemes and methods of definition extension. In experimentation over the Hinoki Sensebank and the Japanese Senseval-2 dictionary task, we demonstrate that sense-sensitive definition extension over hyponyms, hypernyms, and synonyms, combined with definition extension and word tokenization leads to WSD accuracy above both unsupervised and supervised baselines. In doing so, we demonstrate the utility of ontology induction and establish new opportunities for the development of baseline unsupervised WSD methods.
Timothy Baldwin, Su Nam Kim, Francis Bond, Sanae Fujita, David Martínez 0001, Takaaki Tanaka
ACM Trans. Asian Lang. Inf. Process.3
2009 Hypernym Discovery Based on Distributional Similarity and Hierarchical Structures
Ichiro Yamada, Kentaro Torisawa, Jun'ichi Kazama, Kow Kuroda, Masaki Murata, Stijn De Saeger, Francis Bond, Asuka Sumida
EMNLP7
2008 MRD-based Word Sense Disambiguation: Further Extending Lesk
Timothy Baldwin, Su Nam Kim, Francis Bond, Sanae Fujita, David Martínez 0001, Takaaki Tanaka
IJCNLP3
2008 Boot-Strapping a WordNet Using Multiple Existing WordNets
Francis Bond, Hitoshi Isahara, Kyoko Kanzaki, Kiyotaka Uchimoto
LREC1
2008 Development of the Japanese WordNet
Hitoshi Isahara, Francis Bond, Kiyotaka Uchimoto, Masao Utiyama, Kyoko Kanzaki
LREC2
2008 Extraction of Attribute Concepts from Japanese Adjectives
Kyoko Kanzaki, Francis Bond, Noriko Tomuro, Hitoshi Isahara
LREC2
2007 Word Sense Disambiguation Incorporating Lexical and Structural Semantic Information
Takaaki Tanaka, Francis Bond, Timothy Baldwin, Sanae Fujita, Chikara Hashimoto
EMNLP-CoNLL2
2007 A method of creating new valency entries
Sanae Fujita, Francis Bond
Mach. Transl.2
2006 An Implemented Description of Japanese: The Lexeed Dictionary and the Hinoki Treebank
abstract
In this paper we describe the current state of a new Japanese lexical resource: the Hinoki treebank. The treebank is built from dictionary definition sentences, and uses an HPSG based Japanese grammar to encode both syntactic and semantic information. It is combined with an ontology based on the definition sentences to give a detailed sense level description of the most familiar 28,000 words of Japanese.
Sanae Fujita, Takaaki Tanaka, Francis Bond, Hiromi Nakaiwa
ACL3
2005 High Precision Treebanking-Blazing Useful Trees Using POS Information
abstract
In this paper we present a quantitative and qualitative analysis of annotation in the Hinoki treebank of Japanese, and investigate a method of speeding annotation by using part-of-speech tags. The Hinoki treebank is a Redwoods-style treebank of Japanese dictionary definition sentences. 5,000 sentences are annotated by three different annotators and the agreement evaluated. An average agreement of 65.4% was found using strict agreement, and 83.5% using labeled precision. Exploiting POS tags allowed the annotators to choose the best parse with 19.5% fewer decisions.
Takaaki Tanaka, Francis Bond, Stephan Oepen, Sanae Fujita
ACL2
2005 Robust Ontology Acquisition from Machine-Readable Dictionaries
Eric Nichols, Francis Bond, Dan Flickinger
IJCAI2
2005 SEM-I Rational MT: Enriching Deep Grammars with a Semantic Interface for Scalable Machine Translation
abstract
In the LOGON machine translation system where semantic transfer using Minimal Recursion Semantics is being developed in conjunction with two existing broad-coverage grammars of Norwegian and English, we motivate the use of a grammar-specific semantic interface (SEM-I) to facilitate the construction and maintenance of a scalable translation engine. The SEM-I is a theoretically grounded component of each grammar, capturing several classes of lexical regularities while also serving the crucial engineering function of supplying a reliable and complete specification of the elementary predications the grammar can realize. We make extensive use of underspecification and type hierarchies to maximize generality and precision.
Dan Flickinger, Jan Tore Lønning, Helge Dyvik, Stephan Oepen, Francis Bond
MTSummit5
2005 Extracting Representative Arguments from Dictionaries for Resolving Zero Pronouns
abstract
We propose a method to alleviate the problem of referential granularity for Japanese zero pronoun resolution. We use dictionary definition sentences to extract ‘representative’ arguments of predicative definition words; e.g. ‘arrest’ is likely to take police as the subject and criminal as its object. These representative arguments are far more informative than ‘person’ that is provided by other valency dictionaries. They are auto-extracted using both Shallow parsing and Deep parsing for greater quality and quantity. Initial results are highly promising, obtaining more specific information about selectional preferences. An architecture of zero pronoun resolution using these representative arguments is described.
Shigeko Nariyama, Eric Nichols, Francis Bond, Takaaki Tanaka, Hiromi Nakaiwa
MTSummit3
2005 Introduction to the special issue on multiword expressions: Having a crack at a hard nut
Aline Villavicencio, Francis Bond, Anna Korhonen, Diana McCarthy
Comput. Speech Lang.2
2004 Source Language Effect on Translating Korean Honorifics
Kyonghee Paik, Kiyonori Ohtake, Francis Bond, Kazuhide Yamamoto
CICLing3
2004 Acquiring an Ontology for a Fundamental Vocabulary
Francis Bond, Eric Nichols, Sanae Fujita, Takaaki Tanaka
COLING1
2004 The Hinoki Treebank A Treebank for Text Understanding
Francis Bond, Sanae Fujita, Chikara Hashimoto, Kaname Kasahara, Shigeko Nariyama, Eric Nichols, Akira Ohtani, Takaaki Tanaka, Shigeaki Amano
IJCNLP1
2004 A Lexicon Module for a Grammar Development Environment
Ann A. Copestake, Fabre Lambeau, Benjamin Waldron, Francis Bond, Dan Flickinger, Stephan Oepen
LREC4
2003 Learning the Countability of English Nouns from Corpus Data
abstract
This paper describes a method for learning the countability preferences of English nouns from raw text corpora. The method maps the corpus-attested lexico-syntactic properties of each noun onto a feature vector, and uses a suite of memory-based classifiers to predict membership in 4 countability classes. We were able to assign countability to English nouns with a precision of 94.6%.
Timothy Baldwin, Francis Bond
ACL2
2003 A Plethora of Methods for Learning English Countability
Timothy Baldwin, Francis Bond
EMNLP2
2003 Evaluation of a method of creating new valency entries
abstract
Information on subcategorization and selectional restrictions is important for natural language processing tasks such as deep parsing, rule-based machine translation and automatic summarization. In this paper we present a method of adding detailed entries to a bilingual dictionary, based on information in an existing valency dictionary. The method is based on two assumptions: words with similar meaning have similar subcategorization frames and selectional restrictions; and words with the same translations have similar meanings. Based on these assumptions, new valency entries are constructed from words in a plain bilingual dictionary, using entries with similar source-language meaning and the same target-language translations. We evaluate the effects of various measures of similarity in increasing accuracy.
Francis Bond, Sanae Fujita
MTSummit1
2002 Multiword Expressions: A Pain in the Neck for NLP
Ivan A. Sag, Timothy Baldwin, Francis Bond, Ann A. Copestake, Dan Flickinger
CICLing3
2002 Using an Ontology to Determine English Countability
Francis Bond, Caitlin Vatikiotis-Bateson
COLING1
2002 Multiword expressions: linguistic precision and reusability
Ann A. Copestake, Fabre Lambeau, Aline Villavicencio, Francis Bond, Timothy Baldwin, Ivan A. Sag, Dan Flickinger
LREC4
2002 Towards a Thesaurus of Predicates
Satoshi Shirai, Kazuhide Yamamoto, Francis Bond, Hozumi Tanaka
LREC3
2001 Design and construction of a machine-tractable Japanese-Malay dictionary
abstract
We present a method for combining two bilingual dictionaries to make a third, using one language as a pivot. In this case we combine a Japanese-English dictionary with a Malay-English dictionary, to produce a Japanese-Malay dictionary suitable for use in a machine translation system. Our method differs from previous methods in its use of semantic classes to rank translation equivalents: word pairs with compatible semantic classes are preferred to those with dissimilar classes. We also experiment with the use of two pivot languages. We have made a prototype dictionary of over 75,000 pairs.
Francis Bond, Ruhaida Binti Sulong, Takefumi Yamazaki, Kentaro Ogura
MTSummit1
2000 Reusing an ontology to generate numeral classifiers
Francis Bond, Kyonghee Paik
COLING1
1999 ALT-J/M a prototype Japanese-to-Malay translation system
abstract
In this report we introduce ALT-J/M: a prototype Japanese-to-Malay translation system. The system is a semantic transfer based system that uses the same translation engine as ALT-J/E, a Japanese-to-English system.
Kentaro Ogura, Francis Bond, Yoshifumi Ooyama
MTSummit2
1998 Reference in Japanese-English Machine Translation
Francis Bond, Kentaro Ogura
Mach. Transl.1
1996 Classifiers in Japanese-to-English Machine Translation
Francis Bond, Kentaro Ogura, Satoru Ikehara
COLING1
1994 Countability and Number in Japanese to English Machine Translation
Francis Bond, Kentaro Ogura, Satoru Ikehara
COLING1