William Lane 0002

dblp:239/0671-2 · also William Abbott Lane · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 72% Speech recognition and synthesis · 22% Language models and text generation · 6%
Human-computer interaction and pervasive computing
1 paper
Learning and educational technologies · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
morphological analysis
1.232022
Interactive Word Completion for Plains Cree · ACL (1) 2022
Bootstrapping Techniques for Polysynthetic Morphological Analysis · ACL 2020
Local Word Discovery for Interactive Transcription · EMNLP (1) 2021
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
low-resource speech recognition
0.512021
Local Word Discovery for Interactive Transcription · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › word segmentation
word discovery
0.512021
Local Word Discovery for Interactive Transcription · EMNLP (1) 2021
Learning and educational technologies
language learning
0.212022
Interactive Word Completion for Plains Cree · ACL (1) 2022
Natural language and speech › Language models and text generation
low-resource language processing
0.112020
Bootstrapping Techniques for Polysynthetic Morphological Analysis · ACL 2020

Methods — techniques the papers use, named apart from their topics

finite state morphological analyzer · 1.1finite state transducer · 0.9ranking strategy · 0.6ranking strategies · 0.6morphological transducer · 0.5encoder-decoder · 0.4data augmentation · 0.4
YearPublicationVenuePosition
2022 Interactive Word Completion for Plains Cree
abstract
The composition of richly-inflected words in morphologically complex languages can be a challenge for language learners developing literacy.Accordingly, Lane and Bird (2020) proposed a finite state approach which maps prefixes in a language to a set of possible completions up to the next morpheme boundary, for the incremental building of complex words.In this work, we develop an approach to morph-based auto-completion based on a finite state morphological analyzer of Plains Cree (nêhiyawêwin), showing the portability of the concept to a much larger, more complete morphological transducer.Additionally, we propose and compare various novel ranking strategies on the morph auto-complete output.The best weighting scheme ranks the target completion in the top 10 results in 64.9% of queries, and in the top 50 in 73.9% of queries.
William Lane 0002, Atticus Harrigan, Antti Arppe
ACL (1)1
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC17
2021 Local Word Discovery for Interactive Transcription
abstract
Human expertise and the participation of speech communities are essential factors in the success of technologies for low-resource languages.Accordingly, we propose a new computational task which is tuned to the available knowledge and interests in an Indigenous community, and which supports the construction of high quality texts and lexicons.The task is illustrated for Kunwinjku, a morphologicallycomplex Australian language.We combine a finite state implementation of a published grammar with a partial lexicon, and apply this to a noisy phone representation of the signal.We locate known lexemes in the signal and use the morphological transducer to build these out into hypothetical, morphologicallycomplex words for human validation.We show that applying a single iteration of this method results in a relative transcription density gain of 17%.Further, we find that 75% of breath groups in the test set receive at least one correct partial or full-word suggestion.
William Lane 0002, Steven Bird
EMNLP (1)1
2020 Bootstrapping Techniques for Polysynthetic Morphological Analysis
abstract
Polysynthetic languages have exceptionally large and sparse vocabularies, thanks to the number of morpheme slots and combinations in a word.This complexity, together with a general scarcity of written data, poses a challenge to the development of natural language technologies.To address this challenge, we offer linguistically-informed approaches for bootstrapping a neural morphological analyzer, and demonstrate its application to Kunwinjku, a polysynthetic Australian language.We generate data from a finite state transducer to train an encoderdecoder model.We improve the model by "hallucinating" missing linguistic structure into the training data, and by resampling from a Zipf distribution to simulate a more natural distribution of morphemes.The best model accounts for all instances of reduplication in the test set and achieves an accuracy of 94.7% overall, a 10 percentage point improvement over the FST baseline.This process demonstrates the feasibility of bootstrapping a neural morph analyzer from minimal resources.
William Lane 0002, Steven Bird
ACL1
2020 Interactive Word Completion for Morphologically Complex Languages
abstract
Text input technologies for low-resource languages support literacy, content authoring, and language learning.However, tasks such as word completion pose a challenge for morphologically complex languages thanks to the combinatorial explosion of possible words.We have developed a method for morphologically-aware text input in Kunwinjku, a polysynthetic language of northern Australia.We modify an existing finite state recognizer to map input morph prefixes to morph completions, respecting the morphosyntax and morphophonology of the language.We demonstrate the portability of the method by applying it to Turkish.We show that the space of proximal morph completions is many orders of magnitude smaller than the space of full word completions for Kunwinjku.We provide a visualization of the morph completion space to enable the text completion parameters to be fine-tuned.Finally, we report on a web services deployment, along with a web interface which helps users enter morphologically complex words and which retrieves corresponding entries from the lexicon.
William Lane 0002, Steven Bird
COLING1