Francis M. Tyers

dblp:84/8340 · DBLP profile ↗
← Back
38ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0001-6108-2220ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl
abstract
Robert Pugh, Cheyenne Wing, María Ximena Juárez Huerta, Ángeles Márquez Hernandez, Francis Tyers. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Robert Pugh, Cheyenne Wing, María Ximena Juárez Huerta, Angeles Márquez Hernandez, Francis M. Tyers
NAACL (Long Papers)5
2024 Developing a Benchmark for Pronunciation Feedback: Creation of a Phonemically Annotated Speech Corpus of isiZulu Language Learner Speech
abstract
Pronunciation of the phonemic inventory of a new language often presents difficulties to second language (L2) learners. These challenges can be alleviated by the development of pronunciation feedback tools that take speech input from learners and return information about errors in the utterance. This paper presents the development of a corpus designed for use in pronunciation feedback research. The corpus is comprised of gold standard recordings from isiZulu teachers and recordings from isiZulu L2 learners that have been annotated for pronunciation errors. Exploring the potential benefits of word-level versus phoneme-level feedback necessitates a speech corpus that has been annotated for errors on the phoneme-level. To aid in this discussion, this corpus of isiZulu L2 speech has been annotated for phoneme-errors in utterances, as well as suprasegmental errors in tone.
Alexandra O'Neil, Nils Hjortnaes, Francis M. Tyers, Zinhle Nkosi, Thulile Ndlovu, Zanele Mlondo, Ngami Phumzile Pewa
LREC/COLING3
2024 Producing a Parallel Universal Dependencies Treebank of Ancient Hebrew and Ancient Greek via Cross-Lingual Projection
abstract
In this paper we present the initial construction of a treebank of Ancient Greek containing portions of the Septuagint, a translation of the Hebrew Scriptures (1576 sentences, 39K tokens, roughly 7% of the total corpus). We construct the treebank by word-aligning and projecting from the parallel text in Ancient Hebrew before automatically correcting systematic syntactic mismatches and manually correcting other errors.
Daniel G. Swanson, Bryce D. Bussert, Francis M. Tyers
LREC/COLING3
2024 A Universal Dependencies Treebank for Highland Puebla Nahuatl
abstract
Robert Pugh, Francis Tyers. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Robert Pugh, Francis M. Tyers
NAACL-HLT2
2022 Yet Another Format of Universal Dependencies for Korean
abstract
In this study, we propose a morpheme-based scheme for Korean dependency parsing and adopt the proposed scheme to Universal Dependencies. We present the linguistic rationale that illustrates the motivation and the necessity of adopting the morpheme-based format, and develop scripts that convert between the original format used by Universal Dependencies and the proposed morpheme-based format automatically. The effectiveness of the proposed format for Korean dependency parsing is then testified by both statistical and neural models, including UDPipe and Stanza, with our carefully constructed morpheme-based word embedding for Korean. morphUD outperforms parsing results for all Korean UD treebanks, and we also present detailed error analysis.
Eunkyul Leah Jo, Yundong Yao, Miikka Silfverberg, Francis M. Tyers, Jungyeul Park
COLING6
2022 Curriculum Optimization for Low-Resource Speech Recognition
abstract
Modern end-to-end speech recognition models show astonishing results in transcribing audio signals into written text. However, conventional data feeding pipelines may be sub-optimal for low-resource speech recognition, which still remains a challenging task. We propose an automated curriculum learning approach to optimize the sequence of training examples based on both the progress of the model while training and prior knowledge about the difficulty of the training examples. We introduce a new difficulty measure called compression ratio that can be used as a scoring function for raw audio in various noise conditions. The proposed method improves speech recognition Word Error Rate performance by up to 33% relative over the baseline system.
Anastasia Kuznetsova, Jennifer Drexler Fox, Francis M. Tyers
ICASSP4
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC72
2022 A Free/Open-Source Morphological Analyser and Generator for Sakha
abstract
We present, to our knowledge, the first ever published morphological analyser and generator for Sakha, a marginalised language of Siberia. The transducer, developed using HFST, has coverage of solidly above 90%, and high precision. In the development of the analyser, we have expanded linguistic knowledge about Sakha, and developed strategies for complex grammatical patterns. The transducer is already being used in downstream tasks, including computer assisted language learning applications for linguistic maintenance and computational linguistic shared tasks.
Sardana Ivanova, Jonathan Washington, Francis M. Tyers
LREC3
2022 Universal Dependencies for Western Sierra Puebla Nahuatl
abstract
We present a morpho-syntactically-annotated corpus of Western Sierra Puebla Nahuatl that conforms to the annotation guidelines of the Universal Dependencies project. We describe the sources of the texts that make up the corpus, the annotation process, and important annotation decisions made throughout the development of the corpus. As the first indigenous language of Mexico to be added to the Universal Dependencies project, this corpus offers a good opportunity to test and more clearly define annotation guidelines for the Meso-american linguistic area, spontaneous and elicited spoken data, and code-switching.
Robert Pugh, Marivel Huerta Mendez, Mitsuya Sasaki, Francis M. Tyers
LREC4
2022 A Universal Dependencies Treebank of Ancient Hebrew
abstract
In this paper we present the initial construction of a Universal Dependencies treebank with morphological annotations of Ancient Hebrew containing portions of the Hebrew Scriptures (1579 sentences, 27K tokens) for use in comparative study with ancient translations and for analysis of the development of Hebrew syntax. We construct this treebank by applying a rule-based parser (300 rules) to an existing morphologically-annotated corpus with minimal constituency structure and manually verifying the output and present the results of this semi-automated annotation process and some of the annotation decisions made in the process of applying the UD guidelines to a new language.
Daniel G. Swanson, Francis M. Tyers
LREC2
2021 A Large-Scale Study of Machine Translation in Turkic Languages
abstract
Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Bekhzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, Sriram Chellappan. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis M. Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta, Behzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, Sriram Chellappan
EMNLP (1)5
2021 Do RNN States Encode Abstract Phonological Alternations?
abstract
Miikka Silfverberg, Francis Tyers, Garrett Nicolai, Mans Hulden. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Miikka Silfverberg, Francis M. Tyers, Garrett Nicolai, Mans Hulden
NAACL-HLT2
2021 Recent advances in Apertium, a free/open-source rule-based machine translation platform for low-resource languages
abstract
Abstract This paper presents an overview of Apertium, a free and open-source rule-based machine translation platform. Translation in Apertium happens through a pipeline of modular tools, and the platform continues to be improved as more language pairs are added. Several advances have been implemented since the last publication, including some new optional modules: a module that allows rules to process recursive structures at the structural transfer stage, a module that deals with contiguous and discontiguous multi-word expressions, and a module that resolves anaphora to aid translation. Also highlighted is the hybridisation of Apertium through statistical modules that augment the pipeline, and statistical methods that augment existing modules. This includes morphological disambiguation, weighted structural transfer, and lexical selection modules that learn from limited data. The paper also discusses how a platform like Apertium can be a critical part of access to language technology for so-called low-resource languages, which might be ignored or deemed unapproachable by popular corpus-based translation technologies. Finally, the paper presents some of the released and unreleased language pairs, concluding with a brief look at some supplementary Apertium tools that prove valuable to users as well as language developers. All Apertium-related code, including language data, is free/open-source and available at https://github.com/apertium .
Tanmai Khanna, Jonathan Washington, Francis M. Tyers, Sevilay Bayatli, Daniel G. Swanson, Flammie A. Pirinen, Irene Tang, Hèctor Alòs i Font
Mach. Transl.3
2020 Common Voice: A Massively-Multilingual Speech Corpus
abstract
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla’s DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 ± 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, Gregor Weber
LREC9
2020 An Unsupervised Method for Weighting Finite-state Morphological Analyzers
abstract
Morphological analysis is one of the tasks that have been studied for years. Different techniques have been used to develop models for performing morphological analysis. Models based on finite state transducers have proved to be more suitable for languages with low available resources. In this paper, we have developed a method for weighting a morphological analyzer built using finite state transducers in order to disambiguate its results. The method is based on a word2vec model that is trained in a completely unsupervised way using raw untagged corpora and is able to capture the semantic meaning of the words. Most of the methods used for disambiguating the results of a morphological analyzer relied on having tagged corpora that need to manually built. Additionally, the method developed uses information about the token irrespective of its context unlike most of the other techniques that heavily rely on the word’s context to disambiguate its set of candidate analyses.
Amr Keleg, Francis M. Tyers, Nicholas Howell, Flammie A. Pirinen
LREC2
2020 Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
abstract
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the universal guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages.
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic 0001, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster 0001, Francis M. Tyers, Daniel Zeman
LREC8
2020 A Finite-State Morphological Analyser for Evenki
abstract
It has been widely admitted that morphological analysis is an important step in automated text processing for morphologically rich languages. Evenki is a language with rich morphology, therefore a morphological analyser is highly desirable for processing Evenki texts and developing applications for Evenki. Although two morphological analysers for Evenki have already been developed, they are able to analyse less than a half of the available Evenki corpora. The aim of this paper is to create a new morphological analyser for Evenki. It is implemented using the Helsinki Finite-State Transducer toolkit (HFST). The lexc formalism is used to specify the morphotactic rules, which define the valid orderings of morphemes in a word. Morphophonological alternations and orthographic rules are described using the twol formalism. The lexicon is extracted from available machine-readable dictionaries. Since a part of the corpora belongs to texts in Evenki dialects, a version of the analyser with relaxed rules is developed for processing dialectal features. We evaluate the analyser on available Evenki corpora and estimate precision, recall and F-score. We obtain coverage scores of between 61% and 87% on the available Evenki corpora.
Anna Zueva, Anastasia Kuznetsova, Francis M. Tyers
LREC3
2018 Rule-based machine translation from Kazakh to Turkish
abstract
This paper presents a shallow-transfer machine translation (MT) system for translating from Kazakh to Turkish. Background on the differences between the languages is presented, followed by how the system was designed to handle some of these differences. The system is based on the Apertium free/open-source machine translation platform. The structure of the system and how it works is described, along with an evaluation against two competing systems. Linguistic components were developed, including a Kazakh-Turkish bilingual dictionary, Constraint Grammar disambiguation rules, lexical selection rules, and structural transfer rules. With many known issues yet to be addressed, our RBMT system has reached performance comparable to publicly-available corpus-based MT systems between the languages.
Sevilay Bayatli, Sefer Kurnaz, Ilnar Salimzyanov, Jonathan Washington, Francis M. Tyers
EAMT5
2018 Finite-state morphological analysis for Gagauz
Francis M. Tyers, Sevilay Bayatli, Güllü Karanfil, Memduh Gokirmak
LREC1
2018 The ARIEL-CMU situation frame detection pipeline for LoReHLT16: a model translation approach
Patrick Littell, Ruochen Xu, Zaid Sheikh, David R. Mortensen, Lori S. Levin, Francis M. Tyers, Hiroaki Hayashi, Graham Horwood, Steve Sloto, Emily Tagtow, Alan W. Black, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy
Mach. Transl.7
2016 Universal Dependencies for Turkish
abstract
The Universal Dependencies (UD) project was conceived after the substantial recent interest in unifying annotation schemes across languages. With its own annotation principles and abstract inventory for parts of speech, morphosyntactic features and dependency relations, UD aims to facilitate multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. This paper presents the Turkish IMST-UD Treebank, the first Turkish treebank to be in a UD release. The IMST-UD Treebank was automatically converted from the IMST Treebank, which was also recently released. We describe this conversion procedure in detail, complete with mapping tables. We also present our evaluation of the parsing performances of both versions of the IMST Treebank. Our findings suggest that the UD framework is at least as viable for Turkish as the original annotation framework of the IMST Treebank.
Umut Sulubacak, Memduh Gokirmak, Francis M. Tyers, Çagri Çöltekin, Joakim Nivre, Gülsen Eryigit
COLING3
2016 A Finite-State Morphological Analyser for Sindhi
Raveesh Motlani, Francis M. Tyers, Dipti Misra Sharma
LREC2
2016 A Finite-state Morphological Analyser for Tuvan
Francis M. Tyers, Aziyana Bayyr-ool, Aelita Salchak, Jonathan Washington
LREC1
2015 Evaluating machine translation for assimilation via a gap-filling task
Ekaterina Ageeva, Mikel L. Forcada, Francis M. Tyers, Juan Antonio Pérez-Ortiz
EAMT3
2015 Unsupervised training of maximum-entropy models for lexical selection in rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT1
2014 Why Implementation Matters: Evaluation of an Open-source Constraint Grammar Parser
Dávid Márk Nemeskey, Francis M. Tyers, Mans Hulden
COLING2
2014 Finite-state morphological transducers for three Kypchak languages
Jonathan Washington, Ilnar Salimzyanov, Francis M. Tyers
LREC3
2014 Emily M. Bender: Linguistic fundamentals for natural language processing: 100 essentials from morphology and syntax - Morgan-Claypool, San Rafael, CA, USA, 2013, xviii + 166 pp
Francis M. Tyers
Mach. Transl.1
2013 A Free/Open-source Kazakh-Tatar Machine Translation System
Ilnar Salimzyanov, Jonathan Washington, Francis M. Tyers
MTSummit3
2012 Flexible finite-state lexical selection for rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada
EAMT1
2012 Free/Open Source Shallow-Transfer Based Machine Translation for Spanish and Aragonese
Juan Pablo Martínez 0001, Jim O'Regan, Francis M. Tyers
LREC3
2012 A finite-state morphological transducer for Kyrgyz
Jonathan Washington, Mirlan Ipasov, Francis M. Tyers
LREC3
2011 Apertium-IceNLP: A rule-based Icelandic to English machine translation system
Martha Dís Brandt, Hrafn Loftsson, Hlynur Sigurþórsson, Francis M. Tyers
EAMT4
2011 Rapid rule-based machine translation between Dutch and Afrikaans
Pim Otte, Francis M. Tyers
EAMT2
2011 Apertium: a free/open-source platform for rule-based machine translation
Mikel L. Forcada, Mireia Ginestí-Rosell, Jacob Nordfalk, Jim O'Regan, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Gema Ramírez-Sánchez, Francis M. Tyers
Mach. Transl.9
2010 Rule-based Breton to French machine translation
Francis M. Tyers
EAMT1
2009 Rule-Based Augmentation of Training Data in Breton-French Statistical Machine Translation
Francis M. Tyers
EAMT1
2009 Developing Prototypes for Machine Translation between Two Sami Languages
Francis M. Tyers, Linda Wiechetek, Trond Trosterud
EAMT1