VLDB 2026 Research / reviewers in the wild / expert
Francis M. Tyers
dblp:84/8340
· DBLP profile ↗
38ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0001-6108-2220ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla NahuatlabstractRobert Pugh, Cheyenne Wing, María Ximena Juárez Huerta, Ángeles Márquez Hernandez, Francis Tyers. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Robert Pugh, Cheyenne Wing, María Ximena Juárez Huerta, Angeles Márquez Hernandez, Francis M. Tyers |
NAACL (Long Papers) | 5 |
| 2024 | Developing a Benchmark for Pronunciation Feedback: Creation of a Phonemically Annotated Speech Corpus of isiZulu Language Learner SpeechabstractPronunciation of the phonemic inventory of a new language often presents difficulties to second language (L2) learners. These challenges can be alleviated by the development of pronunciation feedback tools that take speech input from learners and return information about errors in the utterance. This paper presents the development of a corpus designed for use in pronunciation feedback research. The corpus is comprised of gold standard recordings from isiZulu teachers and recordings from isiZulu L2 learners that have been annotated for pronunciation errors. Exploring the potential benefits of word-level versus phoneme-level feedback necessitates a speech corpus that has been annotated for errors on the phoneme-level. To aid in this discussion, this corpus of isiZulu L2 speech has been annotated for phoneme-errors in utterances, as well as suprasegmental errors in tone. Alexandra O'Neil, Nils Hjortnaes, Francis M. Tyers, Zinhle Nkosi, Thulile Ndlovu, Zanele Mlondo, Ngami Phumzile Pewa |
LREC/COLING | 3 |
| 2024 | Producing a Parallel Universal Dependencies Treebank of Ancient Hebrew and Ancient Greek via Cross-Lingual ProjectionabstractIn this paper we present the initial construction of a treebank of Ancient Greek containing portions of the Septuagint, a translation of the Hebrew Scriptures (1576 sentences, 39K tokens, roughly 7% of the total corpus). We construct the treebank by word-aligning and projecting from the parallel text in Ancient Hebrew before automatically correcting systematic syntactic mismatches and manually correcting other errors. Daniel G. Swanson, Bryce D. Bussert, Francis M. Tyers |
LREC/COLING | 3 |
| 2024 | A Universal Dependencies Treebank for Highland Puebla NahuatlabstractRobert Pugh, Francis Tyers. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Robert Pugh, Francis M. Tyers |
NAACL-HLT | 2 |
| 2022 | Yet Another Format of Universal Dependencies for KoreanabstractIn this study, we propose a morpheme-based scheme for Korean dependency parsing and adopt the proposed scheme to Universal Dependencies. We present the linguistic rationale that illustrates the motivation and the necessity of adopting the morpheme-based format, and develop scripts that convert between the original format used by Universal Dependencies and the proposed morpheme-based format automatically. The effectiveness of the proposed format for Korean dependency parsing is then testified by both statistical and neural models, including UDPipe and Stanza, with our carefully constructed morpheme-based word embedding for Korean. morphUD outperforms parsing results for all Korean UD treebanks, and we also present detailed error analysis. Eunkyul Leah Jo, Yundong Yao, Miikka Silfverberg, Francis M. Tyers, Jungyeul Park |
COLING | 6 |
| 2022 | Curriculum Optimization for Low-Resource Speech RecognitionabstractModern end-to-end speech recognition models show astonishing results in transcribing audio signals into written text. However, conventional data feeding pipelines may be sub-optimal for low-resource speech recognition, which still remains a challenging task. We propose an automated curriculum learning approach to optimize the sequence of training examples based on both the progress of the model while training and prior knowledge about the difficulty of the training examples. We introduce a new difficulty measure called compression ratio that can be used as a scoring function for raw audio in various noise conditions. The proposed method improves speech recognition Word Error Rate performance by up to 33% relative over the baseline system. Anastasia Kuznetsova, Jennifer Drexler Fox, Francis M. Tyers |
ICASSP | 4 |
| 2022 | UniMorph 4.0: Universal MorphologyabstractThe Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet. Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova |
LREC | 72 |
| 2022 | A Free/Open-Source Morphological Analyser and Generator for SakhaabstractWe present, to our knowledge, the first ever published morphological analyser and generator for Sakha, a marginalised language of Siberia. The transducer, developed using HFST, has coverage of solidly above 90%, and high precision. In the development of the analyser, we have expanded linguistic knowledge about Sakha, and developed strategies for complex grammatical patterns. The transducer is already being used in downstream tasks, including computer assisted language learning applications for linguistic maintenance and computational linguistic shared tasks. Sardana Ivanova, Jonathan Washington, Francis M. Tyers |
LREC | 3 |
| 2022 | Universal Dependencies for Western Sierra Puebla NahuatlabstractWe present a morpho-syntactically-annotated corpus of Western Sierra Puebla Nahuatl that conforms to the annotation guidelines of the Universal Dependencies project. We describe the sources of the texts that make up the corpus, the annotation process, and important annotation decisions made throughout the development of the corpus. As the first indigenous language of Mexico to be added to the Universal Dependencies project, this corpus offers a good opportunity to test and more clearly define annotation guidelines for the Meso-american linguistic area, spontaneous and elicited spoken data, and code-switching. Robert Pugh, Marivel Huerta Mendez, Mitsuya Sasaki, Francis M. Tyers |
LREC | 4 |
| 2022 | A Universal Dependencies Treebank of Ancient HebrewabstractIn this paper we present the initial construction of a Universal Dependencies treebank with morphological annotations of Ancient Hebrew containing portions of the Hebrew Scriptures (1579 sentences, 27K tokens) for use in comparative study with ancient translations and for analysis of the development of Hebrew syntax. We construct this treebank by applying a rule-based parser (300 rules) to an existing morphologically-annotated corpus with minimal constituency structure and manually verifying the output and present the results of this semi-automated annotation process and some of the annotation decisions made in the process of applying the UD guidelines to a new language. Daniel G. Swanson, Francis M. Tyers |
LREC | 2 |
| 2021 | A Large-Scale Study of Machine Translation in Turkic LanguagesabstractJamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Bekhzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, Sriram Chellappan. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis M. Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta, Behzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, Sriram Chellappan |
EMNLP (1) | 5 |
| 2021 | Do RNN States Encode Abstract Phonological Alternations?abstractMiikka Silfverberg, Francis Tyers, Garrett Nicolai, Mans Hulden. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Miikka Silfverberg, Francis M. Tyers, Garrett Nicolai, Mans Hulden |
NAACL-HLT | 2 |
| 2021 | Recent advances in Apertium, a free/open-source rule-based machine translation platform for low-resource languagesabstractAbstract This paper presents an overview of Apertium, a free and open-source rule-based machine translation platform. Translation in Apertium happens through a pipeline of modular tools, and the platform continues to be improved as more language pairs are added. Several advances have been implemented since the last publication, including some new optional modules: a module that allows rules to process recursive structures at the structural transfer stage, a module that deals with contiguous and discontiguous multi-word expressions, and a module that resolves anaphora to aid translation. Also highlighted is the hybridisation of Apertium through statistical modules that augment the pipeline, and statistical methods that augment existing modules. This includes morphological disambiguation, weighted structural transfer, and lexical selection modules that learn from limited data. The paper also discusses how a platform like Apertium can be a critical part of access to language technology for so-called low-resource languages, which might be ignored or deemed unapproachable by popular corpus-based translation technologies. Finally, the paper presents some of the released and unreleased language pairs, concluding with a brief look at some supplementary Apertium tools that prove valuable to users as well as language developers. All Apertium-related code, including language data, is free/open-source and available at https://github.com/apertium . Tanmai Khanna, Jonathan Washington, Francis M. Tyers, Sevilay Bayatli, Daniel G. Swanson, Flammie A. Pirinen, Irene Tang, Hèctor Alòs i Font |
Mach. Transl. | 3 |
| 2020 | Common Voice: A Massively-Multilingual Speech CorpusabstractThe Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla’s DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 ± 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition. Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, Gregor Weber |
LREC | 9 |
| 2020 | An Unsupervised Method for Weighting Finite-state Morphological AnalyzersabstractMorphological analysis is one of the tasks that have been studied for years. Different techniques have been used to develop models for performing morphological analysis. Models based on finite state transducers have proved to be more suitable for languages with low available resources. In this paper, we have developed a method for weighting a morphological analyzer built using finite state transducers in order to disambiguate its results. The method is based on a word2vec model that is trained in a completely unsupervised way using raw untagged corpora and is able to capture the semantic meaning of the words. Most of the methods used for disambiguating the results of a morphological analyzer relied on having tagged corpora that need to manually built. Additionally, the method developed uses information about the token irrespective of its context unlike most of the other techniques that heavily rely on the word’s context to disambiguate its set of candidate analyses. Amr Keleg, Francis M. Tyers, Nicholas Howell, Flammie A. Pirinen |
LREC | 2 |
| 2020 | Universal Dependencies v2: An Evergrowing Multilingual Treebank CollectionabstractUniversal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the universal guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages. Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic 0001, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster 0001, Francis M. Tyers, Daniel Zeman |
LREC | 8 |
| 2020 | A Finite-State Morphological Analyser for EvenkiabstractIt has been widely admitted that morphological analysis is an important step in automated text processing for morphologically rich languages. Evenki is a language with rich morphology, therefore a morphological analyser is highly desirable for processing Evenki texts and developing applications for Evenki. Although two morphological analysers for Evenki have already been developed, they are able to analyse less than a half of the available Evenki corpora. The aim of this paper is to create a new morphological analyser for Evenki. It is implemented using the Helsinki Finite-State Transducer toolkit (HFST). The lexc formalism is used to specify the morphotactic rules, which define the valid orderings of morphemes in a word. Morphophonological alternations and orthographic rules are described using the twol formalism. The lexicon is extracted from available machine-readable dictionaries. Since a part of the corpora belongs to texts in Evenki dialects, a version of the analyser with relaxed rules is developed for processing dialectal features. We evaluate the analyser on available Evenki corpora and estimate precision, recall and F-score. We obtain coverage scores of between 61% and 87% on the available Evenki corpora. Anna Zueva, Anastasia Kuznetsova, Francis M. Tyers |
LREC | 3 |
| 2018 | Rule-based machine translation from Kazakh to TurkishabstractThis paper presents a shallow-transfer machine translation (MT) system for translating from Kazakh to Turkish. Background on the differences between the languages is presented, followed by how the system was designed to handle some of these differences. The system is based on the Apertium free/open-source machine translation platform. The structure of the system and how it works is described, along with an evaluation against two competing systems. Linguistic components were developed, including a Kazakh-Turkish bilingual dictionary, Constraint Grammar disambiguation rules, lexical selection rules, and structural transfer rules. With many known issues yet to be addressed, our RBMT system has reached performance comparable to publicly-available corpus-based MT systems between the languages. Sevilay Bayatli, Sefer Kurnaz, Ilnar Salimzyanov, Jonathan Washington, Francis M. Tyers |
EAMT | 5 |
| 2018 | Finite-state morphological analysis for Gagauz
Francis M. Tyers, Sevilay Bayatli, Güllü Karanfil, Memduh Gokirmak |
LREC | 1 |
| 2018 | The ARIEL-CMU situation frame detection pipeline for LoReHLT16: a model translation approach
Patrick Littell, Ruochen Xu, Zaid Sheikh, David R. Mortensen, Lori S. Levin, Francis M. Tyers, Hiroaki Hayashi, Graham Horwood, Steve Sloto, Emily Tagtow, Alan W. Black, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy |
Mach. Transl. | 7 |
| 2016 | Universal Dependencies for TurkishabstractThe Universal Dependencies (UD) project was conceived after the substantial recent interest in unifying annotation schemes across languages. With its own annotation principles and abstract inventory for parts of speech, morphosyntactic features and dependency relations, UD aims to facilitate multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. This paper presents the Turkish IMST-UD Treebank, the first Turkish treebank to be in a UD release. The IMST-UD Treebank was automatically converted from the IMST Treebank, which was also recently released. We describe this conversion procedure in detail, complete with mapping tables. We also present our evaluation of the parsing performances of both versions of the IMST Treebank. Our findings suggest that the UD framework is at least as viable for Turkish as the original annotation framework of the IMST Treebank. Umut Sulubacak, Memduh Gokirmak, Francis M. Tyers, Çagri Çöltekin, Joakim Nivre, Gülsen Eryigit |
COLING | 3 |
| 2016 | A Finite-State Morphological Analyser for Sindhi
Raveesh Motlani, Francis M. Tyers, Dipti Misra Sharma |
LREC | 2 |
| 2016 | A Finite-state Morphological Analyser for Tuvan
Francis M. Tyers, Aziyana Bayyr-ool, Aelita Salchak, Jonathan Washington |
LREC | 1 |
| 2015 | Evaluating machine translation for assimilation via a gap-filling task
Ekaterina Ageeva, Mikel L. Forcada, Francis M. Tyers, Juan Antonio Pérez-Ortiz |
EAMT | 3 |
| 2015 | Unsupervised training of maximum-entropy models for lexical selection in rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada |
EAMT | 1 |
| 2014 | Why Implementation Matters: Evaluation of an Open-source Constraint Grammar Parser
Dávid Márk Nemeskey, Francis M. Tyers, Mans Hulden |
COLING | 2 |
| 2014 | Finite-state morphological transducers for three Kypchak languages
Jonathan Washington, Ilnar Salimzyanov, Francis M. Tyers |
LREC | 3 |
| 2014 | Emily M. Bender: Linguistic fundamentals for natural language processing: 100 essentials from morphology and syntax - Morgan-Claypool, San Rafael, CA, USA, 2013, xviii + 166 pp
Francis M. Tyers |
Mach. Transl. | 1 |
| 2013 | A Free/Open-source Kazakh-Tatar Machine Translation System
Ilnar Salimzyanov, Jonathan Washington, Francis M. Tyers |
MTSummit | 3 |
| 2012 | Flexible finite-state lexical selection for rule-based machine translation
Francis M. Tyers, Felipe Sánchez-Martínez, Mikel L. Forcada |
EAMT | 1 |
| 2012 | Free/Open Source Shallow-Transfer Based Machine Translation for Spanish and Aragonese
Juan Pablo Martínez 0001, Jim O'Regan, Francis M. Tyers |
LREC | 3 |
| 2012 | A finite-state morphological transducer for Kyrgyz
Jonathan Washington, Mirlan Ipasov, Francis M. Tyers |
LREC | 3 |
| 2011 | Apertium-IceNLP: A rule-based Icelandic to English machine translation system
Martha Dís Brandt, Hrafn Loftsson, Hlynur Sigurþórsson, Francis M. Tyers |
EAMT | 4 |
| 2011 | Rapid rule-based machine translation between Dutch and Afrikaans
Pim Otte, Francis M. Tyers |
EAMT | 2 |
| 2011 | Apertium: a free/open-source platform for rule-based machine translation
Mikel L. Forcada, Mireia Ginestí-Rosell, Jacob Nordfalk, Jim O'Regan, Sergio Ortiz-Rojas, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, Gema Ramírez-Sánchez, Francis M. Tyers |
Mach. Transl. | 9 |
| 2010 | Rule-based Breton to French machine translation
Francis M. Tyers |
EAMT | 1 |
| 2009 | Rule-Based Augmentation of Training Data in Breton-French Statistical Machine Translation
Francis M. Tyers |
EAMT | 1 |
| 2009 | Developing Prototypes for Machine Translation between Two Sami Languages
Francis M. Tyers, Linda Wiechetek, Trond Trosterud |
EAMT | 1 |