Arturo Oncevay

dblp:167/2188 · also Arturo Oncevay-Marcos · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-7675-6208ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 AfroCS-xs: Creating a Compact, High-Quality, Human-Validated Code-Switched Dataset for African Languages
abstract
Kayode Olaleye, Arturo Oncevay, Mathieu Sibue, Nombuyiselo Zondi, Michelle Terblanche, Sibongile Mapikitla, Richard Lastrucci, Charese Smiley, Vukosi Marivate. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kayode Olaleye, Arturo Oncevay, Mathieu Sibue, Nombuyiselo Zondi, Michelle Terblanche, Sibongile Mapikitla, Richard Lastrucci, Charese Smiley, Vukosi Marivate
ACL (1)2
2025 Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education
abstract
Domain-specific multilingual terminology is essential for accurate machine translation (MT) and cross-lingual NLP applications.We present a gold-standard terminology resource for the tax and financial education domains, built from curated governmental publications and covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.Using this resource, we assess various MT systems and LLMs on translation quality and term accuracy.We annotate over 3,000 terms for domain-specificity, facilitating a comparison between domain-specific and general term translations, and observe models' challenges with specialized tax terms.We also analyze the case of terminology-aided translation, and the LLMs' performance in extracting the translated term given the context.Our results highlight model limitations and the value of high-quality terminologies for advancing MT research in specialized contexts.1 * Contribution done while working at JPMorgan. 1 Please contact the author(s) if you want to have access to the terminologies and parallel data.
Arturo Oncevay, Elena Kochkina, Keshav Ramani, Toyin Aguda, Simerjot Kaur, Charese Smiley
EMNLP1
2025 The Impact of Domain-Specific Terminology on Machine Translation for Finance in European Languages
abstract
Arturo Oncevay, Charese Smiley, Xiaomo Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Arturo Oncevay, Charese Smiley, Xiaomo Liu
NAACL (Long Papers)1
2023 Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained Models
abstract
Abteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, Katharina Kann. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Abteen Ebrahimi, Arya McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez Lugo, Rolando Coto-Solano, Katharina Kann
EACL3
2023 KATIE: A System for Key Attributes Identification in Product Knowledge Graph Construction
abstract
We present part of Huawei's efforts in building a Product Knowledge Graph (PKG). We want to identify which product attributes (i.e. properties) are relevant and important in terms of shopping decisions to product categories (i.e. classes). This is particularly challenging when the attributes and their values are mined from online product catalogues, i.e. HTML pages. These web pages contain semi-structured data, which do not follow a concerted format and use diverse vocabulary to designate the same features. We propose a system for key attribute identification (KATIE) based on fine-tuning pre-trained models (e.g., DistilBERT) to predict the applicability and importance of an attribute to a category. We also propose an attribute synonyms identification module that allows us to discover synonymous attributes by considering not only their labels' similarities but also the similarity of their values sets. We have evaluated our approach to Huawei categories taxonomy and a set of internally mined attributes from web pages. KATIE guarantees promising performance results compared to the most recent baselines.
Btissam Er-Rahmadi, Arturo Oncevay, Yuanyi Ji, Jeff Z. Pan
SIGIR2
2022 AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages
abstract
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John E. Ortega, Ricardo Ramos, Annette Rios, Iván V. Meza, Gustavo Giménez Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Ngoc Thang Vu, Katharina Kann
ACL (1)3
2022 Revisiting Syllables in Language Modelling and Their Application on Low-Resource Machine Translation
abstract
Language modelling and machine translation tasks mostly use subword or character inputs, but syllables are seldom used. Syllables provide shorter sequences than characters, require less-specialised extracting rules than morphemes, and their segmentation is not impacted by the corpus size. In this study, we first explore the potential of syllables for open-vocabulary language modelling in 21 languages. We use rule-based syllabification methods for six languages and address the rest with hyphenation, which works as a syllabification proxy. With a comparable perplexity, we show that syllables outperform characters and other subwords. Moreover, we study the importance of syllables on neural machine translation for a non-related and low-resource language-pair (Spanish–Shipibo-Konibo). In pairwise and multilingual systems, syllables outperform unsupervised subwords, and further morphological segmentation methods, when translating into a highly synthetic language with a transparent orthography (Shipibo-Konibo). Finally, we perform some human evaluation, and discuss limitations and opportunities.
Arturo Oncevay, Kervy Rivas Rojas, Liz Karen Chavez Sanchez, Roberto Zariquiey
COLING1
2022 UniMorph 4.0: Universal Morphology
abstract
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieras, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane 0002, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóga, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer C. White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo Maria Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar 0002, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Tucker Prud'hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
LREC24
2022 Building an Endangered Language Resource in the Classroom: Universal Dependencies for Kakataibo
abstract
In this paper, we launch a new Universal Dependencies treebank for an endangered language from Amazonia: Kakataibo, a Panoan language spoken in Peru. We first discuss the collaborative methodology implemented, which proved effective to create a treebank in the context of a Computational Linguistic course for undergraduates. Then, we describe the general details of the treebank and the language-specific considerations implemented for the proposed annotation. We finally conduct some experiments on part-of-speech tagging and syntactic dependency parsing. We focus on monolingual and transfer learning settings, where we study the impact of a Shipibo-Konibo treebank, another Panoan language resource.
Roberto Zariquiey, Claudia Alvarado, Ximena Echevarría, Luisa Gomez, Rosa Gonzales, Mariana Illescas, Sabina Oporto, Frederic Blum, Arturo Oncevay, Javier Vera
LREC9
2022 Quantifying Synthesis and Fusion and their Impact on Machine Translation
abstract
Arturo Oncevay, Duygu Ataman, Niels Van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Arturo Oncevay, Duygu Ataman, Niels van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva
NAACL-HLT1
2020 Efficient Strategies for Hierarchical Text Classification: External Knowledge and Auxiliary Tasks
abstract
In hierarchical text classification, we perform a sequence of inference steps to predict the category of a document from top to bottom of a given class taxonomy.Most of the studies have focused on developing novels neural network architectures to deal with the hierarchical structure, but we prefer to look for efficient ways to strengthen a baseline model.We first define the task as a sequence-to-sequence problem.Afterwards, we propose an auxiliary synthetic task of bottom-up-classification.Then, from external dictionaries, we retrieve textual definitions for the classes of all the hierarchy's layers, and map them into the word vector space.We use the class-definition embeddings as an additional input to condition the prediction of the next layer and in an adapted beam search.Whereas the modified search did not provide large gains, the combination of the auxiliary task and the additional input of classdefinitions significantly enhance the classification accuracy.With our efficient approaches, we outperform previous studies, using a drastically reduced number of parameters, in two well-known English datasets.
Kervy Rivas Rojas, Gina Bustamante, Arturo Oncevay, Marco Antonio Sobrevilla Cabezudo
ACL3
2020 Bridging Linguistic Typology and Multilingual Machine Translation with Multi-View Language Representations
abstract
Sparse language vectors from linguistic typology databases and learned embeddings from tasks like multilingual machine translation have been investigated in isolation, without analysing how they could benefit from each other's language characterisation.We propose to fuse both views using singular vector canonical correlation analysis and study what kind of information is induced from each source.By inferring typological features and language phylogenies, we observe that our representations embed typology and strengthen correlations with language relationships.We then take advantage of our multi-view language vector space for multilingual machine translation, where we achieve competitive overall translation accuracy in tasks that require information about language similarities, such as language clustering and ranking candidates for multilingual transfer.With our method, which is also released as a tool, we can easily project and assess new languages without expensive retraining of massive multilingual or ranking models, which are major disadvantages of related approaches.
Arturo Oncevay, Barry Haddow, Alexandra Birch
EMNLP (1)1
2020 No Data to Crawl? Monolingual Corpus Creation from PDF Files of Truly low-Resource Languages in Peru
abstract
We introduce new monolingual corpora for four indigenous and endangered languages from Peru: Shipibo-konibo, Ashaninka, Yanesha and Yine. Given the total absence of these languages in the web, the extraction and processing of texts from PDF files is relevant in a truly low-resource language scenario. Our procedure for monolingual corpus creation considers language-specific and language-agnostic steps, and focuses on educational PDF files with multilingual sentences, noisy pages and low-structured content. Through an evaluation based on language modelling and character-level perplexity on a subset of manually extracted sentences, we determine that our method allows the creation of clean corpora for the four languages, a key resource for natural language processing tasks nowadays.
Gina Bustamante, Arturo Oncevay, Roberto Zariquiey
LREC2
2018 ChAnot: An Intelligent Annotation Tool for Indigenous and Highly Agglutinative Languages in Peru
Rodolfo Mercado-Gonzales, José Pereira-Noriega, Marco Antonio Sobrevilla Cabezudo, Arturo Oncevay
LREC4
2018 Corpus Building and Evaluation of Aspect-based Opinion Summaries from Tweets in Spanish
Daniel Peñaloza, Juanjosé Tenorio Peña, Arturo Oncevay, Marco Antonio Sobrevilla Cabezudo
LREC5
2018 WordNet-Shp: Towards the Building of a Lexical Database for a Peruvian Minority Language
Diego Maguiño Valencia, Arturo Oncevay, Marco Antonio Sobrevilla Cabezudo
LREC2
2017 SenseDependency-Rank: A Word Sense Disambiguation Method Based on Random Walks and Dependency Trees
Marco Antonio Sobrevilla Cabezudo, Arturo Oncevay, H. Andrés Melgar S.
CICLing (1)2
2015 Leaf-Based Plant Identification Through Morphological Characterization in Digital Images
Arturo Oncevay, Ronald Juárez Chambi, Sofía Khlebnikov-Núñez, César Beltrán Castañón
CAIP (2)1