VLDB 2026 Research / reviewers in the wild / expert
John E. Ortega
dblp:286/4665 · also John Evan Ortega
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Is Peer-Reviewing Worth the Effort?abstractHow effective is peer-reviewing in identifying important papers? We treat this question as a forecasting task. Can we predict which papers will be highly cited in the future based on venue and “early returns” (citations soon after publication)? We show early returns are more predictive than venue. Finally, we end with a constructive suggestion to simplify reviewing. Kenneth Church 0001, Raman Chandrasekar, John E. Ortega, Ibrahim Said Ahmad |
COLING | 3 |
| 2025 | Semantic Role Labeling of NomBank PartitivesabstractThis article is about Semantic Role Labeling for English partitive nouns (5%/REL of the price/ARG1; The price/ARG1 rose 5 percent/REL) in the NomBank annotated corpus. Several systems are described using traditional and transformer-based machine learning, as well as ensembling. Our highest scoring system achieves an F1 of 91.74% using “gold” parses from the Penn Treebank and 91.12% when using the Berkeley Neural parser. This research includes both classroom and experimental settings for system development. Adam Meyers 0001, Advait Pravin Savant, John E. Ortega |
COLING | 3 |
| 2025 | Lexicography Saves Lives (LSL): Automatically Translating Suicide-Related LanguageabstractRecent years have seen a marked increase in research that aims to identify or predict risk, intention or ideation of suicide. The majority of new tasks, datasets, language models and other resources focus on English and on suicide in the context of Western culture. However, suicide is global issue and reducing suicide rate by 2030 is one of the key goals of the UN’s Sustainable Development Goals. Previous work has used English dictionaries related to suicide to translate into different target languages due to lack of other available resources. Naturally, this leads to a variety of ethical tensions (e.g.: linguistic misrepresentation), where discourse around suicide is not present in a particular culture or country. In this work, we introduce the ‘Lexicography Saves Lives Project’ to address this issue and make three distinct contributions. First, we outline ethical consideration and provide overview guidelines to mitigate harm in developing suicide-related resources. Next, we translate an existing dictionary related to suicidal ideation into 200 different languages and conduct human evaluations on a subset of translated dictionaries. Finally, we introduce a public website to make our resources available and enable community participation. Annika Marie Schoene, John E. Ortega, Rodolfo Zevallos, Laura Haaber Ihle |
COLING | 2 |
| 2024 | Evaluating Self-Supervised Speech Representations for Indigenous American LanguagesabstractThe application of self-supervision to speech representation learning has garnered significant interest in recent years, due to its scalability to large amounts of unlabeled data. However, much progress, both in terms of pre-training and downstream evaluation, has remained concentrated in monolingual models that only consider English. Few models consider other languages, and even fewer consider indigenous ones. In this work, benchmark the efficacy of large SSL models on 6 indigenous America languages: Quechua, Guarani , Bribri, Kotiria, Wa’ikhana, and Totonac on low-resource ASR. Our results show surprisingly strong performance by state-of-the-art SSL models, showing the potential generalizability of large-scale models to real-world data. Chih-Chen Chen, Rodolfo Zevallos, John E. Ortega |
LREC/COLING | 4 |
| 2024 | Related Work Is All You NeedabstractIn modern times, generational artificial intelligence is used in several industries and by many people. One use case that can be considered important but somewhat redundant is the act of searching for related work and other references to cite. As an avenue to better ascertain the value of citations and their corresponding locations, we focus on the common “related work” section as a focus of experimentation with the overall objective to generate the section. In this article, we present a corpus with 400k annotations of that distinguish related work from the rest of the references. Additionally, we show that for the papers in our experiments, the related work section represents the paper just as good, and in many cases, better than the rest of the references. We show that this is the case for more than 74% of the articles when using cosine similarity to measure the distance between two common graph neural network algorithms: Prone and Specter. Rodolfo Zevallos, John E. Ortega, Benjamin Irving |
LREC/COLING | 2 |
| 2023 | Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained ModelsabstractAbteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, Katharina Kann. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Abteen Ebrahimi, Arya McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez Lugo, Rolando Coto-Solano, Katharina Kann |
EACL | 4 |
| 2023 | An Example of (Too Much) Hyper-Parameter Tuning In Suicide Ideation DetectionabstractThis work starts with the TWISCO baseline, a benchmark of suicide-related content from Twitter. We find that hyper-parameter tuning can improve this baseline by 9%. We examined 576 combinations of hyper-parameters: learning rate, batch size, epochs and date range of training data. Reasonable settings of learning rate and batch size produce better results than poor settings. Date range is less conclusive. Balancing the date range of the training data to match the benchmark ought to improve performance, but the differences are relatively small. Optimal settings of learning rate and batch size are much better than poor settings, but optimal settings of date range are not that different from poor settings of date range. Finally, we end with concerns about reproducibility. Of the 576 experiments, 10% produced F1 performance above baseline. It is common practice in the literature to run many experiments and report the best, but doing so may be risky, especially given the sensitive nature of Suicide Ideation Detection. Annika Marie Schoene, John E. Ortega, Silvio Amir, Kenneth Church 0001 |
ICWSM | 2 |
| 2023 | Emerging trends: Unfair, biased, addictive, dangerous, deadly, and insanely profitableabstractAbstract There has been considerable work recently in the natural language community and elsewhere on Responsible AI. Much of this work focuses on fairness and biases (henceforth Risks 1.0), following the 2016 best seller:Weapons of Math Destruction. Two books published in 2022, The Chaos MachineandLike, Comment, Subscribe, raise additional risks to public health/safety/security such as genocide, insurrection, polarized politics, vaccinations (henceforth, Risks 2.0). These books suggest that the use of machine learning to maximize engagement in social media has created a Frankenstein Monster that is exploiting human weaknesses with persuasive technology, the illusory truth effect, Pavlovian conditioning, and Skinner’s intermittent variable reinforcement. Just as we cannot expect tobacco companies to sell fewer cigarettes and prioritize public health ahead of profits, so too, it may be asking too much of companies (and countries) to stop trafficking in misinformation given that it is so effective and so insanely profitable (at least in the short term). Eventually, we believe the current chaos will end, like the lawlessness in Wild West, because chaos is bad for business. As computer scientists, this paper will summarize criticisms from other fields and focus on implications for computer science; we will not attempt to contribute to those other fields. There is quite a bit of work in computer science on these risks, especially on Risks 1.0 (bias and fairness), but more work is needed, especially on Risks 2.0 (addictive, dangerous, and deadly). Kenneth Church 0001, Annika Marie Schoene, John E. Ortega, Raman Chandrasekar, Valia Kordoni |
Nat. Lang. Eng. | 3 |
| 2022 | AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource LanguagesabstractAbteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John E. Ortega, Ricardo Ramos, Annette Rios, Iván V. Meza, Gustavo Giménez Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Ngoc Thang Vu, Katharina Kann |
ACL (1) | 7 |
| 2022 | WordNet-QU: Development of a Lexical Database for Quechua VarietiesabstractIn the effort to minimize the risk of extinction of a language, linguistic resources are fundamental. Quechua, a low-resource language from South America, is a language spoken by millions but, despite several efforts in the past, still lacks the resources necessary to build high-performance computational systems. In this article, we present WordNet-QU which signifies the inclusion of Quechua in a well-known lexical database called wordnet. We propose WordNet-QU to be included as an extension to wordnet after demonstrating a manually-curated collection of multiple digital resources for lexical use in Quechua. Our work uses the synset alignment algorithm to compare Quechua to its geographically nearest high-resource language, Spanish. Altogether, we propose a total of 28,582 unique synset IDs divided according to region like so: 20510 for Southern Quechua, 5993 for Central Quechua, 1121 for Northern Quechua, and 958 for Amazonian Quechua. Nelsi Melgarejo, Rodolfo Zevallos, Hector Gomez, John E. Ortega |
COLING | 4 |
| 2022 | Fuzzy-Match Repair Guided by Quality EstimationabstractComputer-aided translation tools based on translation memories are widely used to assist professional translators. A translation memory (TM) consists of a set of translation units (TU) made up of source- and target-language segment pairs. For the translation of a new source segment$s^{\prime }$, these tools search the TM and retrieve the TUs$(s,t)$whose source segments are more similar to$s^{\prime }$. The translator then chooses a TU and edit the target segment$t$to turn it into an adequate translation of$s^{\prime }$.Fuzzy-match repair(FMR) techniques can be used to automatically modify the parts of$t$that need to be edited. We describe a language-independent FMR method that first uses machine translation to generate, given$s^{\prime }$and$(s,t)$, a set of candidate fuzzy-match repaired segments, and then chooses the best one by estimating their quality. An evaluation on three different language pairs shows that the selected candidate is a good approximation to the best (oracle) candidate produced and is closer to reference translations than machine-translated segments and unrepaired fuzzy matches ($t$). In addition, a single quality estimation model trained on a mix of data from all the languages performs well on any of the languages used. John E. Ortega, Mikel L. Forcada, Felipe Sánchez-Martínez |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Neural machine translation with a polysynthetic low resource language
John E. Ortega, Richard Castro Mamani, Kyunghyun Cho |
Mach. Transl. | 1 |
| 2019 | Improving Translations by Combining Fuzzy-Match Repair with Automatic Post-Editing
John E. Ortega, Felipe Sánchez-Martínez, Marco Turchi, Matteo Negri |
MTSummit (1) | 1 |
| 2018 | Letting a Neural Network Decide Which Machine Translation System to Use for Black-Box Fuzzy-Match RepairabstractWhile systems using the Neural Network-based Machine Translation (NMT) paradigm achieve the highest scores on recent shared tasks, phrase-based (PBMT) systems, rule-based (RBMT) systems and other systems may get better results for individual examples. Therefore, combined systems should achieve the best results for MT, particularly if the system combination method can take advantage of the strengths of each paradigm. In this paper, we describe a system that predicts whether a NMT, PBMT or RBMT will get the best Spanish translation result for a particular English sentence in DGT-TM 20161. Then we use fuzzy-match repair (FMR) as a mechanism to show that the combined system outperforms individual systems in a black-box machine translation setting. John E. Ortega, Weiyi Lu, Adam Meyers 0001, Kyunghyun Cho |
EAMT | 1 |
| 2018 | A Comparative Study of Classifying Legal Documents with Neural NetworksabstractIn recent years, deep learning has shown promising results when used in the field of natural language processing (NLP).Neural networks (NNs) such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been used for various NLP tasks including sentiment analysis, information retrieval, and document classification.In this paper, we the present the Supreme Court Classifier (SCC), a system that applies these methods to the problem of document classification of legal court opinions.We compare methods using traditional machine learning with recent NN-based methods.We also present a CNN used with pre-trained word vectors which shows improvements over the state-of-the-art applied to our dataset.We train and evaluate our system using the Washington University School of Law Supreme Court Database (SCDB).Our best system (word2vec + CNN) achieves 72.4% accuracy when classifying the court decisions into 15 broad SCDB categories and 31.9%accuracy when classifying among 279 finer-grained SCDB categories. Samir Undavia, Adam Meyers 0001, John E. Ortega |
FedCSIS | 3 |