VLDB 2026 Research / reviewers in the wild / expert
Parker Riley
dblp:222/9463
· DBLP profile ↗
9ranked-venue papers
6as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Machine translation · 51% Language models and text generation · 34% Information extraction and text analysis · 14% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
machine translation evaluation |
2.7 | 3 | 2026 | MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation · ACL (1) 2026 From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set · ICML 2025 Enhancing Human Evaluation in Machine Translation with Comparative Judgement · ACL (1) 2025 |
Natural language and speech › Language models and text generation › text evaluation
human evaluation |
1.9 | 2 | 2026 | MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation · ACL (1) 2026 Enhancing Human Evaluation in Machine Translation with Comparative Judgement · ACL (1) 2025 |
Natural language and speech › Machine translation › machine translation evaluation
automatic evaluation metrics |
0.9 | 1 | 2025 | From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set · ICML 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set · ICML 2025 |
Natural language and speech › Information extraction and text analysis › data annotation
inter-annotator agreement |
0.9 | 1 | 2025 | Enhancing Human Evaluation in Machine Translation with Comparative Judgement · ACL (1) 2025 |
Natural language and speech › Language models and text generation › controllable text generation
text style transfer |
0.5 | 1 | 2021 | TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted Restyling · ACL/IJCNLP (1) 2021 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.4 | 1 | 2020 | Translationese as a Language in "Multilingual" NMT · ACL 2020 |
Natural language and speech › Machine translation
neural machine translation |
0.4 | 1 | 2020 | Translationese as a Language in "Multilingual" NMT · ACL 2020 |
Natural language and speech › Machine translation › neural machine translation › multilingual neural machine translation
zero-shot translation |
0.4 | 1 | 2020 | Translationese as a Language in "Multilingual" NMT · ACL 2020 |
Natural language and speech › Information extraction and text analysis › data annotation
annotation schemes |
0.3 | 1 | 2026 | MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation · ACL (1) 2026 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.1 | 1 | 2021 | TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted Restyling · ACL/IJCNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
MQM annotation · 1.0side-by-side comparison · 0.9relative ranking · 0.9multidimensional quality metrics · 0.9large language model · 0.9in-context learning · 0.9few-shot learning · 0.5train-data tagging · 0.4sentence-level classifier · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine TranslationabstractParker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni, Juraj Juraska, Markus Freitag. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Parker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni, Juraj Juraska, Markus Freitag |
ACL (1) | 1 |
| 2025 | Enhancing Human Evaluation in Machine Translation with Comparative JudgementabstractHuman evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups—point-wise Multidimensional Quality Metrics (MQM), side-by-side (S×S) MQM, and its simplified version S×S relative ranking (RR). In MQM, annotators mark error spans with categories and severity levels. S×S MQM extends MQM to pairwise error annotation for two translations of the same input, while S×S RR focuses on selecting the better output without labeling errors.Key findings are: (1) the S×S settings achieve higher inter-annotator agreement than MQM; (2) S×S MQM enhances inter-translation error marking consistency compared to MQM by, on average, 38.5% for explicitly compared MT systems and 19.5% for others; (3) all annotation settings return stable system rankings, with S×S RR offering a more efficient alternative to (S×S) MQM; (4) the S×S settings highlight subtle errors overlooked in MQM without altering absolute system evaluations.To spur further research, we will release the triply annotated datasets comprising 377 ZhEn and 104 EnDe annotation examples, each covering 10 systems. Yixiao Song, Parker Riley, Daniel Deutsch, Markus Freitag |
ACL (1) | 2 |
| 2025 | From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test SetabstractAs LLMs continue to become more powerful and versatile, human evaluation has become intractable at scale and reliance on automatic metrics has become the norm. Recently, it has been shown that LLMs are themselves state-of-the-art evaluators for many tasks. These *Autoraters* are typically designed so that they generalize to new systems *and* test sets. In practice, however, evaluation is performed on a small set of fixed, canonical test sets, which are carefully curated to measure the capabilities of interest and are not changed frequently. In this work, we design a method which specializes a prompted Autorater to a given test set, by leveraging historical ratings on the test set to construct in-context learning (ICL) examples. We evaluate our *Specialist* method on the task of fine-grained machine translation evaluation, and show that it dramatically outperforms the state-of-the-art XCOMET metric by 54% and 119% on the WMT'23 and WMT'24 test sets, respectively. We perform extensive analyses to understand the representations learned by our Specialist metrics, and how variability in rater behavior affects their performance. We also verify the generalizability and robustness of our Specialist method across different numbers of ICL examples, LLM backbones, systems to evaluate, and evaluation tasks. Mara Finkelstein, Daniel Deutsch, Parker Riley, Juraj Juraska, Geza Kovacs, Markus Freitag |
ICML | 3 |
| 2024 | Finding Replicable Human Evaluations via Stable Ranking ProbabilityabstractParker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, Markus Freitag. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Parker Riley, Daniel Deutsch, George F. Foster, Viresh Ratnakar, Ali Dabirmoghaddam, Markus Freitag |
NAACL-HLT | 1 |
| 2023 | FRMT: A Benchmark for Few-Shot Region-Aware Machine TranslationabstractAbstract We present FRMT, a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation. The dataset consists of professional translations from English into two regional variants each of Portuguese and Mandarin Chinese. Source documents are selected to enable detailed analysis of phenomena of interest, including lexically distinct terms and distractor terms. We explore automatic evaluation metrics for FRMT and validate their correlation with expert human evaluation across both region-matched and mismatched rating scenarios. Finally, we present a number of baseline models for this task, and offer guidelines for how researchers can train, evaluate, and compare their own models. Our dataset and evaluation code are publicly available: https://bit.ly/frmt-task. Parker Riley, Timothy Dozat, Jan A. Botha, Xavier Garcia, Dan Garrette, Jason Riesa, Orhan Firat, Noah Constant |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted RestylingabstractParker Riley, Noah Constant, Mandy Guo, Girish Kumar, David Uthus, Zarana Parekh. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Parker Riley, Noah Constant, Mandy Guo, David C. Uthus, Zarana Parekh |
ACL/IJCNLP (1) | 1 |
| 2021 | Outside Computation with Superior FunctionsabstractWe show that a general algorithm for efficient computation of outside values under the minimum of superior functions framework proposed by Knuth (1977) would yield a subexponential time algorithm for SAT, violating the Strong Exponential Time Hypothesis (SETH). Parker Riley, Daniel Gildea |
NAACL-HLT | 1 |
| 2020 | Translationese as a Language in "Multilingual" NMTabstractMachine translation has an undesirable propensity to produce "translationese" artifacts, which can lead to higher BLEU scores while being liked less by human raters.Motivated by this, we model translationese and original (i.e.natural) text as separate languages in a multilingual model, and pose the question: can we perform zero-shot translation between original source text and original target text?There is no data with original source and original target, so we train a sentence-level classifier to distinguish translationese from original target text, and use this classifier to tag the training data for an NMT model.Using this technique we bias the model to produce more natural outputs at test time, yielding gains in human evaluation scores on both adequacy and fluency.Additionally, we demonstrate that it is possible to bias the model to produce translationese and game the BLEU score, increasing it while decreasing human-rated quality.We analyze these outputs using metrics measuring the degree of translationese, and present an analysis of the volatility of heuristic-based train-data tagging. Parker Riley, Isaac Caswell, Markus Freitag, David Grangier |
ACL | 1 |
| 2018 | Feature-Based Decipherment for Machine TranslationabstractOrthographic similarities across languages provide a strong signal for unsupervised probabilistic transduction (decipherment) for closely related language pairs. The existing decipherment models, however, are not well suited for exploiting these orthographic similarities. We propose a log-linear model with latent variables that incorporates orthographic similarity features. Maximum likelihood training is computationally expensive for the proposed log-linear model. To address this challenge, we perform approximate inference via Markov chain Monte Carlo sampling and contrastive divergence. Our results show that the proposed log-linear model with contrastive divergence outperforms the existing generative decipherment models by exploiting the orthographic features. The model both scales to large vocabularies and preserves accuracy in low- and no-resource contexts. Iftekhar Naim, Parker Riley, Daniel Gildea |
Comput. Linguistics | 2 |