VLDB 2026 Research / reviewers in the wild / expert
Samuel Läubli
dblp:99/9242
· DBLP profile ↗
7ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0001-5362-4106ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 40% Machine translation · 40% Language models and text generation · 20% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › monolingual data augmentation
back-translation |
0.7 | 1 | 2023 | Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model · ACL (1) 2023 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation |
0.7 | 1 | 2023 | Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model · ACL (1) 2023 |
Natural language and speech › Language models and text generation › text generation
text rewriting |
0.7 | 1 | 2023 | Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model · ACL (1) 2023 |
Natural language and speech › Machine translation
document-level machine translation |
0.3 | 1 | 2018 | Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation · EMNLP 2018 |
Natural language and speech › Machine translation
machine translation evaluation |
0.3 | 1 | 2018 | Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
sequence-to-sequence · 0.7round-trip translation · 0.7pairwise ranking · 0.3human evaluation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A comparison of translation performance between DeepL and SupertextabstractAs strong machine translation (MT) systems are increasingly based on large language models (LLMs), reliable quality benchmarking requires methods that capture their ability to leverage extended context. This study compares two commercial MT systems – DeepL and Supertext – by assessing their performance on unsegmented texts. We evaluate translation quality across four language directions with professional translators assessing segments with full document-level context. While segment-level assessments indicate no strong preference between the systems in most cases, document-level analysis reveals a preference for Supertext in three out of four language directions, suggesting superior consistency across longer texts. We advocate for more context-sensitive evaluation methodologies to ensure that MT quality assessments reflect real-world usability. We release all evaluation data and scripts for further analysis and reproduction at https://github.com/supertext/evaluation_deepl_supertext. Alex Flückiger, Chantal Amrhein, Tim Graf, Frédéric Odermatt, Martin Pömsl, Philippe Schläpfer, Florian Schottmann, Samuel Läubli |
MTSummit (2) | 8 |
| 2023 | Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting ModelabstractNatural language generation models reproduce and often amplify the biases present in their training data.Previous research explored using sequence-to-sequence rewriting models to transform biased model outputs (or original texts) into more gender-fair language by creating pseudo training data through linguistic rules.However, this approach is not practical for languages with more complex morphology than English.We hypothesise that creating training data in the reverse direction, i.e. starting from gender-fair text, is easier for morphologically complex languages and show that it matches the performance of state-of-the-art rewriting models for English.To eliminate the rule-based nature of data creation, we instead propose using machine translation models to create gender-biased text from real gender-fair text via round-trip translation.Our approach allows us to train a rewriting model for German without the need for elaborate handcrafted rules.The outputs of this model increased genderfairness as shown in a human evaluation study. 1 Chantal Amrhein, Florian Schottmann, Rico Sennrich, Samuel Läubli |
ACL (1) | 4 |
| 2020 | What's the Difference Between Professional Human and Machine Translation? A Blind Multi-language Study on Domain-specific MTabstractMachine translation (MT) has been shown to produce a number of errors that require human post-editing, but the extent to which professional human translation (HT) contains such errors has not yet been compared to MT. We compile pre-translated documents in which MT and HT are interleaved, and ask professional translators to flag errors and post-edit these documents in a blind evaluation. We find that the post-editing effort for MT segments is only higher in two out of three language pairs, and that the number of segments with wrong terminology, omissions, and typographical problems is similar in HT. Lukas Fischer 0003, Samuel Läubli |
EAMT | 2 |
| 2020 | A Set of Recommendations for Assessing Human-Machine Parity in Language TranslationabstractThe quality of machine translation has increased remarkably over the past years, to the degree that it was found to be indistinguishable from professional human translation in a number of empirical investigations. We reassess Hassan et al.'s 2018 investigation into Chinese to English news translation, showing that the finding of human–machine parity was owed to weaknesses in the evaluation design—which is currently considered best practice in the field. We show that the professional human translations contained significantly fewer errors, and that perceived quality in human evaluation depends on the choice of raters, the availability of linguistic context, and the creation of reference translations. Our results call for revisiting current best practices to assess strong machine translation systems in general and human–machine parity in particular, for which we offer a set of recommendations based on our empirical findings. Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, Antonio Toral |
J. Artif. Intell. Res. | 1 |
| 2019 | Post-editing Productivity with Neural Machine Translation: An Empirical Assessment of Speed and Quality in the Banking and Finance Domain
Samuel Läubli, Chantal Amrhein, Patrick Düggelin, Beatriz Gonzalez, Alena Zwahlen, Martin Volk 0001 |
MTSummit (1) | 1 |
| 2018 | mtrain: A Convenience Tool for Machine TranslationabstractWe present mtrain, a convenience tool for machine translation. It wraps existing machine translation libraries and scripts to ease their use. mtrain is written purely in Python 3, well-documented, and freely available.1 Samuel Läubli, Mathias Müller 0002, Beat Horat, Martin Volk 0001 |
EAMT | 1 |
| 2018 | Has Machine Translation Achieved Human Parity? A Case for Document-level EvaluationabstractRecent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese-English news translation task.We empirically test this claim with alternative evaluation protocols, contrasting the evaluation of single sentences and entire documents.In a pairwise ranking experiment, human raters assessing adequacy and fluency show a stronger preference for human over machine translation when evaluating documents as compared to isolated sentences.Our findings emphasise the need to shift towards document-level evaluation as machine translation improves to the degree that errors which are hard or impossible to spot at the sentence-level become decisive in discriminating quality of different translation outputs. Samuel Läubli, Rico Sennrich, Martin Volk 0001 |
EMNLP | 1 |