VLDB 2026 Research / reviewers in the wild / expert
Firas Trabelsi
dblp:379/6048
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 78% Machine translation · 22% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
0.9 | 1 | 2025 | Learning from others' mistakes: Finetuning machine translation models with span-level error annotations · ICML 2025 |
Natural language and speech › Language models and text generation
decoding |
0.8 | 1 | 2024 | Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms · NeurIPS 2024 |
Natural language and speech › Language models and text generation › decoding
efficient decoding |
0.8 | 1 | 2024 | Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms · NeurIPS 2024 |
Natural language and speech › Language models and text generation › decoding
minimum bayes risk decoding |
0.8 | 1 | 2024 | Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
training with annotations · 0.9direct preference optimization · 0.9matrix completion · 0.8low-rank approximation · 0.8alternating least squares · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning from others' mistakes: Finetuning machine translation models with span-level error annotationsabstractDespite growing interest in incorporating feedback to improve language models, most efforts focus only on sequence-level annotations. In this work, we explore the potential of utilizing fine-grained span-level annotations from offline datasets to improve model quality. We develop a simple finetuning algorithm, called Training with Annotations (TWA), to directly train machine translation models on such annotated data. TWA utilizes targeted span-level error information while also flexibly learning what to penalize within a span. Moreover, TWA considers the overall trajectory of a sequence when deciding which non-error spans to utilize as positive signals. Experiments on English-German and Chinese-English machine translation show that TWA outperforms baselines such as supervised finetuning on sequences filtered for quality and Direct Preference Optimization on pairs constructed from the same data. Lily H. Zhang, Hamid Dadkhahi, Mara Finkelstein, Firas Trabelsi, Jiaming Luo, Markus Freitag |
ICML | 4 |
| 2024 | Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion AlgorithmsabstractMinimum Bayes Risk (MBR) decoding is a powerful decoding strategy widely used for text generation tasks but its quadratic computational complexity limits its practical application. This paper presents a novel approach for approximating MBR decoding using matrix completion techniques, focusing on a machine translation task. We formulate MBR decoding as a matrix completion problem, where the utility metric scores between candidate hypotheses and reference translations form a low-rank matrix. First we empirically show that the scores matrices indeed have a low-rank structure. Then we exploit this by only computing a random subset of the scores and efficiently recover the missing entries in the matrix by applying the Alternating Least Squares (ALS) algorithm, thereby enabling fast approximation of the MBR decoding process. Our experimental results on machine translation tasks demonstrate that the proposed method requires 1/16 utility metric computations compared to the vanilla MBR decoding while achieving equal translation quality measured by COMET on the WMT22 dataset (en<>de, en<>ru). We also benchmark our method against other approximation methods and we show significant gains in quality. Firas Trabelsi, David Vilar, Mara Finkelstein, Markus Freitag |
NeurIPS | 1 |