VLDB 2026 Research / reviewers in the wild / expert
Avneesh Saluja
dblp:139/1040
· DBLP profile ↗
6ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Machine translation · 60% Language models and text generation · 40% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › monolingual data augmentation
back-translation |
0.4 | 1 | 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020 |
Natural language and speech › Machine translation
low-resource machine translation |
0.4 | 1 | 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.4 | 1 | 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020 |
Natural language and speech › Language models and text generation › text generation › text simplification
sentence simplification |
0.4 | 1 | 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020 |
Natural language and speech › Language models and text generation › language modeling
n-gram language model |
0.2 | 1 | 2014 | Language Modeling with Power Low Rank Ensembles · EMNLP 2014 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.2 | 1 | 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014 |
Natural language and speech › Language models and text generation › language modeling
statistical language modeling |
0.2 | 1 | 2014 | Language Modeling with Power Low Rank Ensembles · EMNLP 2014 |
Natural language and speech › Machine translation
statistical machine translation |
0.2 | 1 | 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014 |
Natural language and speech › Machine translation › syntax-based machine translation
synchronous context-free grammar |
0.2 | 1 | 2014 | Latent-Variable Synchronous CFGs for Hierarchical Translation · EMNLP 2014 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.2 | 1 | 2014 | Latent-Variable Synchronous CFGs for Hierarchical Translation · EMNLP 2014 |
Natural language and speech › Machine translation › statistical machine translation
translation rule extraction |
0.2 | 1 | 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014 |
Natural language and speech › Language models and text generation › text evaluation
human evaluation |
0.1 | 1 | 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020 |
Natural language and speech › Machine translation
machine translation evaluation |
0.1 | 1 | 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
sequence-to-sequence paraphrase model · 0.4back-translation · 0.4spectral method of moments · 0.2monolingual corpora · 0.2low rank matrix ensemble · 0.2kneser-ney smoothing · 0.2graph-based semi-supervised learning · 0.2graph propagation · 0.2expectation-maximization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Simplify-Then-Translate: Automatic Preprocessing for Black-Box TranslationabstractBlack-box machine translation systems have proven incredibly useful for a variety of applications yet by design are hard to adapt, tune to a specific domain, or build on top of. In this work, we introduce a method to improve such systems via automatic pre-processing (APP) using sentence simplification. We first propose a method to automatically generate a large in-domain paraphrase corpus through back-translation with a black-box MT system, which is used to train a paraphrase model that “simplifies” the original sentence to be more conducive for translation. The model is used to preprocess source sentences of multiple low-resource language pairs. We show that this preprocessing leads to better translation performance as compared to non-preprocessed source sentences. We further perform side-by-side human evaluation to verify that translations of the simplified sentences are better than the original ones. Finally, we provide some guidance on recommended language pairs for generating the simplification model corpora by investigating the relationship between ease of translation of a language pair (as measured by BLEU) and quality of the resulting simplification model from back-translations of this language pair (as measured by SARI), and tie this into the downstream task of low-resource translation. Sneha Mehta, Bahareh Azarnoush, Boris Chen, Avneesh Saluja, Vinith Misra, Ballav Bihani, Ritwik Kumar |
AAAI | 4 |
| 2014 | Graph-based Semi-Supervised Learning of Translation Models from Monolingual DataabstractStatistical phrase-based translation learns translation rules from bilingual corpora, and has traditionally only used monolingual evidence to construct features that rescore existing translation candidates.In this work, we present a semi-supervised graph-based approach for generating new translation rules that leverages bilingual and monolingual data.The proposed technique first constructs phrase graphs using both source and target language monolingual corpora.Next, graph propagation identifies translations of phrases that were not observed in the bilingual corpus, assuming that similar phrases have similar translations.We report results on a large Arabic-English system and a medium-sized Urdu-English system.Our proposed approach significantly improves the performance of competitive phrasebased systems, leading to consistent improvements between 1 and 4 BLEU points on standard evaluation sets.Source!Target! el gato! los gatos!un gato! cat! the cat! the cats! a cat! Target!Prob.! the cat! 0.7! cat! 0.15! …! …! felino!canino!el perro!Target!Prob.! canine!0.6!dog! 0.3!…! …! Target!Prob.! the cats!0.8! cats! 0.1!…! … Avneesh Saluja, Hany Hassan, Kristina Toutanova, Chris Quirk |
ACL (1) | 1 |
| 2014 | Language Modeling with Power Low Rank EnsemblesabstractWe present power low rank ensembles (PLRE), a flexible framework for n-gram language modeling where ensembles of low rank matrices and tensors are used to obtain smoothed probability estimates of words in context.Our method can be understood as a generalization of ngram modeling to non-integer n, and includes standard techniques such as absolute discounting and Kneser-Ney smoothing as special cases.PLRE training is efficient and our approach outperforms stateof-the-art modified Kneser Ney baselines in terms of perplexity on large corpora as well as on BLEU score in a downstream machine translation task. Ankur P. Parikh, Avneesh Saluja, Chris Dyer, Eric P. Xing |
EMNLP | 2 |
| 2014 | Latent-Variable Synchronous CFGs for Hierarchical TranslationabstractData-driven refinement of non-terminal categories has been demonstrated to be a reliable technique for improving mono-lingual parsing with PCFGs. In this pa-per, we extend these techniques to learn latent refinements of single-category syn-chronous grammars, so as to improve translation performance. We compare two estimators for this latent-variable model: one based on EM and the other is a spec-tral algorithm based on the method of mo-ments. We evaluate their performance on a Chinese–English translation task. The re-sults indicate that we can achieve signifi-cant gains over the baseline with both ap-proaches, but in particular the moments-based estimator is both faster and performs better than EM. 1 Avneesh Saluja, Chris Dyer, Shay B. Cohen |
EMNLP | 1 |
| 2014 | Online discriminative learning for machine translation with binary-valued feedback
Avneesh Saluja |
Mach. Transl. | 1 |
| 2011 | Context-aware Language Modeling for Conversational Speech Translation
Avneesh Saluja, Ian Lane, Ying Zhang 0048 |
MTSummit | 1 |