Avneesh Saluja

dblp:139/1040 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Machine translation · 60% Language models and text generation · 40%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation › monolingual data augmentation
back-translation
0.412020
Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020
Natural language and speech › Machine translation
low-resource machine translation
0.412020
Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.412020
Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020
Natural language and speech › Language models and text generation › text generation › text simplification
sentence simplification
0.412020
Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.212014
Language Modeling with Power Low Rank Ensembles · EMNLP 2014
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation
0.212014
Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014
Natural language and speech › Language models and text generation › language modeling
statistical language modeling
0.212014
Language Modeling with Power Low Rank Ensembles · EMNLP 2014
Natural language and speech › Machine translation
statistical machine translation
0.212014
Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014
Natural language and speech › Machine translation › syntax-based machine translation
synchronous context-free grammar
0.212014
Latent-Variable Synchronous CFGs for Hierarchical Translation · EMNLP 2014
Natural language and speech › Machine translation
syntax-based machine translation
0.212014
Latent-Variable Synchronous CFGs for Hierarchical Translation · EMNLP 2014
Natural language and speech › Machine translation › statistical machine translation
translation rule extraction
0.212014
Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data · ACL (1) 2014
Natural language and speech › Language models and text generation › text evaluation
human evaluation
0.112020
Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020
Natural language and speech › Machine translation
machine translation evaluation
0.112020
Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation · AAAI 2020

Methods — techniques the papers use, named apart from their topics

sequence-to-sequence paraphrase model · 0.4back-translation · 0.4spectral method of moments · 0.2monolingual corpora · 0.2low rank matrix ensemble · 0.2kneser-ney smoothing · 0.2graph-based semi-supervised learning · 0.2graph propagation · 0.2expectation-maximization · 0.2
YearPublicationVenuePosition
2020 Simplify-Then-Translate: Automatic Preprocessing for Black-Box Translation
abstract
Black-box machine translation systems have proven incredibly useful for a variety of applications yet by design are hard to adapt, tune to a specific domain, or build on top of. In this work, we introduce a method to improve such systems via automatic pre-processing (APP) using sentence simplification. We first propose a method to automatically generate a large in-domain paraphrase corpus through back-translation with a black-box MT system, which is used to train a paraphrase model that “simplifies” the original sentence to be more conducive for translation. The model is used to preprocess source sentences of multiple low-resource language pairs. We show that this preprocessing leads to better translation performance as compared to non-preprocessed source sentences. We further perform side-by-side human evaluation to verify that translations of the simplified sentences are better than the original ones. Finally, we provide some guidance on recommended language pairs for generating the simplification model corpora by investigating the relationship between ease of translation of a language pair (as measured by BLEU) and quality of the resulting simplification model from back-translations of this language pair (as measured by SARI), and tie this into the downstream task of low-resource translation.
Sneha Mehta, Bahareh Azarnoush, Boris Chen, Avneesh Saluja, Vinith Misra, Ballav Bihani, Ritwik Kumar
AAAI4
2014 Graph-based Semi-Supervised Learning of Translation Models from Monolingual Data
abstract
Statistical phrase-based translation learns translation rules from bilingual corpora, and has traditionally only used monolingual evidence to construct features that rescore existing translation candidates.In this work, we present a semi-supervised graph-based approach for generating new translation rules that leverages bilingual and monolingual data.The proposed technique first constructs phrase graphs using both source and target language monolingual corpora.Next, graph propagation identifies translations of phrases that were not observed in the bilingual corpus, assuming that similar phrases have similar translations.We report results on a large Arabic-English system and a medium-sized Urdu-English system.Our proposed approach significantly improves the performance of competitive phrasebased systems, leading to consistent improvements between 1 and 4 BLEU points on standard evaluation sets.Source!Target! el gato! los gatos!un gato! cat! the cat! the cats! a cat! Target!Prob.! the cat! 0.7! cat! 0.15! …! …! felino!canino!el perro!Target!Prob.! canine!0.6!dog! 0.3!…! …! Target!Prob.! the cats!0.8! cats! 0.1!…! …
Avneesh Saluja, Hany Hassan, Kristina Toutanova, Chris Quirk
ACL (1)1
2014 Language Modeling with Power Low Rank Ensembles
abstract
We present power low rank ensembles (PLRE), a flexible framework for n-gram language modeling where ensembles of low rank matrices and tensors are used to obtain smoothed probability estimates of words in context.Our method can be understood as a generalization of ngram modeling to non-integer n, and includes standard techniques such as absolute discounting and Kneser-Ney smoothing as special cases.PLRE training is efficient and our approach outperforms stateof-the-art modified Kneser Ney baselines in terms of perplexity on large corpora as well as on BLEU score in a downstream machine translation task.
Ankur P. Parikh, Avneesh Saluja, Chris Dyer, Eric P. Xing
EMNLP2
2014 Latent-Variable Synchronous CFGs for Hierarchical Translation
abstract
Data-driven refinement of non-terminal categories has been demonstrated to be a reliable technique for improving mono-lingual parsing with PCFGs. In this pa-per, we extend these techniques to learn latent refinements of single-category syn-chronous grammars, so as to improve translation performance. We compare two estimators for this latent-variable model: one based on EM and the other is a spec-tral algorithm based on the method of mo-ments. We evaluate their performance on a Chinese–English translation task. The re-sults indicate that we can achieve signifi-cant gains over the baseline with both ap-proaches, but in particular the moments-based estimator is both faster and performs better than EM. 1
Avneesh Saluja, Chris Dyer, Shay B. Cohen
EMNLP1
2014 Online discriminative learning for machine translation with binary-valued feedback
Avneesh Saluja
Mach. Transl.1
2011 Context-aware Language Modeling for Conversational Speech Translation
Avneesh Saluja, Ian Lane, Ying Zhang 0048
MTSummit1