VLDB 2026 Research / reviewers in the wild / expert
Anoop Sarkar
dblp:s/AnoopSarkar
· DBLP profile ↗
57ranked-venue papers
9as first author
7since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 9 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
22 papers |
Machine translation · 55% Information extraction and text analysis · 25% Language models and text generation · 12% | |
| Theoretical computer science
4 papers |
Automata and formal languages · 100% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
1.3 | 3 | 2022 | CipherDAug: Ciphertext based Data Augmentation for Neural Machine Translation · ACL (1) 2022 Effectively pretraining a speech translation decoder with Machine Translation data · EMNLP (1) 2020 Top-down Tree Structured Decoding with Syntactic Connections for Neural Machine Translation and Parsing · EMNLP 2018 |
Natural language and speech › Machine translation
simultaneous machine translation |
0.8 | 2 | 2021 | Translation-based Supervision for Policy Generation in Simultaneous Neural Machine Translation · EMNLP (1) 2021 Prediction Improves Simultaneous Neural Machine Translation · EMNLP 2018 |
Natural language and speech › Machine translation
statistical machine translation |
0.7 | 5 | 2015 | Improving Statistical Machine Translation with a Multilingual Paraphrase Database · EMNLP 2015 Graph Propagation for Paraphrasing Out-of-Vocabulary Words in Statistical Machine Translation · ACL (1) 2013 Mixing Multiple Translation Models in Statistical Machine Translation · ACL (1) 2012 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.7 | 1 | 2023 | SpEL: Structured Prediction for Entity Linking · EMNLP 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.6 | 1 | 2022 | CipherDAug: Ciphertext based Data Augmentation for Neural Machine Translation · ACL (1) 2022 |
Automata and formal languages
sequence modeling |
0.6 | 1 | 2022 | Sequence Models for Document Structure Identification in an Undeciphered Script · EMNLP 2022 |
Natural language and speech › Machine translation › speech translation
end-to-end speech translation |
0.4 | 1 | 2020 | Effectively pretraining a speech translation decoder with Machine Translation data · EMNLP (1) 2020 |
Natural language and speech › Machine translation
speech translation |
0.4 | 1 | 2020 | Effectively pretraining a speech translation decoder with Machine Translation data · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.4 | 2 | 2018 | Top-down Tree Structured Decoding with Syntactic Connections for Neural Machine Translation and Parsing · EMNLP 2018 Using LTAG Based Features in Parse Reranking · EMNLP 2003 |
Natural language and speech › Machine translation › statistical machine translation
hierarchical phrase-based translation |
0.4 | 2 | 2014 | Two Improvements to Left-to-Right Decoding for Hierarchical Phrase-based Machine Translation · EMNLP 2014 Efficient Left-to-Right Hierarchical Phrase-Based Translation with Improved Reordering · EMNLP 2013 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.3 | 1 | 2018 | Decipherment of Substitution Ciphers with Neural Language Models · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
constituency parsing |
0.3 | 1 | 2018 | Top-down Tree Structured Decoding with Syntactic Connections for Neural Machine Translation and Parsing · EMNLP 2018 |
Natural language and speech › Machine translation
decipherment |
0.3 | 1 | 2018 | Decipherment of Substitution Ciphers with Neural Language Models · EMNLP 2018 |
Natural language and speech › Language models and text generation
neural language model |
0.3 | 1 | 2018 | Decipherment of Substitution Ciphers with Neural Language Models · EMNLP 2018 |
Automata and formal languages
grammar formalisms |
0.3 | 1 | 2018 | Prefix Lexicalization of Synchronous CFGs using Synchronous TAG · ACL (1) 2018 |
Automata and formal languages › formal grammars
synchronous context-free grammar |
0.3 | 1 | 2018 | Prefix Lexicalization of Synchronous CFGs using Synchronous TAG · ACL (1) 2018 |
Natural language and speech › Machine translation
paraphrase-based translation |
0.2 | 1 | 2015 | Improving Statistical Machine Translation with a Multilingual Paraphrase Database · EMNLP 2015 |
Natural language and speech › Information extraction and text analysis
paraphrase extraction |
0.2 | 1 | 2015 | Improving Statistical Machine Translation with a Multilingual Paraphrase Database · EMNLP 2015 |
Knowledge graphs
knowledge base linking |
0.2 | 1 | 2023 | SpEL: Structured Prediction for Entity Linking · EMNLP 2023 |
Natural language and speech › Machine translation
low-resource machine translation |
0.2 | 1 | 2022 | CipherDAug: Ciphertext based Data Augmentation for Neural Machine Translation · ACL (1) 2022 |
Natural language and speech › Machine translation
out-of-vocabulary word handling |
0.2 | 1 | 2013 | Graph Propagation for Paraphrasing Out-of-Vocabulary Words in Statistical Machine Translation · ACL (1) 2013 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.2 | 1 | 2013 | Graph Propagation for Paraphrasing Out-of-Vocabulary Words in Statistical Machine Translation · ACL (1) 2013 |
Natural language and speech › Information extraction and text analysis
bootstrapping |
0.1 | 1 | 2012 | Bootstrapping via Graph Propagation · ACL (1) 2012 |
Machine learning › Graph learning
graph propagation |
0.1 | 1 | 2012 | Bootstrapping via Graph Propagation · ACL (1) 2012 |
Natural language and speech › Machine translation
system combination |
0.1 | 1 | 2012 | Mixing Multiple Translation Models in Statistical Machine Translation · ACL (1) 2012 |
Machine learning › Efficient and distributed learning
active learning |
0.1 | 1 | 2009 | Active Learning for Multilingual Statistical Machine Translation · ACL/IJCNLP 2009 |
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.1 | 1 | 2007 | Experimental Evaluation of LTAG-Based Features for Semantic Role Labeling · EMNLP-CoNLL 2007 |
Natural language and speech › Language models and text generation › decoding
large language model decoding |
0.1 | 1 | 2014 | Two Improvements to Left-to-Right Decoding for Hierarchical Phrase-based Machine Translation · EMNLP 2014 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2005 | Intimate Learning: A Novel Approach for Combining Labelled and Unlabelled Data · IJCAI 2005 |
Natural language and speech › Machine translation
word reordering |
0.0 | 1 | 2013 | Efficient Left-to-Right Hierarchical Phrase-Based Translation with Improved Reordering · EMNLP 2013 |
Methods — techniques the papers use, named apart from their topics
unsupervised neural sequence modeling · 1.7statistical sequence modeling · 1.7structured prediction · 1.3fine-tuning · 1.3context-sensitive prediction aggregation · 1.3reinforcement learning · 0.8beam search · 0.7multi-source training · 0.6co-regularization · 0.6ROT-k ciphertext · 0.6synchronous tree-adjoining grammar transformation · 0.3LTAG features · 0.0lazy table generation · 0.0LR parsing · 0.0two-level model extension · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SpEL: Structured Prediction for Entity LinkingabstractEntity linking is a prominent thread of research focused on structured data creation by linking spans of text to an ontology or knowledge source.We revisit the use of structured prediction for entity linking which classifies each individual input token as an entity, and aggregates the token predictions.Our system, called SPEL (Structured prediction for Entity Linking) is a state-of-the-art entity linking system that uses some new ideas to apply structured prediction to the task of entity linking including: two refined fine-tuning steps; a context sensitive prediction aggregation strategy; reduction of the size of the model's output vocabulary, and; we address a common problem in entity-linking systems where there is a training vs. inference tokenization mismatch.Our experiments show that we can outperform the state-of-the-art on the commonly used AIDA benchmark dataset for entity linking to Wikipedia.Our method is also very compute efficient in terms of number of parameters and speed of inference.https://github.com/shavarani/SpEL Hassan Shavarani, Anoop Sarkar |
EMNLP | 2 |
| 2022 | CipherDAug: Ciphertext based Data Augmentation for Neural Machine TranslationabstractWe propose a novel data-augmentation technique for neural machine translation based on ROT-k ciphertexts.ROT-k is a simple letter substitution cipher that replaces a letter in the plaintext with the kth letter after it in the alphabet.We first generate multiple ROT-k ciphertexts using different values of k for the plaintext which is the source side of the parallel data.We then leverage this enciphered training data along with the original parallel data via multi-source training to improve neural machine translation.Our method, CipherDAug, uses a co-regularization-inspired training procedure, requires no external data sources other than the original training data, and uses a standard Transformer to outperform strong data augmentation techniques on several datasets by a significant margin.This technique combines easily with existing approaches to data augmentation, and yields particularly strong results in low-resource settings.1 Nishant Kambhatla, Logan Born, Anoop Sarkar |
ACL (1) | 3 |
| 2022 | Auxiliary Subword Segmentations as Related Languages for Low Resource Multilingual TranslationabstractWe propose a novel technique that combines alternative subword tokenizations of a single source-target language pair that allows us to leverage multilingual neural translation training methods. These alternate segmentations function like related languages in multilingual translation. Overall this improves translation accuracy for low-resource languages and produces translations that are lexically diverse and morphologically rich. We also introduce a cross-teaching technique which yields further improvements in translation accuracy and cross-lingual transfer between high- and low-resource language pairs. Compared to other strong multilingual baselines, our approach yields average gains of +1.7 BLEU across the four low-resource datasets from the multilingual TED-talks dataset. Our technique does not require additional training data and is a drop-in improvement for any existing neural translation system. Nishant Kambhatla, Logan Born, Anoop Sarkar |
EAMT | 3 |
| 2022 | Sequence Models for Document Structure Identification in an Undeciphered ScriptabstractThis work describes the first thorough analysis of "header" signs in proto-Elamite, an undeciphered script from 3100-2900 BCE.Headers are a category of signs which have been provisionally identified through painstaking manual analysis of this script by domain experts.We use unsupervised neural and statistical sequence modeling techniques to provide new and independent evidence for the existence of headers, without supervision from domain experts.Having affirmed the existence of headers as a legitimate structural feature, we next arrive at a richer understanding of their possible meaning and purpose by (i) examining which features predict their presence; (ii) identifying correlations between these features and other document properties; and (iii) examining cases where these features predict the presence of a header in texts where domain experts do not expect one (or vice versa).We provide more concrete processes for labeling headers in this corpus and a clearer justification for existing intuitions about document structure in proto-Elamite. Logan Born, M. Willis Monroe, Kathryn Kelley, Anoop Sarkar |
EMNLP | 4 |
| 2021 | Measuring and Improving Faithfulness of Attention in Neural Machine TranslationabstractWhile the attention heatmaps produced by neural machine translation (NMT) models seem insightful, there is little evidence that they reflect a model's true internal reasoning.We provide a measure of faithfulness for NMT based on a variety of stress tests where attention weights which are crucial for prediction are perturbed and the model should alter its predictions if the learned weights are a faithful explanation of the predictions.We show that our proposed faithfulness measure for NMT models can be improved using a novel differentiable objective that rewards faithful behaviour by the model through probability divergence.Our experimental results on multiple language pairs show that our objective function is effective in increasing faithfulness and can lead to a useful analysis of NMT model behaviour and more trustworthy attention heatmaps.Our proposed objective improves faithfulness without reducing the translation quality and has a useful regularization effect on the NMT model and can even improve translation quality in some cases. Pooya Moradi, Nishant Kambhatla, Anoop Sarkar |
EACL | 3 |
| 2021 | Better Neural Machine Translation by Extracting Linguistic Information from BERTabstractAdding linguistic information (syntax or semantics) to neural machine translation (NMT) has mostly focused on using point estimates from pre-trained models.Directly using the capacity of massive pre-trained contextual word embedding models such as BERT (Devlin et al., 2019) has been marginally useful in NMT because effective fine-tuning is difficult to obtain for NMT without making training brittle and unreliable.We augment NMT by extracting dense fine-tuned vector-based linguistic information from BERT instead of using point estimates.Experimental results show that our method of incorporating linguistic information helps NMT to generalize better in a variety of training contexts and is no more difficult to train than conventional Transformerbased NMT. Hassan Shavarani, Anoop Sarkar |
EACL | 2 |
| 2021 | Translation-based Supervision for Policy Generation in Simultaneous Neural Machine TranslationabstractIn simultaneous machine translation, finding an agent with the optimal action sequence of reads and writes that maintain a high level of translation quality while minimizing the average lag in producing target tokens remains an extremely challenging problem.We propose a novel supervised learning approach for training an agent that can detect the minimum number of reads required for generating each target token by comparing simultaneous translations against full-sentence translations during training to generate oracle action sequences.These oracle sequences can then be used to train a supervised model for action generation at inference time.Our approach provides an alternative to current heuristic methods in simultaneous translation by introducing a new training objective, which is easier to train than previous attempts at training the agent using reinforcement learning techniques for this task.Our experimental results show that our novel training method for action generation produces much higher quality translations while minimizing the average lag in simultaneous translation. Ashkan Alinejad, Hassan Shavarani, Anoop Sarkar |
EMNLP (1) | 3 |
| 2020 | Effectively pretraining a speech translation decoder with Machine Translation dataabstractDirectly translating from speech to text using an end-to-end approach is still challenging for many language pairs due to insufficient data.Although pretraining the encoder parameters using the Automatic Speech Recognition (ASR) task improves the results in low resource settings, attempting to use pretrained parameters from the Neural Machine Translation (NMT) task has been largely unsuccessful in previous works.In this paper, we will show that by using an adversarial regularizer, we can bring the encoder representations of the ASR and NMT tasks closer even though they are in different modalities, and how this helps us effectively use a pretrained NMT decoder for speech translation. Ashkan Alinejad, Anoop Sarkar |
EMNLP (1) | 2 |
| 2019 | Deconstructing Supertagging into Multi-Task Sequence PredictionabstractSupertagging is a sequence prediction task where each word is assigned a piece of complex syntactic structure called a supertag.We provide a novel approach to multi-task learning for Tree Adjoining Grammar (TAG) supertagging by deconstructing these complex supertags in order to define a set of related but auxiliary sequence prediction tasks.Our multi-task prediction framework is trained over the exactly same training data used to train the original supertagger where each auxiliary task provides an alternative view on the original prediction task.Our experimental results show that our multi-task approach significantly improves TAG supertagging with a new state-of-the-art accuracy score of 91.39% on the Penn treebank supertagging dataset. Zhenqi Zhu, Anoop Sarkar |
CoNLL | 2 |
| 2019 | An analysis of clausal coordination using synchronous tree adjoining grammarabstractThis paper presents a novel analysis of clausal coordination with shared arguments using synchronous tree adjoining grammar (STAG), a pairing of a tree adjoining grammar (TAG) for syntax and a TAG for semantics. In clausal coordination, one or more arguments can be shared by the verbal predicates of the conjuncts, as in Sue likes and Kim hates Pete, where an object argument Pete is shared by likes and hates. As the predicate-argument structure must be represented within each predicative elementary tree in TAG, modelling argument sharing across clauses poses an interesting challenge for TAG. A widely adopted approach within the TAG literature at present is to employ the conjoin operation (Sakar and Joshi, 1996, Proceedings of COLING ’96, 610–615), a non-standard tree-composing operation in TAG. This operation applies across elementary trees to identify and merge the arguments from each clause, yielding a derivation structure in which the shared arguments are combined with multiple elementary trees and a derived tree in which the shared arguments are dominated by multiple verbal projections. In contrast, our STAG analysis pairs a syntactic elementary tree that participates in the derivation of clausal coordination with a semantic elementary tree that includes a |$\lambda$|-term to abstract over the shared argument. This allows the sharing of arguments in coordination to be instantiated in semantics, without being represented in syntax in the form of multiple dominance, utilizing only the standard TAG operations, substitution and adjoining. Chung-hye Han, Sara Williamson, Logan Born, Anoop Sarkar |
J. Log. Comput. | 4 |
| 2018 | Prefix Lexicalization of Synchronous CFGs using Synchronous TAGabstractWe show that an ε-free, chain-free synchronous context-free grammar (SCFG) can be converted into a weakly equivalent synchronous tree-adjoining grammar (STAG) which is prefix lexicalized.This transformation at most doubles the grammar's rank and cubes its size, but we show that in practice the size increase is only quadratic.Our results extend Greibach normal form from CFGs to SCFGs and prove new formal properties about SCFG, a formalism with many applications in natural language processing. Logan Born, Anoop Sarkar |
ACL (1) | 2 |
| 2018 | Prediction Improves Simultaneous Neural Machine TranslationabstractSimultaneous speech translation aims to maintain translation quality while minimizing the delay between reading input and incrementally producing the output.We propose a new general-purpose prediction action which predicts future words in the input to improve quality and minimize delay in simultaneous translation.We train this agent using reinforcement learning with a novel reward function.Our agent with prediction has better translation quality and less delay compared to an agent-based simultaneous translation system without prediction. Ashkan Alinejad, Maryam Siahbani, Anoop Sarkar |
EMNLP | 3 |
| 2018 | Top-down Tree Structured Decoding with Syntactic Connections for Neural Machine Translation and ParsingabstractThe addition of syntax-aware decoding in Neural Machine Translation (NMT) systems requires an effective tree-structured neural network, a syntax-aware attention model and a language generation model that is sensitive to sentence structure.We exploit a top-down tree-structured model called DRNN (Doubly-Recurrent Neural Networks) first proposed by Alvarez-Melis and Jaakola (2017) to create an NMT model called Seq2DRNN that combines a sequential encoder with tree-structured decoding augmented with a syntax-aware attention model.Unlike previous approaches to syntax-based NMT which use dependency parsing models our method uses constituency parsing which we argue provides useful information for translation.In addition, we use the syntactic structure of the sentence to add new connections to the tree-structured decoder neural network (Seq2DRNN+SynC).We compare our NMT model with sequential and state of the art syntax-based NMT models and show that our model produces more fluent translations with better reordering.Since our model is capable of doing translation and constituency parsing at the same time we also compare our parsing accuracy against other neural parsing models. Jetic Gu, Hassan Shavarani, Anoop Sarkar |
EMNLP | 3 |
| 2018 | Decipherment of Substitution Ciphers with Neural Language ModelsabstractDecipherment of homophonic substitution ciphers using language models (LMs) is a wellstudied task in NLP.Previous work in this topic scores short local spans of possible plaintext decipherments using n-gram LMs.The most widely used technique is the use of beam search with n-gram LMs proposed by Nuhn et al. (2013).We propose a beam search algorithm that scores the entire candidate plaintext at each step of the decipherment using a neural LM.We augment beam search with a novel rest cost estimation that exploits the prediction power of a neural LM.We compare against the state of the art n-gram based methods on many different decipherment tasks.On challenging ciphers such as the Beale cipher we provide significantly better error rates with much smaller beam sizes. Nishant Kambhatla, Anahita Mansouri Bigvand, Anoop Sarkar |
EMNLP | 3 |
| 2017 | Evaluating the Value of Lensing Wikipedia During the Information Seeking ProcessabstractWhile Wikipedia is an excellent source of information about entities, discovering relationships among them is not well-supported by its search features. Lensing Wikipedia was designed as an alternate search and summarization interface, providing a set of filtering, visualization, and exploration tools that enable searching among the connections between people, organizations, and locations. In this paper, we present the results of a user study on how these features are used in each of Vakkari's stages of information seeking (pre-focus, focus formulation, and post-focus), and the participants perceptions of the utility of these features to their overall information-seeking goals. Participants primarily used input and control features during the pre-focus stage, informational and personalization features during the focus formulation stage, and personalization features during the post-focus stage. Findings from this study contribute to understanding how people use advanced search, summarization, and visualization tools to aid their information seeking tasks. Orland Hoeber, Anoop Sarkar, Andrei Vacariu, Max Whitney, Manali Gaikwad, Gursimran Kaur |
CHIIR | 2 |
| 2017 | Joint Prediction of Word Alignment with Alignment TypesabstractCurrent word alignment models do not distinguish between different types of alignment links. In this paper, we provide a new probabilistic model for word alignment where word alignments are associated with linguistically motivated alignment types. We propose a novel task of joint prediction of word alignment and alignment types and propose novel semi-supervised learning algorithms for this task. We also solve a sub-task of predicting the alignment type given an aligned word pair. In our experimental results, the generative models we introduce to model alignment types significantly outperform the models without alignment types. Anahita Mansouri Bigvand, Te Bu, Anoop Sarkar |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | What's Hot in Human Language Technology: Highlights from NAACL HLT 2015abstractThis paper shows a few examples to highlight the trends observed at the NAACL HLT 2015 conference. Joyce Y. Chai, Anoop Sarkar, Rada Mihalcea |
AAAI | 2 |
| 2016 | The Challenge of Simultaneous Speech Translation
Anoop Sarkar |
PACLIC | 1 |
| 2015 | Non-Uniform Stochastic Average Gradient Method for Training Conditional Random FieldsabstractWe apply stochastic average gradient (SAG) algorithms for training conditional random fields (CRFs). We describe a practical implementation that uses structure in the CRF gradient to reduce the memory requirement of this linearly-convergent stochastic gradient method, propose a non-uniform sampling scheme that substantially improves practical performance, and analyze the rate of convergence of the SAGA variant under non-uniform sampling. Our experimental results reveal that our method significantly outperforms existing methods in terms of the training objective, and performs as well or better than optimally-tuned stochastic gradient methods in terms of test error. Mark Schmidt 0001, Reza Babanezhad 0001, Mohamed Osama Ahmed, Aaron Defazio, Ann Clifton, Anoop Sarkar |
AISTATS | 6 |
| 2015 | Improving Statistical Machine Translation with a Multilingual Paraphrase DatabaseabstractThe multilingual Paraphrase Database (PPDB) is a freely available automatically created resource of paraphrases in mul-tiple languages. In statistical machine translation, paraphrases can be used to provide translation for out-of-vocabulary (OOV) phrases. In this paper, we show that a graph propagation approach that uses PPDB paraphrases can be used to im-prove overall translation quality. We pro-vide an extensive comparison with previ-ous work and show that our PPDB-based method improves the BLEU score by up to 1.79 percent points. We show that our approach improves on the state of the art in three different settings: when faced with limited amount of parallel training data; a domain shift between training and test data; and handling a morpho-logically complex source language. Our PPDB-based method outperforms the use of distributional profiles from monolin-gual source data. 1 Ramtin Mehdizadeh Seraj, Maryam Siahbani, Anoop Sarkar |
EMNLP | 3 |
| 2014 | Two Improvements to Left-to-Right Decoding for Hierarchical Phrase-based Machine TranslationabstractLeft-to-right (LR) decoding (Watanabe et al., 2006) is promising decoding algorithm for hierarchical phrase-based translation (Hiero) that visits input spans in arbitrary order producing the output translation in left to right order.This leads to far fewer language model calls, but while LR decoding is more efficient than CKY decoding, it is unable to capture some hierarchical phrase alignments reachable using CKY decoding and suffers from lower translation quality as a result.This paper introduces two improvements to LR decoding that make it comparable in translation quality to CKY-based Hiero. Maryam Siahbani, Anoop Sarkar |
EMNLP | 2 |
| 2014 | Incremental translation using hierarchichal phrase-based translation systemabstractHierarchical phrase-based machine translation [1] (Hiero) is a prominent approach for Statistical Machine Translation usually comparable to or better than conventional phrase-based systems. But Hiero typically uses the CKY decoding algorithm which requires the entire input sentence before decoding begins, as it produces the translation in a bottom-up fashion. Left-to-right (LR) decoding [2] is a promising decoding algorithm for Hiero that produces the output translation in left to right order. In this paper we focus on simultaneous translation using the Hiero translation framework. In simultaneous translation, translations are generated incrementally as source language speech input is processed. We propose a novel approach for incremental translation by integrating segmentation and decoding in LR-Hiero. We compare two incremental decoding algorithms for LR-Hiero and present translation quality scores (BLEU) and the latency of generating translations for both decoders on audio lectures from the TED collection. Maryam Siahbani, Ramtin Mehdizadeh Seraj, Baskaran Sankaran, Anoop Sarkar |
SLT | 4 |
| 2013 | Graph Propagation for Paraphrasing Out-of-Vocabulary Words in Statistical Machine Translation
Majid Razmara, Maryam Siahbani, Gholamreza Haffari, Anoop Sarkar |
ACL (1) | 4 |
| 2013 | Efficient Left-to-Right Hierarchical Phrase-Based Translation with Improved ReorderingabstractLeft-to-right (LR) decoding (Watanabe et al., 2006b) is a promising decoding algorithm for hierarchical phrase-based translation (Hiero).It generates the target sentence by extending the hypotheses only on the right edge.LR decoding has complexity O(n 2 b) for input of n words and beam size b, compared to O(n 3 ) for the CKY algorithm.It requires a single language model (LM) history for each target hypothesis rather than two LM histories per hypothesis as in CKY.In this paper we present an augmented LR decoding algorithm that builds on the original algorithm in (Watanabe et al., 2006b).Unlike that algorithm, using experiments over multiple language pairs we show two new results: our LR decoding algorithm provides demonstrably more efficient decoding than CKY Hiero, four times faster; and by introducing new distortion and reordering features for LR decoding, it maintains the same translation quality (as in BLEU scores) obtained phrase-based and CKY Hiero with the same translation model. Maryam Siahbani, Baskaran Sankaran, Anoop Sarkar |
EMNLP | 3 |
| 2013 | An Online Algorithm for Learning over Constrained Latent Representations using Multiple Views
Ann Clifton, Max Whitney, Anoop Sarkar |
IJCNLP | 3 |
| 2013 | Ensemble Triangulation for Statistical Machine Translation
Majid Razmara, Anoop Sarkar |
IJCNLP | 2 |
| 2013 | Scalable Variational Inference for Extracting Hierarchical Phrase-based Translation Rules
Baskaran Sankaran, Gholamreza Haffari, Anoop Sarkar |
IJCNLP | 3 |
| 2013 | Multi-Metric Optimization Using Ensemble Tuning
Baskaran Sankaran, Anoop Sarkar, Kevin Duh |
HLT-NAACL | 2 |
| 2012 | Mixing Multiple Translation Models in Statistical Machine Translation
Majid Razmara, George F. Foster, Baskaran Sankaran, Anoop Sarkar |
ACL (1) | 4 |
| 2012 | Bootstrapping via Graph Propagation
Max Whitney, Anoop Sarkar |
ACL (1) | 2 |
| 2012 | Improved Reordering for Shallow-n Grammar based Hierarchical Phrase-based Translation
Baskaran Sankaran, Anoop Sarkar |
HLT-NAACL | 2 |
| 2011 | Combining Morpheme-based Machine Translation with Post-processing Morpheme Prediction
Ann Clifton, Anoop Sarkar |
ACL | 2 |
| 2011 | Parsing Schemata for Practical Text Analysis Carlos Gómez Rodríguez (University of A Coruña) London: Imperial College Press (Mathematics, computing, language, and life series, edited by Carlos Martin-Vide, volume 1), 2010, xiv+275 pp; hardbound, ISBN 978-1-84816-560-1, $89.00abstractSikkel's definition of parsing schemas (Sikkel 1997) extends deductive systems by formally defining the semantics of items and related concepts used in deductive systems.In particular, items are sets of partial constituency trees that are licensed by Anoop Sarkar |
Comput. Linguistics | 1 |
| 2009 | Active Learning for Multilingual Statistical Machine Translation
Gholamreza Haffari, Anoop Sarkar |
ACL/IJCNLP | 2 |
| 2009 | Active Learning for Statistical Phrase-based Machine Translation
Gholamreza Haffari, Maxim Roy, Anoop Sarkar |
HLT-NAACL | 3 |
| 2008 | Homotopy-Based Semi-Supervised Hidden Markov Models for Sequence Labeling
Gholamreza Haffari, Anoop Sarkar |
COLING | 2 |
| 2008 | Training a Perceptron with Global and Local Features for Chinese Word Segmentation
Dong Song, Anoop Sarkar |
IJCNLP | 2 |
| 2007 | Transductive learning for statistical machine translation
Nicola Ueffing, Gholamreza Haffari, Anoop Sarkar |
ACL | 3 |
| 2007 | Experimental Evaluation of LTAG-Based Features for Semantic Role Labeling
Anoop Sarkar |
EMNLP-CoNLL | 2 |
| 2007 | Analysis of Semi-Supervised Learning with the Yarowsky Algorithm
Gholamreza Haffari, Anoop Sarkar |
UAI | 2 |
| 2007 | Semi-supervised model adaptation for statistical machine translation
Nicola Ueffing, Gholamreza Haffari, Anoop Sarkar |
Mach. Transl. | 3 |
| 2006 | A Clustering Approach for Nearly Unsupervised Recognition of Nonliteral Language
Julia Birke, Anoop Sarkar |
EACL | 2 |
| 2006 | Tutorial on Inductive Semi-supervised Learning Methods: with Applicability to Natural Language Processing
Anoop Sarkar, Gholamreza Haffari |
HLT-NAACL | 1 |
| 2005 | Intimate Learning: A Novel Approach for Combining Labelled and Unlabelled Data
Zhongmin Shi, Anoop Sarkar |
IJCAI | 2 |
| 2004 | A Smorgasbord of Features for Statistical Machine Translation
Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alexander Fraser 0001, Shankar Kumar, Libin Shen, Katherine Eng, Viren Jain, Zhen Jin 0007, Dragomir R. Radev |
HLT-NAACL | 4 |
| 2004 | Discriminative Reranking for Machine Translation
Libin Shen, Anoop Sarkar, Franz Josef Och |
HLT-NAACL | 2 |
| 2003 | Bootstrapping statistical parsers from small datasets
Mark Steedman, Anoop Sarkar, Miles Osborne, Rebecca Hwa, Stephen Clark, Julia Hockenmaier, Paul Ruhlen, Jeremiah Crim |
EACL | 2 |
| 2003 | Using LTAG Based Features in Parse Reranking
Libin Shen, Anoop Sarkar, Aravind K. Joshi |
EMNLP | 2 |
| 2003 | Example Selection for Bootstrapping Statistical Parsers
Mark Steedman, Rebecca Hwa, Stephen Clark, Miles Osborne, Anoop Sarkar, Julia Hockenmaier, Paul Ruhlen, Jeremiah Crim |
HLT-NAACL | 5 |
| 2002 | Learning Verb Argument Structure from Minimally Annotated Corpora
Anoop Sarkar, Woottiporn Tripasai |
COLING | 1 |
| 2002 | A Note on Typing Feature StructuresabstractFeature structures are used to convey linguistic information in a variety of linguistic formalisms. Various definitions of feature structures exist; one dimension of variation is typing: unlike untyped feature structures, typed ones associate a type with every structure and impose appropriateness constraints on the occurrences of features and on the values that they take. This work demonstrates the benefits that typing can carry even for linguistic formalisms that use untyped feature structures. We present a method for validating the consistency of (untyped) feature structure specifications by imposing a type discipline. This method facilitates a great number of compile-time checks: many possible errors can be detected before the grammar is used for parsing. We have constructed a type signature for an existing broad-coverage grammar of English and implemented a type inference algorithm that operates on the feature structure specifications in the grammar and reports incompatibilities with the signature. We have detected a large number of errors in the grammar, some of which are described in the article. Shuly Wintner, Anoop Sarkar |
Comput. Linguistics | 2 |
| 2001 | Applying Co-Training Methods to Statistical Parsing
Anoop Sarkar |
NAACL | 1 |
| 2000 | Automatic Extraction of Subcategorization Frames for Czech
Anoop Sarkar, Daniel Zeman |
COLING | 1 |
| 2000 | Learning Verb Subcategorization from Corpora: Counting Frame Subsets
Daniel Zeman, Anoop Sarkar |
LREC | 2 |
| 1996 | Incremental Parser Generation for Tree Adjoining GrammarsabstractThis paper describes the incremental generation of parse tables for the LR-type parsing of Tree Adjoining Languages (TALs). The algorithm presented handles modifications to the input grammar by updating the parser generated so far. In this paper, a lazy generation of LR-type parsers for TALs is defined in which parse tables are created by need while parsing. We then describe an incremental parser generator for TALs which responds to modification of the input grammar by updating parse tables built so far. Anoop Sarkar |
ACL | 1 |
| 1996 | Coordination in Tree Adjoining Grammars: Formalization and Implementation
Anoop Sarkar, Aravind K. Joshi |
COLING | 1 |
| 1993 | Extending Kimmo's Two-Level Model of MorphologyabstractThis paper describes the problems faced while using Kimmo's two-level model to describe certain Indian languages such as Tamil and Hindi. The two-level model is shown to be descriptively inadequate to address these problems. A simple extension to the basic two-level model is introduced which allows conflicting phonological rules to coexist. The computational complexity of the extension is the same as Kimmo's two-level model. Anoop Sarkar |
ACL | 1 |