EDBT 2026 Demo / reviewers in the wild / expert
Khalil Sima'an
dblp:82/1262
· DBLP profile ↗
46ranked-venue papers
5as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Machine translation · 64% Language models and text generation · 17% Information extraction and text analysis · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Theoretical computer science
2 papers |
Automata and formal languages · 75% Information theory · 25% |
Topics — the 28 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book? · ICLR 2025 |
Natural language and speech › Machine translation
low-resource machine translation |
0.9 | 1 | 2025 | Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book? · ICLR 2025 |
Natural language and speech › Machine translation
statistical machine translation |
0.7 | 6 | 2015 | Reordering Grammar Induction · EMNLP 2015 Latent Domain Phrase-based Models for Adaptation · EMNLP 2014 A Syntactified Direct Translation Model with Linear-time Decoding · EMNLP 2009 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.6 | 4 | 2015 | Reordering Grammar Induction · EMNLP 2015 Latent Domain Phrase-based Models for Adaptation · EMNLP 2014 Syntactically Lexicalized Phrase-Based SMT · IEEE Trans. Speech Audio Process. 2008 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.3 | 1 | 2017 | Graph Convolutional Encoders for Syntax-aware Neural Machine Translation · EMNLP 2017 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2017 | Graph Convolutional Encoders for Syntax-aware Neural Machine Translation · EMNLP 2017 |
Natural language and speech › Machine translation › word reordering
preordering |
0.2 | 1 | 2015 | Reordering Grammar Induction · EMNLP 2015 |
Automata and formal languages
grammatical inference |
0.2 | 1 | 2015 | Reordering Grammar Induction · EMNLP 2015 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.2 | 2 | 2011 | Learning Hierarchical Translation Structure with Linguistic Annotations · ACL 2011 Syntactically Lexicalized Phrase-Based SMT · IEEE Trans. Speech Audio Process. 2008 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2014 | Latent Domain Phrase-based Models for Adaptation · EMNLP 2014 |
Natural language and speech › Machine translation
machine translation evaluation |
0.2 | 1 | 2014 | Fitting Sentence Level Translation Evaluation with Many Dense Features · EMNLP 2014 |
Natural language and speech › Machine translation › machine translation evaluation
sentence-level MT evaluation |
0.2 | 1 | 2014 | Fitting Sentence Level Translation Evaluation with Many Dense Features · EMNLP 2014 |
Information retrieval › ranking
learning to rank |
0.2 | 1 | 2014 | Fitting Sentence Level Translation Evaluation with Many Dense Features · EMNLP 2014 |
Information retrieval
ranking |
0.2 | 1 | 2014 | Fitting Sentence Level Translation Evaluation with Many Dense Features · EMNLP 2014 |
Natural language and speech › Information extraction and text analysis › data annotation
linguistic annotation |
0.1 | 1 | 2011 | Learning Hierarchical Translation Structure with Linguistic Annotations · ACL 2011 |
Natural language and speech › Machine translation
syntax-aware translation |
0.1 | 1 | 2011 | Learning Hierarchical Translation Structure with Linguistic Annotations · ACL 2011 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 2 | 2009 | An Alternative to Head-Driven Approaches for Parsing a (Relatively) Free Word-Order Language · EMNLP 2009 Tree-gram Parsing: Lexical Dependencies and Structural Relations · ACL 2000 |
Natural language and speech › Language models and text generation
decoding |
0.1 | 1 | 2009 | A Syntactified Direct Translation Model with Linear-time Decoding · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › constituency parsing › lexicalized parsing
head-driven parsing |
0.1 | 1 | 2009 | An Alternative to Head-Driven Approaches for Parsing a (Relatively) Free Word-Order Language · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging |
0.1 | 2 | 2007 | Unsupervised estimation for noisy-channel models · ICML 2007 Supertagged Phrase-Based Statistical Machine Translation · ACL 2007 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.1 | 1 | 2017 | Graph Convolutional Encoders for Syntax-aware Neural Machine Translation · EMNLP 2017 |
Natural language and speech › Machine translation › synchronous grammar
inversion transduction grammar |
0.1 | 1 | 2008 | Phrase Translation Probabilities with ITG Priors and Smoothing as Learning Objective · EMNLP 2008 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.1 | 1 | 2008 | Phrase Translation Probabilities with ITG Priors and Smoothing as Learning Objective · EMNLP 2008 |
Information theory › communication channels › channel models › noisy channel
noisy channel model |
0.1 | 1 | 2007 | Unsupervised estimation for noisy-channel models · ICML 2007 |
Programming languages and type systems
grammar formalisms |
0.0 | 1 | 2009 | A Syntactified Direct Translation Model with Linear-time Decoding · EMNLP 2009 |
Programming languages and type systems › grammar formalisms
synchronous grammars |
0.0 | 1 | 2009 | A Syntactified Direct Translation Model with Linear-time Decoding · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
statistical parsing |
0.0 | 1 | 2000 | Tree-gram Parsing: Lexical Dependencies and Structural Relations · ACL 2000 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › syntactic disambiguation
supertagging |
0.0 | 1 | 2007 | Supertagged Phrase-Based Statistical Machine Translation · ACL 2007 |
Methods — techniques the papers use, named apart from their topics
typological feature prompt · 0.9long-context prompting · 0.9fine-tuning · 0.9word alignment · 0.4permutation trees · 0.4latent variable modeling · 0.4expectation-maximization · 0.3graph convolutional network · 0.3attention · 0.3linear model · 0.2learning to rank · 0.2latent domain variable · 0.2feature engineering · 0.2syntactified direct translation · 0.1linear-time decoding · 0.1maximum likelihood estimation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?abstractExtremely low-resource (XLR) languages lack substantial corpora for training NLP models, motivating the use of all available resources such as dictionaries and grammar books. Machine Translation from One Book (Tanzer et al., 2024) suggests that prompting long-context LLMs with one grammar book enables English–Kalamang translation, an XLR language unseen by LLMs—a noteworthy case of linguistics helping an NLP task. We investigate the source of this translation ability, finding almost all improvements stem from the book’s parallel examples rather than its grammatical explanations. We find similar results for Nepali and Guarani, seen low-resource languages, and we achieve performance comparable to an LLM with a grammar book by simply fine-tuning an encoder-decoder translation model. We then investigate where grammar books help by testing two linguistic tasks, grammaticality judgment and gloss prediction, and we explore what kind of grammatical knowledge helps by introducing a typological feature prompt that achieves leading results on these more relevant tasks. We thus emphasise the importance of task-appropriate data for XLR languages: parallel examples for translation, and grammatical data for linguistic tasks. As we find no evidence that long-context LLMs can make effective use of grammatical explanations for XLR translation, we conclude data collection for multilingual XLR tasks such as translation is best focused on parallel data over linguistic description. Seth Aycock, David Stap, Christof Monz, Khalil Sima'an |
ICLR | 5 |
| 2024 | Continual Reinforcement Learning for Controlled Text GenerationabstractControlled Text Generation (CTG) steers the generation of continuations of a given context (prompt) by a Large Language Model (LLM) towards texts possessing a given attribute (e.g., topic, sentiment). In this paper we view CTG as a Continual Learning problem: how to learn at every step to steer next-word generation, without having to wait for end-of-sentence. This continual view is useful for online applications such as CTG for speech, where end-of-sentence is often uncertain. We depart from an existing model, the Plug-and-Play language models (PPLM), which perturbs the context at each step to better predict next-words that posses the desired attribute. While PPLM is intricate and has many hyper-parameters, we provide a proof that the PPLM objective function can be reduced to a Continual Reinforcement Learning (CRL) reward function, thereby simplifying PPLM and endowing it with a better understood learning framework. Subsequently, we present, the first of its kind, CTG algorithm that is fully based on CRL and exhibit promising empirical results. Velizar Shulev, Khalil Sima'an |
LREC/COLING | 2 |
| 2022 | Passing Parser Uncertainty to the Transformer. Labeled Dependency Distributions for Neural Machine TranslationabstractExisting syntax-enriched neural machine translation (NMT) models work either with the single most-likely unlabeled parse or the set of n-best unlabeled parses coming out of an external parser. Passing a single or n-best parses to the NMT model risks propagating parse errors. Furthermore, unlabeled parses represent only syntactic groupings without their linguistically relevant categories. In this paper we explore the question: Does passing both parser uncertainty and labeled syntactic knowledge to the Transformer improve its translation performance? This paper contributes a novel method for infusing the whole labeled dependency distributions (LDD) of the source sentence’s dependency forest into the self-attention mechanism of the encoder of the Transformer. A range of experimental results on three language pairs demonstrate that the proposed approach outperforms both the vanilla Transformer as well as the single best-parse Transformer model across several evaluation metrics. Dongqi Liu 0001, Khalil Sima'an |
EAMT | 2 |
| 2018 | Deep Generative Model for Joint Alignment and Word RepresentationabstractMiguel Rios, Wilker Aziz, Khalil Sima’an. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Miguel Ángel Ríos-Gaona, Wilker Aziz, Khalil Sima'an |
NAACL-HLT | 3 |
| 2017 | Graph Convolutional Encoders for Syntax-aware Neural Machine TranslationabstractWe present a simple and effective approach to incorporating syntactic structure into neural attention-based encoderdecoder models for machine translation.We rely on graph-convolutional networks (GCNs), a recent class of neural networks developed for modeling graph-structured data.Our GCNs use predicted syntactic dependency trees of source sentences to produce representations of words (i.e.hidden states of the encoder) that are sensitive to their syntactic neighborhoods.GCNs take word representations as input and produce word representations as output, so they can easily be incorporated as layers into standard encoders (e.g., on top of bidirectional RNNs or convolutional neural networks).We evaluate their effectiveness with English-German and English-Czech translation experiments for different types of encoders and observe substantial improvements over their syntax-agnostic versions in all the considered setups. Jasmijn Bastings, Ivan Titov 0001, Wilker Aziz, Diego Marcheggiani, Khalil Sima'an |
EMNLP | 5 |
| 2017 | Elastic-substitution decoding for Hierarchical SMT: efficiency, richer search and double labels
Gideon Maillette de Buy Wenniger, Khalil Sima'an, Andy Way |
MTSummit (1) | 2 |
| 2017 | A survey of domain adaptation for statistical machine translation
Cuong Hoang, Khalil Sima'an |
Mach. Transl. | 2 |
| 2017 | Induction of latent domains in heterogeneous corpora: a case study of word alignment
Cuong Hoang, Khalil Sima'an |
Mach. Transl. | 2 |
| 2016 | Universal Reordering via Linguistic TypologyabstractIn this paper we explore the novel idea of building a single universal reordering model from English to a large number of target languages. To build this model we exploit typological features of word order for a large number of target languages together with source (English) syntactic features and we train this model on a single combined parallel corpus representing all (22) involved language pairs. We contribute experimental evidence for the usefulness of linguistically defined typological features for building such a model. When the universal reordering model is used for preordering followed by monotone translation (no reordering inside the decoder), our experiments show that this pipeline gives comparable or improved translation performance with a phrase-based baseline for a large number of language pairs (12 out of 22) from diverse language families. Joachim Daiber, Milos Stanojevic, Khalil Sima'an |
COLING | 3 |
| 2016 | Hierarchical Permutation Complexity for Word Order EvaluationabstractExisting approaches for evaluating word order in machine translation work with metrics computed directly over a permutation of word positions in system output relative to a reference translation. However, every permutation factorizes into a permutation tree (PET) built of primal permutations, i.e., atomic units that do not factorize any further. In this paper we explore the idea that permutations factorizing into (on average) shorter primal permutations should represent simpler ordering as well. Consequently, we contribute Permutation Complexity, a class of metrics over PETs and their extension to forests, and define tight metrics, a sub-class of metrics implementing this idea. Subsequently we define example tight metrics and empirically test them in word order evaluation. Experiments on the WMT13 data sets for ten language pairs show that a tight metric is more often than not better than the baselines. Milos Stanojevic, Khalil Sima'an |
COLING | 2 |
| 2016 | Adapting to All Domains at Once: Rewarding Domain Invariance in SMTabstractExisting work on domain adaptation for statistical machine translation has consistently assumed access to a small sample from the test distribution (target domain) at training time. In practice, however, the target domain may not be known at training time or it may change to match user needs. In such situations, it is natural to push the system to make safer choices, giving higher preference to domain-invariant translations, which work well across domains, over risky domain-specific alternatives. We encode this intuition by (1) inducing latent subdomains from the training data only; (2) introducing features which measure how specialized phrases are to individual induced sub-domains; (3) estimating feature weights on out-of-domain data (rather than on the target domain). We conduct experiments on three language pairs and a number of different domains. We observe consistent improvements over a baseline which does not explicitly reward domain invariance. Cuong Hoang, Khalil Sima'an, Ivan Titov 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2015 | Reordering Grammar InductionabstractWe present a novel approach for unsupervised induction of a Reordering Grammar using a modified form of permutation trees (Zhang and Gildea, 2007), which we apply to preordering in phrase-based machine translation.Unlike previous approaches, we induce in one step both the hierarchical structure and the transduction function over it from word-aligned parallel corpora.Furthermore, our model (1) handles non-ITG reordering patterns (up to 5-ary branching), ( 2) is learned from all derivations by treating not only labeling but also bracketing as latent variable, (3) is entirely unlexicalized at the level of reordering rules, and (4) requires no linguistic annotation.Our model is evaluated both for accuracy in predicting target order, and for its impact on translation quality.We report significant performance gains over phrase reordering, and over two known preordering baselines for English-Japanese. Milos Stanojevic, Khalil Sima'an |
EMNLP | 2 |
| 2015 | Machine translation with source-predicted target morphology
Joachim Daiber, Khalil Sima'an |
MTSummit | 2 |
| 2015 | Latent Domain Word Alignment for Heterogeneous CorporaabstractThis work focuses on the insensitivity of existing word alignment models to domain differences, which often yields suboptimal results on large heterogeneous data.A novel latent domain word alignment model is proposed, which induces domain-conditioned lexical and alignment statistics.We propose to train the model on a heterogeneous corpus under partial supervision, using a small number of seed samples from different domains.The seed samples allow estimating sharper, domain-conditioned word alignment statistics for sentence pairs.Our experiments show that the derived domain-conditioned statistics, once combined together, produce notable improvements both in word alignment accuracy and in translation accuracy of their resulting SMT systems. Cuong Hoang, Khalil Sima'an |
HLT-NAACL | 2 |
| 2015 | Labeling hierarchical phrase-based models without linguistic resourcesabstractLong-range word order differences are a well-known problem for machine translation. Unlike the standard phrase-based models which work with sequential and local phrase reordering, the hierarchical phrase-based model (Hiero) embeds the reordering of phrases within pairs of lexicalized context-free rules. This allows the model to handle long range reordering recursively. However, the Hiero grammar works with a single nonterminal label, which means that the rules are combined together into derivations independently and without reference to context outside the rules themselves. Follow-up work explored remedies involving nonterminal labels obtained from monolingual parsers and taggers. As of yet, no labeling mechanisms exist for the many languages for which there are no good quality parsers or taggers. In this paper we contribute a novel approach for acquiring reordering labels for Hiero grammars directly from the word-aligned parallel training corpus, without use of any taggers or parsers. The new labels represent types of alignment patterns in which a phrase pair is embedded within larger phrase pairs. In order to obtain alignment patterns that generalize well, we propose to decompose word alignments into trees over phrase pairs. Beside this labeling approach, we contribute coarse and sparse features for learning soft, weighted label-substitution as opposed to standard substitution. We report extensive experiments comparing our model to two baselines: Hiero and the known syntax augmented machine translation (SAMT) variant, which labels Hiero rules with nonterminals extracted from monolingual syntactic parses. We also test a simplified labeling scheme based on inversion transduction grammar (ITG). For the Chinese–English task we obtain performance improvement up to 1 BLEU point, whereas for the German–English task, where morphology is an issue, a minor (but statistically significant) improvement of 0.2 BLEU points is reported over SAMT. While ITG labeling does give a performance improvement, it remains sometimes suboptimal relative to our proposed labeling scheme. Gideon Maillette de Buy Wenniger, Khalil Sima'an |
Mach. Transl. | 2 |
| 2014 | Latent Domain Translation Models in Mix-of-Domains Haystack
Cuong Hoang, Khalil Sima'an |
COLING | 2 |
| 2014 | Latent Domain Phrase-based Models for AdaptationabstractPhrase-based models directly trained on mix-of-domain corpora can be sub-optimal.In this paper we equip phrase-based models with a latent domain variable and present a novel method for adapting them to an in-domain task represented by a seed corpus.We derive an EM algorithm which alternates between inducing domain-focused phrase pair estimates, and weights for mix-domain sentence pairs reflecting their relevance for the in-domain task.By embedding our latent domain phrase model in a sentence-level model and training the two in tandem, we are able to adapt all core translation components together -phrase, lexical and reordering.We show experiments on weighing sentence pairs for relevance as well as adapting phrase-based models, showing significant performance improvement in both tasks. Cuong Hoang, Khalil Sima'an |
EMNLP | 2 |
| 2014 | Fitting Sentence Level Translation Evaluation with Many Dense FeaturesabstractSentence level evaluation in MT has turned out far more difficult than corpus level evaluation. Existing sentence level metrics employ a lim-ited set of features, most of which are rather sparse at the sentence level, and their intricate models are rarely trained for ranking. This pa-per presents a simple linear model exploiting 33 relatively dense features, some of which are novel while others are known but seldom used, and train it under the learning-to-rank frame-work. We evaluate our metric on the stan-dard WMT12 data showing that it outperforms the strong baseline METEOR. We also ana-lyze the contribution of individual features and the choice of training data, language-pair vs. target-language data, providing new insights into this task. 1 Milos Stanojevic, Khalil Sima'an |
EMNLP | 2 |
| 2014 | All Fragments Count in Parser Evaluation
Jasmijn Bastings, Khalil Sima'an |
LREC | 2 |
| 2014 | Learning structural dependencies of words in the Zipfian TailabstractThis article uses semi-supervised Expectation Maximization (EM) to learn lexico-syntactic dependencies, i.e. associations between words and the structures that occur with them. Due to Zipfian distributions in language, such dependencies are extremely sparse in labelled data, and unlabelled data are the only source for learning them. Specifically, we learn sparse lexical parameters of a generative parsing model (a Probabilistic Context-Free Grammar, PCFG) that is initially estimated over the Penn Treebank. Our lexical parameters are similar to supertags—they are fine-grained, and encode complex structural information at the pre-terminal level. Our goal is to use unlabelled data to learn these for words that are rare or unseen in the labelled data. We get large error reductions (up to 17.5%) in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies, resulting in a statistically significant improvement in labelled bracketing score of the treebank PCFG. Our semi-supervised method incorporates structural and lexical priors from the labelled data to guide estimation from unlabelled data, and is the first successful use of semi-supervised EM to improve a generative structured model already trained over large labelled data. The method scales well to larger amounts of unlabelled data, and also gives substantial error reductions (up to 11.5%) for models trained on smaller amounts of labelled data, making it relevant to low-resource languages with small treebanks as well. Tejaswini Deoskar, Markos Mylonakis, Khalil Sima'an |
J. Log. Comput. | 3 |
| 2012 | Adjunct Alignment in Translation Data with an Application to Phrase Based Statistical Machine Translation
Sophie Arnoult, Khalil Sima'an |
EAMT | 2 |
| 2012 | Efficient accurate syntactic direct translation models: one tree at a time
Hany Hassan, Khalil Sima'an, Andy Way |
Mach. Transl. | 2 |
| 2012 | Statistical Translation After Source Reordering: Oracles, Context-Aware Models, and Empirical AnalysisabstractAbstract In source reordering the order of the source words is permuted to minimize word order differences with the target sentence and then fed to a translation model. Earlier work highlights the benefits of resolving long-distance reorderings as a pre-processing step to standard phrase-based models. However, the potential performance improvement of source reordering and its impact on the components of the subsequent translation model remain unexplored. In this paper we study both aspects of source reordering. We set up idealized source reordering (oracle) models with/without syntax and present our own syntax-driven model of source reordering. The latter is a statistical model of inversion transduction grammar (ITG)-like tree transductions manipulating a syntactic parse and working with novel conditional reordering parameters. Having set up the models, we report translation experiments showing significant improvement on three language pairs, and contribute an extensive analysis of the impact of source reordering (both oracle and model) on the translation model regarding the quality of its input, phrase-table, and output. Our experiments show that oracle source reordering has untapped potential in improving translation system output. Besides solving difficult reorderings, we find that source reordering creates more monotone parallel training data at the back-end, leading to significantly larger phrase tables with higher coverage of phrase types in unseen data. Unfortunately, this nice property does not carry over to tree-constrained source reordering. Our analysis shows that, from the string-level perspective, tree-constrained reordering might selectively permute word order, leading to larger phrase tables but without increase in phrase coverage in unseen data. Maxim Khalilov, Khalil Sima'an |
Nat. Lang. Eng. | 2 |
| 2011 | Learning Hierarchical Translation Structure with Linguistic Annotations
Markos Mylonakis, Khalil Sima'an |
ACL | 2 |
| 2011 | Context-Sensitive Syntactic Source-Reordering by Statistical Transduction
Maxim Khalilov, Khalil Sima'an |
IJCNLP | 2 |
| 2010 | Learning Probabilistic Synchronous CFGs for Phrase-Based Translation
Markos Mylonakis, Khalil Sima'an |
CoNLL | 2 |
| 2010 | Source reordering using MaxEnt classifiers and supertags
Maxim Khalilov, Khalil Sima'an |
EAMT | 2 |
| 2009 | A Syntactified Direct Translation Model with Linear-time Decoding
Hany Hassan, Khalil Sima'an, Andy Way |
EMNLP | 2 |
| 2009 | An Alternative to Head-Driven Approaches for Parsing a (Relatively) Free Word-Order Language
Reut Tsarfaty, Khalil Sima'an, Remko Scha |
EMNLP | 2 |
| 2008 | Relational-Realizational Parsing
Reut Tsarfaty, Khalil Sima'an |
COLING | 2 |
| 2008 | Phrase Translation Probabilities with ITG Priors and Smoothing as Learning Objective
Markos Mylonakis, Khalil Sima'an |
EMNLP | 2 |
| 2008 | Parsing with subdomain instance weighting from raw corpora
Barbara Plank, Khalil Sima'an |
INTERSPEECH | 2 |
| 2008 | Subdomain Sensitive Statistical Parsing using Raw Corpora
Barbara Plank, Khalil Sima'an |
LREC | 2 |
| 2008 | A syntactic language model based on incremental CCG parsingabstractSyntactically-enriched language models (parsers) constitute a promising component in applications such as machine translation and speech-recognition. To maintain a useful level of accuracy, existing parsers are non-incremental and must span a combinatorially growing space of possible structures as every input word is processed. This prohibits their incorporation into standard linear-time decoders. In this paper, we present an incremental, linear-time dependency parser based on Combinatory Categorial Grammar (CCG) and classification techniques. We devise a deterministic transform of CCG-bank canonical derivations into incremental ones, and train our parser on this data. We discover that a cascaded, incremental version provides an appealing balance between efficiency and accuracy. Hany Hassan, Khalil Sima'an, Andy Way |
SLT | 2 |
| 2008 | Better statistical estimation can benefit all phrases in phrase-based statistical machine translationabstractThe heuristic estimates of conditional phrase translation probabilities are based on frequency counts in a word-aligned parallel corpus. Earlier attempts at more principled estimation using Expectation-Maximization (EM) under perform this heuristic. This paper shows that a recently introduced novel estimator based on smoothing might provide a good alternative. Whenallphrasepairsare estimated (no length cut-off), this estimator slightly outperforms the heuristic estimator. Khalil Sima'an, Markos Mylonakis |
SLT | 1 |
| 2008 | Part-of-speech tagging of Modern Hebrew textabstractAbstract Words in Semitic texts often consist of a concatenation of word segments , each corresponding to a part-of-speech (POS) category. Semitic words may be ambiguous with regard to their segmentation as well as to the POS tags assigned to each segment. When designing POS taggers for Semitic languages, a major architectural decision concerns the choice of the atomic input tokens (terminal symbols). If the tokenization is at the word level, the output tags must be complex, and represent both the segmentation of the word and the POS tag assigned to each word segment. If the tokenization is at the segment level, the input itself must encode the different alternative segmentations of the words, while the output consists of standard POS tags. Comparing these two alternatives is not trivial, as the choice between them may have global effects on the grammatical model. Moreover, intermediate levels of tokenization between these two extremes are conceivable, and, as we aim to show, beneficial. To the best of our knowledge, the problem of tokenization for POS tagging of Semitic languages has not been addressed before in full generality. In this paper, we study this problem for the purpose of POS tagging of Modern Hebrew texts. After extensive error analysis of the two simple tokenization models, we propose a novel, linguistically motivated, intermediate tokenization model that gives better performance for Hebrew over the two initial architectures. Our study is based on the well-known hidden Markov models (HMMs). We start out from a manually devised morphological analyzer and a very small annotated corpus, and describe how to adapt an HMM-based POS tagger for both tokenization architectures. We present an effective technique for smoothing the lexical probabilities using an untagged corpus, and a novel transformation for casting the segment-level tagger in terms of a standard, word-level HMM implementation. The results obtained using our model are on par with the best published results on Modern Standard Arabic, despite the much smaller annotated corpus available for Modern Hebrew. Roy Bar-Haim, Khalil Sima'an, Yoad Winter |
Nat. Lang. Eng. | 2 |
| 2008 | Syntactically Lexicalized Phrase-Based SMTabstractUntil quite recently, extending phrase-based statistical machine translation (PBSMT) with syntactic knowledge caused system performance to deteriorate. The most recent successful enrichments of PBSMT with hierarchical structure either employ nonlinguistically motivated syntax for capturing hierarchical reordering phenomena, or extend the phrase translation table with redundantly ambiguous syntactic structures over phrase pairs. In this paper, we present an extended, harmonized account of our previous work which showed that incorporating linguistically motivated lexical syntactic descriptions, calledsupertags, can yield significantly better PBSMT systems at insignificant extra computational cost. We describe a novel PBSMT model that integrates supertags into the target language model and the target side of the translation model. Two kinds of supertags are employed: those from lexicalized tree-adjoining grammar and combinatory categorial grammar. Despite the differences between the two sets of supertags, they give similar improvements. In addition to integrating the Markov supertagging approach in PBSMT, we explore the utility of a new surface grammaticality measure based on combinatory operators. We perform various experiments on the Arabic-to-English NIST 2005 test set addressing the issues of sparseness, scalability, and the utility of system subcomponents. We show that even when the parallel training data grows very large, the supertagged system retains a relatively stable absolute performance advantage over the unadorned PBSMT system. Arguably, this hints at a performance gap that cannot be bridged by acquiring more phrase pairs. Our best result shows a relative improvement of 6.1% over a state-of-the-art PBSMT model, which compares favorably with the leading systems on the NIST 2005 task. We also demonstrate that the advantages of a supertag-based system carry over to German-English, where improvements of up to 8.9% relative to the baseline system are observed. Hany Hassan, Khalil Sima'an, Andy Way |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Supertagged Phrase-Based Statistical Machine Translation
Hany Hassan, Khalil Sima'an, Andy Way |
ACL | 2 |
| 2007 | Unsupervised estimation for noisy-channel modelsabstractShannon's Noisy-Channel model, which describes how a corrupted message might be reconstructed, has been the corner stone for much work in statistical language and speech processing. The model factors into two components: a language model to characterize the original message and a channel model to describe the channel's corruptive process. The standard approach for estimating the parameters of the channel model is unsupervised Maximum-Likelihood of the observation data, usually approximated using the Expectation-Maximization (EM) algorithm. In this paper we show that it is better to maximize the joint likelihood of the data at both ends of the noisy-channel. We derive a corresponding bi-directional EM algorithm and show that it gives better performance than standard EM on two tasks: (1) translation using a probabilistic lexicon and (2) adaptation of a part-of-speech tagger between related languages. Markos Mylonakis, Khalil Sima'an, Rebecca Hwa |
ICML | 2 |
| 2006 | Syntactic Phrase-Based Statistical Machine TranslationabstractPhrase-based statistical machine translation (PBSMT) systems represent the dominant approach in MT today. However, unlike systems in other paradigms, it has proven difficult to date to incorporate syntactic knowledge in order to improve translation quality. This paper improves on recent research which uses 'syntactified' target language phrases, by incorporating supertags as constraints to better resolve parse tree fragments. In addition, we do not impose any sentence-length limit, and using a log-linear decoder, we outperform a state-of-the-art PBSMT system by over 1.3 BLEU points (or 3.51% relative) on the NIST 2003 Arabic-English test corpus. Hany Hassan, Mary Hearne, Andy Way, Khalil Sima'an |
SLT | 4 |
| 2006 | Wired for Speech: How Voice Activates and Advances the Human-Computer Relationship
Charles B. Callaway, Khalil Sima'an |
Comput. Linguistics | 2 |
| 2003 | Backoff Parameter Estimation for the DOP Model
Khalil Sima'an, Luciano Buratto |
ECML | 1 |
| 2000 | Tree-gram Parsing: Lexical Dependencies and Structural RelationsabstractThis paper explores the kinds of probabilistic relations that are important in syntactic disambiguation. It proposes that two widely used kinds of relations, lexical dependencies and structural relations, have complementary disambiguation capabilities. It presents a new model based on structural relations, the Tree-gram model, and reports experiments showing that structural relations should benefit from enrichment by lexical dependencies. Khalil Sima'an |
ACL | 1 |
| 1999 | A memory-based model of syntactic analysis: data-oriented parsing
Remko Scha, Rens Bod, Khalil Sima'an |
J. Exp. Theor. Artif. Intell. | 3 |
| 1997 | Explanation-Based Learning of Data-Oriented Parsing
Khalil Sima'an |
CoNLL | 1 |
| 1996 | Computational Complexity of Probabilistic Disambiguation by means of Tree-Grammars
Khalil Sima'an |
COLING | 1 |