Markos Mylonakis

dblp:58/809 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Machine translation · 74% Information extraction and text analysis · 26%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › data annotation
linguistic annotation
0.112011
Learning Hierarchical Translation Structure with Linguistic Annotations · ACL 2011
Natural language and speech › Machine translation
syntax-aware translation
0.112011
Learning Hierarchical Translation Structure with Linguistic Annotations · ACL 2011
Natural language and speech › Machine translation
syntax-based machine translation
0.112011
Learning Hierarchical Translation Structure with Linguistic Annotations · ACL 2011
Natural language and speech › Machine translation › synchronous grammar
inversion transduction grammar
0.112008
Phrase Translation Probabilities with ITG Priors and Smoothing as Learning Objective · EMNLP 2008
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.112008
Phrase Translation Probabilities with ITG Priors and Smoothing as Learning Objective · EMNLP 2008
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging
0.112007
Unsupervised estimation for noisy-channel models · ICML 2007
Natural language and speech › Machine translation
statistical machine translation
0.112007
Unsupervised estimation for noisy-channel models · ICML 2007
Information theory › communication channels › channel models › noisy channel
noisy channel model
0.112007
Unsupervised estimation for noisy-channel models · ICML 2007

Methods — techniques the papers use, named apart from their topics

maximum likelihood estimation · 0.1expectation-maximization · 0.1linguistic annotation · 0.1smoothing · 0.1ITG priors · 0.1
YearPublicationVenuePosition
2014 Learning structural dependencies of words in the Zipfian Tail
abstract
This article uses semi-supervised Expectation Maximization (EM) to learn lexico-syntactic dependencies, i.e. associations between words and the structures that occur with them. Due to Zipfian distributions in language, such dependencies are extremely sparse in labelled data, and unlabelled data are the only source for learning them. Specifically, we learn sparse lexical parameters of a generative parsing model (a Probabilistic Context-Free Grammar, PCFG) that is initially estimated over the Penn Treebank. Our lexical parameters are similar to supertags—they are fine-grained, and encode complex structural information at the pre-terminal level. Our goal is to use unlabelled data to learn these for words that are rare or unseen in the labelled data. We get large error reductions (up to 17.5%) in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies, resulting in a statistically significant improvement in labelled bracketing score of the treebank PCFG. Our semi-supervised method incorporates structural and lexical priors from the labelled data to guide estimation from unlabelled data, and is the first successful use of semi-supervised EM to improve a generative structured model already trained over large labelled data. The method scales well to larger amounts of unlabelled data, and also gives substantial error reductions (up to 11.5%) for models trained on smaller amounts of labelled data, making it relevant to low-resource languages with small treebanks as well.
Tejaswini Deoskar, Markos Mylonakis, Khalil Sima'an
J. Log. Comput.2
2011 Learning Hierarchical Translation Structure with Linguistic Annotations
Markos Mylonakis, Khalil Sima'an
ACL1
2010 Learning Probabilistic Synchronous CFGs for Phrase-Based Translation
Markos Mylonakis, Khalil Sima'an
CoNLL1
2008 Phrase Translation Probabilities with ITG Priors and Smoothing as Learning Objective
Markos Mylonakis, Khalil Sima'an
EMNLP1
2008 Better statistical estimation can benefit all phrases in phrase-based statistical machine translation
abstract
The heuristic estimates of conditional phrase translation probabilities are based on frequency counts in a word-aligned parallel corpus. Earlier attempts at more principled estimation using Expectation-Maximization (EM) under perform this heuristic. This paper shows that a recently introduced novel estimator based on smoothing might provide a good alternative. Whenallphrasepairsare estimated (no length cut-off), this estimator slightly outperforms the heuristic estimator.
Khalil Sima'an, Markos Mylonakis
SLT2
2007 Unsupervised estimation for noisy-channel models
abstract
Shannon's Noisy-Channel model, which describes how a corrupted message might be reconstructed, has been the corner stone for much work in statistical language and speech processing. The model factors into two components: a language model to characterize the original message and a channel model to describe the channel's corruptive process. The standard approach for estimating the parameters of the channel model is unsupervised Maximum-Likelihood of the observation data, usually approximated using the Expectation-Maximization (EM) algorithm. In this paper we show that it is better to maximize the joint likelihood of the data at both ends of the noisy-channel. We derive a corresponding bi-directional EM algorithm and show that it gives better performance than standard EM on two tasks: (1) translation using a probabilistic lexicon and (2) adaptation of a part-of-speech tagger between related languages.
Markos Mylonakis, Khalil Sima'an, Rebecca Hwa
ICML1