EDBT 2026 Demo / reviewers in the wild / expert
Kevin Knight
dblp:77/1610
· DBLP profile ↗
120ranked-venue papers
16as first author
0since 2021 · last 2020
0000-0001-9117-1718ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 114 · 14 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-authorDatabases, data management, data science and information retrieval · 2Theory of computation · 2Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
57 papers |
Machine translation · 39% Information extraction and text analysis · 24% Language models and text generation · 18% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% | |
| Theoretical computer science
4 papers |
Automata and formal languages · 74% Information theory · 23% Mathematical optimization · 3% |
Topics — the 30 heaviest of 76, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
decipherment |
1.0 | 7 | 2015 | Unifying Bayesian Inference and Vector Space Models for Improved Decipherment · ACL (1) 2015 Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014 Dependency-Based Decipherment for Resource-Limited Machine Translation · EMNLP 2013 |
Natural language and speech › Machine translation
low-resource machine translation |
0.7 | 6 | 2017 | Deciphering Related Languages · EMNLP 2017 Transfer Learning for Low-Resource Neural Machine Translation · EMNLP 2016 Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014 |
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion |
0.7 | 2 | 2020 | Learning to Pronounce Chinese Without a Pronunciation Dictionary · EMNLP (1) 2020 Grapheme-to-Phoneme Models for (Almost) Any Language · ACL (1) 2016 |
Natural language and speech › Machine translation
statistical machine translation |
0.6 | 8 | 2014 | Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014 Dependency-Based Decipherment for Resource-Limited Machine Translation · EMNLP 2013 Syntactic Re-Alignment Models for Machine Translation · EMNLP-CoNLL 2007 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.6 | 6 | 2015 | Parsing English into Abstract Meaning Representation Using Syntax-Based Machine Translation · EMNLP 2015 Synchronous Tree Adjoining Machine Translation · EMNLP 2009 Binarizing Syntax Trees to Improve Syntax-Based Machine Translation Accuracy · EMNLP-CoNLL 2007 |
Natural language and speech › Machine translation
neural machine translation |
0.5 | 2 | 2016 | Does String-Based Neural MT Learn Source Syntax? · EMNLP 2016 Why Neural Translations are the Right Length · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › dataset construction
corpus filtering |
0.4 | 1 | 2020 | Parallel Corpus Filtering via Pre-trained Language Models · ACL 2020 |
Natural language and speech › Language models and text generation › decoding
large language model decoding |
0.4 | 1 | 2020 | Solving Historical Dictionary Codes with a Neural Language Model · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
neural language model |
0.4 | 1 | 2020 | Solving Historical Dictionary Codes with a Neural Language Model · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.4 | 2 | 2015 | Parsing English into Abstract Meaning Representation Using Syntax-Based Machine Translation · EMNLP 2015 Aligning English Strings with Abstract Meaning Representation Graphs · EMNLP 2014 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 2 | 2019 | Plan-and-Write: Towards Better Automatic Storytelling · AAAI 2019 Two-Level, Many-Path Generation · ACL 1995 |
Natural language and speech › Language models and text generation › text generation
story generation |
0.4 | 1 | 2019 | Plan-and-Write: Towards Better Automatic Storytelling · AAAI 2019 |
Natural language and speech › Machine translation
translationese |
0.4 | 1 | 2019 | Translating Translationese: A Two-Step Approach to Unsupervised Machine Translation · ACL (1) 2019 |
Natural language and speech › Machine translation
unsupervised machine translation |
0.4 | 1 | 2019 | Translating Translationese: A Two-Step Approach to Unsupervised Machine Translation · ACL (1) 2019 |
Knowledge graphs
domain-specific knowledge graph |
0.4 | 1 | 2019 | PaperRobot: Incremental Draft Generation of Scientific Ideas · ACL (1) 2019 |
Knowledge graphs
knowledge graph construction |
0.4 | 1 | 2019 | PaperRobot: Incremental Draft Generation of Scientific Ideas · ACL (1) 2019 |
Knowledge graphs
link prediction |
0.4 | 1 | 2019 | PaperRobot: Incremental Draft Generation of Scientific Ideas · ACL (1) 2019 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.3 | 2 | 2018 | Grapheme-to-Phoneme Models for (Almost) Any Language · ACL (1) 2016 Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word Embedding · EMNLP 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.3 | 1 | 2018 | Modeling Naive Psychology of Characters in Simple Commonsense Stories · ACL (1) 2018 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual word alignment |
0.3 | 1 | 2018 | Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word Embedding · EMNLP 2018 |
Machine learning › Representation and self-supervised learning › word representation
multilingual word embedding |
0.3 | 1 | 2018 | Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word Embedding · EMNLP 2018 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.3 | 1 | 2018 | Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word Embedding · EMNLP 2018 |
Cryptographic primitives and cryptanalysis › cryptanalysis › cipher cryptanalysis
classical cipher cryptanalysis |
0.3 | 2 | 2014 | Cipher Type Detection · EMNLP 2014 Bayesian Inference for Zodiac and Other Homophonic Ciphers · ACL 2011 |
Natural language and speech › Machine translation › low-resource machine translation
related-language translation |
0.3 | 1 | 2017 | Deciphering Related Languages · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › entity linking
cross-lingual entity linking |
0.2 | 1 | 2016 | A Multi-media Approach to Cross-lingual Entity Knowledge Transfer · ACL (1) 2016 |
Natural language and speech › Information extraction and text analysis › named entity processing
entity discovery and linking |
0.2 | 1 | 2016 | A Multi-media Approach to Cross-lingual Entity Knowledge Transfer · ACL (1) 2016 |
Machine learning › Transfer learning and domain adaptation › language adaptation
low-resource language adaptation |
0.2 | 1 | 2016 | Grapheme-to-Phoneme Models for (Almost) Any Language · ACL (1) 2016 |
Machine learning › Transfer learning and domain adaptation › knowledge transfer
parameter transfer |
0.2 | 1 | 2016 | Transfer Learning for Low-Resource Neural Machine Translation · EMNLP 2016 |
Natural language and speech › Language models and text generation › text generation
poetry generation |
0.2 | 1 | 2016 | Generating Topical Poetry · EMNLP 2016 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.2 | 2 | 2013 | Dependency-Based Decipherment for Resource-Limited Machine Translation · EMNLP 2013 What Can Syntax-Based MT Learn from Phrase-Based MT? · EMNLP-CoNLL 2007 |
Methods — techniques the papers use, named apart from their topics
neural language model · 0.9decoding lattice search · 0.9unsupervised alignment · 0.4pre-trained language model · 0.4many-to-many mapping · 0.4GPT · 0.4BERT · 0.4neural machine translation · 0.4memory-attention network · 0.4hierarchical generation · 0.4graph attention · 0.4dictionary induction · 0.4contextual text attention · 0.4sound matching · 0.2image-to-image retrieval · 0.2face recognition · 0.2compression algorithm · 0.2text classification · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Parallel Corpus Filtering via Pre-trained Language ModelsabstractWeb-crawled data provides a good source of parallel corpora for training machine translation models.It is automatically obtained, but extremely noisy, and recent work shows that neural machine translation systems are more sensitive to noise than traditional statistical machine translation methods.In this paper, we propose a novel approach to filter out noisy sentence pairs from web-crawled corpora via pre-trained language models.We measure sentence parallelism by leveraging the multilingual capability of BERT and use the Generative Pre-training (GPT) language model as a domain filter to balance data domains.We evaluate the proposed method on the WMT 2018 Parallel Corpus Filtering shared task, and on our own web-crawled Japanese-Chinese parallel corpus.Our method significantly outperforms baselines and achieves a new stateof-the-art.In an unsupervised setting, our method achieves comparable performance to the top-1 supervised method.We also evaluate on a web-crawled Japanese-Chinese parallel corpus that we make publicly available. Boliang Zhang, Ajay Nagesh, Kevin Knight |
ACL | 3 |
| 2020 | Learning to Pronounce Chinese Without a Pronunciation DictionaryabstractWe demonstrate a program that learns to pronounce Chinese text in Mandarin, without a pronunciation dictionary.From non-parallel streams of Chinese characters and Chinese pinyin syllables, it establishes a many-to-many mapping between characters and pronunciations.Using unsupervised methods, the program effectively deciphers writing into speech.Its token-level character-to-syllable accuracy is 89%, which significantly exceeds the 22% accuracy of prior work. Christopher Chu, Scot Fang, Kevin Knight |
EMNLP (1) | 3 |
| 2020 | Solving Historical Dictionary Codes with a Neural Language ModelabstractWe solve difficult word-based substitution codes by constructing a decoding lattice and searching that lattice with a neural language model. We apply our method to a set of enciphered letters exchanged between US Army General James Wilkinson and agents of the Spanish Crown in the late 1700s and early 1800s, obtained from the US Library of Congress. We are able to decipher 75.1% of the cipher-word tokens correctly. Christopher Chu, Raphael Valenti, Kevin Knight |
EMNLP (1) | 3 |
| 2020 | ReviewRobot: Explainable Paper Review Generation based on Knowledge SynthesisabstractTo assist human review process, we build a novel ReviewRobot to automatically assign a review score and write comments for multiple categories such as novelty and meaningful comparison.A good review needs to be knowledgeable, namely that the comments should be constructive and informative to help improve the paper; and explainable by providing detailed evidence.ReviewRobot achieves these goals via three steps: (1) We perform domainspecific Information Extraction to construct a knowledge graph (KG) from the target paper under review, a related work KG from the papers cited by the target paper, and a background KG from a large collection of previous papers in the domain.(2) By comparing these three KGs, we predict a review score and detailed structured knowledge as evidence for each review category.(3) We carefully select and generalize human review sentences into templates, and apply these templates to transform the review scores and evidence into natural language comments.Experimental results show that our review score predictor reaches 71.4%-100% accuracy.Human assessment by domain experts shows that 41.7%-70.5% of the comments generated by ReviewRobot are valid and constructive, and better than humanwritten ones for 20% of the time.Thus, Re-viewRobot can serve as an assistant for paper reviewers, program chairs and authors. 1 Qingyun Wang 0005, Qi Zeng 0001, Lifu Huang, Kevin Knight, Heng Ji 0001, Nazneen Fatema Rajani |
INLG | 4 |
| 2019 | Plan-and-Write: Towards Better Automatic StorytellingabstractAutomatic storytelling is challenging since it requires generating long, coherent natural language to describes a sensible sequence of events. Despite considerable efforts on automatic story generation in the past, prior work either is restricted in plot planning, or can only generate stories in a narrow domain. In this paper, we explore open-domain story generation that writes stories given a title (topic) as input. We propose a plan-and-write hierarchical generation framework that first plans a storyline, and then generates a story based on the storyline. We compare two planning strategies. The dynamic schema interweaves story planning and its surface realization in text, while the static schema plans out the entire storyline before generating stories. Experiments show that with explicit storyline planning, the generated stories are more diverse, coherent, and on topic than those generated without creating a full plan, according to both automatic and human evaluations. Lili Yao, Nanyun Peng 0001, Ralph M. Weischedel, Kevin Knight, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 4 |
| 2019 | Translating Translationese: A Two-Step Approach to Unsupervised Machine TranslationabstractGiven a rough, word-by-word gloss of a source language sentence, target language natives can uncover the latent, fully-fluent rendering of the translation.In this work we explore this intuition by breaking translation into a two step process: generating a rough gloss by means of a dictionary and then 'translating' the resulting pseudo-translation, or 'Translationese' into a fully fluent translation.We build our Translationese decoder once from a mish-mash of parallel data that has the target language in common and then can build dictionaries on demand using unsupervised techniques, resulting in rapidly generated unsupervised neural MT systems for many source languages.We apply this process to 14 test languages, obtaining better or comparable translation results on high-resource languages than previously published unsupervised MT studies, and obtaining good quality results for low-resource languages that have never been used in an unsupervised MT scenario. Nima Pourdamghani, Nada Aldarrab, Marjan Ghazvininejad, Kevin Knight, Jonathan May |
ACL (1) | 4 |
| 2019 | PaperRobot: Incremental Draft Generation of Scientific IdeasabstractWe present a PaperRobot who performs as an automatic research assistant by (1) conducting deep understanding of a large collection of human-written papers in a target domain and constructing comprehensive background knowledge graphs (KGs); (2) creating new ideas by predicting links from the background KGs, by combining graph attention and contextual text attention; (3) incrementally writing some key elements of a new paper based on memory-attention networks: from the input title along with predicted related entities to generate a paper abstract, from the abstract to generate conclusion and future work, and finally from future work to generate a title for a follow-on paper.Turing Tests, where a biomedical domain expert is asked to compare a system output and a human-authored string, show PaperRobot generated abstracts, conclusion and future work sections, and new titles are chosen over human-written ones up to 30%, 24% and 12% of the time, respectively. 1 keeps almost the same across years.In 2012, US scientists estimated that they read, on average, only 264 papers per year (1 out of 5000 available papers), which is, statistically, not different from what they reported in an identical survey last conducted in 2005.PaperRobot automatically reads existing papers to build background knowledge graphs (KGs), in which nodes are entities/concepts and edges are the relations between these entities (Section 2.2). Qingyun Wang 0005, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji 0001, Mohit Bansal, Yi Luan |
ACL (1) | 4 |
| 2019 | Decipherment of Historical Manuscript ImagesabstractEuropean libraries and archives are filled with enciphered manuscripts from the early modern period. These include military and diplomatic correspondence, records of secret societies, private letters, and so on. Although they are enciphered with classical cryptographic algorithms, their contents are unavailable to working historians. We therefore attack the problem of automatically converting cipher manuscript images into plaintext. We develop unsupervised models for character segmentation, character-image clustering, and decipherment of cluster sequences. We experiment with both pipelined and joint models, and we give empirical results for multiple ciphers. Xusen Yin, Nada Aldarrab, Beáta Megyesi, Kevin Knight |
ICDAR | 4 |
| 2019 | Neighbors helping the poor: improving low-resource machine translation using related languages
Nima Pourdamghani, Kevin Knight |
Mach. Transl. | 2 |
| 2018 | Modeling Naive Psychology of Characters in Simple Commonsense StoriesabstractUnderstanding a narrative requires reading between the lines and reasoning about the unspoken but obvious implications about events and people's mental states -a capability that is trivial for humans but remarkably hard for machines.To facilitate research addressing this challenge, we introduce a new annotation framework to explain naive psychology of story characters as fully-specified chains of mental states with respect to motivations and emotional reactions.Our work presents a new largescale dataset with rich low-level annotations and establishes baseline performance on several new tasks, suggesting avenues for future research. Hannah Rashkin, Antoine Bosselut, Maarten Sap, Kevin Knight, Yejin Choi 0001 |
ACL (1) | 4 |
| 2018 | AMR Beyond the Sentence: the Multi-sentence AMR corpusabstractThere are few corpora that endeavor to represent the semantic content of entire documents. We present a corpus that accomplishes one way of capturing document level semantics, by annotating coreference and similar phenomena (bridging and implicit roles) on top of gold Abstract Meaning Representations of sentence-level semantics. We present a new corpus of this annotation, with analysis of its quality, alongside a plausible baseline for comparison. It is hoped that this Multi-Sentence AMR corpus (MS-AMR) may become a feasible method for developing rich representations of document meaning, useful for tasks such as information extraction and question answering. Tim O'Gorman, Michael Regan, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Martha Palmer |
COLING | 5 |
| 2018 | Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word EmbeddingabstractWe construct a multilingual common semantic space based on distributional semantics, where words from multiple languages are projected into a shared space via which all available resources and knowledge can be shared across multiple languages.Beyond word alignment, we introduce multiple cluster-level alignments and enforce the word clusters to be consistently distributed across multiple languages.We exploit three signals for clustering: (1) neighbor words in the monolingual word embedding space; (2) character-level information; and (3) linguistic properties (e.g., apposition, locative suffix) derived from linguistic structure knowledge bases available for thousands of languages.We introduce a new cluster-consistent correlational neural network to construct the common semantic space by aligning words as well as clusters.Intrinsic evaluation on monolingual and multilingual QVEC tasks shows our approach achieves significantly higher correlation with linguistic features which are extracted from manually crafted lexical resources than state-of-the-art multi-lingual embedding learning methods do.Using low-resource language name tagging as a case study for extrinsic evaluation, our approach achieves up to 14.6% absolute F-score gain over the state of the art on cross-lingual direct transfer.Our approach is also shown to be robust even when the size of bilingual dictionary is small.1 Lifu Huang, Kyunghyun Cho, Boliang Zhang, Heng Ji 0001, Kevin Knight |
EMNLP | 5 |
| 2018 | Describing a Knowledge BaseabstractWe aim to automatically generate natural language descriptions about an input structured knowledge base (KB).We build our generation framework based on a pointer network which can copy facts from the input KB, and add two attention mechanisms: (i) slot-aware attention to capture the association between a slot type and its corresponding slot value; and (ii) a new table position self-attention to capture the inter-dependencies among related slots.For evaluation, besides standard metrics including BLEU, METEOR, and ROUGE, we propose a KB reconstruction based metric by extracting a KB from the generation output and comparing it with the input KB.We also create a new data set which includes 106,216 pairs of structured KBs and their corresponding natural language descriptions for two distinct entity types.Experiments show that our approach significantly outperforms stateof-the-art methods.The reconstructed KB achieves 68.8% -72.6% F-score. 1 Qingyun Wang 0005, Xiaoman Pan, Lifu Huang, Boliang Zhang, Zhiying Jiang, Heng Ji 0001, Kevin Knight |
INLG | 7 |
| 2018 | Abstract Meaning Representation of Constructions: The More We Include, the Better the Representation
Claire Bonial, Bianca Badarau, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Tim O'Gorman, Martha Palmer, Nathan Schneider 0001 |
LREC | 5 |
| 2018 | Recurrent Neural Networks as Weighted Language RecognizersabstractYining Chen, Sorcha Gilroy, Andreas Maletti, Jonathan May, Kevin Knight. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Sorcha Gilroy, Andreas Maletti, Jonathan May, Kevin Knight |
NAACL-HLT | 5 |
| 2018 | Incident-Driven Machine Translation and Name Tagging for Low-resource Languages
Ulf Hermjakob, Daniel Marcu, Jonathan May, Sabrina J. Mielke, Nima Pourdamghani, Michael Pust, Kevin Knight, Tomer Levinboim, Kenton Murray, David Chiang 0001, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Heng Ji 0001 |
Mach. Transl. | 9 |
| 2017 | Cross-lingual Name Tagging and Linking for 282 LanguagesabstractThe ambitious goal of this work is to develop a cross-lingual name tagging and linking framework for 282 languages that exist in Wikipedia.Given a document in any of these languages, our framework is able to identify name mentions, assign a coarse-grained or fine-grained type to each mention, and link it to an English Knowledge Base (KB) if it is linkable.We achieve this goal by performing a series of new KB mining methods: generating "silver-standard" annotations by transferring annotations from English to other languages through crosslingual links and KB properties, refining annotations through self-training and topic selection, deriving language-specific morphology features from anchor links, and mining word translation pairs from crosslingual links.Both name tagging and linking results for 282 languages are promising on Wikipedia data and on-Wikipedia data.All the data sets, resources and systems for 282 languages are made publicly available as a new benchmark 1 . Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, Heng Ji 0001 |
ACL (1) | 5 |
| 2017 | Deciphering Related LanguagesabstractWe present a method for translating texts between close language pairs.The method does not require parallel data, and it does not require the languages to be written in the same script.We show results for six language pairs: Afrikaans/Dutch, Bosnian/Serbian, Danish/Swedish, Macedonian/Bulgarian, Malaysian/Indonesian, and Polish/Belorussian.We report BLEU scores showing our method to outperform others that do not use parallel data. Nima Pourdamghani, Kevin Knight |
EMNLP | 2 |
| 2017 | Embracing Non-Traditional Linguistic Resources for Low-resource Language Name TaggingabstractCurrent supervised name tagging approaches are inadequate for most low-resource languages due to the lack of annotated data and actionable linguistic knowledge. All supervised learning methods (including deep neural networks (DNN)) are sensitive to noise and thus they are not quite portable without massive clean annotations. We found that the F-scores of DNN-based name taggers drop rapidly (20%-30%) when we replace clean manual annotations with noisy annotations in the training data. We propose a new solution to incorporate many non-traditional language universal resources that are readily available but rarely explored in the Natural Language Processing (NLP) community, such as the World Atlas of Linguistic Structure, CIA names, PanLex and survival guides. We acquire and encode various types of non-traditional linguistic resources into a DNN name tagger. Experiments on three low-resource languages show that feeding linguistic knowledge can make DNN significantly more robust to noise, achieving 8%-22% absolute F-score gains on name tagging without using any human annotation Boliang Zhang, Di Lu 0003, Xiaoman Pan, Halidanmu Abudukelimu, Heng Ji 0001, Kevin Knight |
IJCNLP(1) | 7 |
| 2017 | Team ELISA System for DARPA LORELEI Speech Evaluation 2016
Pavlos Papadopoulos, Ruchir Travadi, Colin Vaz, Nikos Malandrakis, Ulf Hermjakob, Nima Pourdamghani, Michael Pust, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Ondrej Glembek, Murali Karthick Baskar, Martin Karafiát, Lukás Burget, Mark Hasegawa-Johnson, Heng Ji 0001, Jonathan May, Kevin Knight, Shri Narayanan |
INTERSPEECH | 19 |
| 2016 | Grapheme-to-Phoneme Models for (Almost) Any LanguageabstractGrapheme-to-phoneme (g2p) models are rarely available in low-resource languages, as the creation of training and evaluation data is expensive and time-consuming.We use Wiktionary to obtain more than 650k word-pronunciation pairs in more than 500 languages.We then develop phoneme and language distance metrics based on phonological and linguistic knowledge; applying those, we adapt g2p models for highresource languages to create models for related low-resource languages.We provide results for models for 229 adapted languages. Aliya Deri, Kevin Knight |
ACL (1) | 2 |
| 2016 | A Multi-media Approach to Cross-lingual Entity Knowledge TransferabstractWhen a large-scale incident or disaster occurs, there is often a great demand for rapidly developing a system to extract detailed and new information from lowresource languages (LLs).We propose a novel approach to discover comparable documents in high-resource languages (HLs), and project Entity Discovery and Linking results from HLs documents back to LLs.We leverage a wide variety of language-independent forms from multiple data modalities, including image processing (image-to-image retrieval, visual similarity and face recognition) and sound matching.We also propose novel methods to learn entity priors from a large-scale HL corpus and knowledge base.Using Hausa and Chinese as the LLs and English as the HL, experiments show that our approach achieves 36.1% higher Hausa name tagging F-score over a costly supervised model, and 9.4% higher Chineseto-English Entity Linking accuracy over state-of-the-art. Di Lu 0003, Xiaoman Pan, Nima Pourdamghani, Shih-Fu Chang, Heng Ji 0001, Kevin Knight |
ACL (1) | 6 |
| 2016 | Generating Topical PoetryabstractWe describe Hafez, a program that generates any number of distinct poems on a usersupplied topic.Poems obey rhythmic and rhyme constraints.We describe the poetrygeneration algorithm, give experimental data concerning its parameters, and show its generality with respect to language and poetic form. Marjan Ghazvininejad, Yejin Choi 0001, Kevin Knight |
EMNLP | 4 |
| 2016 | Why Neural Translations are the Right Length
Kevin Knight, Deniz Yuret |
EMNLP | 2 |
| 2016 | Does String-Based Neural MT Learn Source Syntax?
Inkit Padhi, Kevin Knight |
EMNLP | 3 |
| 2016 | Transfer Learning for Low-Resource Neural Machine TranslationabstractThe encoder-decoder framework for neural machine translation (NMT) has been shown effective in large data scenarios, but is much less effective for low-resource languages.We present a transfer learning method that significantly improves BLEU scores across a range of low-resource languages.Our key idea is to first train a high-resource language pair (the parent model), then transfer some of the learned parameters to the low-resource pair (the child model) to initialize and constrain training.Using our transfer learning method we improve baseline NMT models by an average of 5.6 BLEU on four low-resource language pairs.Ensembling and unknown word replacement add another 2 BLEU which brings the NMT performance on low-resource machine translation close to a strong syntax based machine translation (SBMT) system, exceeding its performance on one language pair.Additionally, using the transfer learning model for re-scoring, we can improve the SBMT system by an average of 1.3 BLEU, improving the state-of-the-art on low-resource machine translation. Barret Zoph, Deniz Yuret, Jonathan May, Kevin Knight |
EMNLP | 4 |
| 2016 | Generating English from Abstract Meaning RepresentationsabstractWe present a method for generating English sentences from Abstract Meaning Representation (AMR) graphs, exploiting a parallel corpus of AMRs and English sentences.We treat AMR-to-English generation as phrase-based machine translation (PBMT).We introduce a method that learns to linearize tokens of AMR graphs into an English-like order.Our linearization reduces the amount of distortion in PBMT and increases generation quality.We report a Bleu score of 26.8 on the standard AMR/English test set. Nima Pourdamghani, Kevin Knight, Ulf Hermjakob |
INLG | 2 |
| 2016 | Extracting Structured Scholarly Information from the Machine Translation Literature
Eunsol Choi, Matic Horvat, Jonathan May, Kevin Knight, Daniel Marcu |
LREC | 4 |
| 2016 | Name Tagging for Low-resource Incident Languages based on Expectation-driven LearningabstractBoliang Zhang, Xiaoman Pan, Tianlu Wang, Ashish Vaswani, Heng Ji, Kevin Knight, Daniel Marcu. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Boliang Zhang, Xiaoman Pan, Ashish Vaswani, Heng Ji 0001, Kevin Knight, Daniel Marcu |
HLT-NAACL | 6 |
| 2016 | Multi-Source Neural TranslationabstractWe build a multi-source machine translation model and train it to maximize the probability of a target English string given French and German sources.Using the neural encoderdecoder framework, we explore several combination methods and report up to +4.8 Bleu increases on top of a very strong attentionbased neural translation model. Barret Zoph, Kevin Knight |
HLT-NAACL | 2 |
| 2016 | Simple, Fast Noise-Contrastive Estimation for Large RNN VocabulariesabstractBarret Zoph, Ashish Vaswani, Jonathan May, Kevin Knight. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Barret Zoph, Ashish Vaswani, Jonathan May, Kevin Knight |
HLT-NAACL | 4 |
| 2016 | From Image to Translation: Processing the Endangered Nyushu ScriptabstractThe lack of computational support has significantly slowed down automatic understanding of endangered languages. In this paper, we take Nyushu (simplified Chinese: 女书; literally: “women’s writing”) as a case study to present the first computational approach that combines Computer Vision and Natural Language Processing techniques to deeply understand an endangered language. We developed an end-to-end system to read a scanned hand-written Nyushu article, segment it into characters, link them to standard characters, and then translate the article into Mandarin Chinese. We propose several novel methods to address the new challenges introduced by noisy input and low resources, including Nyushu-specific feature selection for character segmentation and linking, and character linking lattice based Machine Translation. The end-to-end system performance indicates that the system is a promising approach and can serve as a standard benchmark. Tongtao Zhang, Aritra Chowdhury, Nimit Dhulekar, Jinjing Xia, Kevin Knight, Heng Ji 0001, Bülent Yener |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2015 | Unifying Bayesian Inference and Vector Space Models for Improved DeciphermentabstractQing Dou, Ashish Vaswani, Kevin Knight, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Qing Dou, Ashish Vaswani, Kevin Knight, Chris Dyer |
ACL (1) | 3 |
| 2015 | Context-aware Entity Morph DecodingabstractBoliang Zhang, Hongzhao Huang, Xiaoman Pan, Sujian Li, Chin-Yew Lin, Heng Ji, Kevin Knight, Zhen Wen, Yizhou Sun, Jiawei Han, Bulent Yener. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Boliang Zhang, Hongzhao Huang, Xiaoman Pan, Sujian Li, Chin-Yew Lin, Heng Ji 0001, Kevin Knight, Yizhou Sun, Jiawei Han 0001, Bülent Yener |
ACL (1) | 7 |
| 2015 | Parsing English into Abstract Meaning Representation Using Syntax-Based Machine TranslationabstractWe present a parser for Abstract Meaning Representation (AMR).We treat Englishto-AMR conversion within the framework of string-to-tree, syntax-based machine translation (SBMT).To make this work, we transform the AMR structure into a form suitable for the mechanics of SBMT and useful for modeling.We introduce an AMR-specific language model and add data and features drawn from semantic resources.Our resulting AMR parser significantly improves upon state-of-the-art results. Michael Pust, Ulf Hermjakob, Kevin Knight, Daniel Marcu, Jonathan May |
EMNLP | 3 |
| 2015 | How Much Information Does a Human Translator Add to the Original?abstractWe ask how much information a human translator adds to an original text, and we provide a bound.We address this question in the context of bilingual text compression: given a source text, how many bits of additional information are required to specify the target text produced by a human translator?We develop new compression algorithms and establish a benchmark task. Barret Zoph, Marjan Ghazvininejad, Kevin Knight |
EMNLP | 3 |
| 2015 | How to Make a Frenemy: Multitape FSTs for Portmanteau GenerationabstractA portmanteau is a type of compound word that fuses the sounds and meanings of two component words; for example, "frenemy" (friend + enemy) or "smog" (smoke + fog).We develop a system, including a novel multitape FST, that takes an input of two words and outputs possible portmanteaux.Our system is trained on a list of known portmanteaux and their component words, and achieves 45% exact matches in cross-validated experiments. Aliya Deri, Kevin Knight |
HLT-NAACL | 2 |
| 2015 | How to Memorize a Random 60-Bit StringabstractUser-generated passwords tend to be memorable, but not secure.A random, computergenerated 60-bit string is much more secure.However, users cannot memorize random 60bit strings.In this paper, we investigate methods for converting arbitrary bit strings into English word sequences (both prose and poetry), and we study their memorability and other properties. Marjan Ghazvininejad, Kevin Knight |
HLT-NAACL | 2 |
| 2015 | Unsupervised Entity Linking with Abstract Meaning RepresentationabstractXiaoman Pan, Taylor Cassidy, Ulf Hermjakob, Heng Ji, Kevin Knight. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Xiaoman Pan, Taylor Cassidy, Ulf Hermjakob, Heng Ji 0001, Kevin Knight |
HLT-NAACL | 5 |
| 2015 | Building and Using a Knowledge Graph to Combat Human Trafficking
Pedro A. Szekely, Craig A. Knoblock, Jason Slepicka, Andrew Philpot, Chengye Yin, Dipsy Kapoor, Premkumar Natarajan, Daniel Marcu, Kevin Knight, David Stallard, Subessware S. Karunamoorthy, Rajagopal Bojanapalli, Steven Minton, Brian Amanatullah, Todd Hughes, Mike Tamayo, David Flynt, Rachel Artiss, Shih-Fu Chang, Tao Chen 0015, Gerald Hiebel, Lidia Silva Ferreira |
ISWC (2) | 10 |
| 2014 | Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine TranslationabstractInspired by previous work, where decipherment is used to improve machine translation, we propose a new idea to combine word alignment and decipherment into a single learning process.We use EM to estimate the model parameters, not only to maximize the probability of parallel corpus, but also the monolingual corpus.We apply our approach to improve Malagasy-English machine translation, where only a small amount of parallel data is available.In our experiments, we observe gains of 0.9 to 2.1 Bleu over a strong baseline. Qing Dou, Ashish Vaswani, Kevin Knight |
EMNLP | 3 |
| 2014 | Cipher Type DetectionabstractManual analysis and decryption of enciphered documents is a tedious and error prone work.Often-even after spending large amounts of time on a particular cipher-no decipherment can be found.Automating the decryption of various types of ciphers makes it possible to sift through the large number of encrypted messages found in libraries and archives, and to focus human effort only on a small but potentially interesting subset of them.In this work, we train a classifier that is able to predict which encipherment method has been used to generate a given ciphertext.We are able to distinguish 50 different cipher types (specified by the American Cryptogram Association) with an accuracy of 58.5%.This is a 11.2% absolute improvement over the best previously published classifier. Malte Nuhn, Kevin Knight |
EMNLP | 2 |
| 2014 | Aligning English Strings with Abstract Meaning Representation GraphsabstractWe align pairs of English sentences and corresponding Abstract Meaning Repre-sentations (AMR), at the token level. Such alignments will be useful for downstream extraction of semantic interpretation and generation rules. Our method involves linearizing AMR structures and perform-ing symmetrized EM training. We obtain 86.5 % and 83.1 % alignment F score on de-velopment and test sets. 1 Nima Pourdamghani, Ulf Hermjakob, Kevin Knight |
EMNLP | 4 |
| 2014 | Aligning context-based statistical models of language with brain activity during readingabstractMany statistical models for natural language processing exist, including context-based neural networks that (1) model the previously seen context as a latent feature vector, (2) integrate successive words into the context using some learned representation (embedding), and (3) compute output probabilities for incoming words given the context.On the other hand, brain imaging studies have suggested that during reading, the brain (a) continuously builds a context from the successive words and every time it encounters a word it (b) fetches its properties from memory and (c) integrates it with the previous context with a degree of effort that is inversely proportional to how probable the word is.This hints to a parallelism between the neural networks and the brain in modeling context (1 and a), representing the incoming words (2 and b) and integrating it (3 and c).We explore this parallelism to better understand the brain processes and the neural networks representations.We study the alignment between the latent vectors used by neural networks and brain activity observed via Magnetoencephalography (MEG) when subjects read a story.For that purpose we apply the neural network to the same text the subjects are reading, and explore the ability of these three vector representations to predict the observed word-by-word brain activity.Our novel results show that: before a new word i is read, brain activity is well predicted by the neural network latent representation of context and the predictability decreases as the brain integrates the word and changes its own representation of context.Secondly, the neural network embedding of word i can predict the MEG activity when word i is presented to the subject, revealing that it is correlated with the brain's own representation of word i.Moreover, we obtain that the activity is predicted in different regions of the brain with varying delay.The delay is consistent with the placement of each region on the processing pathway that starts in the visual cortex and moves to higher level regions.Finally, we show that the output probability computed by the neural networks agrees with the brain's own assessment of the probability of word i, as it can be used to predict the brain activity after the word i's properties have been fetched from memory and the brain is in the process of integrating it into the context. Leila Wehbe, Ashish Vaswani, Kevin Knight, Tom M. Mitchell |
EMNLP | 3 |
| 2014 | Mapping Between English Strings and Reentrant Semantic Graphs
Fabienne Braune, Daniel Bauer 0002, Kevin Knight |
LREC | 3 |
| 2013 | Parsing Graphs with Hyperedge Replacement Grammars
David Chiang 0001, Jacob Andreas, Daniel Bauer 0002, Karl Moritz Hermann, Bevan K. Jones, Kevin Knight |
ACL (1) | 6 |
| 2013 | Dependency-Based Decipherment for Resource-Limited Machine TranslationabstractWe introduce dependency relations into deciphering foreign languages and show that dependency relations help improve the state-ofthe-art deciphering accuracy by over 500%.We learn a translation lexicon from large amounts of genuinely non parallel data with decipherment to improve a phrase-based machine translation system trained with limited parallel data.In experiments, we observe BLEU gains of 1.2 to 1.8 across three different test sets. Qing Dou, Kevin Knight |
EMNLP | 2 |
| 2013 | Curating and contextualizing Twitter stories to assist with social newsgatheringabstractWhile journalism is evolving toward a rather open-minded participatory paradigm, social media presents overwhelming streams of data that make it difficult to identify the information of a journalist's interest. Given the increasing interest of journalists in broadening and democratizing news by incorporating social media sources, we have developed TweetGathering, a prototype tool that provides curated and contextualized access to news stories on Twitter. This tool was built with the aim of assisting journalists both with gathering and with researching news stories as users comment on them. Five journalism professionals who tested the tool found helpful characteristics that could assist them with gathering additional facts on breaking news, as well as facilitating discovery of potential information sources such as witnesses in the geographical locations of news. Arkaitz Zubiaga, Heng Ji 0001, Kevin Knight |
IUI | 3 |
| 2012 | Semantics-Based Machine Translation with Hyperedge Replacement Grammars
Bevan K. Jones, Jacob Andreas, Daniel Bauer 0002, Karl Moritz Hermann, Kevin Knight |
COLING | 5 |
| 2012 | Large Scale Decipherment for Out-of-Domain Machine Translation
Qing Dou, Kevin Knight |
EMNLP-CoNLL | 2 |
| 2011 | Deciphering Foreign Language
Sujith Ravi, Kevin Knight |
ACL | 2 |
| 2011 | Bayesian Inference for Zodiac and Other Homophonic Ciphers
Sujith Ravi, Kevin Knight |
ACL | 2 |
| 2010 | Efficient Inference through Cascades of Weighted Tree Transducers
Jonathan May, Kevin Knight, Heiko Vogler |
ACL | 2 |
| 2010 | Minimized Models and Grammar-Informed Initialization for Supertagging with Highly Ambiguous Lexicons
Sujith Ravi, Jason Baldridge, Kevin Knight |
ACL | 3 |
| 2010 | A Statistical Model for Lost Language Decipherment
Benjamin Snyder, Regina Barzilay, Kevin Knight |
ACL | 3 |
| 2010 | Fast, Greedy Model Minimization for Unsupervised Tagging
Sujith Ravi, Ashish Vaswani, Kevin Knight, David Chiang 0001 |
COLING | 3 |
| 2010 | Automatic Analysis of Rhythmic Poetry with Applications to Generation and Translation
Erica Greene, Tugba Bodrumlu, Kevin Knight |
EMNLP | 3 |
| 2010 | Bayesian Inference for Finite-State Transducers
David Chiang 0001, Jonathan Graehl, Kevin Knight, Adam Pauls, Sujith Ravi |
HLT-NAACL | 3 |
| 2010 | Unsupervised Syntactic Alignment with Inversion Transduction Grammars
Adam Pauls, Daniel Klein 0001, David Chiang 0001, Kevin Knight |
HLT-NAACL | 4 |
| 2010 | Does GIZA++ Make Search Errors?abstractWord alignment is a critical procedure within statistical machine translation (SMT). Brown et al. (1993) have provided the most popular word alignment algorithm to date, one that has been implemented in the GIZA (Al-Onaizan et al., 1999) and GIZA++ (Och and Ney 2003) software and adopted by nearly every SMT project. In this article, we investigate whether this algorithm makes search errors when it computes Viterbi alignments, that is, whether it returns alignments that are sub-optimal according to a trained model. Sujith Ravi, Kevin Knight |
Comput. Linguistics | 2 |
| 2010 | Re-structuring, Re-labeling, and Re-aligning for Syntax-Based Machine TranslationabstractThis article shows that the structure of bilingual material from standard parsing and alignment tools is not optimal for training syntax-based statistical machine translation (SMT) systems. We present three modifications to the MT training data to improve the accuracy of a state-of-the-art syntax MT system: re-structuring changes the syntactic structure of training parse trees to enable reuse of substructures; re-labeling alters bracket labels to enrich rule application context; and re-aligning unifies word alignment across sentences to remove bad word alignments and refine good ones. Better structures, labels, and word alignments are learned by the EM algorithm. We show that each individual technique leads to improvement as measured by BLEU, and we also show that the greatest improvement is achieved by combining them. We report an overall 1.48 BLEU improvement on the NIST08 evaluation set over a strong baseline in Chinese/English translation. Wei Wang 0006, Jonathan May, Kevin Knight, Daniel Marcu |
Comput. Linguistics | 3 |
| 2009 | Fast Consensus Decoding over Translation Forests
John DeNero, David Chiang 0001, Kevin Knight |
ACL/IJCNLP | 3 |
| 2009 | Minimized Models for Unsupervised Part-of-Speech Tagging
Sujith Ravi, Kevin Knight |
ACL/IJCNLP | 2 |
| 2009 | Synchronous Tree Adjoining Machine Translation
Steve DeNeefe, Kevin Knight |
EMNLP | 2 |
| 2009 | 11,001 New Features for Statistical Machine Translation
David Chiang 0001, Kevin Knight, Wei Wang 0006 |
HLT-NAACL | 2 |
| 2009 | Learning Phoneme Mappings for Transliteration without Parallel Data
Sujith Ravi, Kevin Knight |
HLT-NAACL | 2 |
| 2009 | Binarization of Synchronous Context-Free GrammarsabstractSystems based on synchronous grammars and tree transducers promise to improve the quality of statistical machine translation output, but are often very computationally intensive. The complexity is exponential in the size of individual grammar rules due to arbitrary re-orderings between the two languages. We develop a theory of binarization for synchronous context-free grammars and present a linear-time algorithm for binarizing synchronous rules when possible. In our large-scale experiments, we found that almost all rules are binarizable and the resulting binarized rule set significantly improves the speed and accuracy of a state-of-the-art syntax-based machine translation system. We also discuss the more general, and computationally more difficult, problem of finding good parsing strategies for non-binarizable rules, and present an approximate polynomial-time algorithm for this problem. Liang Huang 0001, Hao Zhang 0010, Daniel Gildea, Kevin Knight |
Comput. Linguistics | 4 |
| 2009 | The Power of Extended Top-Down Tree TransducersabstractExtended top-down tree transducers (transducteurs généralisés descendants; see [A. Arnold and M. Dauchet, Bi-transductions de forêts, in Proceedings of the 3rd International Colloquium on Automata, Languages and Programming, Edinburgh University Press, Edinburgh, 1976, pp. 74–86]) received renewed interest in the field of natural language processing. Here those transducers are extensively and systematically studied. Their main properties are identified and their relation to classical top-down tree transducers is exactly characterized. The obtained properties completely explain the Hasse diagram of the induced classes of tree transformations. In addition, it is shown that most interesting classes of transformations computed by extended top-down tree transducers are not closed under composition. Andreas Maletti, Jonathan Graehl, Mark Hopkins, Kevin Knight |
SIAM J. Comput. | 4 |
| 2008 | Name Translation in Statistical Machine Translation - Learning When to Transliterate
Ulf Hermjakob, Kevin Knight, Hal Daumé III |
ACL | 2 |
| 2008 | Attacking Decipherment Problems Optimally with Low-Order N-gram Models
Sujith Ravi, Kevin Knight |
EMNLP | 2 |
| 2008 | Automatic Prediction of Parser Accuracy
Sujith Ravi, Kevin Knight, Radu Soricut |
EMNLP | 2 |
| 2008 | Training Tree TransducersabstractMany probabilistic models for natural language are now written in terms of hierarchical tree structure. Tree-based modeling still lacks many of the standard tools taken for granted in (finite-state) string-based modeling. The theory of tree transducer automata provides a possible framework to draw on, as it has been worked out in an extensive literature. We motivate the use of tree transducers for natural language and address the training problem for probabilistic tree-to-tree and tree-to-string transducers. Jonathan Graehl, Kevin Knight, Jonathan May |
Comput. Linguistics | 2 |
| 2007 | What Can Syntax-Based MT Learn from Phrase-Based MT?
Steve DeNeefe, Kevin Knight, Wei Wang 0006, Daniel Marcu |
EMNLP-CoNLL | 2 |
| 2007 | Syntactic Re-Alignment Models for Machine Translation
Jonathan May, Kevin Knight |
EMNLP-CoNLL | 2 |
| 2007 | Binarizing Syntax Trees to Improve Syntax-Based Machine Translation Accuracy
Wei Wang 0006, Kevin Knight, Daniel Marcu |
EMNLP-CoNLL | 2 |
| 2007 | Capturing practical natural language transformations
Kevin Knight |
Mach. Transl. | 1 |
| 2006 | Scalable Inference and Training of Context-Rich Syntactic Translation ModelsabstractStatistical MT has made great progress in the last few years, but current translation models are weak on re-ordering and target language fluency. Syntactic approaches seek to remedy these problems. In this paper, we take the framework for acquiring multi-level syntactic translation rules of (Galley et al., 2004) from aligned tree-string pairs, and present two main extensions of their approach: first, instead of merely computing a single derivation that minimally explains a sentence pair, we construct a large number of derivations that include contextually richer rules, and account for multiple interpretations of unaligned words. Second, we propose probability estimates and a training procedure for weighting these rules. We contrast different approaches on real examples, show that our estimates based on multiple derivations favor phrasal re-orderings that are linguistically better motivated, and establish that our larger rules provide a 3.63 BLEU point increase over minimal rules. Michel Galley, Jonathan Graehl, Kevin Knight, Daniel Marcu, Steve DeNeefe, Wei Wang 0006, Ignacio Thayer |
ACL | 3 |
| 2006 | Unsupervised Analysis for Decipherment Problems
Kevin Knight, Anish Nair, Nishit Rathod, Kenji Yamada |
ACL | 1 |
| 2006 | SPMT: Statistical Machine Translation with Syntactified Target Language Phrases
Daniel Marcu, Wei Wang 0006, Abdessamad Echihabi, Kevin Knight |
EMNLP | 4 |
| 2006 | Building an English-iraqi Arabic machine translation system for spoken utterances with limited resourcesabstractThis paper presents an English-Iraqi Arabic speech-to-speech statistical machine translation system using limited resources. In it, we explore the constraints involved, how we endeavored to mitigate such problems as a non-standard orthography and a highly inflected grammar, and discuss leveraging existing plentiful resources for Modern Standard Arabic to assist in this task. These combined techniques yield a reduction in unknown words at translation time by over 40 % and a +3.65 increase in BLEU score over a previous state-of-the-art system using the same parallel training corpus of spoken utterances. Index Terms: speech translation, limited resources, Arabic 1. Jason Riesa, Behrang Mohit, Kevin Knight, Daniel Marcu |
INTERSPEECH | 3 |
| 2006 | Relabeling Syntax Trees to Improve Syntax-Based Machine Translation Quality
Bryant Huang, Kevin Knight |
HLT-NAACL | 2 |
| 2006 | A Better N-Best List: Practical Determinization of Weighted Finite Tree Automata
Jonathan May, Kevin Knight |
HLT-NAACL | 2 |
| 2006 | Capitalizing Machine Translation
Wei Wang 0006, Kevin Knight, Daniel Marcu |
HLT-NAACL | 2 |
| 2006 | Synchronous Binarization for Machine Translation
Hao Zhang 0010, Liang Huang 0001, Daniel Gildea, Kevin Knight |
HLT-NAACL | 4 |
| 2006 | No More Strings, pleaseabstractSummary form only given. In natural language research, many (grammar) trees were felled in 1992, to make room for the highly successful string-based HMM industry. A small literature survived on parsing (putting a tree on a string) and syntactic language modeling (putting a weight on a string). However, trees are making a comeback. Tree transformations are turning out to be very useful in large-scale machine translation (MT), and we will cover recent developments in this area. Most of the tree techniques used in MT turn out to be generic, leading to tools and software for manipulating tree automata in general. Tree acceptors and transducers generalize HMM techniques to the world of trees, raising many interesting theoretical and practical problems. Kevin Knight |
SLT | 1 |
| 2006 | Tiburon: A Weighted Tree Automata Toolkit
Jonathan May, Kevin Knight |
CIAA | 2 |
| 2006 | Discovering the linear writing order of a two-dimensional ancient hieroglyphic script
Shou-De Lin, Kevin Knight |
Artif. Intell. | 2 |
| 2005 | Transonics: A Practical Speech-to-Speech Translator for English-Farsi Medical Dialogs
Robert S. Belvin, Emil Ettelaie, Sudeep Gandhe, Panayiotis G. Georgiou, Kevin Knight, Daniel Marcu, Scott Millward, Shri Narayanan, Howard Neely, David R. Traum |
ACL | 5 |
| 2005 | Interactively Exploring a Machine Translation Model
Steve DeNeefe, Kevin Knight, Hayward H. Chan |
ACL | 2 |
| 2005 | An Overview of Probabilistic Tree Transducers for Natural Language Processing
Kevin Knight, Jonathan Graehl |
CICLing | 1 |
| 2005 | Machine translation in the year 2004abstractMachine translation (MT) accuracy has recently increased, due to better techniques and to the availability of larger parallel training sets. Statistical MT systems are now able to translate across a wide variety of language pairs. This paper covers the basic elements of state-of-the-art, statistical MT, including modeling, decoding, evaluation, and data preparation. Kevin Knight, Daniel Marcu |
ICASSP (5) | 1 |
| 2004 | What's in a translation rule?
Michel Galley, Mark Hopkins, Kevin Knight, Daniel Marcu |
HLT-NAACL | 3 |
| 2004 | Training Tree Transducers
Jonathan Graehl, Kevin Knight |
HLT-NAACL | 2 |
| 2004 | Fast and optimal decoding for machine translation
Ulrich Germann, Michael Jahr, Kevin Knight, Daniel Marcu, Kenji Yamada |
Artif. Intell. | 3 |
| 2003 | Feature-Rich Statistical Translation of Noun PhrasesabstractWe define noun phrase translation as a subtask of machine translation. This enables us to build a dedicated noun phrase translation subsystem that improves over the currently best general statistical machine translation methods by incorporating special modeling and special features. We achieved 65.5% translation accuracy in a German-English translation task vs. 53.2% with IBM Model 4. Philipp Koehn, Kevin Knight |
ACL | 2 |
| 2003 | Empirical Methods for Compound Splitting
Philipp Koehn, Kevin Knight |
EACL | 2 |
| 2003 | Syntax-based language models for statistical machine translationabstractWe present a syntax-based language model for use in noisy-channel machine translation. In particular, a language model based upon that described in (Cha01) is combined with the syntax based translation-model described in (YK01). The resulting system was used to translate 347 sentences from Chinese to English and compared with the results of an IBM-model-4-based system, as well as that of (YK02), all trained on the same data. The translations were sorted into four groups: good/bad syntax crossed with good/bad meaning. While the total number of translations that preserved meaning were the same for (YK02) and the syntax-based system (and both higher than the IBM-model-4-based system), the syntax based system had 45% more translations that also had good syntax than did (YK02) (and approximately 70% more than IBM Model 4). The number of translations that did not preserve meaning, but at least had good grammar, also increased, though to less avail. Eugene Charniak, Kevin Knight, Kenji Yamada |
MTSummit | 2 |
| 2003 | What's New in Statistical Machine Translation
Kevin Knight, Philipp Koehn |
HLT-NAACL | 1 |
| 2003 | Cognates Can Improve Statistical Translation Models
Grzegorz Kondrak, Daniel Marcu, Kevin Knight |
HLT-NAACL | 3 |
| 2003 | Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight |
HLT-NAACL | 12 |
| 2003 | Syntax-based Alignment of Multiple Translations: Extracting Paraphrases and Generating New Sentences
Bo Pang 0001, Kevin Knight, Daniel Marcu |
HLT-NAACL | 2 |
| 2002 | Translating Named Entities Using Monolingual and Bilingual ResourcesabstractNamed entity phrases are some of the most difficult phrases to translate because new phrases can appear from nowhere, and because many are domain specific, not to be found in bilingual dictionaries. We present a novel algorithm for translating named entity phrases using easily obtainable monolingual and bilingual resources. We report on the application and evaluation of this algorithm in translating Arabic named entities to English. We also compare our results with the results obtained from human translations and a commercial system for the same task. Yaser Al-Onaizan, Kevin Knight |
ACL | 2 |
| 2002 | A Decoder for Syntax-based Statistical MTabstractThis paper describes a decoding algorithm for a syntax-based translation model (Yamada and Knight, 2001). The model has been extended to incorporate phrasal translations as presented here. In contrast to a conventional word-to-word statistical model, a decoder for the syntax-based model builds up an English parse tree given a sentence in a foreign language. As the model size becomes huge in a practical setting, and the decoder considers multiple syntactic structures for each word alignment, several pruning techniques are necessary. We tested our decoder in a Chinese-to-English translation system, and obtained better results than IBM Model 4. We also discuss issues concerning the relation between this decoder and a language model. Kenji Yamada, Kevin Knight |
ACL | 2 |
| 2002 | The Importance of Lexicalized Syntax Models for Natural Language Generation Tasks
Hal Daumé III, Kevin Knight, Irene Langkilde-Geary, Daniel Marcu, Kenji Yamada |
INLG | 2 |
| 2002 | Summarization beyond sentence extraction: A probabilistic approach to sentence compression
Kevin Knight, Daniel Marcu |
Artif. Intell. | 1 |
| 2002 | Translation with Scarce Bilingual Resources
Yaser Al-Onaizan, Ulrich Germann, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Daniel Marcu, Kenji Yamada |
Mach. Transl. | 4 |
| 2001 | Fast Decoding and Optimal Decoding for Machine TranslationabstractA good decoding algorithm is critical to the success of any statistical machine translation system. The decoder's job is to find the translation that is most likely according to set of previously learned parameters (and a formula for combining them). Since the space of possible translations is extremely large, typical decoding algorithms are only able to examine a portion of it, thus risking to miss good solutions. In this paper, we compare the speed and output quality of a traditional stack-based decoding algorithm with two new decoders: a fast greedy decoder and a slow but optimal decoder that treats decoding as an integer-programming optimization problem. Ulrich Germann, Michael Jahr, Kevin Knight, Daniel Marcu, Kenji Yamada |
ACL | 3 |
| 2001 | A Syntax-based Statistical Translation ModelabstractWe present a syntax-based statistical translation model.Our model transforms a source-language parse tree into a target-language string by applying stochastic operations at each node.These operations capture linguistic differences such as word order and case marking.Model parameters are estimated in polynomial time using an EM algorithm.The model produces word alignments that are better than those produced by IBM Model 5. Kenji Yamada, Kevin Knight |
ACL | 2 |
| 2001 | Knowledge Sources for Word-Level Translation Models
Philipp Koehn, Kevin Knight |
EMNLP | 2 |
| 1999 | Decoding Complexity in Word-Replacement Translation Models
Kevin Knight |
Comput. Linguistics | 1 |
| 1998 | The Practical Value Of N-Grams Is In Generation
Irene Langkilde-Geary, Kevin Knight |
INLG | 2 |
| 1998 | Machine Transliteration
Kevin Knight, Jonathan Graehl |
Comput. Linguistics | 1 |
| 1997 | Machine TransliterationabstractIt is challenging to translate names and technical terms across languages with different alphabets and sound inventories. These items are commonly transliterated, i.e., replaced with approximate phonetic equivalents. For example, computer in English comes out as (konpyuutaa) in Japanese. Translating such items from Japanese back to English is even more challenging, and of practical interest, as transliterated items make up the bulk of text phrases not found in bilingual dictionaries. We describe and evaluate a method for performing backwards transliterations by machine. This method uses a generative model, incorporating several distinct stages in the transliteration process. Kevin Knight, Jonathan Graehl |
ACL | 1 |
| 1995 | Two-Level, Many-Path GenerationabstractLarge-scale natural language generation requires the integration of vast amounts of knowledge: lexical, grammatical, and conceptual. A robust generator must be able to operate well even when pieces of knowledge are missing. It must also be robust against incomplete or inaccurate inputs. To attack these problems, we have built a hybrid generator, in which gaps in symbolic knowledge are filled by statistical methods. We describe algorithms and show experimental results. We also discuss how the hybrid generation model can be used to simplify current generators and enhance their portability, even when perfect knowledge is in principle obtainable. Kevin Knight, Vasileios Hatzivassiloglou |
ACL | 1 |
| 1995 | Unification-Based Glossing
Vasileios Hatzivassiloglou, Kevin Knight |
IJCAI | 2 |
| 1995 | Filling Knowledge Gaps in a Broad-Coverage Machine Translation System
Kevin Knight, Ishwar Chander, Matthew Haines, Vasileios Hatzivassiloglou, Eduard H. Hovy, Masayo Iida, Steve K. Luk, Richard Whitney, Kenji Yamada |
IJCAI | 1 |
| 1994 | Automated Postediting of Documents
Kevin Knight, Ishwar Chander |
AAAI | 1 |
| 1994 | Building a Large-Scale Knowledge Base for Machine Translation
Kevin Knight, Steve K. Luk |
AAAI | 1 |
| 1993 | Are Many Reactive Agents Better Than a Few Deliberative Ones?
Kevin Knight |
IJCAI | 1 |
| 1992 | Integrating knowledge acquisition and language acquisition
Kevin Knight |
Appl. Intell. | 1 |