Colin Cherry

dblp:99/6601 · DBLP profile ↗
← Back
61ranked-venue papers
17as first author
15since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 13 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
abstract
Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Senyu Li, Jiayi Wang 0010, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, David Ifeoluwa Adelani
EMNLP4
2025 Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
abstract
Data contamination—the accidental consumption of evaluation examples within the pre-training data—can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task. Starting from a carefully decontaminated train-test split, we systematically introduce contamination at various stages, scales, and data formats to isolate its effect and measure its impact on performance metrics. Our experiments reveal that contamination with both source and target substantially inflates BLEU scores, and this inflation is 2.5 times larger (up to 30 BLEU points) for 8B compared to 1B models. In contrast, source-only and target-only contamination generally produce smaller, less consistent over-estimations. Finally, we study how the temporal distribution and frequency of contaminated samples influence performance over-estimation across languages with varying degrees of data resources.
Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo, Colin Cherry, Markus Freitag
ICML5
2024 Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
abstract
Christian Tomani, David Vilar, Markus Freitag, Colin Cherry, Subhajit Naskar, Mara Finkelstein, Xavier Garcia, Daniel Cremers. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Christian Tomani, David Vilar, Markus Freitag, Colin Cherry, Subhajit Naskar, Mara Finkelstein, Xavier Garcia, Daniel Cremers
ACL (1)4
2024 When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
abstract
While large language models (LLMs) often adopt finetuning to unlock their capabilities for downstream applications, our understanding on the inductive biases (especially the scaling properties) of different finetuning methods is still limited. To fill this gap, we conduct systematic experiments studying whether and how different scaling factors, including LLM model size, pretraining data size, new finetuning parameter size and finetuning data size, affect the finetuning performance. We consider two types of finetuning – full-model tuning (FMT) and parameter efficient tuning (PET, including prompt tuning and LoRA), and explore their scaling behaviors in the data-limited regime where the LLM model size substantially outweighs the finetuning data size. Based on two sets of pretrained bilingual LLMs from 1B to 16B and experiments on bilingual machine translation and multilingual summarization benchmarks, we find that 1) LLM finetuning follows a powerbased multiplicative joint scaling law between finetuning data size and each other scaling factor; 2) LLM finetuning benefits more from LLM model scaling than pretraining data scaling, and PET parameter scaling is generally ineffective; and 3) the optimal finetuning method is highly task- and finetuning data-dependent. We hope our findings could shed light on understanding, selecting and developing LLM finetuning methods.
Biao Zhang 0006, Zhongtao Liu, Colin Cherry, Orhan Firat
ICLR3
2024 To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation
abstract
Abstract We conduct a large-scale fine-grained comparative analysis of machine translations (MTs) against human translations (HTs) through the lens of morphosyntactic divergence. Across three language pairs and two types of divergence defined as the structural difference between the source and the target, MT is consistently more conservative than HT, with less morphosyntactic diversity, more convergent patterns, and more one-to-one alignments. Through analysis on different decoding algorithms, we attribute this discrepancy to the use of beam search that biases MT towards more convergent patterns. This bias is most amplified when the convergent pattern appears around 50% of the time in training data. Lastly, we show that for a majority of morphosyntactic divergences, their presence in HT is correlated with decreased MT performance, presenting a greater challenge for MT systems.
Jiaming Luo, Colin Cherry, George F. Foster
Trans. Assoc. Comput. Linguistics2
2023 Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability
abstract
Large, multilingual language models exhibit surprisingly good zero-or few-shot machine translation capabilities, despite having never seen the intentionally-included translation examples provided to typical neural translation systems.We investigate the role of incidental bilingualism-the unintentional consumption of bilingual signals, including translation examples-in explaining the translation capabilities of large language models, taking the Pathways Language Model (PaLM) as a case study.We introduce a mixed-method approach to measure and understand incidental bilingualism at scale.We show that PaLM is exposed to over 30 million translation pairs across at least 44 languages.Furthermore, the amount of incidental bilingual content is highly correlated with the amount of monolingual in-language content for non-English languages.We relate incidental bilingual content to zero-shot prompts and show that it can be used to mine new prompts to improve PaLM's out-of-English zero-shot translation quality.Finally, in a series of small-scale ablations, we show that its presence has a substantial impact on translation capabilities, although this impact diminishes with model scale.
Eleftheria Briakou, Colin Cherry, George F. Foster
ACL (1)2
2023 Prompting PaLM for Translation: Assessing Strategies and Performance
abstract
David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, George Foster. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, George F. Foster
ACL (1)3
2023 The Unreasonable Effectiveness of Few-shot Learning for Machine Translation
abstract
We demonstrate the potential of few-shot translation systems, trained with unpaired language data, for both high and low-resource language pairs. We show that with only 5 examples of high-quality translation data shown at inference, a transformer decoder-only model trained solely with self-supervised learning, is able to match specialized supervised state-of-the-art models as well as more general commercial translation systems. In particular, we outperform the best performing system on the WMT'21 English-Chinese news translation task by only using five examples of English-Chinese parallel data at inference. Furthermore, the resulting models are two orders of magnitude smaller than state-of-the-art language models. We then analyze the factors which impact the performance of few-shot translation systems, and highlight that the quality of the few-shot demonstrations heavily determines the quality of the translations generated by our models. Finally, we show that the few-shot paradigm also provides a way to control certain attributes of the translation --- we show that we are able to control for regional varieties and formality using only a five examples at inference, paving the way towards controllable machine translation systems.
Xavier Garcia, Yamini Bansal, Colin Cherry, George F. Foster, Maxim Krikun, Melvin Johnson, Orhan Firat
ICML3
2022 Scaling Laws for Neural Machine Translation
Behrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna, Maxim Krikun, Xavier Garcia, Ciprian Chelba, Colin Cherry
ICLR8
2022 Data Scaling Laws in NMT: The Effect of Noise and Architecture
abstract
In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.
Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang 0006, Colin Cherry, Behnam Neyshabur, Orhan Firat
ICML5
2022 XTREME-S: Evaluating Cross-lingual Speech Representations
abstract
We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages.XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval.Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning.This paper describes the new benchmark and establishes the first speech-only and speechtext baselines using XLS-R and mSLAM on all downstream tasks.We motivate the design choices and detail how to use the benchmark.Datasets and fine-tuning scripts are made easily accessible through the HuggingFace platform. 1
Alexis Conneau, Ankur Bapna, Yu Zhang 0033, Patrick von Platen, Anton Lozhkov, Colin Cherry, Ye Jia, Clara Rivera, Mihir Kale, Daan van Esch, Vera Axelrod, Simran Khanuja, Jonathan H. Clark, Orhan Firat, Michael Auli, Sebastian Ruder, Jason Riesa, Melvin Johnson
INTERSPEECH7
2022 Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation
abstract
End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research.Recent works have demonstrated that the performance of such direct S2ST systems is approaching that of conventional cascade S2ST when trained on comparable datasets.However, in practice, the performance of direct S2ST is bounded by the availability of paired S2ST training data.In this work, we explore multiple approaches for leveraging much more widely available unsupervised and weakly-supervised speech and text data to improve the performance of direct S2ST based on Translatotron 2. With our most effective approaches, the average translation quality of direct S2ST on 21 language pairs on the CVSS-C corpus is improved by +13.6 BLEU (or +113% relatively), as compared to the previous state-of-the-art trained without additional data.The improvements on low-resource language are even more significant (+398% relatively on average).Our comparative studies suggest future research directions for S2ST and speech representation learning.
Ye Jia, Yifan Ding 0004, Ankur Bapna, Colin Cherry, Yu Zhang 0033, Alexis Conneau, Nobuyuki Morioka
INTERSPEECH4
2021 Sentence Boundary Augmentation for Neural Machine Translation Robustness
abstract
Neural Machine Translation (NMT) models have demonstrated strong state of the art performance on translation tasks where well-formed training and evaluation data are pro-vided, but they remain sensitive to inputs that include errors of various types. Specifically, in the context of long-form speech translation systems, where the input transcripts come from Automatic Speech Recognition (ASR), the NMT models have to handle errors including phoneme substitutions, grammatical structure, and sentence boundaries, all of which pose challenges to NMT robustness. Through in-depth error analysis, we show that sentence boundary segmentation has the largest impact on quality, and we develop a simple data augmentation strategy to improve segmentation robustness.
Te I, Naveen Arivazhagan, Colin Cherry, Dirk Padfield
ICASSP4
2021 Subtitle Translation as Markup Translation
Colin Cherry, Naveen Arivazhagan, Dirk Padfield, Maxim Krikun
Interspeech1
2021 Assessing Reference-Free Peer Evaluation for Machine Translation
abstract
Sweta Agrawal, George Foster, Markus Freitag, Colin Cherry. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Sweta Agrawal, George F. Foster, Markus Freitag, Colin Cherry
NAACL-HLT4
2020 Inference Strategies for Machine Translation with Conditional Masking
abstract
Conditional masked language model (CMLM) training has proven successful for nonautoregressive and semi-autoregressive sequence generation tasks, such as machine translation.Given a trained CMLM, however, it is not clear what the best inference strategy is.We formulate masked inference as a factorization of conditional probabilities of partial sequences, show that this does not harm performance, and investigate a number of simple heuristics motivated by this perspective.We identify a thresholding strategy that has advantages over the standard "mask-predict" algorithm, and provide analyses of its behavior on machine translation tasks.
Julia Kreutzer, George F. Foster, Colin Cherry
EMNLP (1)3
2020 Re-Translation Strategies for Long Form, Simultaneous, Spoken Language Translation
abstract
We investigate the problem of simultaneous machine translation of long-form speech content. We target a continuous speech-to-text scenario, generating translated captions for a live audio feed, such as a lecture or play-by-play commentary. As this scenario allows for revisions to our incremental translations, we adopt a re-translation approach to simultaneous translation, where the source is repeatedly translated from scratch as it grows. This approach naturally exhibits very low latency and high final quality, but at the cost of incremental instability as the output is continuously refined. We experiment with a pipeline of industry-grade speech recognition and translation tools, augmented with simple inference heuristics to improve stability. We use TED Talks as a source of multilingual test data, developing our techniques on English-to-German spoken language translation. Our minimalist approach to simultaneous translation allows us to scale our final evaluation to several other target languages, dramatically improving incremental stability for all of them.
Naveen Arivazhagan, Colin Cherry, Te I, Wolfgang Macherey, Pallavi Baljekar, George F. Foster
ICASSP2
2020 Shaping the Narrative Arc: Information-Theoretic Collaborative DialoguePaper type: Technical Paper
Kory W. Mathewson, Pablo Samuel Castro, Colin Cherry, George F. Foster, Marc G. Bellemare
ICCC3
2019 Monotonic Infinite Lookback Attention for Simultaneous Machine Translation
abstract
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, Colin Raffel. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li 0133, Colin Raffel
ACL (1)2
2018 Revisiting Character-Based Neural Machine Translation with Capacity and Compression
abstract
Translating characters instead of words or word-fragments has the potential to simplify the processing pipeline for neural machine translation (NMT), and improve results by eliminating hyper-parameters and manual feature engineering.However, it results in longer sequences in which each symbol contains less information, creating both modeling and computational challenges.In this paper, we show that the modeling problem can be solved by standard sequence-to-sequence architectures of sufficient depth, and that deep models operating at the character level outperform identical models operating over word fragments.This result implies that alternative architectures for handling character input are better viewed as methods for reducing computation time than as improved ways of modeling longer sequences.From this perspective, we evaluate several techniques for characterlevel NMT, verify that they do not match the performance of our deep character baseline model, and evaluate the performance versus computation time tradeoffs they offer.Within this framework, we also perform the first evaluation for NMT of conditional computation over time, in which the model learns which timesteps can be skipped, rather than having them be dictated by a fixed schedule specified before training begins.
Colin Cherry, George F. Foster, Ankur Bapna, Orhan Firat, Wolfgang Macherey
EMNLP1
2017 A Challenge Set Approach to Evaluating Machine Translation
abstract
Neural machine translation represents an exciting leap forward in translation quality.But what longstanding weaknesses does it resolve, and which remain?We address these questions with a challenge set approach to translation evaluation and error analysis.A challenge set consists of a small set of sentences, each hand-designed to probe a system's capacity to bridge a particular structural divergence between languages.To exemplify this approach, we present an English-French challenge set, and use it to analyze phrase-based and neural systems.The resulting analysis provides not only a more fine-grained picture of the strengths of neural systems, but also insight into which linguistic phenomena remain out of reach.
Pierre Isabelle, Colin Cherry, George F. Foster
EMNLP2
2016 A Dataset for Detecting Stance in Tweets
Saif M. Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu 0001, Colin Cherry
LREC5
2016 An Empirical Evaluation of Noise Contrastive Estimation for the Neural Network Joint Model of Translation
abstract
The neural network joint model of translation or NNJM (Devlin et al., 2014) combines source and target context to produce a powerful translation feature.However, its softmax layer necessitates a sum over the entire output vocabulary, which results in very slow maximum likelihood (MLE) training.This has led some groups to train using Noise Contrastive Estimation (NCE), which side-steps this sum.We carry out the first direct comparison of MLE and NCE training objectives for the NNJM, showing that NCE is significantly outperformed by MLE on large-scale Arabic-English and Chinese-English translation tasks.We also show that this drop can be avoided by using a recently proposed translation noise distribution.
Colin Cherry
HLT-NAACL1
2016 Integrating Morphological Desegmentation into Phrase-based Decoding
Mohammad Salameh, Colin Cherry, Grzegorz Kondrak
HLT-NAACL2
2015 The Unreasonable Effectiveness of Word Representations for Twitter Named Entity Recognition
abstract
Named entity recognition (NER) systems trained on newswire perform very badly when tested on Twitter. Signals that were reliable in copy-edited text disappear almost entirely in Twitter’s informal chatter, requiring the construction of specialized models. Using well understood techniques, we set out to improve Twitter NER performance when given a small set of annotated training tweets. To leverage unlabeled tweets, we build Brown clusters and word vectors, enabling generalizations across distributionally similar words. To leverage annotated newswire data, we employ an importance weighting scheme. Taken all together, we establish a new state-of-the-art on two common test sets. Though it is wellknown that word representations are useful for NER, supporting experiments have thus far focused on newswire data. We emphasize the effectiveness of representations on Twitter NER, and demonstrate that their inclusion can improve performance by up to 20 F1.
Colin Cherry
HLT-NAACL1
2015 Inflection Generation as Discriminative String Transduction
abstract
We approach the task of morphological inflection generation as discriminative string transduction.Our supervised system learns to generate word-forms from lemmas accompanied by morphological tags, and refines them by referring to the other forms within a paradigm.Results of experiments on six diverse languages with varying amounts of training data demonstrate that our approach improves the state of the art in terms of predicting inflected word-forms.
Garrett Nicolai, Colin Cherry, Grzegorz Kondrak
HLT-NAACL2
2014 Lattice Desegmentation for Statistical Machine Translation
abstract
Morphological segmentation is an effec-tive sparsity reduction strategy for statis-tical machine translation (SMT) involv-ing morphologically complex languages. When translating into a segmented lan-guage, an extra step is required to deseg-ment the output; previous studies have de-segmented the 1-best output from the de-coder. In this paper, we expand our trans-lation options by desegmenting n-best lists or lattices. Our novel lattice desegmenta-tion algorithm effectively combines both segmented and desegmented views of the target language for a large subspace of possible translation outputs, which allows for inclusion of features related to the de-segmentation process, as well as an un-segmented language model (LM). We in-vestigate this technique in the context of English-to-Arabic and English-to-Finnish translation, showing significant improve-ments in translation quality over deseg-mentation of 1-best decoder outputs. 1
Mohammad Salameh, Colin Cherry, Grzegorz Kondrak
ACL (1)2
2013 Regularized Minimum Error Rate Training
abstract
Minimum Error Rate Training (MERT) remains one of the preferred methods for tuning linear parameters in machine translation systems, yet it faces significant issues.First, MERT is an unregularized learner and is therefore prone to overfitting.Second, it is commonly used on a noisy, non-convex loss function that becomes more difficult to optimize as the number of parameters increases.To address these issues, we study the addition of a regularization term to the MERT objective function.Since standard regularizers such as ℓ 2 are inapplicable to MERT due to the scale invariance of its objective function, we turn to two regularizers-ℓ 0 and a modification of ℓ 2and present methods for efficiently integrating them during search.To improve search in large parameter spaces, we also present a new direction finding algorithm that uses the gradient of expected BLEU to orient MERT's exact line searches.Experiments with up to 3600 features show that these extensions of MERT yield results comparable to PRO, a learner often used with large feature sets.
Michel Galley, Chris Quirk, Colin Cherry, Kristina Toutanova
EMNLP3
2013 Improved Reordering for Phrase-Based Translation using Sparse Features
Colin Cherry
HLT-NAACL1
2013 Reversing Morphological Tokenization in English-to-Arabic SMT
Mohammad Salameh, Colin Cherry, Grzegorz Kondrak
HLT-NAACL2
2013 À la Recherche du Temps Perdu: extracting temporal relations from medical text in the 2012 i2b2 NLP challenge
abstract
OBJECTIVE: An analysis of the timing of events is critical for a deeper understanding of the course of events within a patient record. The 2012 i2b2 NLP challenge focused on the extraction of temporal relationships between concepts within textual hospital discharge summaries. MATERIALS AND METHODS: The team from the National Research Council Canada (NRC) submitted three system runs to the second track of the challenge: typifying the time-relationship between pre-annotated entities. The NRC system was designed around four specialist modules containing statistical machine learning classifiers. Each specialist targeted distinct sets of relationships: local relationships, 'sectime'-type relationships, non-local overlap-type relationships, and non-local causal relationships. RESULTS: The best NRC submission achieved a precision of 0.7499, a recall of 0.6431, and an F1 score of 0.6924, resulting in a statistical tie for first place. Post hoc improvements led to a precision of 0.7537, a recall of 0.6455, and an F1 score of 0.6954, giving the highest scores reported on this task to date. DISCUSSION AND CONCLUSIONS: Methods for general relation extraction extended well to temporal relations, and gave top-ranked state-of-the-art results. Careful ordering of predictions within result sets proved critical to this success.
Colin Cherry, Xiaodan Zhu 0001, Joel D. Martin, Berry de Bruijn
J. Am. Medical Informatics Assoc.1
2013 Detecting concept relations in clinical text: Insights from a state-of-the-art model
Xiaodan Zhu 0001, Colin Cherry, Svetlana Kiritchenko, Joel D. Martin, Berry de Bruijn
J. Biomed. Informatics2
2013 A Graph-Partitioning Framework for Aligning Hierarchical Topic Structures to Presentations
abstract
This paper studies the problem of imposing an existing hierarchical semantic structure onto a corresponding spoken document in which the structures are embedded, with the goal of indexing such documents for easier access. We propose a graph-partitioning framework to solve a semantic tree-to-string alignment problem through optimizing a normalized-cut criterion. We present models with different modeling capabilities and time complexities in this framework and provide experimental evidence of their performance. We relate graph partitioning to conventional dynamic time warping (DTW) as it applies to this problem, and show that the proposed framework can naturally include topic segmentation to accommodate cohesion constraints.
Xiaodan Zhu 0001, Colin Cherry, Gerald Penn
IEEE Trans. Speech Audio Process.2
2012 Paraphrasing for Style
Wei Xu 0004, Alan Ritter, William B. Dolan, Ralph Grishman, Colin Cherry
COLING5
2012 Batch Tuning Strategies for Statistical Machine Translation
Colin Cherry, George F. Foster
HLT-NAACL1
2012 MSR SPLAT, a language analysis toolkit
Chris Quirk, Pallavi Choudhury, Jianfeng Gao 0001, Hisami Suzuki, Kristina Toutanova, Michael Gamon, Scott Yih, Colin Cherry, Lucy Vanderwende
HLT-NAACL8
2011 Lexically-Triggered Hidden Markov Models for Clinical Document Coding
Svetlana Kiritchenko, Colin Cherry
ACL2
2011 Data-Driven Response Generation in Social Media
Alan Ritter, Colin Cherry, William B. Dolan
EMNLP2
2011 Indexing Spoken Documents with Hierarchical Semantic Structures: Semantic Tree-to-string Alignment Models
Xiaodan Zhu 0001, Colin Cherry, Gerald Penn
IJCNLP2
2011 Machine-learned solutions for three stages of clinical information extraction: the state of the art at i2b2 2010
abstract
OBJECTIVE: As clinical text mining continues to mature, its potential as an enabling technology for innovations in patient care and clinical research is becoming a reality. A critical part of that process is rigid benchmark testing of natural language processing methods on realistic clinical narrative. In this paper, the authors describe the design and performance of three state-of-the-art text-mining applications from the National Research Council of Canada on evaluations within the 2010 i2b2 challenge. DESIGN: The three systems perform three key steps in clinical information extraction: (1) extraction of medical problems, tests, and treatments, from discharge summaries and progress notes; (2) classification of assertions made on the medical problems; (3) classification of relations between medical concepts. Machine learning systems performed these tasks using large-dimensional bags of features, as derived from both the text itself and from external sources: UMLS, cTAKES, and Medline. MEASUREMENTS: Performance was measured per subtask, using micro-averaged F-scores, as calculated by comparing system annotations with ground-truth annotations on a test set. RESULTS: The systems ranked high among all submitted systems in the competition, with the following F-scores: concept extraction 0.8523 (ranked first); assertion detection 0.9362 (ranked first); relationship detection 0.7313 (ranked second). CONCLUSION: For all tasks, we found that the introduction of a wide range of features was crucial to success. Importantly, our choice of machine learning algorithms allowed us to be versatile in our feature design, and to introduce a large number of features without overfitting and without encountering computing-resource bottlenecks.
Berry de Bruijn, Colin Cherry, Svetlana Kiritchenko, Joel D. Martin, Xiaodan Zhu 0001
J. Am. Medical Informatics Assoc.2
2010 Fast and Accurate Arc Filtering for Dependency Parsing
Shane Bergsma, Colin Cherry
COLING2
2010 Integrating Joint n-gram Features into a Discriminative Training Framework
Sittichai Jiampojamarn, Colin Cherry, Grzegorz Kondrak
HLT-NAACL2
2010 Unsupervised Modeling of Twitter Conversations
Alan Ritter, Colin Cherry, William B. Dolan
HLT-NAACL2
2010 Statistical Machine Translation Philipp Koehn (University of Edinburgh) Cambridge University Press, 2010, xii+433 pp; ISBN 978-0-521-87415-1, $60.00
Colin Cherry
Comput. Linguistics1
2009 A global model for joint lemmatization and part-of-speech prediction
Kristina Toutanova, Colin Cherry
ACL/IJCNLP2
2009 Discriminative Substring Decoding for Transliteration
Colin Cherry, Hisami Suzuki
EMNLP1
2009 On the Syllabification of Phonemes
Susan Bartlett, Grzegorz Kondrak, Colin Cherry
HLT-NAACL3
2009 Unsupervised Morphological Segmentation with Log-Linear Models
Hoifung Poon, Colin Cherry, Kristina Toutanova
HLT-NAACL2
2008 Automatic Syllabification with Structured SVMs for Letter-to-Phoneme Conversion
Susan Bartlett, Grzegorz Kondrak, Colin Cherry
ACL3
2008 Cohesive Phrase-Based Decoding for Statistical Machine Translation
Colin Cherry
ACL1
2008 Joint Processing and Discriminative Training for Letter-to-Phoneme Conversion
Sittichai Jiampojamarn, Colin Cherry, Grzegorz Kondrak
ACL2
2006 Soft Syntactic Constraints for Word Alignment through Discriminative Training
Colin Cherry, Dekang Lin
ACL1
2006 Improved Large Margin Dependency Parsing via Local Constraints and Laplacian Regularization
Qin Iris Wang, Colin Cherry, Daniel J. Lizotte, Dale Schuurmans
CoNLL2
2006 A Comparison of Syntactically Motivated Word Alignment Spaces
Colin Cherry, Dekang Lin
EACL1
2005 Dependency Treelet Translation: Syntactically Informed Phrasal SMT
abstract
We describe a novel approach to statistical machine translation that combines syntactic information in the source language with recent advances in phrasal translation. This method requires a source-language dependency parser, target language word segmentation and an unsupervised word alignment component. We align a parallel corpus, project the source dependency parse onto the target sentence, extract dependency treelet translation pairs, and train a tree-based ordering model. We describe an efficient decoder and show that using these tree-based models in combination with conventional SMT models provides a promising approach that incorporates the power of phrasal SMT with the linguistic generality available in a parser.
Chris Quirk, Arul Menezes, Colin Cherry
ACL3
2005 An Expectation Maximization Approach to Pronoun Resolution
Colin Cherry, Shane Bergsma
CoNLL1
2003 A Probability Model to Improve Word Alignment
abstract
Word alignment plays a crucial role in statistical machine translation. Word-aligned corpora have been found to be an excellent source of translation-related knowledge. We present a statistical model for computing the probability of an alignment given a sentence pair. This model allows easy integration of context-specific features. Our experiments show that this model can be an effective tool for improving an existing word alignment.
Colin Cherry, Dekang Lin
ACL1
2003 Word Alignment with Cohesion Constraint
Dekang Lin, Colin Cherry
HLT-NAACL2
1975 Celebration of the 25th Anniversary of Norbert Wiener's Cybernetics
Colin Cherry
IEEE Trans. Syst. Man Cybern.1
1963 Review: Book Review
abstract
Modulation and Coding in Information Systems Gordon M. Russell 1962 ; 260. (London : Prentice-Hall International Inc. , 42s.)
Colin Cherry
Comput. J.1
1961 A New Type of Computer for Problems in Propositional Logic, with Greatly Reduced Scanning Procedures
Colin Cherry, Peter K. T. Vaswani
Inf. Control.1