Adam Lopez

dblp:65/5274 · DBLP profile ↗
← Back
39ranked-venue papers
4as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
16 papers
Information extraction and text analysis · 41% Language models and text generation · 15% Deep learning architectures and training · 14%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.222023
Cross-lingual Transfer Can Worsen Bias in Sentiment Analysis · EMNLP 2023
Intrinsic Bias Metrics Do Not Correlate with Application Bias · ACL/IJCNLP (1) 2021
Machine learning › Learning paradigms
multi-label classification
0.812024
Taming the Sigmoid Bottleneck: Provably Argmaxable Sparse Multi-Label Classification · AAAI 2024
Machine learning › Deep learning architectures and training
output layer
0.812024
Taming the Sigmoid Bottleneck: Provably Argmaxable Sparse Multi-Label Classification · AAAI 2024
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.722019
A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages · EMNLP/IJCNLP (1) 2019
What do character-level models learn about morphology? The case of dependency parsing · EMNLP 2018
Machine learning › Deep learning architectures and training
encoder-decoder architecture
0.412020
Inflecting When There's No Majority: Limitations of Encoder-Decoder Neural Networks as Cognitive Models for German Plurals · ACL 2020
Natural language and speech › Information extraction and text analysis › computational morphology
morphological inflection
0.412020
Inflecting When There's No Majority: Limitations of Encoder-Decoder Neural Networks as Cognitive Models for German Plurals · ACL 2020
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
low-resource dependency parsing
0.412019
A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis › semantic parsing
semantic graph parsing
0.412019
Semantic graph parsing with recurrent neural network DAG grammars · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis
semantic parsing
0.412019
Semantic graph parsing with recurrent neural network DAG grammars · EMNLP/IJCNLP (1) 2019
Automata and formal languages
graph grammars
0.412019
Semantic graph parsing with recurrent neural network DAG grammars · EMNLP/IJCNLP (1) 2019
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.432011
Training a Log-Linear Parser with Loss Functions via Softmax-Margin · EMNLP 2011
Efficient CCG Parsing: A* versus Adaptive Supertagging · ACL 2011
A Comparison of Loopy Belief Propagation and Dual Decomposition for Integrated CCG Supertagging and Parsing · ACL 2011
Natural language and speech › Information extraction and text analysis
morphological analysis
0.312018
What do character-level models learn about morphology? The case of dependency parsing · EMNLP 2018
Natural language and speech › Language models and text generation › language modeling
character-level language modeling
0.312017
From Characters to Words to in Between: Do We Capture Morphology? · ACL (1) 2017
Natural language and speech › Language models and text generation
language modeling
0.312017
From Characters to Words to in Between: Do We Capture Morphology? · ACL (1) 2017
Machine learning › Representation and self-supervised learning › word representation
morphological representation
0.312017
From Characters to Words to in Between: Do We Capture Morphology? · ACL (1) 2017
Machine learning › Representation and self-supervised learning › word representation
subword representation
0.312017
From Characters to Words to in Between: Do We Capture Morphology? · ACL (1) 2017
Natural language and speech › Information extraction and text analysis › syntactic parsing › grammar-based parsing
combinatory categorial grammar parsing
0.222011
Efficient CCG Parsing: A* versus Adaptive Supertagging · ACL 2011
A Comparison of Loopy Belief Propagation and Dual Decomposition for Integrated CCG Supertagging and Parsing · ACL 2011
Natural language and speech › Information extraction and text analysis › semantic analysis
negation scope detection
0.212016
Neural Networks For Negation Scope Detection · ACL (1) 2016
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.212016
N-gram language models for massively parallel devices · ACL (1) 2016
Natural language and speech › Information extraction and text analysis › syntactic parsing › syntactic disambiguation
supertagging
0.222011
Efficient CCG Parsing: A* versus Adaptive Supertagging · ACL 2011
A Comparison of Loopy Belief Propagation and Dual Decomposition for Integrated CCG Supertagging and Parsing · ACL 2011
GPUs and heterogeneous computing › GPU computing
GPU implementation
0.212016
N-gram language models for massively parallel devices · ACL (1) 2016
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.212023
Cross-lingual Transfer Can Worsen Bias in Sentiment Analysis · EMNLP 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
maximum entropy models
0.112011
Training a Log-Linear Parser with Loss Functions via Softmax-Margin · EMNLP 2011
Natural language and speech › Language models and text generation
natural language understanding
0.112016
Neural Networks For Negation Scope Detection · ACL (1) 2016
Natural language and speech › Machine translation › statistical machine translation
hierarchical phrase-based translation
0.112007
Hierarchical Phrase-Based Translation with Suffix Arrays · EMNLP-CoNLL 2007
Natural language and speech › Machine translation
statistical machine translation
0.112007
Hierarchical Phrase-Based Translation with Suffix Arrays · EMNLP-CoNLL 2007
Natural language and speech › Information extraction and text analysis › text mining
web text mining
0.012013
Dirt Cheap Web-Scale Parallel Text from the Common Crawl · ACL (1) 2013

Methods — techniques the papers use, named apart from their topics

recurrent neural network · 0.8discrete fourier transform output layer · 0.8cross-lingual transfer learning · 0.7counterfactual evaluation · 0.7unargmaxable token detection · 0.6low-rank softmax analysis · 0.6correlation analysis · 0.5encoder-decoder neural network · 0.4systematic comparison · 0.4character-level models · 0.3b-tree · 0.2
YearPublicationVenuePosition
2024 Taming the Sigmoid Bottleneck: Provably Argmaxable Sparse Multi-Label Classification
abstract
Sigmoid output layers are widely used in multi-label classification (MLC) tasks, in which multiple labels can be assigned to any input. In many practical MLC tasks, the number of possible labels is in the thousands, often exceeding the number of input features and resulting in a low-rank output layer. In multi-class classification, it is known that such a low-rank output layer is a bottleneck that can result in unargmaxable classes: classes which cannot be predicted for any input. In this paper, we show that for MLC tasks, the analogous sigmoid bottleneck results in exponentially many unargmaxable label combinations. We explain how to detect these unargmaxable outputs and demonstrate their presence in three widely used MLC datasets. We then show that they can be prevented in practice by introducing a Discrete Fourier Transform (DFT) output layer, which guarantees that all sparse label combinations with up to k active labels are argmaxable. Our DFT layer trains faster and is more parameter efficient, matching the F1@k score of a sigmoid layer while using up to 50% fewer trainable parameters. Our code is publicly available at https://github.com/andreasgrv/sigmoid-bottleneck.
Andreas Grivas, Antonio Vergari, Adam Lopez
AAAI3
2024 Human Temporal Inferences Go Beyond Aspectual Class
abstract
Past work in NLP has proposed the task of classifying English verb phrases into situation aspect categories, assuming that these categories play an important role in tasks requiring temporal reasoning. We investigate this assumption by gathering crowd-sourced judgements about aspectual entailments from non-expert, native English participants. The results suggest that aspectual class alone is not sufficient to explain the response patterns of the participants. We propose that looking at scenarios which can feasibly accompany an action description contributes towards a better explanation of the participants' answers. A further experiment using GPT-3.5 shows that its outputs follow different patterns than human answers, suggesting that such conceivable scenarios cannot be fully accounted for in the language alone. We release our dataset to support further research.
Katarzyna Prus, Mark Steedman, Adam Lopez
EACL (1)3
2024 First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models
abstract
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez
NAACL-HLT4
2023 Cross-lingual Transfer Can Worsen Bias in Sentiment Analysis
abstract
Sentiment analysis (SA) systems are widely deployed in many of the world's languages, and there is well-documented evidence of demographic bias in these systems.In languages beyond English, scarcer training data is often supplemented with transfer learning using pretrained models, including multilingual models trained on other languages.In some cases, even supervision data comes from other languages.Does cross-lingual transfer also import new biases?To answer this question, we use counterfactual evaluation to test whether gender or racial biases are imported when using cross-lingual transfer, compared to a monolingual transfer setting.Across five languages, we find that systems using cross-lingual transfer usually become more biased than their monolingual counterparts.We also find racial biases to be much more prevalent than gender biases.To spur further research on this topic, we release the sentiment models we used for this study, and the intermediate checkpoints throughout training, yielding 1,525 distinct models; we also release our evaluation code. 1
Seraphina Goldfarb-Tarrant, Björn Ross, Adam Lopez
EMNLP3
2022 Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in Practice
abstract
Classifiers in natural language processing (NLP) often have a large number of output classes.For example, neural language models (LMs) and machine translation (MT) models both predict tokens from a vocabulary of thousands.The Softmax output layer of these models typically receives as input a dense feature representation, which has much lower dimensionality than the output.In theory, the result is some words may be impossible to be predicted via argmax, irrespective of input features, and empirically, there is evidence this happens in small language models (Demeter et al., 2020).In this paper we ask whether it can happen in practical large language models and translation models.To do so, we develop algorithms to detect such unargmaxable tokens in public models.We find that 13 out of 150 models do indeed have such tokens; however, they are very infrequent and unlikely to impact model quality.We release our algorithms and code so that others can test their models.1
Andreas Grivas, Nikolay Bogoychev, Adam Lopez
ACL (1)3
2022 Regularization or lexical probability-matching? How German speakers generalize plural morphology
Kate McCurdy, Sharon Goldwater, Adam Lopez
CogSci3
2021 Intrinsic Bias Metrics Do Not Correlate with Application Bias
abstract
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, Adam Lopez. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, Adam Lopez
ACL/IJCNLP (1)5
2020 Inflecting When There's No Majority: Limitations of Encoder-Decoder Neural Networks as Cognitive Models for German Plurals
abstract
Can artificial neural networks learn to represent inflectional morphology and generalize to new words as human speakers do?Kirov and Cotterell (2018) argue that the answer is yes: modern Encoder-Decoder (ED) architectures learn human-like behavior when inflecting English verbs, such as extending the regular past tense form /-(e)d/ to novel words.However, their work does not address the criticism raised by Marcus et al. (1995): that neural models may learn to extend not the regular, but the most frequent class -and thus fail on tasks like German number inflection, where infrequent suffixes like /-s/ can still be productively generalized.To investigate this question, we first collect a new dataset from German speakers (production and ratings of plural forms for novel nouns) that is designed to avoid sources of information unavailable to the ED model.The speaker data show high variability, and two suffixes evince 'regular' behavior, appearing more often with phonologically atypical inputs.Encoder-decoder models do generalize the most frequently produced plural class, but do not show human-like variability or 'regular' extension of these other plural markers.We conclude that modern neural models may still struggle with minority-class generalization.
Kate McCurdy, Sharon Goldwater, Adam Lopez
ACL3
2020 Cross-Lingual Topic Prediction For Speech Using Translations
abstract
Given a large amount of unannotated speech in a low-resource language, can we classify the speech utterances by topicƒ We consider this question in the setting where a small amount of speech in the low-resource language is paired with text translations in a high-resource language. We develop an effective cross-lingual topic classifier by training on just 20 hours of translated speech, using a recent model for direct speech-to-text translation. While the translations are poor, they are still good enough to correctly classify the topic of 1-minute speech segments over 70% of the time—a 20% improvement over a majority-class baseline. Such a system could be useful for humanitarian applications like crisis response, where incoming speech in a foreign low-resource language must be quickly assessed for further action.
Sameer Bansal, Herman Kamper, Adam Lopez, Sharon Goldwater
ICASSP3
2019 I Wanna Talk Like You: Speaker Adaptation to Dialogue Style in L2 Practice Conversation
Arabella Sinclair, Rafael Ferreira Leite de Mello, Dragan Gasevic, Christopher G. Lucas, Adam Lopez
AIED (2)5
2019 Tutorbot Corpus: Evidence of Human-Agent Verbal Alignment in Second Language Learner Dialogues
Arabella Sinclair, Kate McCurdy, Christopher G. Lucas, Adam Lopez, Dragan Gasevic
EDM4
2019 Semantic graph parsing with recurrent neural network DAG grammars
abstract
Federico Fancellu, Sorcha Gilroy, Adam Lopez, Mirella Lapata. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Federico Fancellu, Sorcha Gilroy, Adam Lopez, Mirella Lapata
EMNLP/IJCNLP (1)3
2019 A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages
abstract
Clara Vania, Yova Kementchedjhieva, Anders Søgaard, Adam Lopez. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Clara Vania, Yova Kementchedjhieva, Anders Søgaard, Adam Lopez
EMNLP/IJCNLP (1)4
2018 What do character-level models learn about morphology? The case of dependency parsing
abstract
When parsing morphologically-rich languages with neural models, it is beneficial to model input at the character level, and it has been claimed that this is because character-level models learn morphology.We test these claims by comparing character-level models to an oracle with access to explicit morphological analysis on twelve languages with varying morphological typologies.Our results highlight many strengths of character-level models, but also show that they are poor at disambiguating some words, particularly in the face of case syncretism.We then demonstrate that explicitly modeling morphological case improves our best model, showing that characterlevel models can benefit from targeted forms of explicit morphological modeling.
Clara Vania, Andreas Grivas, Adam Lopez
EMNLP3
2018 Low-Resource Speech-to-Text Translation
abstract
Speech-to-text translation has many potential applications for low-resource languages, but the typical approach of cascading speech recognition with machine translation is often impossible, since the transcripts needed to train a speech recognizer are usually not available for low-resource languages. Recent work has found that neural encoder-decoder models can learn to directly translate foreign speech in high-resource scenarios, without the need for intermediate transcription. We investigate whether this approach also works in settings where both data and computation are limited. To make the approach efficient, we make several architectural changes, including a change from character-level to word-level decoding. We find that this choice yields crucial speed improvements that allow us to train with fewer computational resources, yet still performs well on frequent words. We explore models trained on between 20 and 160 hours of data, and find that although models trained on less data have considerably lower BLEU scores, they can still predict words with relatively high precision and recall---around 50% for a model trained on 50 hours of data, versus around 60% for the full 160 hour model. Thus, they may still be useful for some low-resource scenarios.
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, Sharon Goldwater
INTERSPEECH4
2018 A Structured Syntax-Semantics Interface for English-AMR Alignment
abstract
Ida Szubert, Adam Lopez, Nathan Schneider. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Ida Szubert, Adam Lopez, Nathan Schneider 0001
NAACL-HLT2
2018 Does Ability Affect Alignment in Second Language Tutorial Dialogue?
abstract
The role of alignment between interlocutors in second language learning is different to that in fluent conversational dialogue.Learners gain linguistic skill through increased alignment, yet the extent to which they can align will be constrained by their ability.Tutors may use alignment to teach and encourage the student, yet still must push the student and correct their errors, decreasing alignment.To understand how learner ability interacts with alignment, we measure the influence of ability on lexical priming, an indicator of alignment.We find that lexical priming in learner-tutor dialogues differs from that in conversational and task-based dialogues, and we find evidence that alignment increases with ability and with word complexity.
Arabella Sinclair, Adam Lopez, Christopher G. Lucas, Dragan Gasevic
SIGDIAL Conference2
2018 Weighted DAG Automata for Semantic Graphs
abstract
Graphs have a variety of uses in natural language processing, particularly as representations of linguistic meaning. A deficit in this area of research is a formal framework for creating, combining, and using models involving graphs that parallels the frameworks of finite automata for strings and finite tree automata for trees. A possible starting point for such a framework is the formalism of directed acyclic graph (DAG) automata, defined by Kamimura and Slutzki and extended by Quernheim and Knight. In this article, we study the latter in depth, demonstrating several new results, including a practical recognition algorithm that can be used for inference and learning with models defined on DAG automata. We also propose an extension to graphs with unbounded node degree and show that our results carry over to the extended formalism.
David Chiang 0001, Frank Drewes, Daniel Gildea, Adam Lopez, Giorgio Satta
Comput. Linguistics4
2017 From Characters to Words to in Between: Do We Capture Morphology?
abstract
Words can be represented by composing the representations of subword units such as word segments, characters, and/or character n-grams.While such representations are effective and may capture the morphological regularities of words, they have not been systematically compared, and it is not understood how they interact with different morphological typologies.On a language modeling task, we present experiments that systematically vary (1) the basic unit of representation, (2) the composition of these representations, and (3) the morphological typology of the language modeled.Our results extend previous findings that character representations are effective across typologies, and we find that a previously unstudied combination of character trigram representations composed with bi-LSTMs outperforms most others.But we also find room for improvement: none of the character-level models match the predictive accuracy of a model with access to true morphological analyses, even when learned from an order of magnitude more data.
Clara Vania, Adam Lopez
ACL (1)2
2017 Weakly supervised spoken term discovery using cross-lingual side information
abstract
Recent work on unsupervised term discovery (UTD) aims to identify and cluster repeated word-like units from audio alone. These systems are promising for some very low-resource languages where transcribed audio is unavailable, or where no written form of the language exists. However, in some cases it may still be feasible (e.g., through crowdsourcing) to obtain (possibly noisy) text translations of the audio. If so, this information could be used as a source of side information to improve UTD. Here, we present a simple method for rescoring the output of a UTD system using text translations, and test it on a corpus of Spanish audio with English translations. We show that it greatly improves the average precision of the results over a wide range of system configurations and data preprocessing methods.
Sameer Bansal, Herman Kamper, Sharon Goldwater, Adam Lopez
ICASSP4
2016 N-gram language models for massively parallel devices
abstract
For many applications, the query speed of N -gram language models is a computational bottleneck.Although massively parallel hardware like GPUs offer a potential solution to this bottleneck, exploiting this hardware requires a careful rethinking of basic algorithms and data structures.We present the first language model designed for such hardware, using B-trees to maximize data parallelism and minimize memory footprint and latency.Compared with a single-threaded instance of KenLM (Heafield, 2011), a highly optimized CPUbased language model, our GPU implementation produces identical results with a smaller memory footprint and a sixfold increase in throughput on a batch query task.When we saturate both devices, the GPU delivers nearly twice the throughput per hardware dollar even when the CPU implementation uses faster data structures.
Nikolay Bogoychev, Adam Lopez
ACL (1)2
2016 Neural Networks For Negation Scope Detection
abstract
Automatic negation scope detection is a task that has been tackled using different classifiers and heuristics.Most systems are however 1) highly-engineered, 2) English-specific, and 3) only tested on the same genre they were trained on.We start by addressing 1) and 2) using a neural network architecture.Results obtained on data from the *SEM2012 shared task on negation scope detection show that even a simple feed-forward neural network using word-embedding features alone, performs on par with earlier classifiers, with a bi-directional LSTM outperforming all of them.We then address 3) by means of a specially-designed synthetic test set; in doing so, we explore the problem of detecting the negation scope more in depth and show that performance suffers from genre effects and differs with the type of negation considered.
Federico Fancellu, Adam Lopez, Bonnie L. Webber
ACL (1)2
2015 AMRICA: an AMR Inspector for Cross-language Alignments
abstract
Meaning Representation (AMR), an annotation scheme for natural language semantics, has drawn attention for its simplicity and representational power.Because AMR annotations are not designed for human readability, we present AMRICA, a visual aid for exploration of AMR annotations.AMRICA can visualize an AMR or the difference between two AMRs to help users diagnose interannotator disagreement or errors from an AMR parser.AMRICA can also automatically align and visualize the AMRs of a sentence and its translation in a parallel text.We believe AMRICA will simplify and streamline exploratory research on cross-lingual AMR corpora.
Naomi Saphra, Adam Lopez
HLT-NAACL2
2015 Gappy Pattern Matching on GPUs for On-Demand Extraction of Hierarchical Translation Grammars
abstract
Grammars for machine translation can be materialized on demand by finding source phrases in an indexed parallel corpus and extracting their translations. This approach is limited in practical applications by the computational expense of online lookup and extraction. For phrase-based models, recent work has shown that on-demand grammar extraction can be greatly accelerated by parallelization on general purpose graphics processing units (GPUs), but these algorithms do not work for hierarchical models, which require matching patterns that contain gaps. We address this limitation by presenting a novel GPU algorithm for on-demand hierarchical grammar extraction that is at least an order of magnitude faster than a comparable CPU algorithm when processing large batches of sentences. In terms of end-to-end translation, with decoding on the CPU, we increase throughput by roughly two thirds on a standard MT evaluation dataset. The GPU necessary to achieve these improvements increases the cost of a server by about a third. We believe that GPU-based extraction of hierarchical grammars is an attractive proposition, particularly for MT applications that demand high throughput.
Jimmy Lin, Adam Lopez
Trans. Assoc. Comput. Linguistics3
2013 Dirt Cheap Web-Scale Parallel Text from the Common Crawl
Jason Smith 0006, Herve Saint-Amand, Magdalena Plamada, Philipp Koehn, Chris Callison-Burch, Adam Lopez
ACL (1)6
2013 Massively Parallel Suffix Array Queries and On-Demand Phrase Extraction for Statistical Machine Translation Using GPUs
Jimmy Lin, Adam Lopez
HLT-NAACL3
2013 Learning to translate with products of novices: a suite of open-ended challenge problems for teaching MT
abstract
Machine translation (MT) draws from several different disciplines, making it a complex subject to teach. There are excellent pedagogical texts, but problems in MT and current algorithms for solving them are best learned by doing. As a centerpiece of our MT course, we devised a series of open-ended challenges for students in which the goal was to improve performance on carefully constrained instances of four key MT tasks: alignment, decoding, evaluation, and reranking. Students brought a diverse set of techniques to the problems, including some novel solutions which performed remarkably well. A surprising and exciting outcome was that student solutions or their combinations fared competitively on some tasks, demonstrating that even newcomers to the field can help improve the state-of-the-art on hard NLP problems while simultaneously learning a great deal. The problems, baseline code, and results are freely available.
Adam Lopez, Matt Post, Chris Callison-Burch, Jonathan Weese, Juri Ganitkevitch, Narges Ahmidi, Olivia Buzek, Leah Hanson, Beaniesh Jamil, Matthias A. Lee, Ya-Ting Lin, Henry Pao, Fatima Rivera, Leili Shahriyari, Debu Sinha, Adam R. Teichert, Stephen Wampler, Michael Weinberger, Daguang Xu, Lin Yang 0002, Shang Zhao 0002
Trans. Assoc. Comput. Linguistics1
2012 Semi-supervised discriminative language modeling for Turkish ASR
abstract
We present our work on semi-supervised learning of discriminative language models where the negative examples for sentences in a text corpus are generated using confusion models for Turkish at various granularities, specifically, word, sub-word, syllable and phone levels. We experiment with different language models and various sampling strategies to select competing hypotheses for training with a variant of the perceptron algorithm. We find that morph-based confusion models with a sample selection strategy aiming to match the error distribution of the baseline ASR system gives the best performance. We also observe that substituting half of the supervised training examples with those obtained in a semi-supervised manner gives similar results.
Arda Çelebi, Hasim Sak, Erinç Dikici, Murat Saraclar, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Kenji Sagae, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP20
2012 Hallucinated n-best lists for discriminative language modeling
abstract
This paper investigates semi-supervised methods for discriminative language modeling, whereby n-best lists are “hallucinated” for given reference text and are then used for training n-gram language models using the perceptron algorithm. We perform controlled experiments on a very strong baseline English CTS system, comparing three methods for simulating ASR output, and compare the results with training with “real” n-best list output from the baseline recognizer. We find that methods based on extracting phrasal cohorts - similar to methods from machine translation for extracting phrase tables - yielded the largest gains of our three methods, achieving over half of the WER reduction of the fully supervised methods.
Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Damianos Karakos, Sanjeev Khudanpur, Brian Roark, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP17
2012 Continuous space discriminative language modeling
abstract
Discriminative language modeling is a structured classification problem. Log-linear models have been previously used to address this problem. In this paper, the standard dot-product feature representation used in log-linear models is replaced by a non-linear function parameterized by a neural network. Embeddings are learned for each word and features are extracted automatically through the use of convolutional layers. Experimental results show that as a stand-alone model the continuous space model yields significantly lower word error rate (1% absolute), while having a much more compact parameterization (60%-90% smaller). If the baseline scores are combined, our approach performs equally well.
Puyang Xu, Sanjeev Khudanpur, Maider Lehr, Emily Tucker Prud'hommeaux, Nathan Glenn, Damianos Karakos, Brian Roark, Kenji Sagae, Murat Saraclar, Izhak Shafran, Dan Bikel, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
ICASSP17
2012 Deriving conversation-based features from unlabeled speech for discriminative language modeling
Damianos Karakos, Brian Roark, Izhak Shafran, Kenji Sagae, Maider Lehr, Emily Tucker Prud'hommeaux, Puyang Xu, Nathan Glenn, Sanjeev Khudanpur, Murat Saraclar, Dan Bikel, Mark Dredze, Chris Callison-Burch, Yuan Cao 0007, Keith B. Hall, Eva Hasler, Philipp Koehn, Adam Lopez, Matt Post, Darcey Riley
INTERSPEECH18
2011 A Comparison of Loopy Belief Propagation and Dual Decomposition for Integrated CCG Supertagging and Parsing
Michael Auli, Adam Lopez
ACL2
2011 Efficient CCG Parsing: A* versus Adaptive Supertagging
Michael Auli, Adam Lopez
ACL2
2011 Training a Log-Linear Parser with Loss Functions via Softmax-Margin
Michael Auli, Adam Lopez
EMNLP2
2010 Monte Carlo techniques for phrase-based translation
Abhishek Arun, Barry Haddow, Philipp Koehn, Adam Lopez, Chris Dyer, Phil Blunsom
Mach. Transl.4
2009 Monte Carlo inference and maximization for phrase-based translation
Abhishek Arun, Chris Dyer, Barry Haddow, Phil Blunsom, Adam Lopez, Philipp Koehn
CoNLL5
2009 Translation as Weighted Deduction
Adam Lopez
EACL1
2008 Tera-Scale Translation Models via Pattern Matching
Adam Lopez
COLING1
2007 Hierarchical Phrase-Based Translation with Suffix Arrays
Adam Lopez
EMNLP-CoNLL1