Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Myle Ott

dblp:92/9767 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
7since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Machine translation · 29% Efficient and distributed learning · 19% Language models and text generation · 16%
Databases, data mining, and information retrieval
2 papers
Web and social media mining · 72% Data mining · 21% Recommender systems · 7%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
neural machine translation
1.032018
Analyzing Uncertainty in Neural Machine Translation · ICML 2018
Phrase-Based & Neural Unsupervised Machine Translation · EMNLP 2018
Understanding Back-Translation at Scale · EMNLP 2018
Machine learning › Generative modeling
energy-based model
0.922021
Residual Energy-Based Models for Text · J. Mach. Learn. Res. 2021
Residual Energy-Based Models for Text Generation · ICLR 2020
Machine learning › Efficient and distributed learning
distributed training
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training › data parallel training
fully sharded data parallel
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training
large model training
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.612022
Few-shot Learning with Multilingual Generative Language Models · EMNLP 2022
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.612022
Few-shot Learning with Multilingual Generative Language Models · EMNLP 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.612022
Efficient Large Scale Language Modeling with Mixtures of Experts · EMNLP 2022
Natural language and speech › Language models and text generation
multilingual language models
0.612022
Few-shot Learning with Multilingual Generative Language Models · EMNLP 2022
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning
0.412020
Unsupervised Cross-lingual Representation Learning at Scale · ACL 2020
Natural language and speech › Machine translation
machine translation evaluation
0.412020
On The Evaluation of Machine Translation SystemsTrained With Back-Translation · ACL 2020
Natural language and speech › Language models and text generation
masked language modeling
0.412020
Unsupervised Cross-lingual Representation Learning at Scale · ACL 2020
Natural language and speech › Language models and text generation
text generation
0.412020
Residual Energy-Based Models for Text Generation · ICLR 2020
Natural language and speech › Machine translation › controllable machine translation
diverse machine translation
0.412019
Mixture Models for Diverse Machine Translation: Tricks of the Trade · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412019
Mixture Models for Diverse Machine Translation: Tricks of the Trade · ICML 2019
Natural language and speech › Machine translation
low-resource machine translation
0.412019
The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English · EMNLP/IJCNLP (1) 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.412019
Mixture Models for Diverse Machine Translation: Tricks of the Trade · ICML 2019
Natural language and speech › Machine translation › monolingual data augmentation
back-translation
0.312018
Understanding Back-Translation at Scale · EMNLP 2018
Machine learning › Deep learning architectures and training
data augmentation
0.312018
Understanding Back-Translation at Scale · EMNLP 2018
Machine learning › Trustworthy machine learning › calibration
model calibration
0.312018
Analyzing Uncertainty in Neural Machine Translation · ICML 2018
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation
0.312018
Phrase-Based & Neural Unsupervised Machine Translation · EMNLP 2018
Natural language and speech › Machine translation
synthetic parallel data
0.312018
Understanding Back-Translation at Scale · EMNLP 2018
Natural language and speech › Machine translation
unsupervised machine translation
0.312018
Phrase-Based & Neural Unsupervised Machine Translation · EMNLP 2018
Natural language and speech › Information extraction and text analysis › misinformation detection
deceptive review detection
0.322014
Towards a General Rule for Identifying Deceptive Opinion Spam · ACL (1) 2014
Finding Deceptive Opinion Spam by Any Stretch of the Imagination · ACL 2011
Web and social media mining › online review analysis
fake review detection
0.322013
Identifying Manipulated Offerings on Review Portals · EMNLP 2013
Estimating the prevalence of deception in online review communities · WWW 2012
Operating systems › resource management
memory management
0.212023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Natural language and speech › Language models and text generation › neural language model
autoregressive language model
0.112021
Residual Energy-Based Models for Text · J. Mach. Learn. Res. 2021
Web and social media mining
online review analysis
0.112012
Estimating the prevalence of deception in online review communities · WWW 2012
Data mining › statistical analysis › statistical estimation
quantification
0.112012
Estimating the prevalence of deception in online review communities · WWW 2012
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search
0.112018
Analyzing Uncertainty in Neural Machine Translation · ICML 2018

Methods — techniques the papers use, named apart from their topics

sharding · 1.3back-translation · 0.8data-parallel training · 0.7data parallel training · 0.7sparse routing · 0.6mixture of experts · 0.6generative language model · 0.6few-shot prompting · 0.6perplexity evaluation · 0.5energy-based model · 0.5discriminative training · 0.5semi-supervised manifold ranking · 0.2statistical analysis · 0.1
YearPublicationVenuePosition
2023 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
abstract
It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.
Yanli Zhao, Andrew Gu, Rohan Varma, Chien-Chin Huang, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews
Proc. VLDB Endow.9
2022 Efficient Large Scale Language Modeling with Mixtures of Experts
abstract
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov
EMNLP5
2022 Few-shot Learning with Multilingual Generative Language Models
abstract
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal 0001, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, Veselin Stoyanov, Xian Li 0003
EMNLP7
2021 Analyzing the Forgetting Problem in Pretrain-Finetuning of Open-domain Dialogue Response Models
abstract
Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, Fuchun Peng. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Tianxing He, Kyunghyun Cho, Myle Ott, Bing Liu 0024, James R. Glass, Fuchun Peng
EACL4
2021 Recipes for Building an Open-Domain Chatbot
abstract
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Stephen Roller, Emily Dinan, Naman Goyal 0001, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu 0014, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston
EACL8
2021 The Source-Target Domain Mismatch Problem in Machine Translation
abstract
Jiajun Shen, Peng-Jen Chen, Matthew Le, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Peng-Jen Chen, Matt Le 0001, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc'Aurelio Ranzato
EACL6
2021 Residual Energy-Based Models for Text
abstract
Current large-scale auto-regressive language models display impressive fluency and can generate convincing text. In this work we start by asking the question: Can the generations of these models be reliably distinguished from real text by statistical discriminators? We find experimentally that the answer is affirmative when we have access to the training data for the model, and guardedly affirmative even if we do not. This suggests that the auto-regressive models can be improved by incorporating the (globally normalized) discriminators into the generative process. We give a formalism for this using the Energy-Based Model framework, and show that it indeed improves the results of the generative models, measured both in terms of perplexity and in terms of human evaluation.
Anton Bakhtin, Yuntian Deng, Sam Gross, Myle Ott, Marc'Aurelio Ranzato, Arthur Szlam
J. Mach. Learn. Res.4
2020 Unsupervised Cross-lingual Representation Learning at Scale
abstract
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Alexis Conneau, Kartikay Khandelwal, Naman Goyal 0001, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, Veselin Stoyanov
ACL8
2020 On The Evaluation of Machine Translation SystemsTrained With Back-Translation
abstract
Back-translation is a widely used data augmentation technique which leverages target monolingual data.However, its effectiveness has been challenged since automatic metrics such as BLEU only show significant improvements for test examples where the source itself is a translation, or translationese.This is believed to be due to translationese inputs better matching the back-translated training data.In this work, we show that this conjecture is not empirically supported and that backtranslation improves translation quality of both naturally occurring text as well as translationese according to professional human translators.We provide empirical evidence to support the view that back-translation is preferred by humans because it produces more fluent outputs.BLEU cannot capture human preferences because references are translationese when source sentences are natural text.We recommend complementing BLEU with a language model score to measure fluency.
Sergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael Auli
ACL2
2020 Residual Energy-Based Models for Text Generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, Marc'Aurelio Ranzato
ICLR3
2019 The FLORES Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English
abstract
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc’Aurelio Ranzato. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino 0001, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, Marc'Aurelio Ranzato
EMNLP/IJCNLP (1)3
2019 Mixture Models for Diverse Machine Translation: Tricks of the Trade
abstract
Mixture models trained via EM are among the simplest, most widely used and well understood latent variable models in the machine learning literature. Surprisingly, these models have been hardly explored in text generation applications such as machine translation. In principle, they provide a latent variable to control generation and produce a diverse set of hypotheses. In practice, however, mixture models are prone to degeneracies—often only one component gets trained or the latent variable is simply ignored. We find that disabling dropout noise in responsibility computation is critical to successful training. In addition, the design choices of parameterization, prior distribution, hard versus soft EM and online versus offline assignment can dramatically affect model performance. We develop an evaluation protocol to assess both quality and diversity of generations against multiple references, and provide an extensive empirical study of several mixture model variants. Our analysis shows that certain types of mixture models are more robust and offer the best trade-off between translation quality and diversity compared to variational models and diverse decoding approaches.\footnote{Code to reproduce the results in this paper is available at \url{https://github.com/pytorch/fairseq}}
Tianxiao Shen, Myle Ott, Michael Auli, Marc'Aurelio Ranzato
ICML2
2018 Understanding Back-Translation at Scale
abstract
An effective method to improve neural machine translation with monolingual data is to augment the parallel training corpus with back-translations of target language sentences.This work broadens the understanding of back-translation and investigates a number of methods to generate synthetic source sentences.We find that in all but resource poor settings back-translations obtained via sampling or noised beam outputs are most effective.Our analysis shows that sampling or noisy synthetic data gives a much stronger training signal than data generated by beam or greedy search.We also compare how synthetic data compares to genuine bitext and study various domain effects.Finally, we scale to hundreds of millions of monolingual sentences and achieve a new state of the art of 35 BLEU on the WMT'14 English-German test set.
Sergey Edunov, Myle Ott, Michael Auli, David Grangier
EMNLP2
2018 Phrase-Based & Neural Unsupervised Machine Translation
abstract
Machine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentences, which hinders their applicability to the majority of language pairs.This work investigates how to learn to translate when having access to only large monolingual corpora in each language.We propose two model variants, a neural and a phrase-based model.Both versions leverage a careful initialization of the parameters, the denoising effect of language models and automatic generation of parallel data by iterative back-translation.These models are significantly better than methods from the literature, while being simpler and having fewer hyper-parameters.On the widely used WMT'14 English-French and WMT'16 German-English benchmarks, our models respectively obtain 28.1 and 25.2 BLEU points without using a single parallel sentence, outperforming the state of the art by more than 11 BLEU points.On low-resource languages like English-Urdu and English-Romanian, our methods achieve even better results than semisupervised and supervised approaches leveraging the paucity of available bitexts.Our code for NMT and PBSMT is publicly available.
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, Marc'Aurelio Ranzato
EMNLP2
2018 Analyzing Uncertainty in Neural Machine Translation
abstract
Machine translation is a popular test bed for research in neural sequence-to-sequence models but despite much recent research, there is still a lack of understanding of these models. Practitioners report performance degradation with large beams, the under-estimation of rare words and a lack of diversity in the final translations. Our study relates some of these issues to the inherent uncertainty of the task, due to the existence of multiple valid translations for a single source sentence, and to the extrinsic uncertainty caused by noisy training data. We propose tools and metrics to assess how uncertainty in the data is captured by the model distribution and how it affects search strategies that generate translations. Our results show that search works remarkably well but that the models tend to spread too much probability mass over the hypothesis space. Next, we propose tools to assess model calibration and show how to easily fix some shortcomings of current models. We release both code and multiple human reference translations for two popular benchmarks.
Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato
ICML1
2018 Classical Structured Prediction Losses for Sequence to Sequence Learning
abstract
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc’Aurelio Ranzato. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Sergey Edunov, Myle Ott, Michael Auli, David Grangier, Marc'Aurelio Ranzato
NAACL-HLT2
2014 Towards a General Rule for Identifying Deceptive Opinion Spam
abstract
Consumers' purchase decisions are increasingly influenced by user-generated online reviews.Accordingly, there has been growing concern about the potential for posting deceptive opinion spamfictitious reviews that have been deliberately written to sound authentic, to deceive the reader.In this paper, we explore generalized approaches for identifying online deceptive opinion spam based on a new gold standard dataset, which is comprised of data from three different domains (i.e.Hotel, Restaurant, Doctor), each of which contains three types of reviews, i.e. customer generated truthful reviews, Turker generated deceptive reviews and employee (domain-expert) generated deceptive reviews.Our approach tries to capture the general difference of language usage between deceptive and truthful reviews, which we hope will help customers when making purchase decisions and review portal operators, such as TripAdvisor or Yelp, investigate possible fraudulent activity on their sites.1
Jiwei Li 0001, Myle Ott, Claire Cardie, Eduard H. Hovy
ACL (1)2
2013 Identifying Manipulated Offerings on Review Portals
abstract
Recent work has developed supervised methods for detecting deceptive opinion spamfake reviews written to sound authentic and deliberately mislead readers.And whereas past work has focused on identifying individual fake reviews, this paper aims to identify offerings (e.g., hotels) that contain fake reviews.We introduce a semi-supervised manifold ranking algorithm for this task, which relies on a small set of labeled individual reviews for training.Then, in the absence of gold standard labels (at an offering level), we introduce a novel evaluation procedure that ranks artificial instances of real offerings, where each artificial offering contains a known number of injected deceptive reviews.Experiments on a novel dataset of hotel reviews show that the proposed method outperforms state-of-art learning baselines.
Jiwei Li 0001, Myle Ott, Claire Cardie
EMNLP2
2013 Properties, Prediction, and Prevalence of Useful User-Generated Comments for Descriptive Annotation of Social Media Objects
Elaheh Momeni, Claire Cardie, Myle Ott
ICWSM3
2013 Negative Deceptive Opinion Spam
Myle Ott, Claire Cardie, Jeffrey T. Hancock
HLT-NAACL1
2012 Estimating the prevalence of deception in online review communities
abstract
Consumers' purchase decisions are increasingly influenced by user-generated online reviews. Accordingly, there has been growing concern about the potential for posting deceptive opinion spam---fictitious reviews that have been deliberately written to sound authentic, to deceive the reader. But while this practice has received considerable public attention and concern, relatively little is known about the actual prevalence, or rate, of deception in online review communities, and less still about the factors that influence it.
Myle Ott, Claire Cardie, Jeffrey T. Hancock
WWW1
2011 Finding Deceptive Opinion Spam by Any Stretch of the Imagination
Myle Ott, Yejin Choi 0001, Claire Cardie, Jeffrey T. Hancock
ACL1