VLDB 2026 Research / reviewers in the wild / expert
Kyunghyun Cho
dblp:41/9736 · also KyungHyun Cho
· DBLP profile ↗
170ranked-venue papers
13as first author
64since 2021 · last 2026
0000-0003-1669-3211ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 156 · 12 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain TextsabstractThe history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored in NLP due to a lack of accessible historical corpora. To address this gap, we introduce the Open Korean Historical Corpus, a large-scale, openly licensed dataset spanning 1,300 years and 6 languages, as well as under-represented writing systems like Korean-style Sinitic (Idu) and Hanja-Hangul mixed script. This corpus contains 17.7 million documents and 5.1 billion tokens from 19 sources, ranging from the 7th century to 2025. We leverage this resource to quantitatively analyze major linguistic shifts: (1) Idu usage peaked in the 1860s before declining sharply; (2) the transition from Hanja to Hangul was a rapid transformation starting around 1890; and (3) North Korea's lexical divergence causes modern tokenizers to produce up to 51 times higher out-of-vocabulary rates. This work provides a foundational resource for quantitative diachronic analysis by capturing the history of the Korean language. Moreover, it can serve as a pre-training corpus for large language models, potentially improving their understanding of Sino-Korean vocabulary in modern Hangul as well as archaic writing systems. Seyoung Song 0001, Nawon Kim, Songeun Chae, Kiwoong Park, Jiho Jin, Haneul Yoo, Kyunghyun Cho, Alice Oh |
LREC | 7 |
| 2025 | Semiparametric conformal predictionabstractMany risk-sensitive applications require well-calibrated prediction sets over multiple, potentially correlated target variables, for which the prediction algorithm may report correlated errors. In this work, we aim to construct the conformal prediction set accounting for the joint correlation structure of the vector-valued non-conformity scores. Drawing from the rich literature on multivariate quantiles and semiparametric statistics, we propose an algorithm to estimate the $1-\alpha$ quantile of the scores, where $\alpha$ is the user-specified miscoverage rate. In particular, we flexibly estimate the joint cumulative distribution function (CDF) of the scores using nonparametric vine copulas and improve the asymptotic efficiency of the quantile estimate using its influence function. The vine decomposition allows our method to scale well to a large number of targets. As well as guaranteeing asymptotically exact coverage, our method yields desired coverage and competitive efficiency on a range of real-world regression problems, including those with missing-at-random labels in the calibration set. Ji Won Park, Kyunghyun Cho |
AISTATS | 2 |
| 2025 | Language Models as Causal Effect GeneratorsabstractIn this work, we present sequence-driven structural causal models (SD-SCMs), a framework for specifying causal models with user-defined structure and language-model-defined mechanisms.We characterize how an SD-SCM enables sampling from observational, interventional, and counterfactual distributions according to the desired causal structure.We then leverage this procedure to propose a new type of benchmark for causal inference methods, generating individual-level counterfactual data to test treatment effect estimation.We create an example benchmark consisting of thousands of datasets, and test a suite of popular estimation methods for average, conditional average, and individual treatment effect estimation.We find under this benchmark that (1) causal methods outperform non-causal methods and that (2) even state-of-the-art methods struggle with individualized effect estimation, suggesting this benchmark captures some inherent difficulties in causal estimation.Apart from generating data, this same technique can underpin the auditing of language models for (un)desirable causal effects, such as misinformation or discrimination.We believe SD-SCMs can serve as a useful tool in any application that would benefit from sequential data with controllable causal structure. Lucius Bynum, Kyunghyun Cho |
EMNLP | 2 |
| 2025 | Following Length Constraints in InstructionsabstractAligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral. Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014 |
EMNLP | 4 |
| 2025 | Aioli: A Unified Optimization Framework for Language Model Data MixingabstractLanguage model performance depends on identifying the optimal mixture of data groups to train on (e.g., law, code, math). Prior work has proposed a diverse set of methods to efficiently learn mixture proportions, ranging from fitting regression models over training runs to dynamically updating proportions throughout training. Surprisingly, we find that no existing method consistently outperforms a simple stratified sampling baseline in terms of average test perplexity. To understand this inconsistency, we unify existing methods into a standard framework, showing they are equivalent to solving a common optimization problem: minimize average loss subject to a method-specific mixing law---an implicit assumption on the relationship between loss and mixture proportions. This framework suggests that measuring the fidelity of a method's mixing law can offer insights into its performance. Empirically, we find that existing methods set their mixing law parameters inaccurately, resulting in the inconsistent mixing performance we observe. Using this insight, we derive a new online method named Aioli, which directly estimates the mixing law parameters throughout training and uses them to dynamically adjust proportions. Empirically, Aioli outperforms stratified sampling on 6 out of 6 datasets by an average of 0.27 test perplexity points, whereas existing methods fail to consistently beat stratified sampling, doing up to 6.9 points worse. Moreover, in a practical setting where proportions are learned on shorter runs due to computational constraints, Aioli can dynamically adjust these proportions over the full training run, consistently improving performance over existing methods by up to 12.012 test perplexity points. Mayee F. Chen, Michael Y. Hu, Nicholas Lourie, Kyunghyun Cho, Christopher Ré |
ICLR | 4 |
| 2025 | Concept Bottleneck Language Models For Protein DesignabstractWe introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our architecture offers three key benefits: i) Control: We can intervene on concept values to precisely control the properties of generated proteins, achieving a 3$\times$ larger change in desired concept values compared to baselines. ii) Interpretability: A linear mapping between concept values and predicted tokens allows transparent analysis of the model's decision-making process. iii) Debugging: This transparency facilitates easy debugging of trained models. Our models achieve pre-training perplexity and downstream task performance comparable to traditional masked protein language models, demonstrating that interpretability does not compromise performance. While adaptable to any language model, we focus on masked protein language models due to their importance in drug discovery and the ability to validate our model's capabilities through real-world experiments and expert knowledge. We scale our CB-pLM from 24 million to 3 billion parameters, making them the largest Concept Bottleneck Models trained and the first capable of generative language modeling. Aya Abdelsalam Ismail, Tuomas P. Oikarinen, Amy Wang, Julius Adebayo, Samuel Stanton, Héctor Corrada Bravo, Kyunghyun Cho, Nathan C. Frey |
ICLR | 7 |
| 2025 | X-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs
Vlad Sobal, Mark Ibrahim, Randall Balestriero, Vivien Cabannes, Diane Bouchacourt, Pietro Astolfi, Kyunghyun Cho, Yann LeCun |
ICLR | 7 |
| 2025 | Generalists vs. Specialists: Evaluating LLMs on Highly-Constrained Biophysical Sequence Optimization TasksabstractAlthough large language models (LLMs) have shown promise in biomolecule optimization problems, they incur heavy computational costs and struggle to satisfy precise constraints. On the other hand, specialized solvers like LaMBO-2 offer efficiency and fine-grained control but require more domain expertise.
Comparing these approaches is challenging due to expensive laboratory validation and inadequate synthetic benchmarks.
We address this by introducing Ehrlich functions, a synthetic test suite that captures the geometric structure of biophysical sequence optimization problems.
With prompting alone, off-the-shelf LLMs struggle to optimize Ehrlich functions.
In response, we propose LLOME (Language Model Optimization with Margin Expectation), a bilevel optimization routine for online black-box optimization.
When combined with a novel preference learning loss, we find LLOME can not only learn to solve some Ehrlich functions, but
can even perform as well as or better than LaMBO-2 on moderately difficult Ehrlich variants.
However, LLMs also exhibit some likelihood-reward miscalibration and struggle without explicit rewards.
Our results indicate LLMs can occasionally provide significant benefits, but specialized solvers are still competitive and incur less overhead. Angelica Chen, Samuel Stanton, Frances Ding, Robert G. Alberstein, Andrew M. Watkins, Richard Bonneau, Vladimir Gligorijevic, Kyunghyun Cho, Nathan C. Frey |
ICML | 8 |
| 2025 | Why Knowledge Distillation Works in Generative Models: A Minimal Working ExplanationabstractKnowledge distillation (KD) is a core component in the training and deployment of modern generative models, particularly large language models (LLMs). While its empirical benefits are well documented---enabling smaller student models to emulate the performance of much larger teachers---the underlying mechanisms by which KD improves generative quality remain poorly understood. In this work, we present a minimal working explanation of KD in generative modeling. Using a controlled simulation with mixtures of Gaussians, we demonstrate that distillation induces a trade-off between precision and recall in the student model. As the teacher distribution becomes more selective, the student concentrates more probability mass on high-likelihood regions at the expense of coverage, which is a behavior modulated by a single entropy-controlling parameter. We then validate this effect in a large-scale language modeling setup using the SmolLM2 family of models. Empirical results reveal the same precision-recall dynamics observed in simulation, where precision corresponds to sample quality and recall to distributional coverage. This precision-recall trade-off in LLMs is found to be especially beneficial in scenarios where sample quality is more important than diversity, such as instruction tuning or downstream generation. Our analysis provides a simple and general explanation for the effectiveness of KD in generative modeling. Sungmin Cha, Kyunghyun Cho |
NeurIPS | 2 |
| 2025 | Test Time Scaling for Neural ProcessesabstractUncertainty-aware meta-learning aims not only for rapid adaptation to new tasks but also for reliable uncertainty estimation under limited supervision. Neural Processes (NPs) offer a flexible solution by learning implicit stochastic processes directly from data, often using a global latent variable to capture functional uncertainty. However, we empirically find that variational posteriors for this global latent variable are frequently miscalibrated, limiting both predictive accuracy and the reliability of uncertainty estimates. To address this issue, we propose Test Time Scaling for Neural Processes (TTSNPs), a sequential inference framework based on Sequential Monte Carlo Sampler (SMCS) that refines latent samples at test time without modifying the pre-trained NP model. TTSNPs iteratively transform variational samples into better approximations of the true posterior using neural transition kernels, significantly improving both prediction quality and uncertainty calibration. This makes NPs more robust and trustworthy, extending applicability to various scenarios requiring well-calibrated uncertainty estimates. Hyungi Lee, Moonseok Choi, Hyunsu Kim, Kyunghyun Cho, Rajesh Ranganath |
NeurIPS | 4 |
| 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference schemeabstractConditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability. Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho |
NeurIPS | 5 |
| 2025 | Efficient semantic uncertainty quantification in language models via diversity-steered samplingabstractAccurately estimating *semantic* aleatoric and epistemic uncertainties in large language models (LLMs) is particularly challenging in free-form question answering (QA), where obtaining stable estimates often requires many expensive generations. We introduce a **diversity-steered sampler** that discourages semantically redundant outputs during decoding, covers both autoregressive and masked diffusion paradigms, and yields substantial sample-efficiency gains. The key idea is to inject a continuous semantic-similarity penalty into the model’s proposal distribution using a natural language inference (NLI) model lightly finetuned on partial prefixes or intermediate diffusion states. We debias downstream uncertainty estimates with importance reweighting and shrink their variance with control variates. Across four QA benchmarks, our method matches or surpasses baselines while covering more semantic clusters with the same number of samples. Being modular and requiring no gradient access to the base LLM, the framework promises to serve as a drop-in enhancement for uncertainty estimation in risk-sensitive model deployments. Ji Won Park, Kyunghyun Cho |
NeurIPS | 2 |
| 2025 | AION-1: Omnimodal Foundation Model for Astronomical SciencesabstractWhile foundation models have shown promise across a variety of fields, astronomy lacks a unified framework for joint modeling across its highly diverse data modalities. In this paper, we present AION-1, the first large-scale multimodal foundation family of models for astronomy. AION-1 enables arbitrary transformations between heterogeneous data types using a two-stage architecture: modality-specific tokenization followed by transformer-based masked modeling of cross-modal token sequences. Trained on over 200M astronomical objects, AION-1 demonstrates strong performance across regression, classification, generation, and object retrieval tasks. Beyond astronomy, AION-1 provides a scalable blueprint for multimodal scientific foundation models that can seamlessly integrate heterogeneous combinations of real-world observations. Our model release is entirely open source, including the dataset, training script, and weights. Liam Holden Parker, François Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Pierre Cornette, Keiya Hirashima, Géraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Kyunghyun Cho, Miles D. Cranmer, Shirley Ho |
NeurIPS | 23 |
| 2025 | Generative property enhancer: implicit guided generation through conditional density estimationabstractGenerative modeling is increasingly important for data-driven computational design. Conventional approaches pair a generative model with a discriminative model to select or guide samples toward optimized designs. Yet discriminative models often struggle in data-scarce settings, common in scientific applications, and are unreliable in the tails of the distribution where optimal designs typically lie. We introduce generative property enhancer (GPE), an approach that implicitly guides generation by matching samples with lower property values to higher-value ones. Formulated as conditional density estimation, our framework defines a target distribution with improved properties, compelling the generative model to produce enhanced, diverse designs without auxiliary predictors. GPE is simple, scalable, end-to-end, modality-agnostic, and integrates seamlessly with diverse generative model architectures and losses. We demonstrate competitive empirical results on standard _in silico_ offline (non-sequential) protein fitness optimization benchmarks. Finally, we propose iterative training on a combination of limited real data and self-generated synthetic data, enabling extrapolation beyond the original property ranges. Pedro O. Pinheiro, Pan Kessel, Aya Abdelsalam Ismail, Sai Pooja Mahajan, Kyunghyun Cho, Saeed Saremi, Natasa Tagasovska |
NeurIPS | 5 |
| 2025 | Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsabstractA long-standing goal in AI is to develop agents capable of solving diverse tasks across a range of environments, including those never seen during training. Two dominant paradigms address this challenge: (i) reinforcement learning (RL), which learns policies via trial and error, and (ii) optimal control, which plans actions using a known or learned dynamics model. However, their comparative strengths in the offline setting—where agents must learn from reward-free trajectories—remain underexplored. In this work, we systematically evaluate RL and control-based methods on a suite of navigation tasks, using offline datasets of varying quality. On the RL side, we consider goal-conditioned and zero-shot methods. On the control side, we train a latent dynamics model using the Joint Embedding Predictive Architecture (JEPA) and employ it for planning. We investigate how factors such as data diversity, trajectory quality, and environment variability influence the performance of these approaches. Our results show that model-free RL benefits most from large amounts of high-quality data, whereas model-based planning generalizes better to unseen layouts and is more data-efficient, while achieving trajectory stitching performance comparable to leading model-free methods. Notably, planning with a latent dynamics model proves to be a strong approach for handling suboptimal offline data and adapting to diverse environments. Uladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero, Tim G. J. Rudner, Yann LeCun |
NeurIPS | 3 |
| 2025 | NaturalReasoning: Reasoning in the Wild with 2.8M Challenging QuestionsabstractScaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding. Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003 |
NeurIPS | 8 |
| 2024 | System-Level Natural Language FeedbackabstractNatural language (NL) feedback offers rich insights into user experience.While existing studies focus on an instance-level approach, where feedback is used to refine specific examples, we introduce a framework for system-level use of NL feedback.We show how to use feedback to formalize system-level design decisions in a human-in-the-loop-process -in order to produce better models.In particular this is done through: (i) metric design for tasks; and (ii) language model prompt design for refining model responses.We conduct two case studies of this approach for improving search query and dialog response generation, demonstrating the effectiveness of system-level feedback.We show the combination of system-level and instancelevel feedback brings further gains, and that human written instance-level feedback results in more grounded refinements than GPT-3.5 written ones, underlying the importance of human feedback for building systems.We release our code and data at https://github.com/ yyy-Apple/Sys-NL-Feedback. Weizhe Yuan, Kyunghyun Cho, Jason Weston |
EACL (1) | 2 |
| 2024 | Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsabstractMost interpretability research in NLP focuses on understanding the behavior and features of a fully trained model. However, certain insights into model behavior may only be accessible by observing the trajectory of the training process. We present a case study of syntax acquisition in masked language models (MLMs) that demonstrates how analyzing the evolution of interpretable artifacts throughout training deepens our understanding of emergent behavior. In particular, we study Syntactic Attention Structure (SAS), a naturally emerging property of MLMs wherein specific Transformer heads tend to focus on specific syntactic relations. We identify a brief window in pretraining when models abruptly acquire SAS, concurrent with a steep drop in loss. This breakthrough precipitates the subsequent acquisition of linguistic capabilities. We then examine the causal role of SAS by manipulating SAS during training, and demonstrate that SAS is necessary for the development of grammatical capabilities. We further find that SAS competes with other beneficial traits during training, and that briefly suppressing SAS improves model quality. These findings offer an interpretation of a real-world example of both simplicity bias and breakthrough training dynamics. Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, Naomi Saphra |
ICLR | 3 |
| 2024 | Protein Discovery with Discrete Walk-Jump SamplingabstractWe resolve difficulties in training and sampling from a discrete generative model by learning a smoothed energy function, sampling from the smoothed data manifold with Langevin Markov chain Monte Carlo (MCMC), and projecting back to the true data manifold with one-step denoising. Our $\textit{Discrete Walk-Jump Sampling}$ formalism combines the contrastive divergence training of an energy-based model and improved sample quality of a score-based model, while simplifying training and sampling by requiring only a single noise level. We evaluate the robustness of our approach on generative modeling of antibody proteins and introduce the $\textit{distributional conformity score}$ to benchmark protein generative models. By optimizing and sampling from our models for the proposed distributional conformity score, 97-100\% of generated samples are successfully expressed and purified and 70\% of functional designs show equal or improved binding affinity compared to known functional antibodies on the first attempt in a single round of laboratory experiments. We also report the first demonstration of long-run fast-mixing MCMC chains where diverse antibody protein classes are visited in a single MCMC chain. Nathan C. Frey, Daniel Berenberg, Karina Zadorozhny, Joseph Kleinhenz, Julien Lafrance-Vanasse, Isidro Hötzel, Yan Wu 0027, Stephen Ra, Richard Bonneau, Kyunghyun Cho, Andreas Loukas, Vladimir Gligorijevic, Saeed Saremi |
ICLR | 10 |
| 2024 | Concept Bottleneck Generative ModelsabstractWe introduce a generative model with an intrinsically interpretable layer---a concept bottleneck layer---that constrains the model to encode human-understandable concepts. The concept bottleneck layer partitions the generative model into three parts: the pre-concept bottleneck portion, the CB layer, and the post-concept bottleneck portion. To train CB generative models, we complement the traditional task-based loss function for training generative models with a concept loss and an orthogonality loss. The CB layer and these loss terms are model agnostic, which we demonstrate by applying the CB layer to three different families of generative models: generative adversarial networks, variational autoencoders, and diffusion models. On multiple datasets across different types of generative models, steering a generative model, with the CB layer, outperforms all baselines---in some cases, it is \textit{10 times} more effective. In addition, we show how the CB layer can be used to interpret the output of the generative model and debug the model during or post training. Aya Abdelsalam Ismail, Julius Adebayo, Héctor Corrada Bravo, Stephen Ra, Kyunghyun Cho |
ICLR | 5 |
| 2024 | Regularizing with Pseudo-Negatives for Continual Self-Supervised LearningabstractWe introduce a novel Pseudo-Negative Regularization (PNR) framework for effective continual self-supervised learning (CSSL). Our PNR leverages pseudo-negatives obtained through model-based augmentation in a way that newly learned representations may not contradict what has been learned in the past. Specifically, for the InfoNCE-based contrastive learning methods, we define symmetric pseudo-negatives obtained from current and previous models and use them in both main and regularization loss terms. Furthermore, we extend this idea to non-contrastive learning methods which do not inherently rely on negatives. For these methods, a pseudo-negative is defined as the output from the previous model for a differently augmented version of the anchor sample and is asymmetrically applied to the regularization term. Extensive experimental results demonstrate that our PNR framework achieves state-of-the-art performance in representation learning during CSSL by effectively balancing the trade-off between plasticity and stability. Sungmin Cha, Kyunghyun Cho, Taesup Moon |
ICML | 2 |
| 2024 | Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial OptimizationabstractActive learning is increasingly adopted for expensive multi-objective combinatorial optimization problems, but it involves a challenging subset selection problem, optimizing the batch acquisition score that quantifies the goodness of a batch for evaluation. Due to the excessively large search space of the subset selection problem, prior methods optimize the batch acquisition on the latent space, which has discrepancies with the actual space, or optimize individual acquisition scores without considering the dependencies among candidates in a batch instead of directly optimizing the batch acquisition. To manage the vast search space, a simple and effective approach is the greedy method, which decomposes the problem into smaller subproblems, yet it has difficulty in parallelization since each subproblem depends on the outcome from the previous ones. To this end, we introduce a novel greedy-style subset selection algorithm that optimizes batch acquisition directly on the combinatorial space by sequential greedy sampling from the greedy policy, specifically trained to address all greedy subproblems concurrently. Notably, our experiments on the red fluorescent proteins design task show that our proposed method achieves the baseline performance in 1.69x fewer queries, demonstrating its efficiency. Deokjae Lee, Hyun Oh Song, Kyunghyun Cho |
ICML | 3 |
| 2024 | BOtied: Multi-objective Bayesian optimization with tied multivariate ranksabstractMany scientific and industrial applications require the joint optimization of multiple, potentially competing objectives. Multi-objective Bayesian optimization (MOBO) is a sample-efficient framework for identifying Pareto-optimal solutions. At the heart of MOBO is the acquisition function, which determines the next candidate to evaluate by navigating the best compromises among the objectives. Acquisition functions that rely on integrating over the objective space scale poorly to a large number of objectives. In this paper, we show a natural connection between the non-dominated solutions and the highest multivariate rank, which coincides with the extreme level line of the joint cumulative distribution function (CDF). Motivated by this link, we propose the CDF indicator, a Pareto-compliant metric for evaluating the quality of approximate Pareto sets, that can complement the popular hypervolume indicator. We then introduce an acquisition function based on the CDF indicator, called BOtied. BOtied can be implemented efficiently with copulas, a statistical tool for modeling complex, high-dimensional distributions. Our experiments on a variety of synthetic and real-world experiments demonstrate that BOtied outperforms state-of-the-art MOBO algorithms while being computationally efficient for many objectives. Ji Won Park, Natasa Tagasovska, Michael Maser, Stephen Ra, Kyunghyun Cho |
ICML | 5 |
| 2024 | Self-Rewarding Language ModelsabstractWe posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these reward models require additional human preferences data to further improve.In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training, not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes. Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 0003, Sainbayar Sukhbaatar, Jing Xu 0014, Jason Weston |
ICML | 3 |
| 2024 | Show Your Work with Confidence: Confidence Bands for Tuning CurvesabstractNicholas Lourie, Kyunghyun Cho, He He. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nicholas Lourie, Kyunghyun Cho, He He 0001 |
NAACL-HLT | 2 |
| 2024 | First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language ModelsabstractNaomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Naomi Saphra, Eve Fleisig, Kyunghyun Cho, Adam Lopez |
NAACL-HLT | 3 |
| 2024 | Preference Learning Algorithms Do Not Learn Preference RankingsabstractPreference learning algorithms (e.g., RLHF and DPO) are frequently used to steer LLMs to produce generations that are more preferred by humans, but our understanding of their inner workings is still limited. In this work, we study the conventional wisdom that preference learning trains models to assign higher likelihoods to more preferred outputs than less preferred outputs, measured via *ranking accuracy*.
Surprisingly, we find that most state-of-the-art preference-tuned models achieve a ranking accuracy of less than 60% on common preference datasets. We furthermore derive the *idealized ranking accuracy* that a preference-tuned LLM would achieve if it optimized the DPO or RLHF objective perfectly. We demonstrate that existing models exhibit a significant *alignment gap* -- *i.e.*, a gap between the observed and idealized ranking accuracies.
We attribute this discrepancy to the DPO objective, which is empirically and theoretically ill-suited to correct even mild ranking errors in the reference model, and derive a simple and efficient formula for quantifying the difficulty of learning a given preference datapoint.
Finally, we demonstrate that ranking accuracy strongly correlates with the empirically popular win rate metric when the model is close to the reference model used in the objective, shedding further light on the differences between on-policy (e.g., RLHF) and off-policy (e.g., DPO) preference learning algorithms. Angelica Chen, Sadhika Malladi, Lily H. Zhang, Xinyi Chen 0001, Qiuyi Zhang 0001, Rajesh Ranganath, Kyunghyun Cho |
NeurIPS | 7 |
| 2024 | Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal LearningabstractSupervised multi-modal learning involves mapping multiple modalities to a target label. Previous studies in this field have concentrated on capturing in isolation either the inter-modality dependencies (the relationships between different modalities and the label) or the intra-modality dependencies (the relationships within a single modality and the label). We argue that these conventional approaches that rely solely on either inter- or intra-modality dependencies may not be optimal in general. We view the multi-modal learning problem from the lens of generative models where we consider the target as a source of multiple modalities and the interaction between them. Towards that end, we propose inter- \& intra-modality modeling (I2M2) framework, which captures and integrates both the inter- and intra-modality dependencies, leading to more accurate predictions. We evaluate our approach using real-world healthcare and vision-and-language datasets with state-of-the-art models, demonstrating superior performance over traditional methods focusing only on one type of modality dependency. The code is available at https://github.com/divyam3897/I2M2. Divyam Madaan, Taro Makino, Sumit Chopra, Kyunghyun Cho |
NeurIPS | 4 |
| 2024 | Multiple Physics Pretraining for Spatiotemporal Surrogate ModelsabstractWe introduce multiple physics pretraining (MPP), an autoregressive task-agnostic pretraining approach for physical surrogate modeling of spatiotemporal systems with transformers. In MPP, rather than training one model on a specific physical system, we train a backbone model to predict the dynamics of multiple heterogeneous physical systems simultaneously in order to learn features that are broadly useful across systems and facilitate transfer. In order to learn effectively in this setting, we introduce a shared embedding and normalization strategy that projects the fields of multiple systems into a shared embedding space. We validate the efficacy of our approach on both pretraining and downstream tasks over a broad fluid mechanics-oriented benchmark. We show that a single MPP-pretrained transformer is able to match or outperform task-specific baselines on all pretraining sub-tasks without the need for finetuning. For downstream tasks, we demonstrate that finetuning MPP-trained models results in more accurate predictions across multiple time-steps on systems with previously unseen physical components or higher dimensional systems compared to training from scratch or finetuning pretrained video foundation models. We open-source our code and model weights trained at multiple scales for reproducibility. Michael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana, Miles D. Cranmer, Alberto Bietti, Michael Eickenberg, Siavash Golkar, Géraud Krawezik, François Lanusse, Mariel Pettee, Tiberiu Tesileanu, Kyunghyun Cho, Shirley Ho |
NeurIPS | 13 |
| 2024 | Iterative Reasoning Preference OptimizationabstractIterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks. In this work we develop an iterative approach that optimizes the preference between competing generated Chain-of-Thought (CoT) candidates by optimizing for winning vs. losing reasoning steps. We train using a modified DPO loss with an additional negative log-likelihood term, which we find to be crucial. We show reasoning improves across repeated iterations of this scheme. While only relying on examples in the training set, our approach results in increasing accuracy on GSM8K, MATH, and ARC-Challenge for Llama-2-70B-Chat, outperforming other Llama-2-based models not relying on additionally sourced datasets. For example, we see a large improvement from 55.6% to 81.6% on GSM8K and an accuracy of 88.7% with majority voting out of 32 samples. Richard Yuanzhe Pang, Weizhe Yuan, He He 0001, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston |
NeurIPS | 4 |
| 2024 | Implicitly Guided Design with PropEn: Match your Data to Follow the GradientabstractAcross scientific domains, generating new models or optimizing existing ones while meeting specific criteria is crucial. Traditional machine learning frameworks for guided design use a generative model and a surrogate model (discriminator), requiring large datasets. However, real-world scientific applications often have limited data and complex landscapes, making data-hungry models inefficient or impractical. We propose a new framework, PropEn, inspired by ``matching'', which enables implicit guidance without training a discriminator. By matching each sample with a similar one that has a better property value, we create a larger training dataset that inherently indicates the direction of improvement. Matching, combined with an encoder-decoder architecture, forms a domain-agnostic generative framework for property enhancement. We show that training with a matched dataset approximates the gradient of the property of interest while remaining within the data distribution, allowing efficient design optimization. Extensive evaluations in toy problems and scientific applications, such as therapeutic protein design and airfoil optimization, demonstrate PropEn's advantages over common baselines. Notably, the protein design results are validated with wet lab experiments, confirming the competitiveness and effectiveness of our approach. Our code is available at https://github.com/prescient-design/propen. Natasa Tagasovska, Vladimir Gligorijevic, Kyunghyun Cho, Andreas Loukas |
NeurIPS | 3 |
| 2024 | Non-convolutional graph neural networksabstractRethink convolution-based graph neural networks (GNN)---they characteristically suffer from limited expressiveness, over-smoothing, and over-squashing, and require specialized sparse kernels for efficient computation.
Here, we design a simple graph learning module entirely free of convolution operators, coined _random walk with unifying memory_ (RUM) neural network, where an RNN merges the topological and semantic graph features along the random walks terminating at each node.
Relating the rich literature on RNN behavior and graph topology, we theoretically show and experimentally verify that RUM attenuates the aforementioned symptoms and is more expressive than the Weisfeiler-Lehman (WL) isomorphism test.
On a variety of node- and graph-level classification and regression tasks, RUM not only achieves competitive performance, but is also robust, memory-efficient, scalable, and faster than the simplest convolutional GNNs. Kyunghyun Cho |
NeurIPS | 2 |
| 2024 | : Visualization of AI-Assisted Task Guidance in ARabstractThe concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi-year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine-tune the components of the AR assistant. Sonia Castelo Quispe, João Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Irán R. Román, Roque Lopez, Ethan Brewer, Chen Zhao 0013, Kyunghyun Cho, He He 0001, Qi Sun 0003, Huy T. Vo, Juan Pablo Bello, Michael Krone, Cláudio T. Silva |
IEEE Trans. Vis. Comput. Graph. | 12 |
| 2023 | On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationabstractTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tianxing He, Tianle Wang 0003, Sachin Kumar 0009, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov |
ACL (1) | 5 |
| 2023 | A Transformer-based Function Symbol Name Inference Model from an Assembly Language for Binary ReversingabstractReverse engineering of a stripped binary has a wide range of applications, yet it is challenging mainly due to the lack of contextually useful information within. Once debugging symbols (e.g., variable names, types, function names) are discarded, recovering such information is not technically viable with traditional approaches like static or dynamic binary analysis. We focus on a function symbol name recovery, which allows a reverse engineer to gain a quick overview of an unseen binary. The key insight is that a well-developed program labels a meaningful function name that describes its underlying semantics well. In this paper, we present AsmDepictor, the Transformer-based framework that generates a function symbol name from a set of assembly codes (i.e., machine instructions), which consists of three major components: binary code refinement, model training, and inference. To this end, we conduct systematic experiments on the effectiveness of code refinement that can enhance an overall performance. We introduce the per-layer positional embedding and Unique-softmax for AsmDepictor so that both can aid to capture a better relationship between tokens. Lastly, we devise a novel evaluation metric tailored for a short description length, the Jaccard* score. Our empirical evaluation shows that the performance of AsmDepictor by far surpasses that of the state-of-the-art models up to around 400%. The best AsmDepictor model achieves an F1 of 71.5 and Jaccard* of 75.4. JinYeong Bak, Kyunghyun Cho, Hyungjoon Koo |
AsiaCCS | 3 |
| 2023 | A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at ScaleabstractUnpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploying these methods to improve production ASR systems that are trained on very large supervised corpora and with realistic requirements like a constrained model size and CPU budget, streaming capability, and a rich lattice for rescoring and for downstream NLU tasks. In this work, we compare three state-of-the-art semi-supervised methods encompassing both unpaired text and audio as well as several of their combinations in a controlled setting using joint training. We find that in our setting these methods offer many improvements beyond raw WER, including substantial gains in tail-word WER, decoder computation during inference, and lattice density. Cal Peyser, Michael Picheny, Kyunghyun Cho, Rohit Prabhavalkar, W. Ronny Huang, Tara N. Sainath |
ICASSP | 3 |
| 2023 | A Non-monotonic Self-terminating Language Model
Eugene Choi, Kyunghyun Cho, Cheolhyoung Lee |
ICLR | 2 |
| 2023 | Linear Connectivity Reveals Generalization Strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, Naomi Saphra |
ICLR | 3 |
| 2023 | Towards Understanding and Improving GFlowNet TrainingabstractGenerative flow networks (GFlowNets) are a family of algorithms that learn a generative policy to sample discrete objects $x$ with non-negative reward $R(x)$. Learning objectives guarantee the GFlowNet samples $x$ from the target distribution $p^*(x) \propto R(x)$ when loss is globally minimized over all states or trajectories, but it is unclear how well they perform with practical limits on training resources. We introduce an efficient evaluation strategy to compare the learned sampling distribution to the target reward distribution. As flows can be underdetermined given training data, we clarify the importance of learned flows to generalization and matching $p^*(x)$ in practice. We investigate how to learn better flows, and propose (i) prioritized replay training of high-reward $x$, (ii) relative edge flow policy parametrization, and (iii) a novel guided trajectory balance objective, and show how it can solve a substructure credit assignment problem. We substantially improve sample efficiency on biochemical design tasks. Max W. Shen, Emmanuel Bengio, Ehsan Hajiramezanali, Andreas Loukas, Kyunghyun Cho, Tommaso Biancalani |
ICML | 5 |
| 2023 | Improving Joint Speech-Text Representations Without Alignment
Cal Peyser, Zhong Meng, Rohit Prabhavalkar, Andrew Rosenberg, Tara N. Sainath, Michael Picheny, Kyunghyun Cho |
INTERSPEECH | 7 |
| 2023 | Protein Design with Guided Discrete DiffusionabstractA popular approach to protein design is to combine a generative model with a discriminative model for conditional sampling. The generative model samples plausible sequences while the discriminative model guides a search for sequences with high fitness. Given its broad success in conditional sampling, classifier-guided diffusion modeling is a promising foundation for protein design, leading many to develop guided diffusion models for structure with inverse folding to recover sequences. In this work, we propose diffusioN Optimized Sampling (NOS), a guidance method for discrete diffusion models that follows gradients in the hidden states of the denoising network. NOS makes it possible to perform design directly in sequence space, circumventing significant limitations of structure-based methods, including scarce data and challenging inverse design. Moreover, we use NOS to generalize LaMBO, a Bayesian optimization procedure for sequence design that facilitates multiple objectives and edit-based constraints. The resulting method, LaMBO-2, enables discrete diffusions and stronger performance with limited edits through a novel application of saliency maps. We apply LaMBO-2 to a real-world protein design task, optimizing antibodies for higher expression yield and binding affinity to several therapeutic targets under locality and developability constraints, attaining a 99\% expression rate and 40\% binding rate in exploratory in vitro experiments. Nate Gruver, Samuel Stanton, Nathan C. Frey, Tim G. J. Rudner, Isidro Hötzel, Julien Lafrance-Vanasse, Arvind Rajpal, Kyunghyun Cho, Andrew Gordon Wilson |
NeurIPS | 8 |
| 2023 | AbDiffuser: full-atom generation of in-vitro functioning antibodiesabstractWe introduce AbDiffuser, an equivariant and physics-informed diffusion model for the joint generation of antibody 3D structures and sequences. AbDiffuser is built on top of a new representation of protein structure, relies on a novel architecture for aligned proteins, and utilizes strong diffusion priors to improve the denoising process. Our approach improves protein diffusion by taking advantage of domain knowledge and physics-based constraints; handles sequence-length changes; and reduces memory complexity by an order of magnitude, enabling backbone and side chain generation. We validate AbDiffuser in silico and in vitro. Numerical experiments showcase the ability of AbDiffuser to generate antibodies that closely track the sequence and structural properties of a reference set. Laboratory experiments confirm that all 16 HER2 antibodies discovered were expressed at high levels and that 57.1% of the selected designs were tight binders. Karolis Martinkus, Jan Ludwiczak, Wei-Ching Liang, Julien Lafrance-Vanasse, Isidro Hötzel, Arvind Rajpal, Yan Wu 0027, Kyunghyun Cho, Richard Bonneau, Vladimir Gligorijevic, Andreas Loukas |
NeurIPS | 8 |
| 2022 | DEEP: DEnoising Entity Pre-training for Neural Machine TranslationabstractIt has been shown that machine translation models usually generate poor translations for named entities that are infrequent in the training corpus.Earlier named entity translation methods mainly focus on phonetic transliteration, which ignores the sentence context for translation and is limited in domain and language coverage.To address this limitation, we propose DEEP, a DEnoising Entity Pretraining method that leverages large amounts of monolingual data and a knowledge base to improve named entity translation accuracy within sentences.Besides, we investigate a multi-task learning strategy that finetunes a pre-trained neural machine translation model on both entity-augmented monolingual data and parallel data to further improve entity translation.Experimental results on three language pairs demonstrate that DEEP results in significant improvements over strong denoising autoencoding baselines, with a gain of up to 1.3 BLEU and up to 9.2 entity accuracy points for English-Russian translation. 1 Junjie Hu 0001, Hiroaki Hayashi, Kyunghyun Cho, Graham Neubig |
ACL (1) | 3 |
| 2022 | Translation between Molecules and Natural LanguageabstractWe present MolT5 -a self-supervised learning framework for pretraining models on a vast amount of unlabeled natural language text and molecule strings.MolT5 allows for new, useful, and challenging analogs of traditional vision-language tasks, such as molecule captioning and text-based de novo molecule generation (altogether: translation between molecules and language), which we explore for the first time.Since MolT5 pretrains models on single-modal data, it helps overcome the chemistry domain shortcoming of data scarcity.Furthermore, we consider several metrics, including a new cross-modal embedding-based metric, to evaluate the tasks of molecule captioning and text-based molecule generation.Our results show that MolT5-based models are able to generate outputs, both molecules and captions, which in many cases are high quality 1 . Carl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, Heng Ji 0001 |
EMNLP | 5 |
| 2022 | Chemical-Reaction-Aware Molecule Representation Learning
Weijiang Li, Xiaomeng Jin, Kyunghyun Cho, Heng Ji 0001, Jiawei Han 0001, Martin D. Burke |
ICLR | 4 |
| 2022 | Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural NetworksabstractWe hypothesize that due to the greedy nature of learning in multi-modal deep neural networks, these models tend to rely on just one modality while under-fitting the other modalities. Such behavior is counter-intuitive and hurts the models’ generalization, as we observe empirically. To estimate the model’s dependence on each modality, we compute the gain on the accuracy when the model has access to it in addition to another modality. We refer to this gain as the conditional utilization rate. In the experiments, we consistently observe an imbalance in conditional utilization rates between modalities, across multiple tasks and architectures. Since conditional utilization rate cannot be computed efficiently during training, we introduce a proxy for it based on the pace at which the model learns from each modality, which we refer to as the conditional learning speed. We propose an algorithm to balance the conditional learning speeds between modalities during training and demonstrate that it indeed addresses the issue of greedy learning. The proposed algorithm improves the model’s generalization on three datasets: Colored MNIST, ModelNet40, and NVIDIA Dynamic Hand Gesture. Nan Wu 0008, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. Geras |
ICML | 3 |
| 2022 | Towards Disentangled Speech RepresentationsabstractThe careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks.Increasingly, such approaches have emphasized "disentanglement", where a representation contains only parts of the speech signal relevant to transcription while discarding irrelevant information.In this paper, we construct a representation learning task based on joint modeling of ASR and TTS, and seek to learn a representation of audio that disentangles that part of the speech signal that is relevant to transcription from that part which is not.We present empirical evidence that successfully finding such a representation is tied to the randomness inherent in training.We then make the observation that these desired, disentangled solutions to the optimization problem possess unique statistical properties.Finally, we show that enforcing these properties during training improves WER by 24.5% relative on average for our joint modeling task.These observations motivate a novel approach to learning effective audio representations. Cal Peyser, W. Ronny Huang, Andrew Rosenberg, Tara N. Sainath, Michael Picheny, Kyunghyun Cho |
INTERSPEECH | 6 |
| 2022 | On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language ModelabstractSeongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woomyoung Park, Jung-Woo Ha, Nako Sung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Seongjin Shin, Sang-Woo Lee 0001, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woo-Myoung Park, Jung-Woo Ha 0001, Nako Sung |
NAACL-HLT | 7 |
| 2022 | Generative multitask learning mitigates target-causing confoundingabstractWe propose generative multitask learning (GMTL), a simple and scalable approach to causal machine learning in the multitask setting. Our approach makes a minor change to the conventional multitask inference objective, and improves robustness to target shift. Since GMTL only modifies the inference objective, it can be used with existing multitask learning methods without requiring additional training. The improvement in robustness comes from mitigating unobserved confounders that cause the targets, but not the input. We refer to them as \emph{target-causing confounders}. These confounders induce spurious dependencies between the input and targets. This poses a problem for conventional multitask learning, due to its assumption that the targets are conditionally independent given the input. GMTL mitigates target-causing confounding at inference time, by removing the influence of the joint target distribution, and predicting all targets jointly. This removes the spurious dependencies between the input and targets, where the degree of removal is adjustable via a single hyperparameter. This flexibility is useful for managing the trade-off between in- and out-of-distribution generalization. Our results on the Attributes of People and Taskonomy datasets reflect an improved robustness to target shift across four multitask learning methods. Taro Makino, Krzysztof J. Geras, Kyunghyun Cho |
NeurIPS | 3 |
| 2022 | Dual Learning for Large Vocabulary On-Device ASRabstractDual learning is a paradigm for semi-supervised machine learning that seeks to leverage unsupervised data by solving two opposite tasks at once. In this scheme, each model is used to generate pseudo-labels for unlabeled examples that are used to train the other model. Dual learning has seen some use in speech processing by pairing ASR and TTS as dual tasks. However, these results mostly address only the case of using unpaired examples to compensate for very small supervised datasets, and mostly on large, non-streaming models. Dual learning has not yet been proven effective for using unsupervised data to improve realistic on-device streaming models that are already trained on large supervised corpora. We provide this missing piece though an analysis of an on-device-sized streaming conformer trained on the entirety of Librispeech, showing relative WER improvements of 10.7%/5.2% without an LM and 11.7%/16.4% with an LM. Cal Peyser, W. Ronny Huang, Tara N. Sainath, Rohit Prabhavalkar, Michael Picheny, Kyunghyun Cho |
SLT | 6 |
| 2022 | NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site predictionabstractMOTIVATION: Machine learning models for predicting cell-type-specific transcription factor (TF) binding sites have become increasingly more accurate thanks to the increased availability of next-generation sequencing data and more standardized model evaluation criteria. However, knowledge transfer from data-rich to data-limited TFs and cell types remains crucial for improving TF binding prediction models because available binding labels are highly skewed towards a small collection of TFs and cell types. Transfer prediction of TF binding sites can potentially benefit from a multitask learning approach; however, existing methods typically use shallow single-task models to generate low-resolution predictions. Here, we propose NetTIME, a multitask learning framework for predicting cell-type-specific TF binding sites with base-pair resolution. RESULTS: We show that the multitask learning strategy for TF binding prediction is more efficient than the single-task approach due to the increased data availability. NetTIME trains high-dimensional embedding vectors to distinguish TF and cell-type identities. We show that this approach is critical for the success of the multitask learning strategy and allows our model to make accurate transfer predictions within and beyond the training panels of TFs and cell types. We additionally train a linear-chain conditional random field (CRF) to classify binding predictions and show that this CRF eliminates the need for setting a probability threshold and reduces classification noise. We compare our method's predictive performance with two state-of-the-art methods, Catchitt and Leopard, and show that our method outperforms previous methods under both supervised and transfer learning settings. AVAILABILITY AND IMPLEMENTATION: NetTIME is freely available at https://github.com/ryi06/NetTIME and the code is also archived at https://doi.org/10.5281/zenodo.6994897. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ren Yi, Kyunghyun Cho, Richard Bonneau |
Bioinform. | 2 |
| 2021 | MLE-Guided Parameter Search for Task Loss Minimization in Neural Sequence ModelingabstractNeural autoregressive sequence models are used to generate sequences in a variety of natural language processing (NLP) tasks, where they are evaluated according to sequence-level task losses. These models are typically trained with maximum likelihood estimation, which ignores the task loss, yet empirically performs well as a surrogate objective. Typical approaches to directly optimizing the task loss such as policy gradient and minimum risk training are based around sampling in the sequence space to obtain candidate update directions that are scored based on the loss of a single sequence. In this paper, we develop an alternative method based on random search in the parameter space that leverages access to the maximum likelihood gradient. We propose maximum likelihood guided parameter search (MGS), which samples from a distribution over update directions that is a mixture of random search around the current parameters and around the maximum likelihood gradient, with each direction weighted by its improvement in the task loss. MGS shifts sampling to the parameter space, and scores candidates using losses that are pooled from multiple sequences. Our experiments show that MGS is capable of optimizing sequence-level losses, with substantial reductions in repetition and non-termination in sequence completion, and similar improvements to those of minimum risk training in machine translation. Sean Welleck, Kyunghyun Cho |
AAAI | 2 |
| 2021 | Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with SearchabstractGyuwan Kim, Kyunghyun Cho. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Gyuwan Kim, Kyunghyun Cho |
ACL/IJCNLP (1) | 2 |
| 2021 | Comparing Test Sets with Item Response TheoryabstractClara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Clara Vania, Phu Mon Htut, William Huang, Dhara A. Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, Samuel R. Bowman |
ACL/IJCNLP (1) | 8 |
| 2021 | Analyzing the Forgetting Problem in Pretrain-Finetuning of Open-domain Dialogue Response ModelsabstractTianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, Fuchun Peng. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Tianxing He, Kyunghyun Cho, Myle Ott, Bing Liu 0024, James R. Glass, Fuchun Peng |
EACL | 3 |
| 2021 | AdapterFusion: Non-Destructive Task Composition for Transfer LearningabstractJonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, Iryna Gurevych. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, Iryna Gurevych |
EACL | 4 |
| 2021 | The Future is not One-dimensional: Complex Event Schema Induction by Graph Modeling for Event PredictionabstractEvent schemas encode knowledge of stereotypical structures of events and their connections.As events unfold, schemas are crucial to act as a scaffolding.Previous work on event schema induction focuses either on atomic events or linear temporal event sequences, ignoring the interplay between events via arguments and argument relations.We introduce a new concept of Temporal Complex Event Schema: a graph-based schema representation that encompasses events, arguments, temporal connections and argument relations.In addition, we propose a Temporal Event Graph Model that predicts event instances following the temporal complex event schema.To build and evaluate such schemas, we release a new schema learning corpus containing 6,399 documents accompanied with event graphs, and we have manually constructed gold-standard schemas.Intrinsic evaluations by schema matching and instance graph perplexity, prove the superior quality of our probabilistic graph schema library compared to linear representations.Extrinsic evaluation on schema-guided future event prediction further demonstrates the predictive power of our event graph model, significantly outperforming human schemas and baselines by more than 23.8% on HITS@1. 1 Manling Li, Zhenhailong Wang, Lifu Huang, Kyunghyun Cho, Heng Ji 0001, Jiawei Han 0001, Clare R. Voss |
EMNLP (1) | 5 |
| 2021 | Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule
Shuhei Kurita, Kyunghyun Cho |
ICLR | 2 |
| 2021 | Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationabstractThe early phase of training a deep neural network has a dramatic effect on the local curvature of the loss function. For instance, using a small learning rate does not guarantee stable optimization because the optimization trajectory has a tendency to steer towards regions of the loss surface with increasing local curvature. We ask whether this tendency is connected to the widely observed phenomenon that the choice of the learning rate strongly influences generalization. We first show that stochastic gradient descent (SGD) implicitly penalizes the trace of the Fisher Information Matrix (FIM), a measure of the local curvature, from the start of training. We argue it is an implicit regularizer in SGD by showing that explicitly penalizing the trace of the FIM can significantly improve generalization. We highlight that poor final generalization coincides with the trace of the FIM attaining a large value early in training, to which we refer as catastrophic Fisher explosion. Finally, to gain insight into the regularization effect of penalizing the trace of the FIM, we show that it limits memorization by reducing the learning speed of examples with noisy labels more than that of the examples with clean labels. Stanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg, Huan Wang 0016, Caiming Xiong, Richard Socher, Kyunghyun Cho, Krzysztof J. Geras |
ICML | 8 |
| 2021 | Rissanen Data Analysis: Examining Dataset Characteristics via Description LengthabstractWe introduce a method to determine if a certain capability helps to achieve an accurate model of given data. We view labels as being generated from the inputs by a program composed of subroutines with different capabilities, and we posit that a subroutine is useful if and only if the minimal program that invokes it is shorter than the one that does not. Since minimum program length is uncomputable, we instead estimate the labels’ minimum description length (MDL) as a proxy, giving us a theoretically-grounded method for analyzing dataset characteristics. We call the method Rissanen Data Analysis (RDA) after the father of MDL, and we showcase its applicability on a wide variety of settings in NLP, ranging from evaluating the utility of generating subquestions before answering a question, to analyzing the value of rationales and explanations, to investigating the importance of different parts of speech, and uncovering dataset gender bias. Ethan Perez, Douwe Kiela, Kyunghyun Cho |
ICML | 3 |
| 2021 | True Few-Shot Learning with Language ModelsabstractPretrained language models (LMs) perform well on many tasks even when learning from a few examples, but prior work uses many held-out examples to tune various aspects of learning, such as hyperparameters, training objectives, and natural language templates ("prompts"). Here, we evaluate the few-shot ability of LMs when such held-out examples are unavailable, a setting we call true few-shot learning. We test two model selection criteria, cross-validation and minimum description length, for choosing LM prompts and hyperparameters in the true few-shot setting. On average, both marginally outperform random selection and greatly underperform selection based on held-out examples. Moreover, selection criteria often prefer models that perform significantly worse than randomly-selected ones. We find similar results even when taking into account our uncertainty in a model's true performance during selection, as well as when varying the amount of computation and number of examples used for selection. Overall, our findings suggest that prior work significantly overestimated the true few-shot ability of LMs given the difficulty of few-shot model selection. Ethan Perez, Douwe Kiela, Kyunghyun Cho |
NeurIPS | 3 |
| 2021 | NetQuilt: deep multispecies network-based protein function prediction using homology-informed network similarityabstractMOTIVATION: Transferring knowledge between species is challenging: different species contain distinct proteomes and cellular architectures, which cause their proteins to carry out different functions via different interaction networks. Many approaches to protein functional annotation use sequence similarity to transfer knowledge between species. These approaches cannot produce accurate predictions for proteins without homologues of known function, as many functions require cellular context for meaningful prediction. To supply this context, network-based methods use protein-protein interaction (PPI) networks as a source of information for inferring protein function and have demonstrated promising results in function prediction. However, most of these methods are tied to a network for a single species, and many species lack biological networks. RESULTS: In this work, we integrate sequence and network information across multiple species by computing IsoRank similarity scores to create a meta-network profile of the proteins of multiple species. We use this integrated multispecies meta-network as input to train a maxout neural network with Gene Ontology terms as target labels. Our multispecies approach takes advantage of more training examples, and consequently leads to significant improvements in function prediction performance compared to two network-based methods, a deep learning sequence-based method and the BLAST annotation method used in the Critial Assessment of Functional Annotation. We are able to demonstrate that our approach performs well even in cases where a species has no network information available: when an organism's PPI network is left out we can use our multi-species method to make predictions for the left-out organism with good performance. AVAILABILITY AND IMPLEMENTATION: The code is freely available at https://github.com/nowittynamesleft/NetQuilt. The data, including sequences, PPI networks and GO annotations are available at https://string-db.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Meet Barot, Vladimir Gligorijevic, Kyunghyun Cho, Richard Bonneau |
Bioinform. | 3 |
| 2021 | An interpretable classifier for high-resolution breast cancer screening images utilizing weakly supervised localizationabstractMedical images differ from natural images in significantly higher resolutions and smaller regions of interest. Because of these differences, neural network architectures that work well for natural images might not be applicable to medical image analysis. In this work, we propose a novel neural network model to address these unique properties of medical images. This model first uses a low-capacity, yet memory-efficient, network on the whole image to identify the most informative regions. It then applies another higher-capacity network to collect details from chosen regions. Finally, it employs a fusion module that aggregates global and local information to make a prediction. While existing methods often require lesion segmentation during training, our model is trained with only image-level labels and can generate pixel-level saliency maps indicating possible malignant findings. We apply the model to screening mammography interpretation: predicting the presence or absence of benign and malignant lesions. On the NYU Breast Cancer Screening Dataset, our model outperforms (AUC = 0.93) ResNet-34 and Faster R-CNN in classifying breasts with malignant findings. On the CBIS-DDSM dataset, our model achieves performance (AUC = 0.858) on par with state-of-the-art approaches. Compared to ResNet-34, our model is 4.1x faster for inference while using 78.4% less GPU memory. Furthermore, we demonstrate, in a reader study, that our model surpasses radiologist-level AUC by a margin of 0.11. Yiqiu Shen, Nan Wu 0008, Jason Phang, Jungkyu Park 0001, Kangning Liu, Sudarshini Tyagi, Laura Heacock, Sungheon Gene Kim, Linda Moy, Kyunghyun Cho, Krzysztof J. Geras |
Medical Image Anal. | 10 |
| 2021 | Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics dataabstractThe analysis of single-cell genomics data presents several statistical challenges, and extensive efforts have been made to produce methods for the analysis of this data that impute missing values, address sampling issues and quantify and correct for noise. In spite of such efforts, no consensus on best practices has been established and all current approaches vary substantially based on the available data and empirical tests. The k-Nearest Neighbor Graph (kNN-G) is often used to infer the identities of, and relationships between, cells and is the basis of many widely used dimensionality-reduction and projection methods. The kNN-G has also been the basis for imputation methods using, e.g., neighbor averaging and graph diffusion. However, due to the lack of an agreed-upon optimal objective function for choosing hyperparameters, these methods tend to oversmooth data, thereby resulting in a loss of information with regard to cell identity and the specific gene-to-gene patterns underlying regulatory mechanisms. In this paper, we investigate the tuning of kNN- and diffusion-based denoising methods with a novel non-stochastic method for optimally preserving biologically relevant informative variance in single-cell data. The framework, Denoising Expression data with a Weighted Affinity Kernel and Self-Supervision (DEWÄKSS), uses a self-supervised technique to tune its parameters. We demonstrate that denoising with optimal parameters selected by our objective function (i) is robust to preprocessing methods using data from established benchmarks, (ii) disentangles cellular identity and maintains robust clusters over dimension-reduction methods, (iii) maintains variance along several expression dimensions, unlike previous heuristic-based methods that tend to oversmooth data variance, and (iv) rarely involves diffusion but rather uses a fixed weighted kNN graph for denoising. Together, these findings provide a new understanding of kNN- and diffusion-based denoising methods. Code and example data for DEWÄKSS is available at https://gitlab.com/Xparx/dewakss/-/tree/Tjarnberg2020branch. Andreas Tjärnberg, Omar Mahmood, Christopher A. Jackson, Giuseppe-Antonio Saldi, Kyunghyun Cho, Lionel A. Christiaen, Richard Bonneau |
PLoS Comput. Biol. | 5 |
| 2020 | Learning to Learn Morphological Inflection for Resource-Poor LanguagesabstractWe propose to cast the task of morphological inflection—mapping a lemma to an indicated inflected form—for resource-poor languages as a meta-learning problem. Treating each language as a separate task, we use data from high-resource source languages to learn a set of model parameters that can serve as a strong initialization point for fine-tuning on a resource-poor target language. Experiments with two model architectures on 29 target languages from 3 families show that our suggested approach outperforms all baselines. In particular, it obtains a 31.7% higher absolute accuracy than a previously proposed cross-lingual transfer model and outperforms the previous state of the art by 1.7% absolute accuracy on average over languages. Katharina Kann, Samuel R. Bowman, Kyunghyun Cho |
AAAI | 3 |
| 2020 | Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorabstractAlthough neural machine translation models reached high translation quality, the autoregressive nature makes inference difficult to parallelize and leads to high translation latency. Inspired by recent refinement-based approaches, we propose LaNMT, a latent-variable non-autoregressive model with continuous latent variables and deterministic inference procedure. In contrast to existing approaches, we use a deterministic inference algorithm to find the target sequence that maximizes the lowerbound to the log-probability. During inference, the length of translation automatically adapts itself. Our experiments show that the lowerbound can be greatly increased by running the inference algorithm, resulting in significantly improved translation quality. Our proposed model closes the performance gap between non-autoregressive and autoregressive approaches on ASPEC Ja-En dataset with 8.6x faster decoding. On WMT'14 En-De dataset, our model narrows the gap with autoregressive baseline to 2.0 BLEU points with 12.5x speedup. By decoding multiple initial latent variables in parallel and rescore using a teacher model, the proposed model further brings the gap down to 1.0 BLEU point on WMT'14 En-De task with 6.8x speedup. Raphael Shu, Jason Lee 0002, Hideki Nakayama, Kyunghyun Cho |
AAAI | 4 |
| 2020 | Neural Machine Translation with Byte-Level SubwordsabstractAlmost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can unnecessarily take up vocabulary slots and limit its compactness. Representing text at the level of bytes and using the 256 byte set as vocabulary is a potential solution to this issue. High computational cost has however prevented it from being widely deployed or used in practice. In this paper, we investigate byte-level subwords, specifically byte-level BPE (BBPE), which is compacter than character vocabulary and has no out-of-vocabulary tokens, but is more efficient than using pure bytes only is. We claim that contextualizing BBPE embeddings is necessary, which can be implemented by a convolutional or recurrent layer. Our experiments show that BBPE has comparable performance to BPE while its size is only 1/8 of that for BPE. In the multilingual setting, BBPE maximizes vocabulary sharing across many languages and achieves better translation quality. Moreover, we show that BBPE enables transferring models between languages with non-overlapping character sets. Changhan Wang, Kyunghyun Cho, Jiatao Gu |
AAAI | 2 |
| 2020 | Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood TrainingabstractGenerative dialogue models currently suffer from a number of problems which standard maximum likelihood training does not address.They tend to produce generations that (i) rely too much on copying from the context, (ii) contain repetitions within utterances, (iii) overuse frequent words, and (iv) at a deeper level, contain logical flaws.In this work we show how all of these problems can be addressed by extending the recently introduced unlikelihood loss (Welleck et al., 2019a) to these cases.We show that appropriate loss functions which regularize generated outputs to match human distributions are effective for the first three issues.For the last important general issue, we show applying unlikelihood to collected data of what a model should not do is effective for improving logical consistency, potentially paving the way to generative models with greater reasoning ability.We demonstrate the efficacy of our approach across several dialogue tasks. Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, Jason Weston |
ACL | 6 |
| 2020 | Asking and Answering Questions to Evaluate the Factual Consistency of SummariesabstractPractical applications of abstractive summarization models are limited by frequent factual inconsistencies with respect to their input.Existing automatic evaluation metrics for summarization are largely insensitive to such errors.We propose QAGS, 1 an automatic evaluation protocol that is designed to identify factual inconsistencies in a generated summary.QAGS is based on the intuition that if we ask questions about a summary and its source, we will receive similar answers if the summary is factually consistent with the source.To evaluate QAGS, we collect human judgments of factual consistency on model-generated summaries for the CNN/DailyMail (Hermann et al., 2015) and XSUM (Narayan et al., 2018) summarization datasets.QAGS has substantially higher correlations with these judgments than other automatic evaluation metrics.Also, QAGS offers a natural form of interpretability: The answers and questions generated while computing QAGS indicate which tokens of a summary are inconsistent and why.We believe QAGS is a promising tool in automatically generating usable and factually consistent text.Code for QAGS will be available at https://github. com/W4ngatang/qags.Article: On Friday, 28-year-old Usman Khan stabbed reportedly several people at Fishmongers' Hall in London with a large knife, then fled up London Bridge.Members of the public confronted him; one man sprayed Khan with a fire extinguisher, others struck him with their fists and took his knife, and another, a Polish chef named ukasz, harried him with a five-foot narwhal tusk.[. . .] Summary : On Friday afternoon , a man named Faisal Khan entered a Cambridge University building and started attacking people with a knife and a fire extinguisher .Question 1: What did the attacker have ?Article answer: a large knife Summary answer: a knife and a fire extinguisher Question 2: When did the attack take place ? Kyunghyun Cho, Mike Lewis |
ACL | 2 |
| 2020 | Improving Conversational Question Answering Systems after Deployment using Feedback-Weighted LearningabstractThe interaction of conversational systems with users poses an exciting opportunity for improving them after deployment, but little evidence has been provided of its feasibility.In most applications, users are not able to provide the correct answer to the system, but they are able to provide binary (correct, incorrect) feedback.In this paper we propose feedback-weighted learning based on importance sampling to improve upon an initial supervised system using binary user feedback.We perform simulated experiments on document classification (for development) and Conversational Question Answering datasets like QuAC and DoQA, where binary user feedback is derived from gold annotations.The results show that our method is able to improve over the initial supervised system, getting close to a fully-supervised system that has access to the same labeled examples in in-domain experiments (QuAC), and even matching in out-of-domain experiments (DoQA).Our work opens the prospect to exploit interactions with real users and improve conversational systems after deployment. Jon Ander Campos, Kyunghyun Cho, Arantxa Otegi, Aitor Soroa, Eneko Agirre, Gorka Azkune |
COLING | 2 |
| 2020 | Learning Non-Monotonic Automatic Post-Editing of Translations from Human OrderingsabstractRecent research in neural machine translation has explored flexible generation orders, as an alternative to left-to-right generation. However, training non-monotonic models brings a new complication: how to search for a good ordering when there is a combinatorial explosion of orderings arriving at the same final result? Also, how do these automatic orderings compare with the actual behaviour of human translators? Current models rely on manually built biases or are left to explore all possibilities on their own. In this paper, we analyze the orderings produced by human post-editors and use them to train an automatic post-editing system. We compare the resulting system with those trained with left-to-right and random post-editing orderings. We observe that humans tend to follow a nearly left-to-right order, but with interesting deviations, such as preferring to start by correcting punctuation or verbs. António Góis, Kyunghyun Cho, André F. T. Martins |
EAMT | 2 |
| 2020 | Iterative Refinement in the Continuous Space for Non-Autoregressive Neural Machine TranslationabstractWe propose an efficient inference procedure for non-autoregressive machine translation that iteratively refines translation purely in the continuous space.Given a continuous latent variable model for machine translation (Shu et al., 2020), we train an inference network to approximate the gradient of the marginal log probability of the target sentence, using only the latent variable as input.This allows us to use gradient-based optimization to find the target sentence at inference time that approximately maximizes its marginal probability.As each refinement step only involves computation in the latent space of low dimensionality (we use 8 in our experiments), we avoid computational overhead incurred by existing non-autoregressive inference procedures that often refine in token space.We compare our approach to a recently proposed EM-like inference procedure (Shu et al., 2020) that optimizes in a hybrid space, consisting of both discrete and continuous variables.We evaluate our approach on WMT'14 En→De, WMT'16 Ro→En and IWSLT'16 De→En, and observe two advantages over the EM-like inference: (1) it is computationally efficient, i.e. each refinement step is twice as fast, and (2) it is more effective, resulting in higher marginal probabilities and BLEU scores with the same number of refinement steps.On WMT'14 En→De, for instance, our approach is able to decode 6.2 times faster than the autoregressive model with minimal degradation to translation quality (0.9 BLEU). Jason Lee 0002, Raphael Shu, Kyunghyun Cho |
EMNLP (1) | 3 |
| 2020 | Connecting the Dots: Event Graph Schema Induction with Path Language ModelingabstractManling Li, Qi Zeng, Ying Lin, Kyunghyun Cho, Heng Ji, Jonathan May, Nathanael Chambers, Clare Voss. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Manling Li, Qi Zeng 0001, Kyunghyun Cho, Heng Ji 0001, Jonathan May, Nathanael Chambers, Clare R. Voss |
EMNLP (1) | 4 |
| 2020 | SSMBA: Self-Supervised Manifold Based Data Augmentation for Improving Out-of-Domain RobustnessabstractModels that perform well on a training domain often fail to generalize to out-of-domain (OOD) examples.Data augmentation is a common method used to prevent overfitting and improve OOD generalization.However, in natural language, it is difficult to generate new examples that stay on the underlying data manifold.We introduce SSMBA, a data augmentation method for generating synthetic training examples by using a pair of corruption and reconstruction functions to move randomly on a data manifold.We investigate the use of SSMBA in the natural language domain, leveraging the manifold assumption to reconstruct corrupted text with masked language models.In experiments on robustness benchmarks across 3 tasks and 9 datasets, SSMBA consistently outperforms existing data augmentation methods and baseline models on both in-domain and OOD data, achieving gains of 0.8% accuracy on OOD Amazon reviews, 1.8% accuracy on OOD MNLI, and 1.4 BLEU on in-domain IWSLT14 German-English. 1 Nathan Ng 0001, Kyunghyun Cho, Marzyeh Ghassemi |
EMNLP (1) | 2 |
| 2020 | Unsupervised Question Decomposition for Question AnsweringabstractWe aim to improve question answering (QA) by decomposing hard questions into simpler sub-questions that existing QA systems are capable of answering.Since labeling questions with decompositions is cumbersome, we take an unsupervised approach to produce sub-questions, also enabling us to leverage millions of questions from the internet.Specifically, we propose an algorithm for One-to-N Unsupervised Sequence transduction (ONUS) that learns to map one hard, multi-hop question to many simpler, singlehop sub-questions.We answer sub-questions with an off-the-shelf QA model and give the resulting answers to a recomposition model that combines them into a final answer.We show large QA improvements on HOTPOTQA over a strong baseline on the original, out-ofdomain, and multi-hop dev sets.ONUS automatically learns to decompose different kinds of questions, while matching the utility of supervised and heuristic decomposition methods for QA and exceeding those methods in fluency.Qualitatively, we find that using subquestions is promising for shedding light on why a QA system makes a prediction. 1 * KC was a part-time research scientist at Facebook AI Research while working on this paper.1 Our code, data, and pretrained models are available at https://github.com/facebookresearch/ UnsupervisedDecomposition. What profession do H. L. Mencken and Albert Camus have in common? Ethan Perez, Patrick S. H. Lewis, Scott Yih, Kyunghyun Cho, Douwe Kiela |
EMNLP (1) | 4 |
| 2020 | Consistency of a Recurrent Language Model With Respect to Incomplete DecodingabstractDespite strong performance on a variety of tasks, neural sequence models trained with maximum likelihood have been shown to exhibit issues such as length bias and degenerate repetition.We study the related issue of receiving infinite-length sequences from a recurrent language model when using common decoding algorithms.To analyze this issue, we first define inconsistency of a decoding algorithm, meaning that the algorithm can yield an infinite-length sequence that has zero probability under the model.We prove that commonly used incomplete decoding algorithms -greedy search, beam search, top-k sampling, and nucleus sampling -are inconsistent, despite the fact that recurrent language models are trained to produce sequences of finite length.Based on these insights, we propose two remedies which address inconsistency: consistent variants of top-k and nucleus sampling, and a selfterminating recurrent language model.Empirical results show that inconsistency occurs in practice, and that the proposed methods prevent inconsistency. Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, Kyunghyun Cho |
EMNLP (1) | 5 |
| 2020 | The Break-Even Point on Optimization Trajectories of Deep Neural Networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, Krzysztof J. Geras |
ICLR | 6 |
| 2020 | Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
Cheolhyoung Lee, Kyunghyun Cho, Wanmo Kang |
ICLR | 2 |
| 2020 | Neural Text Generation With Unlikelihood Training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, Jason Weston |
ICLR | 5 |
| 2020 | Dynamics-Aware Embeddings
William F. Whitney, Rajat Agarwal, Kyunghyun Cho, Abhinav Gupta 0001 |
ICLR | 3 |
| 2020 | Classifier-agnostic saliency map extractionabstractCurrently available methods for extracting saliency maps identify parts of the input which are the most important to a specific fixed classifier. We show that this strong dependence on a given classifier hinders their performance. To address this problem, we propose classifier-agnostic saliency map extraction, which finds all parts of the image that any classifier could use, not just one given in advance. We observe that the proposed approach extracts higher quality saliency maps than prior work while being conceptually simple and easy to implement. The method sets the new state of the art result for localization task on the ImageNet data, outperforming all existing weakly-supervised localization techniques, despite not using the ground truth labels at the inference time. The code reproducing the results is available at https://github.com/kondiz/casme. Konrad Zolna, Krzysztof J. Geras, Kyunghyun Cho |
Comput. Vis. Image Underst. | 3 |
| 2020 | A Unified Framework of Online Learning Algorithms for Training Recurrent Neural NetworksabstractWe present a framework for compactly summarizing many recent results in efficient and/or biologically plausible online training of recurrent neural networks (RNN). The framework organizes algorithms according to several criteria: (a) past vs. future facing, (b) tensor structure, (c) stochastic vs. deterministic, and (d) closed form vs. numerical. These axes reveal latent conceptual connections among several recent advances in online learning. Furthermore, we provide novel mathematical intuitions for their degree of success. Testing these algorithms on two parametric task families shows that performances cluster according to our criteria. Although a similar clustering is also observed for pairwise gradient alignment, alignment with exact methods does not explain ultimate performance. This suggests the need for better comparison metrics. Owen Marschall, Kyunghyun Cho, Cristina Savin |
J. Mach. Learn. Res. | 2 |
| 2020 | Neural machine translation with a polysynthetic low resource language
John E. Ortega, Richard Castro Mamani, Kyunghyun Cho |
Mach. Transl. | 3 |
| 2020 | Deep Neural Networks Improve Radiologists' Performance in Breast Cancer ScreeningabstractWe present a deep convolutional neural network for breast cancer screening exam classification, trained, and evaluated on over 200000 exams (over 1000000 images). Our network achieves an AUC of 0.895 in predicting the presence of cancer in the breast, when tested on the screening population. We attribute the high accuracy to a few technical advances. 1) Our network's novel two-stage architecture and training procedure, which allows us to use a high-capacity patch-level network to learn from pixel-level labels alongside a network learning from macroscopic breast-level labels. 2) A custom ResNet-based network used as a building block of our model, whose balance of depth and width is optimized for high-resolution medical images. 3) Pretraining the network on screening BI-RADS classification, a related task with more noisy labels. 4) Combining multiple input views in an optimal way among a number of possible choices. To validate our model, we conducted a reader study with 14 readers, each reading 720 screening mammogram exams, and show that our model is as accurate as experienced radiologists when presented with the same data. We also show that a hybrid model, averaging the probability of malignancy predicted by a radiologist with a prediction of our neural network, is more accurate than either of the two separately. To further understand our results, we conduct a thorough analysis of our network's performance on different subpopulations of the screening population, the model's design, training procedure, errors, and properties of its internal representations. Our best models are publicly available at https://github.com/nyukat/breast_cancer_classifier. Nan Wu 0008, Jason Phang, Jungkyu Park 0001, Yiqiu Shen, Zhe Huang 0005, Masha Zorin, Stanislaw Jastrzebski, Thibault Févry, Joe Katsnelson, Eric Kim, Stacey Wolfson, Ujas Parikh, Sushma Gaddam, Leng Leng Young Lin, Kara Ho, Joshua D. Weinstein, Beatriu Reig, Yiming Gao 0003, Hildegard Toth, Kristine Pysarenko, Alana Lewin, Jiyon Lee, Krystal Airola, Eralda Mema, Stephanie Chung, Esther Hwang, Naziya Samreen, Sungheon Gene Kim, Laura Heacock, Linda Moy, Kyunghyun Cho, Krzysztof J. Geras |
IEEE Trans. Medical Imaging | 31 |
| 2019 | Classifier-Agnostic Saliency Map ExtractionabstractExtracting saliency maps, which indicate parts of the image important to classification, requires many tricks to achieve satisfactory performance when using classifier-dependent methods. Instead, we propose classifier-agnostic saliency map extraction. This allows to find all parts of the image that any classifier could use, not just one given in advance. This way we extract much higher quality saliency maps. Konrad Zolna, Krzysztof J. Geras, Kyunghyun Cho |
AAAI | 3 |
| 2019 | Improved Zero-shot Neural Machine Translation via Ignoring Spurious CorrelationsabstractZero-shot translation, translating between language pairs on which a Neural Machine Translation (NMT) system has never been trained, is an emergent property when training the system in multilingual settings.However, naïve training for zero-shot NMT easily fails, and is sensitive to hyper-parameter setting.The performance typically lags far behind the more conventional pivot-based approach which translates twice using a third language as a pivot.In this work, we address the degeneracy problem due to capturing spurious correlations by quantitatively analyzing the mutual information between language IDs of the source and decoded sentences.Inspired by this analysis, we propose to use two simple but effective approaches: (1) decoder pre-training; (2) backtranslation.These methods show significant improvement (4 ∼ 22 BLEU points) over the vanilla zero-shot translation on three challenging multilingual datasets, and achieve similar or better results than the pivot-based approach. Jiatao Gu, Yong Wang 0032, Kyunghyun Cho, Victor O. K. Li |
ACL (1) | 3 |
| 2019 | Generating Diverse Translations with Sentence CodesabstractUsers of machine translation systems may desire to obtain multiple candidates translated in different ways.In this work, we attempt to obtain diverse translations by using sentence codes to condition the sentence generation.We describe two methods to extract the codes, either with or without the help of syntax information.For diverse generation, we sample multiple candidates, each of which conditioned on a unique code.Experiments show that the sampled translations have much higher diversity scores when using reasonable sentence codes, where the translation quality is still on par with the baselines even under strong constraint imposed by the codes.In qualitative analysis, we show that our method is able to generate paraphrase translations with drastically different structures.The proposed approach can be easily adopted to existing translation systems as no modification to the model is required. Raphael Shu, Hideki Nakayama, Kyunghyun Cho |
ACL (1) | 3 |
| 2019 | Dialogue Natural Language InferenceabstractConsistency is a long standing issue faced by dialogue models. In this paper, we frame the consistency of dialogue agents as natural language inference (NLI) and create a new natural language inference dataset called Dialogue NLI. We propose a method which demonstrates that a model trained on Dialogue NLI can be used to improve the consistency of a dialogue model, and evaluate the method with human evaluation and with automatic metrics on a suite of evaluation sets designed to measure a dialogue model’s consistency. Sean Welleck, Jason Weston, Arthur Szlam, Kyunghyun Cho |
ACL (1) | 4 |
| 2019 | Retrieval-Augmented Convolutional Neural Networks Against Adversarial ExamplesabstractWe propose a retrieval-augmented convolutional network (RaCNN) and propose to train it with local mixup, a novel variant of the recently proposed mixup algorithm. The proposed hybrid architecture combining a convolutional network and an off-the-shelf retrieval engine was designed to mitigate the adverse effect of off-manifold adversarial examples, while the proposed local mixup addresses on-manifold ones by explicitly encouraging the classifier to locally behave linearly on the data manifold. Our evaluation of the proposed approach against seven readilyavailable adversarial attacks on three datasets-CIFAR-10, SVHN and ImageNet-demonstrate the improved robustness compared to a vanilla convolutional network, and comparable performance with the state-of-the-art reactive defense approaches. Jake Zhao, Kyunghyun Cho |
CVPR | 2 |
| 2019 | Emergent Linguistic Phenomena in Multi-Agent Communication GamesabstractLaura Harding Graesser, Kyunghyun Cho, Douwe Kiela. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Laura Graesser, Kyunghyun Cho, Douwe Kiela |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Towards Realistic Practices In Low-Resource Natural Language Processing: The Development SetabstractKatharina Kann, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Katharina Kann, Kyunghyun Cho, Samuel R. Bowman |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Countering Language Drift via Visual GroundingabstractJason Lee, Kyunghyun Cho, Douwe Kiela. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jason Lee 0002, Kyunghyun Cho, Douwe Kiela |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Finding Generalizable Evidence by Learning to Convince Q&A ModelsabstractEthan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, Kyunghyun Cho. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ethan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, Kyunghyun Cho |
EMNLP/IJCNLP (1) | 6 |
| 2019 | DialogWAE: Multimodal Response Generation with Conditional Wasserstein Auto-Encoder
Xiaodong Gu 0002, Kyunghyun Cho, Jung-Woo Ha 0001, Sunghun Kim 0001 |
ICLR (Poster) | 2 |
| 2019 | Non-Monotonic Sequential Text GenerationabstractStandard sequential generation methods assume a pre-specified generation order, such as text generation methods which generate words from left to right. In this work, we propose a framework for training models of text generation that operate in non-monotonic orders; the model directly learns good orders, without any additional annotation. Our framework operates by generating a word at an arbitrary position, and then recursively generating words to its left and then words to its right, yielding a binary tree. Learning is framed as imitation learning, including a coaching method which moves from imitating an oracle to reinforcing the policy’s own preferences. Experimental results demonstrate that using the proposed method, it is possible to learn policies which generate text without pre-specifying a generation order, while achieving competitive performance with conventional left-to-right generation. Sean Welleck, Kianté Brantley, Hal Daumé III, Kyunghyun Cho |
ICML | 4 |
| 2019 | Importance of Search and Evaluation Strategies in Neural Dialogue ModelingabstractWe investigate the impact of search strategies in neural dialogue modeling.We first compare two standard search algorithms, greedy and beam search, as well as our newly proposed iterative beam search which produces a more diverse set of candidate responses.We evaluate these strategies in realistic full conversations with humans and propose a modelbased Bayesian calibration to address annotator bias.These conversations are analyzed using two automatic metrics: log-probabilities assigned by the model and utterance diversity.Our experiments reveal that better search algorithms lead to higher rated conversations.However, finding the optimal selection mechanism to choose from a more diverse set of candidates is still an open question. Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, Jason Weston |
INLG | 3 |
| 2019 | Can Unconditional Language Models Recover Arbitrary Sentences?abstractNeural network-based generative language models like ELMo and BERT can work effectively as general purpose sentence encoders in text classification without further fine-tuning. Is it possible to adapt them in a similar way for use as general-purpose decoders? For this to be possible, it would need to be the case that for any target sentence of interest, there is some continuous representation that can be passed to the language model to cause it to reproduce that sentence. We set aside the difficult problem of designing an encoder that can produce such representations and, instead, ask directly whether such representations exist at all. To do this, we introduce a pair of effective, complementary methods for feeding representations into pretrained unconditional language models and a corresponding set of methods to map sentences into and out of this representation space, the reparametrized sentence space. We then investigate the conditions under which a language model can be made to generate a sentence through the identification of a point in such a space and find that it is possible to recover arbitrary sentences nearly perfectly with language models and representations of moderate size. Nishant Subramani, Samuel R. Bowman, Kyunghyun Cho |
NeurIPS | 3 |
| 2019 | Insertion-based Decoding with Automatically Inferred Generation OrderabstractConventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. We extend Transformer, a state-of-the-art sequence generation model, to efficiently implement the proposed approach, enabling it to be trained with either a pre-defined generation order or adaptive orders obtained from beam-search. Experiments on four real-world tasks, including word order recovery, machine translation, image caption, and code generation, demonstrate that our algorithm can generate sequences following arbitrary orders, while achieving competitive or even better performance compared with the conventional left-to-right generation. The generated sequences show that InDIGO adopts adaptive generation orders based on input information. Jiatao Gu, Qi Liu 0049, Kyunghyun Cho |
Trans. Assoc. Comput. Linguistics | 3 |
| 2018 | Search Engine Guided Neural Machine TranslationabstractIn this paper, we extend an attention-based neural machine translation (NMT) model by allowing it to access an entire training set of parallel sentence pairs even after training. The proposed approach consists of two stages. In the first stage –retrieval stage–, an off-the-shelf, black-box search engine is used to retrieve a small subset of sentence pairs from a training set given a source sentence. These pairs are further filtered based on a fuzzy matching score based on edit distance. In the second stage–translation stage–, a novel translation model, called search engine guided NMT (SEG-NMT), seamlessly uses both the source sentence and a set of retrieved sentence pairs to perform the translation. Empirical evaluation on three language pairs (En-Fr, En-De, and En-Es) shows that the proposed approach significantly outperforms the baseline approach and the improvement is more significant when more relevant sentence pairs were retrieved. Jiatao Gu, Yong Wang 0032, Kyunghyun Cho, Victor O. K. Li |
AAAI | 3 |
| 2018 | Zero-Shot Transfer Learning for Event ExtractionabstractMost previous supervised event extraction methods have relied on features derived from manual annotations, and thus cannot be applied to new event types without extra annotation effort.We take a fresh look at event extraction and model it as a generic grounding problem: mapping each event mention to a specific type in a target event ontology.We design a transferable architecture of structural and compositional neural networks to jointly represent and map event mentions and types into a shared semantic space.Based on this new framework, we can select, for each event mention, the event type which is semantically closest in this space as its type.By leveraging manual annotations available for a small set of existing event types, our framework can be applied to new unseen event types without additional manual annotations.When tested on 23 unseen event types, this zeroshot framework, without manual annotations, achieves performance comparable to a supervised model trained from 3,000 sentences annotated with 500 event mentions.1 Lifu Huang, Heng Ji 0001, Kyunghyun Cho, Ido Dagan, Sebastian Riedel 0001, Clare R. Voss |
ACL (1) | 3 |
| 2018 | Letting a Neural Network Decide Which Machine Translation System to Use for Black-Box Fuzzy-Match RepairabstractWhile systems using the Neural Network-based Machine Translation (NMT) paradigm achieve the highest scores on recent shared tasks, phrase-based (PBMT) systems, rule-based (RBMT) systems and other systems may get better results for individual examples. Therefore, combined systems should achieve the best results for MT, particularly if the system combination method can take advantage of the strengths of each paradigm. In this paper, we describe a system that predicts whether a NMT, PBMT or RBMT will get the best Spanish translation result for a particular English sentence in DGT-TM 20161. Then we use fuzzy-match repair (FMR) as a mechanism to show that the combined system outperforms individual systems in a black-box machine translation setting. John E. Ortega, Weiyi Lu, Adam Meyers 0001, Kyunghyun Cho |
EAMT | 4 |
| 2018 | A Stable and Effective Learning Strategy for Trainable Greedy DecodingabstractBeam search is a widely used approximate search strategy for neural network decoders, and it generally outperforms simple greedy decoding on tasks like machine translation.However, this improvement comes at substantial computational cost.In this paper, we propose a flexible new method that allows us to reap nearly the full benefits of beam search with nearly no additional computational cost.The method revolves around a small neural network actor that is trained to observe and manipulate the hidden state of a previouslytrained decoder.To train this actor network, we introduce the use of a pseudo-parallel corpus built using the output of beam search on a base model, ranked by a target quality metric like BLEU.Our method is inspired by earlier work on this problem, but requires no reinforcement learning, and can be trained reliably on a range of models.Experiments on three parallel corpora and three architectures show that the method yields substantial improvements in translation quality and speed over each base system. Yun Chen 0007, Victor O. K. Li, Kyunghyun Cho, Samuel R. Bowman |
EMNLP | 3 |
| 2018 | Meta-Learning for Low-Resource Neural Machine TranslationabstractIn this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm (MAML, Finn et al., 2017) for lowresource neural machine translation (NMT).We frame low-resource translation as a metalearning problem, and we learn to adapt to low-resource languages based on multilingual high-resource language tasks.We use the universal lexical representation (Gu et al., 2018b) to overcome the input-output mismatch across different languages.We evaluate the proposed meta-learning strategy using eighteen European languages (Bg, Cs, Da, De, El, Es, Et, Fr, Hu, It, Lt, Nl, Pl, Pt, Sk, Sl, Sv and Ru) as source tasks and five diverse languages (Ro, Lv, Fi, Tr and Ko) as target tasks.We show that the proposed approach significantly outperforms the multilingual, transfer learning based approach (Zoph et al., 2016) and enables us to train a competitive NMT system with only a fraction of training examples.For instance, the proposed approach can achieve as high as 22.04 BLEU on Romanian-English WMT'16 by seeing only 16,000 translated words (⇠ 600 parallel sentences). Jiatao Gu, Yong Wang 0032, Yun Chen 0007, Victor O. K. Li, Kyunghyun Cho |
EMNLP | 5 |
| 2018 | Conditional Word Embedding and Hypothesis Testing via Bayes-by-BackpropabstractConventional word embedding models do not leverage information from document metadata, and they do not model uncertainty.We address these concerns with a model that incorporates document covariates to estimate conditional word embedding distributions.Our model allows for (a) hypothesis tests about the meanings of terms, (b) assessments as to whether a word is near or far from another conditioned on different covariate values, and (c) assessments as to whether estimated differences are statistically significant. Rujun Han, Michael Gill, Arthur Spirling, Kyunghyun Cho |
EMNLP | 4 |
| 2018 | Grammar Induction with Neural Language Models: An Unusual ReplicationabstractA substantial thread of recent work on latent tree learning has attempted to develop neural network models with parse-valued latent variables and train them on non-parsing tasks, in the hope of having them discover interpretable tree structure.In a recent paper, Shen et al. (2018) introduce such a model and report nearstate-of-the-art results on the target task of language modeling, and the first strong latent tree learning result on constituency parsing.In an attempt to reproduce these results, we discover issues that make the original results hard to trust, including tuning and even training on what is effectively the test set.Here, we attempt to reproduce these results in a fair experiment and to extend them to two new datasets.We find that the results of this work are robust: All variants of the model under study outperform all latent tree learning baselines, and perform competitively with symbolic grammar induction systems.We find that this model represents the first empirical success for latent tree learning, and that neural network language modeling warrants further study as a setting for grammar induction. Phu Mon Htut, Kyunghyun Cho, Samuel R. Bowman |
EMNLP | 2 |
| 2018 | Multi-lingual Common Semantic Space Construction via Cluster-Consistent Word EmbeddingabstractWe construct a multilingual common semantic space based on distributional semantics, where words from multiple languages are projected into a shared space via which all available resources and knowledge can be shared across multiple languages.Beyond word alignment, we introduce multiple cluster-level alignments and enforce the word clusters to be consistently distributed across multiple languages.We exploit three signals for clustering: (1) neighbor words in the monolingual word embedding space; (2) character-level information; and (3) linguistic properties (e.g., apposition, locative suffix) derived from linguistic structure knowledge bases available for thousands of languages.We introduce a new cluster-consistent correlational neural network to construct the common semantic space by aligning words as well as clusters.Intrinsic evaluation on monolingual and multilingual QVEC tasks shows our approach achieves significantly higher correlation with linguistic features which are extracted from manually crafted lexical resources than state-of-the-art multi-lingual embedding learning methods do.Using low-resource language name tagging as a case study for extrinsic evaluation, our approach achieves up to 14.6% absolute F-score gain over the state of the art on cross-lingual direct transfer.Our approach is also shown to be robust even when the size of bilingual dictionary is small.1 Lifu Huang, Kyunghyun Cho, Boliang Zhang, Heng Ji 0001, Kevin Knight |
EMNLP | 2 |
| 2018 | Dynamic Meta-Embeddings for Improved Sentence RepresentationsabstractWhile one of the first steps in many NLP systems is selecting what pre-trained word embeddings to use, we argue that such a step is better left for neural networks to figure out by themselves.To that end, we introduce dynamic meta-embeddings, a simple yet effective method for the supervised learning of embedding ensembles, which leads to stateof-the-art performance within the same model class on a variety of tasks.We subsequently show how the technique can be used to shed new light on the usage of word embeddings in NLP systems. Douwe Kiela, Changhan Wang, Kyunghyun Cho |
EMNLP | 3 |
| 2018 | Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative RefinementabstractWe propose a conditional non-autoregressive neural sequence model based on iterative refinement.The proposed model is designed based on the principles of latent variable models and denoising autoencoders, and is generally applicable to any sequence generation task.We extensively evaluate the proposed model on machine translation (En$De and En$Ro) and image caption generation, and observe that it significantly speeds up decoding while maintaining the generation quality comparable to the autoregressive counterpart. Jason Lee 0002, Elman Mansimov, Kyunghyun Cho |
EMNLP | 3 |
| 2018 | Breast Density Classification with Deep Convolutional Neural NetworksabstractBreast density classification is an essential part of breast cancer screening. Although a lot of prior work considered this problem as a task for learning algorithms, to our knowledge, all of them used small and not clinically realistic data both for training and evaluation of their models. In this work, we explored the limits of this task with a data set coming from over 200,000 breast cancer screening exams. We used this data to train and evaluate a strong convolutional neural network classifier. In a reader study, we found that our model can perform this task comparably to a human expert. Nan Wu 0008, Krzysztof J. Geras, Yiqiu Shen, Jingyi Su, Sungheon Gene Kim, Eric Kim, Stacey Wolfson, Linda Moy, Kyunghyun Cho |
ICASSP | 9 |
| 2018 | Unsupervised Neural Machine Translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, Kyunghyun Cho |
ICLR (Poster) | 4 |
| 2018 | Emergent Communication in a Multi-Modal, Multi-Step Referential Game
Katrina Evtimova, Andrew Drozdov, Douwe Kiela, Kyunghyun Cho |
ICLR (Poster) | 4 |
| 2018 | Boundary Seeking GANs
R. Devon Hjelm, Athul Paul Jacob, Adam Trischler, Gerry Che, Kyunghyun Cho, Yoshua Bengio |
ICLR (Poster) | 5 |
| 2018 | Emergent Translation in Multi-Agent Communication
Jason Lee 0002, Kyunghyun Cho, Jason Weston, Douwe Kiela |
ICLR (Poster) | 2 |
| 2018 | Loss Functions for Multiset PredictionabstractWe study the problem of multiset prediction. The goal of multiset prediction is to train a predictor that maps an input to a multiset consisting of multiple items. Unlike existing problems in supervised learning, such as classification, ranking and sequence generation, there is no known order among items in a target multiset, and each item in the multiset may appear more than once, making this problem extremely challenging. In this paper, we propose a novel multiset loss function by viewing this problem from the perspective of sequential decision making. The proposed multiset loss function is empirically evaluated on two families of datasets, one synthetic and the other real, with varying levels of difficulty, against various baseline loss functions including reinforcement learning, sequence, and aggregated distribution matching loss functions. The experiments reveal the effectiveness of the proposed loss function over the others. Sean Welleck, Zixin Yao, Yu Gai, Jialin Mao, Zheng Zhang 0001, Kyunghyun Cho |
NeurIPS | 6 |
| 2018 | Fine-grained attention mechanism for neural machine translation
Heeyoul Choi, Kyunghyun Cho, Yoshua Bengio |
Neurocomputing | 2 |
| 2018 | Dynamic Neural Turing Machine with Continuous and Discrete Addressing SchemesabstractWe extend the neural Turing machine (NTM) model into a dynamic neural Turing machine (D-NTM) by introducing trainable address vectors. This addressing scheme maintains for each memory cell two separate vectors, content and address vectors. This allows the D-NTM to learn a wide variety of location-based addressing strategies, including both linear and nonlinear ones. We implement the D-NTM with both continuous and discrete read and write mechanisms. We investigate the mechanisms and effects of learning to read and write into a memory through experiments on Facebook bAbI tasks using both a feedforward and GRU controller. We provide extensive analysis of our model and compare different variations of neural Turing machines on this task. We show that our model outperforms long short-term memory and NTM variants. We provide further experimental results on the sequential [Formula: see text]MNIST, Stanford Natural Language Inference, associative recall, and copy tasks. Caglar Gulcehre, Sarath Chandar, Kyunghyun Cho, Yoshua Bengio |
Neural Comput. | 3 |
| 2017 | Query-Efficient Imitation Learning for End-to-End Simulated DrivingabstractOne way to approach end-to-end autonomous driving is to learn a policy that maps from a sensory input, such as an image frame from a front-facing camera, to a driving action, by imitating an expert driver, or a reference policy. This can be done by supervised learning, where a policy is tuned to minimize the difference between the predicted and ground-truth actions. A policy trained in this way however is known to suffer from unexpected behaviours due to the mismatch between the states reachable by the reference policy and trained policy. More advanced algorithms for imitation learning, such as DAgger, addresses this issue by iteratively collecting training examples from both reference and trained policies. These algorithms often require a large number of queries to a reference policy, which is undesirable as the reference policy is often expensive. In this paper, we propose an extension of the DAgger, called SafeDAgger, that is query-efficient and more suitable for end-to-end autonomous driving. We evaluate the proposed SafeDAgger in a car racing simulator and show that it indeed requires less queries to a reference policy. We observe a significant speed up in convergence, which we conjecture to be due to the effect of automated curriculum learning. Jiakai Zhang, Kyunghyun Cho |
AAAI | 2 |
| 2017 | Learning to Translate in Real-time with Neural Machine TranslationabstractTranslating in real-time, a.k.a.simultaneous translation, outputs translation words before the input sentence ends, which is a challenging problem for conventional machine translation methods.We propose a neural machine translation (NMT) framework for simultaneous translation in which an agent learns to make decisions on when to translate from the interaction with a pre-trained NMT environment.To trade off quality and delay, we extensively explore various targets for delay and design a method for beam-search applicable in the simultaneous MT setting.Experiments against state-of-the-art baselines on two language pairs demonstrate the efficacy of the proposed framework both quantitatively and qualitatively. 1 Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O. K. Li |
EACL (1) | 3 |
| 2017 | Trainable Greedy Decoding for Neural Machine TranslationabstractRecent research in neural machine translation has largely focused on two aspects; neural network architectures and end-toend learning algorithms.The problem of decoding, however, has received relatively little attention from the research community.In this paper, we solely focus on the problem of decoding given a trained neural machine translation model.Instead of trying to build a new decoding algorithm for any specific decoding objective, we propose the idea of trainable decoding algorithm in which we train a decoding algorithm to find a translation that maximizes an arbitrary decoding objective.More specifically, we design an actor that observes and manipulates the hidden state of the neural machine translation decoder and propose to train it using a variant of deterministic policy gradient.We extensively evaluate the proposed algorithm using four language pairs and two decoding objectives, and show that we can indeed train a trainable greedy decoder that generates a better translation (in terms of a target decoding objective) with minimal computational overhead. Jiatao Gu, Kyunghyun Cho, Victor O. K. Li |
EMNLP | 2 |
| 2017 | Task-Oriented Query Reformulation with Reinforcement LearningabstractSearch engines play an important role in our everyday lives by assisting us in finding the information we need.When we input a complex query, however, results are often far from satisfactory.In this work, we introduce a query reformulation system based on a neural network that rewrites a query to maximize the number of relevant documents returned.We train this neural network with reinforcement learning.The actions correspond to selecting terms to build a reformulated query, and the reward is the document recall.We evaluate our approach on three datasets against strong baselines and show a relative improvement of 5-20% in terms of recall.Furthermore, we present a simple method to estimate a conservative upperbound performance of a model in a particular environment and verify that there is still large room for improvements. Rodrigo Nogueira 0001, Kyunghyun Cho |
EMNLP | 2 |
| 2017 | Convolutional recurrent neural networks for music classificationabstractWe introduce a convolutional recurrent neural network (CRNN) for music tagging. CRNNs take advantage of convolutional neural networks (CNNs) for local feature extraction and recurrent neural networks for temporal summarisation of the extracted features. We compare CRNN with three CNN structures that have been used for music tagging while controlling the number of parameters with respect to their performance and training time per sample. Overall, we found that CRNNs show a strong performance with respect to the number of parameter and training time, indicating the effectiveness of its hybrid structure in music feature extraction and feature summarisation. Keunwoo Choi, György Fazekas, Mark B. Sandler, Kyunghyun Cho |
ICASSP | 4 |
| 2017 | Saliency-based Sequential Image Attention with Multiset PredictionabstractHumans process visual scenes selectively and sequentially using attention. Central to models of human visual attention is the saliency map. We propose a hierarchical visual architecture that operates on a saliency map and uses a novel attention mechanism to sequentially focus on salient regions and take additional glimpses within those regions. The architecture is motivated by human visual attention, and is used for multi-label image classification on a novel multiset task, demonstrating that it achieves high precision and recall while localizing objects with its attention. Unlike conventional multi-label image classification models, the model supports multiset prediction due to a reinforcement-learning based training process that allows for arbitrary label permutation and multiple instances per label. Sean Welleck, Jialin Mao, Kyunghyun Cho, Zheng Zhang 0001 |
NIPS | 3 |
| 2017 | Context-dependent word representation for neural machine translation
Heeyoul Choi, Kyunghyun Cho, Yoshua Bengio |
Comput. Speech Lang. | 2 |
| 2017 | Introduction to the special issue on deep learning approaches for machine translation
Marta R. Costa-jussà, Alexandre Allauzen, Loïc Barrault, Kyunghyun Cho, Holger Schwenk |
Comput. Speech Lang. | 4 |
| 2017 | Multi-way, multilingual neural machine translation
Orhan Firat, Kyunghyun Cho, Baskaran Sankaran, Fatos T. Yarman-Vural, Yoshua Bengio |
Comput. Speech Lang. | 2 |
| 2017 | On integrating a language model into neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Yoshua Bengio |
Comput. Speech Lang. | 4 |
| 2017 | The representational geometry of word meanings acquired by neural machine translation models
Felix Hill, Kyunghyun Cho, Sébastien Jean, Yoshua Bengio |
Mach. Transl. | 2 |
| 2017 | Fully Character-Level Neural Machine Translation without Explicit SegmentationabstractMost existing machine translation systems operate at the level of words, relying on explicit segmentation to extract tokens. We introduce a neural machine translation (NMT) model that maps a source character sequence to a target character sequence without any segmentation. We employ a character-level convolutional network with max-pooling at the encoder to reduce the length of source representation, allowing the model to be trained at a speed comparable to subword-level models while capturing local regularities. Our character-to-character model outperforms a recently proposed baseline with a subword-level encoder on WMT’15 DE-EN and CS-EN, and gives comparable performance on FI-EN and RU-EN. We then demonstrate that it is possible to share a single character-level encoder across multiple languages by training a model on a many-to-one translation task. In this multilingual setting, the character-level encoder significantly outperforms the subword-level encoder on all the language pairs. We observe that on CS-EN, FI-EN and RU-EN, the quality of the multilingual character-level translation even surpasses the models specifically trained on that language pair alone, both in terms of the BLEU score and human judgment. Jason Lee 0002, Kyunghyun Cho, Thomas Hofmann 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | A Character-level Decoder without Explicit Segmentation for Neural Machine TranslationabstractThe existing machine translation systems, whether phrase-based or neural, have relied almost exclusively on word-level modelling with explicit segmentation.In this paper, we ask a fundamental question: can neural machine translation generate a character sequence without any explicit segmentation?To answer this question, we evaluate an attention-based encoderdecoder with a subword-level encoder and a character-level decoder on four language pairs-En-Cs, En-De, En-Ru and En-Fiusing the parallel corpora from WMT'15.Our experiments show that the models with a character-level decoder outperform the ones with a subword-level decoder on all of the four language pairs.Furthermore, the ensembles of neural models with a character-level decoder outperform the state-of-the-art non-neural machine translation systems on En-Cs, En-De and En-Fi and perform comparably on En-Ru. Junyoung Chung, Kyunghyun Cho, Yoshua Bengio |
ACL (1) | 2 |
| 2016 | Larger-Context Language Modelling with Recurrent Neural NetworkabstractIn this work, we propose a novel method to incorporate corpus-level discourse information into language modelling.We call this larger-context language model.We introduce a late fusion approach to a recurrent language model based on long shortterm memory units (LSTM), which helps the LSTM unit keep intra-sentence dependencies and inter-sentence dependencies separate from each other.Through the evaluation on four corpora (IMDB, BBC, Penn TreeBank, and Fil9), we demonstrate that the proposed model improves perplexity significantly.In the experiments, we evaluate the proposed approach while varying the number of context sentences and observe that the proposed late fusion is superior to the usual way of incorporating additional inputs to the LSTM.By analyzing the trained larger-context language model, we discover that content words, including nouns, adjectives and verbs, benefit most from an increasing number of context sentences.This analysis suggests that larger-context language model improves the unconditional language model by capturing the theme of a document better and more easily. *Recently, (Ji et al., 2015) independently proposed a similar approach. Kyunghyun Cho |
ACL (1) | 2 |
| 2016 | Oracle Performance for Visual Captioning
Nicolas Ballas, Kyunghyun Cho, John R. Smith, Yoshua Bengio |
BMVC | 3 |
| 2016 | A Correlational Encoder Decoder Architecture for Pivot Based Sequence GenerationabstractInterlingua based Machine Translation (MT) aims to encode multiple languages into a common linguistic representation and then decode sentences in multiple target languages from this representation. In this work we explore this idea in the context of neural encoder decoder architectures, albeit on a smaller scale and without MT as the end goal. Specifically, we consider the case of three languages or modalities X, Z and Y wherein we are interested in generating sequences in Y starting from information available in X. However, there is no parallel training data available between X and Y but, training data is available between X & Z and Z & Y (as is often the case in many real world applications). Z thus acts as a pivot/bridge. An obvious solution, which is perhaps less elegant but works very well in practice is to train a two stage model which first converts from X to Z and then from Z to Y. Instead we explore an interlingua inspired solution which jointly learns to do the following (i) encode X and Z to a common representation and (ii) decode Y from this common representation. We evaluate our model on two tasks: (i) bridge transliteration and (ii) bridge captioning. We report promising results in both these applications and believe that this is a right step towards truly interlingua inspired encoder decoder architectures. Amrita Saha, Mitesh M. Khapra, Sarath Chandar, Janarthanan Rajendran, Kyunghyun Cho |
COLING | 5 |
| 2016 | Zero-Resource Translation with Multi-Lingual Neural Machine TranslationabstractIn this paper, we propose a novel finetuning algorithm for the recently introduced multiway, multilingual neural machine translate that enables zero-resource machine translation.When used together with novel manyto-one translation strategies, we empirically show that this finetuning algorithm allows the multi-way, multilingual model to translate a zero-resource language pair (1) as well as a single-pair neural translation model trained with up to 1M direct parallel sentences of the same language pair and (2) better than pivotbased translation strategy, while keeping only one additional copy of attention-related parameters. Orhan Firat, Baskaran Sankaran, Yaser Al-Onaizan, Fatos T. Yarman-Vural, Kyunghyun Cho |
EMNLP | 5 |
| 2016 | Gated Word-Character Recurrent Language ModelabstractWe introduce a recurrent neural network language model (RNN-LM) with long shortterm memory (LSTM) units that utilizes both character-level and word-level inputs.Our model has a gate that adaptively finds the optimal mixture of the character-level and wordlevel inputs.The gate creates the final vector representation of a word by combining two distinct representations of the word.The character-level inputs are converted into vector representations of words using a bidirectional LSTM.The word-level inputs are projected into another high-dimensional space by a word lookup table.The final vector representations of words are used in the LSTM language model which predicts the next word given all the preceding words.Our model with the gating mechanism effectively utilizes the character-level inputs for rare and out-ofvocabulary words and outperforms word-level language models on several English corpora. Yasumasa Onoe, Kyunghyun Cho |
EMNLP | 2 |
| 2016 | Multi-Way, Multilingual Neural Machine Translation with a Shared Attention MechanismabstractWe propose multi-way, multilingual neural machine translation.The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages.This is made possible by having a single attention mechanism that is shared across all language pairs.We train the proposed multiway, multilingual model on ten language pairs from WMT'15 simultaneously and observe clear performance improvements over models trained on only one language pair.In particular, we observe that the proposed model significantly improves the translation quality of low-resource language pairs. Orhan Firat, Kyunghyun Cho, Yoshua Bengio |
HLT-NAACL | 2 |
| 2016 | Learning Distributed Representations of Sentences from Unlabelled DataabstractUnsupervised methods for learning distributed representations of words are ubiquitous in today's NLP research, but far less is known about the best ways to learn distributed phrase or sentence representations from unlabelled data. This paper is a systematic comparison of models that learn such representations. We find that the optimal approach depends critically on the intended application. Deeper, more complex models are preferable for representations to be used in supervised systems, but shallow log-linear models work best for building representation spaces that can be decoded with simple spatial distance metrics. We also propose two new unsupervised representation-learning objectives designed to optimise the trade-off between training time, domain portability and performance. Felix Hill, Kyunghyun Cho, Anna Korhonen |
HLT-NAACL | 2 |
| 2016 | Joint Event Extraction via Recurrent Neural NetworksabstractEvent extraction is a particularly challenging problem in information extraction.The stateof-the-art models for this problem have either applied convolutional neural networks in a pipelined framework (Chen et al., 2015) or followed the joint architecture via structured prediction with rich local and global features (Li et al., 2013).The former is able to learn hidden feature representations automatically from data based on the continuous and generalized representations of words.The latter, on the other hand, is capable of mitigating the error propagation problem of the pipelined approach and exploiting the inter-dependencies between event triggers and argument roles via discrete structures.In this work, we propose to do event extraction in a joint framework with bidirectional recurrent neural networks, thereby benefiting from the advantages of the two models as well as addressing issues inherent in the existing approaches.We systematically investigate different memory features for the joint model and demonstrate that the proposed model achieves the state-of-the-art performance on the ACE 2005 dataset. Thien Huu Nguyen, Kyunghyun Cho, Ralph Grishman |
HLT-NAACL | 2 |
| 2016 | Iterative Refinement of the Approximate Posterior for Directed Belief NetworksabstractVariational methods that rely on a recognition network to approximate the posterior of directed graphical models offer better inference and learning than previous methods. Recent advances that exploit the capacity and flexibility in this approach have expanded what kinds of models can be trained. However, as a proposal for the posterior, the capacity of the recognition network is limited, which can constrain the representational power of the generative model and increase the variance of Monte Carlo estimates. To address these issues, we introduce an iterative refinement procedure for improving the approximate posterior of the recognition network and show that training with the refined posterior is competitive with state-of-the-art methods. The advantages of refinement are further evident in an increased effective sample size, which implies a lower variance of gradient estimates. R. Devon Hjelm, Ruslan Salakhutdinov, Kyunghyun Cho, Nebojsa Jojic, Vince D. Calhoun, Junyoung Chung |
NIPS | 3 |
| 2016 | End-to-End Goal-Driven Web NavigationabstractWe propose a goal-driven web navigation as a benchmark task for evaluating an agent with abilities to understand natural language and plan on partially observed environments. In this challenging task, an agent navigates through a website, which is represented as a graph consisting of web pages as nodes and hyperlinks as directed edges, to find a web page in which a query appears. The agent is required to have sophisticated high-level reasoning based on natural languages and efficient sequential decision-making capability to succeed. We release a software tool, called WebNav, that automatically transforms a website into this goal-driven web navigation task, and as an example, we make WikiNav, a dataset constructed from the English Wikipedia. We extensively evaluate different variants of neural net based artificial agents on WikiNav and observe that the proposed goal-driven web navigation well reflects the advances in models, making it a suitable benchmark for evaluating future progress. Furthermore, we extend the WikiNav with question-answer pairs from Jeopardy! and test the proposed agent based on recurrent neural networks against strong inverted index based search engines. The artificial agents trained on WikiNav outperforms the engined based approaches, demonstrating the capability of the proposed goal-driven navigation as a good proxy for measuring the progress in real-world tasks such as focused crawling and question-answering. Rodrigo Nogueira 0001, Kyunghyun Cho |
NIPS | 2 |
| 2016 | Learning to Understand Phrases by Embedding the DictionaryabstractDistributional models that learn rich semantic word representations are a success story of recent NLP research. However, developing models that learn useful representations of phrases and sentences has proved far harder. We propose using the definitions found in everyday dictionaries as a means of bridging this gap between lexical and phrasal semantics. Neural language embedding models can be effectively trained to map dictionary definitions (phrases) to (lexical) representations of the words defined by those definitions. We present two applications of these architectures: reverse dictionaries that return the name of a concept given a definition or description and general-knowledge crossword question answerers. On both tasks, neural language embedding models trained on definitions from a handful of freely-available lexical resources perform as well or better than existing commercial systems that rely on significant task-specific engineering. The results highlight the effectiveness of both neural embedding architectures and definition-based training for developing models that understand phrases and sentences. Felix Hill, Kyunghyun Cho, Anna Korhonen, Yoshua Bengio |
Trans. Assoc. Comput. Linguistics | 2 |
| 2015 | On Using Very Large Target Vocabulary for Neural Machine TranslationabstractSébastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Sébastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio |
ACL (1) | 2 |
| 2015 | Describing Videos by Exploiting Temporal StructureabstractRecent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description. However, while images are static, working with videos requires modeling their dynamic temporal structure and then properly integrating that information into a natural language description model. In this context, we propose an approach that successfully takes into account both the local and global temporal structure of videos to produce descriptions. First, our approach incorporates a spatial temporal 3-D convolutional neural network (3-D CNN) representation of the short temporal dynamics. The 3-D CNN representation is trained on video action recognition tasks, so as to produce a representation that is tuned to human motion and behavior. Second we propose a temporal attention mechanism that allows to go beyond local temporal modeling and learns to automatically select the most relevant temporal segments given the text-generating RNN. Our approach exceeds the current state-of-art for both BLEU and METEOR metrics on the Youtube2Text dataset. We also present results on a new, larger and more challenging dataset of paired video and natural language descriptions. Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Joseph Pal, Hugo Larochelle, Aaron C. Courville |
ICCV | 3 |
| 2015 | Gated Feedback Recurrent Neural NetworksabstractIn this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper recurrent layers to lower layers using a global gating unit for each pair of layers. The recurrent signals exchanged between layers are gated adaptively based on the previous hidden states and the current input. We evaluated the proposed GF-RNN with different types of recurrent units, such as tanh, long short-term memory and gated recurrent units, on the tasks of character-level language modeling and Python program evaluation. Our empirical evaluation of different RNN units, revealed that in both tasks, the GF-RNN outperforms the conventional approaches to build deep stacked RNNs. We suggest that the improvement arises because the GF-RNN can adaptively assign different layers to different timescales and layer-to-layer interactions (including the top-down ones which are not usually present in a stacked RNN) by learning to gate these interactions. Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, Yoshua Bengio |
ICML | 3 |
| 2015 | Show, Attend and Tell: Neural Image Caption Generation with Visual AttentionabstractInspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images. We describe how we can train this model in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound. We also show through visualization how the model is able to automatically learn to fix its gaze on salient objects while generating the corresponding words in the output sequence. We validate the use of attention with state-of-the-art performance on three benchmark datasets: Flickr8k, Flickr30k and MS COCO. Kelvin Xu, Jimmy Ba, Jamie Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, Yoshua Bengio |
ICML | 4 |
| 2015 | A study of the recurrent neural network encoder-decoder for large vocabulary speech recognitionabstractDeep neural networks have advanced the state-of-the-art in automatic speech recognition, when combined with hidden Markov models (HMMs). Recently there has been interest in using systems based on recurrent neural networks (RNNs) to perform sequence modelling directly, without the requirement of an HMM superstructure. In this paper, we study the RNN encoder-decoder approach for large vocabulary end-toend speech recognition, whereby an encoder transforms a sequence of acoustic vectors into a sequence of feature representations, from which a decoder recovers a sequence of words. We investigated this approach on the Switchboard corpus using a training set of around 300 hours of transcribed audio data. Without the use of an explicit language model or pronunciation lexicon, we achieved promising recognition accuracy, demonstrating that this approach warrants further investigation. Index Terms: end-to-end speech recognition, deep neural networks, recurrent neural networks, encoder-decoder. Liang Lu 0001, Xingxing Zhang 0002, Kyunghyun Cho, Steve Renals |
INTERSPEECH | 3 |
| 2015 | Attention-Based Models for Speech RecognitionabstractRecurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks including machine translation, handwriting synthesis and image caption generation. We extend the attention-mechanism with features needed for speech recognition. We show that while an adaptation of the model used for machine translation reaches a competitive 18.6\% phoneme error rate (PER) on the TIMIT phoneme recognition task, it can only be applied to utterances which are roughly as long as the ones it was trained on. We offer a qualitative explanation of this failure and propose a novel and generic method of adding location-awareness to the attention mechanism to alleviate this issue. The new method yields a model that is robust to long inputs and achieves 18\% PER in single utterances and 20\% in 10-times longer (repeated) utterances. Finally, we propose a change to the attention mechanism that prevents it from concentrating too much on single frames, which further reduces PER to 17.6\% level. Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, Yoshua Bengio |
NIPS | 4 |
| 2015 | Learning Distributed Representations from Reviews for Collaborative FilteringabstractRecent work has shown that collaborative filter-based recommender systems can be improved by incorporating side information, such as natural language reviews, as a way of regularizing the derived product representations. Motivated by the success of this approach, we introduce two different models of reviews and study their effect on collaborative filtering performance. While the previous state-of-the-art approach is based on a latent Dirichlet allocation (LDA) model of reviews, the models we explore are neural network based: a bag-of-words product-of-experts model and a recurrent neural network. We demonstrate that the increased flexibility offered by the product-of-experts model allowed it to achieve state-of-the-art performance on the Amazon review dataset, outperforming the LDA-based approach. However, interestingly, the greater modeling power offered by the recurrent neural network appears to undermine the model's ability to act as a regularizer of the product representations. Amjad Almahairi, Kyle Kastner, Kyunghyun Cho, Aaron C. Courville |
RecSys | 3 |
| 2015 | Measuring the usefulness of hidden units in Boltzmann machines with mutual information
Mathias Berglund, Tapani Raiko, Kyunghyun Cho |
Neural Networks | 3 |
| 2015 | Two-layer contractive encodings for learning stable nonlinear features
Hannes Schulz, Kyunghyun Cho, Tapani Raiko, Sven Behnke |
Neural Networks | 2 |
| 2015 | Describing Multimedia Content Using Attention-Based Encoder-Decoder NetworksabstractWhereas deep neural networks were first mostly used for classification tasks, they are rapidly expanding in the realm of structured output problems, where the observed target is composed of multiple random variables that have a rich joint distribution, given the input. In this paper we focus on the case where the input also has a rich structure and the input and output structures are somehow related. We describe systems that learn to attend to different places in the input, for each element of the output, for a variety of tasks: machine translation, image caption generation, video clip description, and speech recognition. All these systems are based on a shared set of building blocks: gated recurrent neural networks and convolutional neural networks, along with trained attention mechanisms. We report on experimental results with these systems, showing impressively good performance and the advantage of the attention mechanism. Kyunghyun Cho, Aaron C. Courville, Yoshua Bengio |
IEEE Trans. Multim. | 1 |
| 2014 | Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine TranslationabstractKyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014. Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio |
EMNLP | 1 |
| 2014 | Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, Yoshua Bengio |
NIPS | 4 |
| 2014 | On the Number of Linear Regions of Deep Neural Networks
Guido Montúfar, Razvan Pascanu, Kyunghyun Cho, Yoshua Bengio |
NIPS | 3 |
| 2014 | Iterative Neural Autoregressive Distribution Estimator NADE-k
Tapani Raiko, Kyunghyun Cho, Yoshua Bengio |
NIPS | 3 |
| 2014 | Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks
Caglar Gulcehre, Kyunghyun Cho, Razvan Pascanu, Yoshua Bengio |
ECML/PKDD (1) | 2 |
| 2014 | On the Equivalence between Deep NADE and Generative Stochastic Networks
Sherjil Ozair, Kyunghyun Cho, Yoshua Bengio |
ECML/PKDD (3) | 3 |
| 2013 | Boltzmann Machines for Image Denoising
Kyunghyun Cho |
ICANN | 1 |
| 2013 | A Two-Stage Pretraining Algorithm for Deep Boltzmann Machines
Kyunghyun Cho, Tapani Raiko, Alexander Ilin, Juha Karhunen |
ICANN | 1 |
| 2013 | Gaussian-Bernoulli restricted Boltzmann machines and automatic feature extraction for noise robust missing data mask estimationabstractA missing data mask estimation method based on Gaussian-Bernoulli restricted Boltzmann machine (GRBM) trained on cross-correlation representation of the audio signal is presented in the study. The automatically learned features by the GRBM are utilized in dividing the time-frequency units of the spectrographic mask into noise and speech dominant. The system is evaluated against two baseline mask estimation methods in a reverberant multisource environment speech recognition task. The proposed system is shown to provide a performance improvement in the speech recognition accuracy over the previous multifeature approaches. Sami Keronen, Kyunghyun Cho, Tapani Raiko, Alexander Ilin, Kalle J. Palomäki |
ICASSP | 2 |
| 2013 | Simple Sparsification Improves Sparse Denoising Autoencoders in Denoising Highly Corrupted ImagesabstractRecently Burger et al. (2012) and Xie et al. (2012) proposed to use a denoising autoencoder (DAE) for denoising noisy images. They showed that a plain, deep DAE can denoise noisy images as well as the conventional methods such as BM3D and KSVD. Both of them approached image denoising by denoising small, image patches of a larger image and combining them to form a clean image. In this setting, it is usual to use the encoder of the DAE to obtain the latent representation and subsequently apply the decoder to get the clean patch. We propose that a simple sparsification of the latent representation found by the encoder improves denoising performance, when the DAE was trained with sparsity regularization. The experiments confirm that the proposed sparsification indeed helps both denoising a small image patch and denoising a larger image consisting of those patches. Furthermore, it is found out that the proposed method improves even classification performance when test samples are corrupted with noise. Kyunghyun Cho |
ICML (3) | 1 |
| 2013 | Measuring the Usefulness of Hidden Units in Boltzmann Machines with Mutual Information
Mathias Berglund, Tapani Raiko, Kyunghyun Cho |
ICONIP (1) | 3 |
| 2013 | Understanding Dropout: Training Multi-Layer Perceptrons with Auxiliary Independent Stochastic Neurons
Kyunghyun Cho |
ICONIP (1) | 1 |
| 2013 | Two-Layer Contractive Encodings with Shortcuts for Semi-supervised Learning
Hannes Schulz, Kyunghyun Cho, Tapani Raiko, Sven Behnke |
ICONIP (1) | 2 |
| 2013 | Gaussian-Bernoulli deep Boltzmann machineabstractIn this paper, we study a model that we call Gaussian-Bernoulli deep Boltzmann machine (GDBM) and discuss potential improvements in training the model. GDBM is designed to be applicable to continuous data and it is constructed from Gaussian-Bernoulli restricted Boltzmann machine (GRBM) by adding multiple layers of binary hidden neurons. The studied improvements of the learning algorithm for GDBM include parallel tempering, enhanced gradient, adaptive learning rate and layer-wise pretraining. We empirically show that they help avoid some of the common difficulties found in training deep Boltzmann machines such as divergence of learning, the difficulty in choosing right learning rate scheduling, and the existence of meaningless higher layers. Kyunghyun Cho, Tapani Raiko, Alexander Ilin |
IJCNN | 1 |
| 2013 | Enhanced Gradient for Training Restricted Boltzmann MachinesabstractRestricted Boltzmann machines (RBMs) are often used as building blocks in greedy learning of deep networks. However, training this simple model can be laborious. Traditional learning algorithms often converge only with the right choice of metaparameters that specify, for example, learning rate scheduling and the scale of the initial weights. They are also sensitive to specific data representation. An equivalent RBM can be obtained by flipping some bits and changing the weights and biases accordingly, but traditional learning rules are not invariant to such transformations. Without careful tuning of these training settings, traditional algorithms can easily get stuck or even diverge. In this letter, we present an enhanced gradient that is derived to be invariant to bit-flipping transformations. We experimentally show that the enhanced gradient yields more stable training of RBMs both when used with a fixed learning rate and an adaptive one. Kyunghyun Cho, Tapani Raiko, Alexander Ilin |
Neural Comput. | 1 |
| 2012 | Tikhonov-Type Regularization for Restricted Boltzmann Machines
Kyunghyun Cho, Alexander Ilin, Tapani Raiko |
ICANN (1) | 1 |
| 2012 | An iterative algorithm for singular value decomposition on noisy incomplete matricesabstractIn this paper, we propose a simple iterative algorithm, called iSVD, for estimating the singular value decomposition (SVD) of a noisy incomplete given matrix. The iSVD relies on first order optimization over orthogonal manifolds and automatically estimates the rank of the SVD. The main goal here is to estimate the singular vectors through optimization in the right space, which is the space of the orthogonal matrix manifolds. The rank estimation is based on the ratio between estimated large singular values and the sum of all singular values. We empirically evaluate the iSVD on synthetic matrices and image reconstruction tasks. The evaluation shows that the iSVD is comparable to the recently introduced methods for matrix completion such as singular value thresholding (SVT) and fixed-point iteration with approximate SVD (FPCA). Kyunghyun Cho, Nima Reyhani |
IJCNN | 1 |
| 2011 | Improved Learning of Gaussian-Bernoulli Restricted Boltzmann Machines
Kyunghyun Cho, Alexander Ilin, Tapani Raiko |
ICANN (1) | 1 |
| 2011 | Enhanced Gradient and Adaptive Learning Rate for Training Restricted Boltzmann Machines
Kyunghyun Cho, Tapani Raiko, Alexander Ilin |
ICML | 1 |
| 2010 | Parallel tempering is efficient for learning restricted Boltzmann machinesabstractA new interest towards restricted Boltzmann machines (RBMs) has risen due to their usefulness in greedy learning of deep neural networks. While contrastive divergence learning has been considered an efficient way to learn an RBM, it has a drawback due to a biased approximation in the learning gradient. We propose to use an advanced Monte Carlo method called parallel tempering instead, and show experimentally that it works efficiently. Kyunghyun Cho, Tapani Raiko, Alexander Ilin |
IJCNN | 1 |