Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dami Choi

dblp:209/9687 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 24% Optimization for machine learning · 22% Probabilistic and Bayesian machine learning · 16%
Network and information security
1 paper
Security and privacy of machine learning · 77% Cryptographic primitives and cryptanalysis · 23%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
gradient estimation
0.822020
Gradient Estimation with Stochastic Softmax Tricks · NeurIPS 2020
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation · ICLR (Poster) 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
probabilistic regression
0.812024
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language · NeurIPS 2024
Natural language and speech › Language models and text generation
prompting
0.812024
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language · NeurIPS 2024
Machine learning › Learning paradigms
imbalanced learning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Trustworthy machine learning
model auditing
0.712023
Tools for Verifying Neural Models' Training Data · NeurIPS 2023
Natural language and speech › Language models and text generation
multilingual language models
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Natural language and speech › Language models and text generation
multilingual learning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
multilingual transfer
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Learning paradigms
multi-task learning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pre-training and fine-tuning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Security and privacy of machine learning › model intellectual property protection
model provenance
0.712023
Tools for Verifying Neural Models' Training Data · NeurIPS 2023
Machine learning › Optimization for machine learning
black-box optimization
0.622024
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation · ICLR (Poster) 2018
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning
gumbel-softmax
0.412020
Gradient Estimation with Stochastic Softmax Tricks · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412020
Gradient Estimation with Stochastic Softmax Tricks · NeurIPS 2020
Machine learning › Optimization for machine learning › evolutionary computation
evolution strategies
0.412019
Guided evolutionary strategies: augmenting random search with surrogate gradients · ICML 2019
Machine learning › Deep learning architectures and training › spiking neural network
surrogate gradient
0.412019
Guided evolutionary strategies: augmenting random search with surrogate gradients · ICML 2019
Machine learning › Optimization for machine learning › black-box optimization
zeroth-order optimization
0.412019
Guided evolutionary strategies: augmenting random search with surrogate gradients · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
control variates
0.312018
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation · ICLR (Poster) 2018
Machine learning › Optimization for machine learning
variance reduction
0.312018
Backpropagation through the Void: Optimizing control variates for black-box gradient estimation · ICLR (Poster) 2018
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.212024
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language · NeurIPS 2024
Natural language and speech › Machine translation
neural machine translation
0.212023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Cryptographic primitives and cryptanalysis › cryptographic foundations
cryptographic commitments
0.212023
Tools for Verifying Neural Models' Training Data · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

fine-tuning · 1.4overfitting detection · 1.3membership inference · 1.3cryptographic commitment · 1.3prompting · 0.8large language model · 0.8chain-of-thought · 0.8task weighting · 0.7pre-training · 0.7gumbel-max trick · 0.4
YearPublicationVenuePosition
2024 LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language
abstract
Machine learning practitioners often face significant challenges in formally integrating their prior knowledge and beliefs into predictive models, limiting the potential for nuanced and context-aware analyses. Moreover, the expertise needed to integrate this prior knowledge into probabilistic modeling typically limits the application of these models to specialists. Our goal is to build a regression model that can process numerical data and make probabilistic predictions at arbitrary locations, guided by natural language text which describes a user's prior knowledge. Large Language Models (LLMs) provide a useful starting point for designing such a tool since they 1) provide an interface where users can incorporate expert insights in natural language and 2) provide an opportunity for leveraging latent problem-relevant knowledge encoded in LLMs that users may not have themselves. We start by exploring strategies for eliciting explicit, coherent numerical predictive distributions from LLMs. We examine these joint predictive distributions, which we call LLM Processes, over arbitrarily-many quantities in settings such as forecasting, multi-dimensional regression, black-box optimization, and image modeling. We investigate the practical details of prompting to elicit coherent predictive distributions, and demonstrate their effectiveness at regression. Finally, we demonstrate the ability to usefully incorporate text into numerical predictions, improving predictive performance and giving quantitative structure that reflects qualitative descriptions. This lets us begin to explore the rich, grounded hypothesis space that LLMs implicitly encode.
James Requeima, John Bronskill, Dami Choi, Richard E. Turner, David Duvenaud
NeurIPS3
2024 Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
abstract
One way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data. While this removes the explicit information, implicit information can remain scattered across various training documents. Could an LLM infer the censored knowledge by piecing together these implicit hints? As a step towards answering this question, we study inductive out-of-context reasoning (OOCR), a type of generalization in which LLMs infer latent information from evidence distributed across training documents and apply it to downstream tasks without in-context learning. Using a suite of five tasks, we demonstrate that frontier LLMs can perform inductive OOCR. In one experiment we finetune an LLM on a corpus consisting only of distances between an unknown city and other known cities. Remarkably, without in-context examples or Chain of Thought, the LLM can verbalize that the unknown city is Paris and use this fact to answer downstream questions. Further experiments show that LLMs trained only on individual coin flip outcomes can verbalize whether the coin is biased, and those trained only on pairs $(x,f(x))$ can articulate a definition of $f$ and compute inverses. While OOCR succeeds in a range of cases, we also show that it is unreliable, particularly for smaller LLMs learning complex structures. Overall, the ability of LLMs to "connect the dots" without explicit in-context learning poses a potential obstacle to monitoring and controlling the knowledge acquired by LLMs.
Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger B. Grosse, Owain Evans
NeurIPS2
2023 Tools for Verifying Neural Models' Training Data
abstract
It is important that consumers and regulators can verify the provenance of large neural models to evaluate their capabilities and risks. We introduce the concept of a "Proof-of-Training-Data": any protocol that allows a model trainer to convince a Verifier of the training data that produced a set of model weights. Such protocols could verify the amount and kind of data and compute used to train the model, including whether it was trained on specific harmful or beneficial data sources. We explore efficient verification strategies for Proof-of-Training-Data that are compatible with most current large-model training procedures. These include a method for the model-trainer to verifiably pre-commit to a random seed used in training, and a method that exploits models' tendency to temporarily overfit to training data in order to detect whether a given data-point was included in training. We show experimentally that our verification procedures can catch a wide variety of attacks, including all known attacks from the Proof-of-Learning literature.
Dami Choi, Yonadav Shavit, David Duvenaud
NeurIPS1
2023 Order Matters in the Presence of Dataset Imbalance for Multilingual Learning
abstract
In this paper, we empirically study the optimization dynamics of multi-task learning, particularly focusing on those that govern a collection of tasks with significant data imbalance. We present a simple yet effective method of pre-training on high-resource tasks, followed by fine-tuning on a mixture of high/low-resource tasks. We provide a thorough empirical study and analysis of this method's benefits showing that it achieves consistent improvements relative to the performance trade-off profile of standard static weighting. We analyze under what data regimes this method is applicable and show its improvements empirically in neural machine translation (NMT) and multi-lingual language modeling.
Dami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer, Ankush Garg, Orhan Firat, Chih-Kuan Yeh, Andrew M. Dai, Behrooz Ghorbani
NeurIPS1
2020 Gradient Estimation with Stochastic Softmax Tricks
abstract
The Gumbel-Max trick is the basis of many relaxed gradient estimators. These estimators are easy to implement and low variance, but the goal of scaling them comprehensively to large combinatorial distributions is still outstanding. Working within the perturbation model framework, we introduce stochastic softmax tricks, which generalize the Gumbel-Softmax trick to combinatorial spaces. Our framework is a unified perspective on existing relaxed estimators for perturbation models, and it contains many novel relaxations. We design structured relaxations for subset selection, spanning trees, arborescences, and others. When compared to less structured baselines, we find that stochastic softmax tricks can be used to train latent variable models that perform better and discover more latent structure.
Max B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 0001, Chris J. Maddison
NeurIPS2
2019 Guided evolutionary strategies: augmenting random search with surrogate gradients
abstract
Many applications in machine learning require optimizing a function whose true gradient is unknown or computationally expensive, but where surrogate gradient information, directions that may be correlated with the true gradient, is cheaply available. For example, this occurs when an approximate gradient is easier to compute than the full gradient (e.g. in meta-learning or unrolled optimization), or when a true gradient is intractable and is replaced with a surrogate (e.g. in reinforcement learning or training networks with discrete variables). We propose Guided Evolutionary Strategies (GES), a method for optimally using surrogate gradient directions to accelerate random search. GES defines a search distribution for evolutionary strategies that is elongated along a subspace spanned by the surrogate gradients and estimates a descent direction which can then be passed to a first-order optimizer. We analytically and numerically characterize the tradeoffs that result from tuning how strongly the search distribution is stretched along the guiding subspace and use this to derive a setting of the hyperparameters that works well across problems. We evaluate GES on several example problems, demonstrating an improvement over both standard evolutionary strategies and first-order methods that directly follow the surrogate gradient.
Niru Maheswaranathan, Luke Metz, George Tucker, Dami Choi, Jascha Sohl-Dickstein
ICML4
2018 Backpropagation through the Void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoffrey Roeder, David Duvenaud
ICLR (Poster)2