Alexander H. Miller

dblp:190/7117 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-4139-4208ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 54% Question answering and dialogue systems · 14% Planning, search and constraint satisfaction · 12%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing
0.712023
Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning · ICLR 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning · ICLR 2023
Machine learning › Reinforcement learning
exploration
0.412020
The NetHack Learning Environment · NeurIPS 2020
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
procedural environment generation
0.412020
The NetHack Learning Environment · NeurIPS 2020
Machine learning › Reinforcement learning › exploration › novelty-based exploration
random network distillation
0.412020
The NetHack Learning Environment · NeurIPS 2020
Machine learning › Reinforcement learning
reinforcement learning environment
0.412020
The NetHack Learning Environment · NeurIPS 2020
Natural language and speech › Language models and text generation › large language model › knowledge in language models
language models as knowledge bases
0.412019
Language Models as Knowledge Bases? · EMNLP/IJCNLP (1) 2019
Computer vision › Vision and language
grounded language learning
0.312018
Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent · ICLR (Poster) 2018
Machine learning › Reinforcement learning › reinforcement learning from human feedback
learning from human feedback
0.312018
Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent · ICLR (Poster) 2018
Natural language and speech › Question answering and dialogue systems
question asking
0.312017
Learning through Dialogue Interactions by Asking Questions · ICLR (Poster) 2017
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
document question answering
0.212016
Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016
Machine learning › Deep learning architectures and training › memory-augmented neural networks
memory network
0.212016
Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.212016
Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning
0.112020
The NetHack Learning Environment · NeurIPS 2020
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning
0.112020
The NetHack Learning Environment · NeurIPS 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction
0.112019
Language Models as Knowledge Bases? · EMNLP/IJCNLP (1) 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base
0.112016
Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.7planning · 0.7random network distillation · 0.4distributed deep reinforcement learning · 0.4knowledge probing · 0.4cloze-style querying · 0.4mechanical turker descent · 0.3human-in-the-loop · 0.3dialogue learning · 0.3dialogue interaction · 0.3
YearPublicationVenuePosition
2025 AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
abstract
AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a challenging benchmark where agents compete in Kaggle competitions to solve real-world machine learning problems. We formalize AI research agents as search policies that navigate a space of candidate solutions, iteratively modifying them using operators. By designing and systematically varying different operator sets and search policies (Greedy, MCTS, Evolutionary), we show that their interplay is critical for achieving high performance. Our best pairing of search strategy and operator set achieves a state-of-the-art result on MLE-bench lite, increasing the success rate of achieving a Kaggle medal from 39.6% to 47.7%. Our investigation underscores the importance of jointly considering the search strategy, operator design, and evaluation methodology in advancing automated machine learning.
Edan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra, Nicolas Mario Baldwin, Alexis Audran-Reiss, Michael Kuchnik, Despoina Magka, Minqi Jiang, Alisia Maria Lupidi, Andrei Lupu, Roberta Raileanu, Tatiana Shavrina, Kelvin Niu, Jean-Christophe Gagnon-Audet, Michael Shvartsman, Shagun Sodhani, Alexander H. Miller, Abhishek Charnalia, Derek Dunfield, Carole-Jean Wu, Pontus Stenetorp, Nicola Cancedda, Jakob N. Foerster, Yoram Bachrach
NeurIPS18
2025 The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
abstract
Rapidly improving large language models (LLMs) have the potential to assist in scientific progress. One critical skill in this endeavor is the ability to faithfully reproduce existing work. To evaluate the capability of AI agents to reproduce complex code in an active research area, we introduce the Automated LLM Speedrunning Benchmark, leveraging the research community's contributions to the $\textit{NanoGPT speedrun}$, a competition to train a GPT-2 model in the shortest time. Each of the 19 speedrun tasks provides the agent with the previous record's training script, optionally paired with one of three hint formats, ranging from pseudocode to paper-like descriptions of the new record's improvements. Records execute quickly by design and speedrun improvements encompass diverse code-level changes, ranging from high-level algorithmic advancements to hardware-aware optimizations. These features make the benchmark both accessible and realistic for the frontier problem of improving LLM training. We find that recent frontier reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints. Our benchmark thus provides a simple, non-saturated measure of an LLM's ability to automate scientific reproduction, a necessary (but not sufficient) skill for an autonomous research agent.
Bingchen Zhao, Despoina Magka, Minqi Jiang, Xian Li 0003, Roberta Raileanu, Tatiana Shavrina, Jean-Christophe Gagnon-Audet, Kelvin Niu, Shagun Sodhani, Michael Shvartsman, Andrei Lupu, Alisia Maria Lupidi, Karen Hambardzumyan, Martin Josifoski, Edan Toledo, Thomas Foster, Lucia Cipolina-Kun, Derek Dunfield, Abhishek Charnalia, Alexander H. Miller, Oisin Mac Aodha, Jakob Foerster, Yoram Bachrach
NeurIPS20
2023 Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
Anton Bakhtin, David J. Wu 0002, Adam Lerer, Jonathan Gray, Athul Paul Jacob, Gabriele Farina, Alexander H. Miller, Noam Brown
ICLR7
2020 The NetHack Learning Environment
abstract
Progress in Reinforcement Learning (RL) algorithms goes hand-in-hand with the development of challenging environments that test the limits of current methods. While existing RL environments are either sufficiently complex or based on fast simulation, they are rarely both. Here, we present the NetHack Learning Environment (NLE), a scalable, procedurally generated, stochastic, rich, and challenging environment for RL research based on the popular single-player terminal-based roguelike game, NetHack. We argue that NetHack is sufficiently complex to drive long-term research on problems such as exploration, planning, skill acquisition, and language-conditioned RL, while dramatically reducing the computational resources required to gather a large amount of experience. We compare NLE and its task suite to existing alternatives, and discuss why it is an ideal medium for testing the robustness and systematic generalization of RL agents. We demonstrate empirical success for early stages of the game using a distributed Deep RL baseline and Random Network Distillation exploration, alongside qualitative analysis of various agents trained in the environment. NLE is open source and available at https://github.com/facebookresearch/nle.
Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, Tim Rocktäschel
NeurIPS3
2019 Language Models as Knowledge Bases?
abstract
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander Miller. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel 0001, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller
EMNLP/IJCNLP (1)7
2019 Importance of Search and Evaluation Strategies in Neural Dialogue Modeling
abstract
We investigate the impact of search strategies in neural dialogue modeling.We first compare two standard search algorithms, greedy and beam search, as well as our newly proposed iterative beam search which produces a more diverse set of candidate responses.We evaluate these strategies in realistic full conversations with humans and propose a modelbased Bayesian calibration to address annotator bias.These conversations are analyzed using two automatic metrics: log-probabilities assigned by the model and utterance diversity.Our experiments reveal that better search algorithms lead to higher rated conversations.However, finding the optimal selection mechanism to choose from a more diverse set of candidates is still an open question.
Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, Jason Weston
INLG2
2018 Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent
Zhilin Yang 0001, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller, Arthur Szlam, Douwe Kiela, Jason Weston
ICLR (Poster)5
2017 Dialogue Learning With Human-in-the-Loop
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston
ICLR (Poster)2
2017 Learning through Dialogue Interactions by Asking Questions
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston
ICLR (Poster)2
2016 Key-Value Memory Networks for Directly Reading Documents
abstract
Directly reading documents and being able to answer questions from them is an unsolved challenge.To avoid its inherent difficulty, question answering (QA) has been directed towards using Knowledge Bases (KBs) instead, which has proven effective.Unfortunately KBs often suffer from being too restrictive, as the schema cannot support certain types of answers, and too sparse, e.g.Wikipedia contains much more information than Freebase.In this work we introduce a new method, Key-Value Memory Networks, that makes reading documents more viable by utilizing different encodings in the addressing and output stages of the memory read operation.To compare using KBs, information extraction or Wikipedia documents directly in a single framework we construct an analysis tool, WIKIMOVIES, a QA dataset that contains raw text alongside a preprocessed KB, in the domain of movies.Our method reduces the gap between all three settings.It also achieves state-of-the-art results on the existing WIKIQA benchmark.
Alexander H. Miller, Adam Fisch, Jesse Dodge, Antoine Bordes, Jason Weston
EMNLP1