VLDB 2026 Research / reviewers in the wild / expert
Alexander H. Miller
dblp:190/7117
· DBLP profile ↗
10ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-4139-4208ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 54% Question answering and dialogue systems · 14% Planning, search and constraint satisfaction · 12% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing |
0.7 | 1 | 2023 | Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning · ICLR 2023 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.7 | 1 | 2023 | Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning · ICLR 2023 |
Machine learning › Reinforcement learning
exploration |
0.4 | 1 | 2020 | The NetHack Learning Environment · NeurIPS 2020 |
Machine learning › Reinforcement learning › reinforcement learning environment › environment design
procedural environment generation |
0.4 | 1 | 2020 | The NetHack Learning Environment · NeurIPS 2020 |
Machine learning › Reinforcement learning › exploration › novelty-based exploration
random network distillation |
0.4 | 1 | 2020 | The NetHack Learning Environment · NeurIPS 2020 |
Machine learning › Reinforcement learning
reinforcement learning environment |
0.4 | 1 | 2020 | The NetHack Learning Environment · NeurIPS 2020 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
language models as knowledge bases |
0.4 | 1 | 2019 | Language Models as Knowledge Bases? · EMNLP/IJCNLP (1) 2019 |
Computer vision › Vision and language
grounded language learning |
0.3 | 1 | 2018 | Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent · ICLR (Poster) 2018 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
learning from human feedback |
0.3 | 1 | 2018 | Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent · ICLR (Poster) 2018 |
Natural language and speech › Question answering and dialogue systems
question asking |
0.3 | 1 | 2017 | Learning through Dialogue Interactions by Asking Questions · ICLR (Poster) 2017 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
document question answering |
0.2 | 1 | 2016 | Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016 |
Machine learning › Deep learning architectures and training › memory-augmented neural networks
memory network |
0.2 | 1 | 2016 | Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.2 | 1 | 2016 | Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
language-conditioned reinforcement learning |
0.1 | 1 | 2020 | The NetHack Learning Environment · NeurIPS 2020 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning |
0.1 | 1 | 2020 | The NetHack Learning Environment · NeurIPS 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base construction |
0.1 | 1 | 2019 | Language Models as Knowledge Bases? · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base |
0.1 | 1 | 2016 | Key-Value Memory Networks for Directly Reading Documents · EMNLP 2016 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.7planning · 0.7random network distillation · 0.4distributed deep reinforcement learning · 0.4knowledge probing · 0.4cloze-style querying · 0.4mechanical turker descent · 0.3human-in-the-loop · 0.3dialogue learning · 0.3dialogue interaction · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-benchabstractAI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a challenging benchmark where agents compete in Kaggle competitions to solve real-world machine learning problems. We formalize AI research agents as search policies that navigate a space of candidate solutions, iteratively modifying them using operators. By designing and systematically varying different operator sets and search policies (Greedy, MCTS, Evolutionary), we show that their interplay is critical for achieving high performance. Our best pairing of search strategy and operator set achieves a state-of-the-art result on MLE-bench lite, increasing the success rate of achieving a Kaggle medal from 39.6% to 47.7%. Our investigation underscores the importance of jointly considering the search strategy, operator design, and evaluation methodology in advancing automated machine learning. Edan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra, Nicolas Mario Baldwin, Alexis Audran-Reiss, Michael Kuchnik, Despoina Magka, Minqi Jiang, Alisia Maria Lupidi, Andrei Lupu, Roberta Raileanu, Tatiana Shavrina, Kelvin Niu, Jean-Christophe Gagnon-Audet, Michael Shvartsman, Shagun Sodhani, Alexander H. Miller, Abhishek Charnalia, Derek Dunfield, Carole-Jean Wu, Pontus Stenetorp, Nicola Cancedda, Jakob N. Foerster, Yoram Bachrach |
NeurIPS | 18 |
| 2025 | The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT ImprovementsabstractRapidly improving large language models (LLMs) have the potential to assist in scientific progress. One critical skill in this endeavor is the ability to faithfully reproduce existing work. To evaluate the capability of AI agents to reproduce complex code in an active research area, we introduce the Automated LLM Speedrunning Benchmark, leveraging the research community's contributions to the $\textit{NanoGPT speedrun}$, a competition to train a GPT-2 model in the shortest time. Each of the 19 speedrun tasks provides the agent with the previous record's training script, optionally paired with one of three hint formats, ranging from pseudocode to paper-like descriptions of the new record's improvements. Records execute quickly by design and speedrun improvements encompass diverse code-level changes, ranging from high-level algorithmic advancements to hardware-aware optimizations. These features make the benchmark both accessible and realistic for the frontier problem of improving LLM training. We find that recent frontier reasoning LLMs combined with SoTA scaffolds struggle to reimplement already-known innovations in our benchmark, even when given detailed hints. Our benchmark thus provides a simple, non-saturated measure of an LLM's ability to automate scientific reproduction, a necessary (but not sufficient) skill for an autonomous research agent. Bingchen Zhao, Despoina Magka, Minqi Jiang, Xian Li 0003, Roberta Raileanu, Tatiana Shavrina, Jean-Christophe Gagnon-Audet, Kelvin Niu, Shagun Sodhani, Michael Shvartsman, Andrei Lupu, Alisia Maria Lupidi, Karen Hambardzumyan, Martin Josifoski, Edan Toledo, Thomas Foster, Lucia Cipolina-Kun, Derek Dunfield, Abhishek Charnalia, Alexander H. Miller, Oisin Mac Aodha, Jakob Foerster, Yoram Bachrach |
NeurIPS | 20 |
| 2023 | Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
Anton Bakhtin, David J. Wu 0002, Adam Lerer, Jonathan Gray, Athul Paul Jacob, Gabriele Farina, Alexander H. Miller, Noam Brown |
ICLR | 7 |
| 2020 | The NetHack Learning EnvironmentabstractProgress in Reinforcement Learning (RL) algorithms goes hand-in-hand with the development of challenging environments that test the limits of current methods. While existing RL environments are either sufficiently complex or based on fast simulation, they are rarely both. Here, we present the NetHack Learning Environment (NLE), a scalable, procedurally generated, stochastic, rich, and challenging environment for RL research based on the popular single-player terminal-based roguelike game, NetHack. We argue that NetHack is sufficiently complex to drive long-term research on problems such as exploration, planning, skill acquisition, and language-conditioned RL, while dramatically reducing the computational resources required to gather a large amount of experience. We compare NLE and its task suite to existing alternatives, and discuss why it is an ideal medium for testing the robustness and systematic generalization of RL agents. We demonstrate empirical success for early stages of the game using a distributed Deep RL baseline and Random Network Distillation exploration, alongside qualitative analysis of various agents trained in the environment. NLE is open source and available at https://github.com/facebookresearch/nle. Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, Tim Rocktäschel |
NeurIPS | 3 |
| 2019 | Language Models as Knowledge Bases?abstractFabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander Miller. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Fabio Petroni, Tim Rocktäschel, Sebastian Riedel 0001, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Importance of Search and Evaluation Strategies in Neural Dialogue ModelingabstractWe investigate the impact of search strategies in neural dialogue modeling.We first compare two standard search algorithms, greedy and beam search, as well as our newly proposed iterative beam search which produces a more diverse set of candidate responses.We evaluate these strategies in realistic full conversations with humans and propose a modelbased Bayesian calibration to address annotator bias.These conversations are analyzed using two automatic metrics: log-probabilities assigned by the model and utterance diversity.Our experiments reveal that better search algorithms lead to higher rated conversations.However, finding the optimal selection mechanism to choose from a more diverse set of candidates is still an open question. Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, Jason Weston |
INLG | 2 |
| 2018 | Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent
Zhilin Yang 0001, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller, Arthur Szlam, Douwe Kiela, Jason Weston |
ICLR (Poster) | 5 |
| 2017 | Dialogue Learning With Human-in-the-Loop
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston |
ICLR (Poster) | 2 |
| 2017 | Learning through Dialogue Interactions by Asking Questions
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston |
ICLR (Poster) | 2 |
| 2016 | Key-Value Memory Networks for Directly Reading DocumentsabstractDirectly reading documents and being able to answer questions from them is an unsolved challenge.To avoid its inherent difficulty, question answering (QA) has been directed towards using Knowledge Bases (KBs) instead, which has proven effective.Unfortunately KBs often suffer from being too restrictive, as the schema cannot support certain types of answers, and too sparse, e.g.Wikipedia contains much more information than Freebase.In this work we introduce a new method, Key-Value Memory Networks, that makes reading documents more viable by utilizing different encodings in the addressing and output stages of the memory read operation.To compare using KBs, information extraction or Wikipedia documents directly in a single framework we construct an analysis tool, WIKIMOVIES, a QA dataset that contains raw text alongside a preprocessed KB, in the domain of movies.Our method reduces the gap between all three settings.It also achieves state-of-the-art results on the existing WIKIQA benchmark. Alexander H. Miller, Adam Fisch, Jesse Dodge, Antoine Bordes, Jason Weston |
EMNLP | 1 |