VLDB 2026 Research / reviewers in the wild / expert
Nicholas Corrado
dblp:340/2322 · also Nicholas E. Corrado
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 34% Language models and text generation · 34% Efficient and distributed learning · 17% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model training
data mixing |
0.9 | 1 | 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › data curation
training data curation |
0.9 | 1 | 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.8 | 1 | 2024 | Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024 |
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
replay ratio |
0.8 | 1 | 2024 | Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024 |
Machine learning › Reinforcement learning
sample efficiency |
0.8 | 1 | 2024 | Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024 |
Machine learning › Reinforcement learning › sparse reward reinforcement learning
sparse reward tasks |
0.2 | 1 | 2024 | Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
adaptive data mixing · 0.9model-free reinforcement learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMsabstractNicholas E. Corrado, Julian Katz-Samuels, Adithya M Devraj, Hyokun Yun, Chao Zhang, Yi Xu, Yi Pan, Bing Yin, Trishul Chilimbi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nicholas Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun, Yi Xu 0011, Trishul Chilimbi |
ACL (1) | 1 |
| 2025 | When Can Model-Free Reinforcement Learning be Enough for Thinking?abstractRecent work on large language models has demonstrated the use of model-free reinforcement learning (RL) to train reasoning-like capabilities. The emergence of "thinking" through model-free RL is interesting as thinking actions neither produce reward nor change the external world state to one where the agent is more likely to get reward. This paper seeks to build a domain-independent understanding of when model-free RL will lead to such "thinking" as a strategy for reward maximization. To build this understanding, we first introduce a theoretical model which we call a thought Markov decision process (MDP). Thought MDPs minimally extend the classical MDP model to include an abstract notion of thought state and thought action. Using the thought MDP model, we prove the importance of policy initialization in determining whether or not thinking emerges and show formally that thought actions are equivalent to the agent choosing to perform a step of policy improvement before continuing to act. We then show that open-source LLMs satisfy the conditions that our theory predicts are necessary for model-free RL to produce thinking-like behavior. Finally, we hypothesize sufficient conditions that would enable thinking to be learned outside of language generation and introduce a toy domain where a combination of multi-task pre-training and designated thought actions enable more data-efficient RL compared to non-thinking agents. Josiah Hanna, Nicholas Corrado |
NeurIPS | 2 |
| 2024 | Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning UpdatesabstractRecently, data augmentation (DA) has emerged as a method for leveraging domain knowledge to inexpensively generate additional data in reinforcement learning (RL) tasks, often yielding substantial improvements in data efficiency.
While prior work has demonstrated the utility of incorporating augmented data directly into model-free RL updates,
it is not well-understood when a particular DA strategy will improve data efficiency.
In this paper, we seek to identify general aspects of DA responsible for observed learning improvements.
Our study focuses on sparse-reward tasks with dynamics-invariant data augmentation functions, serving as an initial step towards a more general understanding of DA and its integration into RL training.
Experimentally, we isolate three relevant aspects of DA: state-action coverage, reward density, and the number of augmented transitions generated per update (the augmented replay ratio).
From our experiments, we draw two conclusions: (1) increasing state-action coverage often has a much greater impact on data efficiency than increasing reward density, and (2) decreasing the augmented replay ratio substantially improves data efficiency.
In fact, certain tasks in our empirical study are solvable only when the replay ratio is sufficiently low. Nicholas Corrado, Josiah Hanna |
ICLR | 1 |