Nicholas Corrado

dblp:340/2322 · also Nicholas E. Corrado · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 34% Language models and text generation · 34% Efficient and distributed learning · 17%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model training
data mixing
0.912025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025
Machine learning › Efficient and distributed learning › data curation
training data curation
0.912025
AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs · ACL (1) 2025
Machine learning › Deep learning architectures and training
data augmentation
0.812024
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024
Machine learning › Reinforcement learning › off-policy reinforcement learning › experience replay
replay ratio
0.812024
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024
Machine learning › Reinforcement learning
sample efficiency
0.812024
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024
Machine learning › Reinforcement learning › sparse reward reinforcement learning
sparse reward tasks
0.212024
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates · ICLR 2024

Methods — techniques the papers use, named apart from their topics

adaptive data mixing · 0.9model-free reinforcement learning · 0.8
YearPublicationVenuePosition
2025 AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
abstract
Nicholas E. Corrado, Julian Katz-Samuels, Adithya M Devraj, Hyokun Yun, Chao Zhang, Yi Xu, Yi Pan, Bing Yin, Trishul Chilimbi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nicholas Corrado, Julian Katz-Samuels, Adithya M. Devraj, Hyokun Yun, Yi Xu 0011, Trishul Chilimbi
ACL (1)1
2025 When Can Model-Free Reinforcement Learning be Enough for Thinking?
abstract
Recent work on large language models has demonstrated the use of model-free reinforcement learning (RL) to train reasoning-like capabilities. The emergence of "thinking" through model-free RL is interesting as thinking actions neither produce reward nor change the external world state to one where the agent is more likely to get reward. This paper seeks to build a domain-independent understanding of when model-free RL will lead to such "thinking" as a strategy for reward maximization. To build this understanding, we first introduce a theoretical model which we call a thought Markov decision process (MDP). Thought MDPs minimally extend the classical MDP model to include an abstract notion of thought state and thought action. Using the thought MDP model, we prove the importance of policy initialization in determining whether or not thinking emerges and show formally that thought actions are equivalent to the agent choosing to perform a step of policy improvement before continuing to act. We then show that open-source LLMs satisfy the conditions that our theory predicts are necessary for model-free RL to produce thinking-like behavior. Finally, we hypothesize sufficient conditions that would enable thinking to be learned outside of language generation and introduce a toy domain where a combination of multi-task pre-training and designated thought actions enable more data-efficient RL compared to non-thinking agents.
Josiah Hanna, Nicholas Corrado
NeurIPS2
2024 Understanding when Dynamics-Invariant Data Augmentations Benefit Model-free Reinforcement Learning Updates
abstract
Recently, data augmentation (DA) has emerged as a method for leveraging domain knowledge to inexpensively generate additional data in reinforcement learning (RL) tasks, often yielding substantial improvements in data efficiency. While prior work has demonstrated the utility of incorporating augmented data directly into model-free RL updates, it is not well-understood when a particular DA strategy will improve data efficiency. In this paper, we seek to identify general aspects of DA responsible for observed learning improvements. Our study focuses on sparse-reward tasks with dynamics-invariant data augmentation functions, serving as an initial step towards a more general understanding of DA and its integration into RL training. Experimentally, we isolate three relevant aspects of DA: state-action coverage, reward density, and the number of augmented transitions generated per update (the augmented replay ratio). From our experiments, we draw two conclusions: (1) increasing state-action coverage often has a much greater impact on data efficiency than increasing reward density, and (2) decreasing the augmented replay ratio substantially improves data efficiency. In fact, certain tasks in our empirical study are solvable only when the replay ratio is sufficiently low.
Nicholas Corrado, Josiah Hanna
ICLR1