Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jake Grigsby

dblp:276/6109 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 71% Deep learning architectures and training · 20% Transfer learning and domain adaptation · 9%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning
1.522024
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024
Machine learning › Reinforcement learning
meta-reinforcement learning
1.522024
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024
Machine learning › Reinforcement learning
multi-task reinforcement learning
1.422024
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024
Cross-Episodic Curriculum for Transformer Agents · NeurIPS 2023
Machine learning › Deep learning architectures and training
transformer
1.022024
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.712023
PGrad: Learning Principal Gradients For Domain Generalization · ICLR 2023
Machine learning › Reinforcement learning
imitation learning
0.712023
Cross-Episodic Curriculum for Transformer Agents · NeurIPS 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.212024
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024

Methods — techniques the papers use, named apart from their topics

transformer · 1.4off-policy reinforcement learning · 0.8hindsight relabeling · 0.8classification-based actor-critic · 0.8principal gradient · 0.7meta-learning · 0.7cross-episodic curriculum · 0.7
YearPublicationVenuePosition
2024 AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
abstract
We introduce AMAGO, an in-context Reinforcement Learning (RL) agent that uses sequence models to tackle the challenges of generalization, long-term memory, and meta-learning. Recent works have shown that off-policy learning can make in-context RL with recurrent policies viable. Nonetheless, these approaches require extensive tuning and limit scalability by creating key bottlenecks in agents' memory capacity, planning horizon, and model size. AMAGO revisits and redesigns the off-policy in-context approach to successfully train long-sequence Transformers over entire rollouts in parallel with end-to-end RL. Our agent is scalable and applicable to a wide range of problems, and we demonstrate its strong performance empirically in meta-RL and long-term memory domains. AMAGO's focus on sparse rewards and off-policy data also allows in-context learning to extend to goal-conditioned problems with challenging exploration. When combined with a multi-goal hindsight relabeling scheme, AMAGO can solve a previously difficult category of open-world domains, where agents complete many possible instructions in procedurally generated environments.
Jake Grigsby, Linxi Fan, Yuke Zhu
ICLR1
2024 Launchpad: Learning to Schedule Using Offline and Online RL Methods
Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi
JSSPP2
2024 AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers
abstract
Language models trained on diverse datasets unlock generalization by in-context learning. Reinforcement Learning (RL) policies can achieve a similar effect by meta-learning within the memory of a sequence model. However, meta-RL research primarily focuses on adapting to minor variations of a single task. It is difficult to scale towards more general behavior without confronting challenges in multi-task optimization, and few solutions are compatible with meta-RL's goal of learning from large training sets of unlabeled tasks. To address this challenge, we revisit the idea that multi-task RL is bottlenecked by imbalanced training losses created by uneven return scales across different tasks. We build upon recent advancements in Transformer-based (in-context) meta-RL and evaluate a simple yet scalable solution where both an agent's actor and critic objectives are converted to classification terms that decouple optimization from the current scale of returns. Large-scale comparisons in Meta-World ML45, Multi-Game Procgen, Multi-Task POPGym, Multi-Game Atari, and BabyAI find that this design unlocks significant progress in online multi-task adaptation and memory problems without explicit task labels.
Jake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi, Amy Zhang 0001, Yuke Zhu
NeurIPS1
2023 PGrad: Learning Principal Gradients For Domain Generalization
Zhe Wang 0025, Jake Grigsby, Yanjun Qi
ICLR2
2023 Cross-Episodic Curriculum for Transformer Agents
abstract
We present a new algorithm, Cross-Episodic Curriculum (CEC), to boost the learning efficiency and generalization of Transformer agents. Central to CEC is the placement of cross-episodic experiences into a Transformer’s context, which forms the basis of a curriculum. By sequentially structuring online learning trials and mixed-quality demonstrations, CEC constructs curricula that encapsulate learning progression and proficiency increase across episodes. Such synergy combined with the potent pattern recognition capabilities of Transformer models delivers a powerful cross-episodic attention mechanism. The effectiveness of CEC is demonstrated under two representative scenarios: one involving multi-task reinforcement learning with discrete control, such as in DeepMind Lab, where the curriculum captures the learning progression in both individual and progressively complex settings; and the other involving imitation learning with mixed-quality data for continuous control, as seen in RoboMimic, where the curriculum captures the improvement in demonstrators' expertise. In all instances, policies resulting from CEC exhibit superior performance and strong generalization. Code is open-sourced on the project website https://cec-agent.github.io/ to facilitate research on Transformer agent learning.
Lucy Xiaoyang Shi, Yunfan Jiang 0001, Jake Grigsby, Linxi Fan, Yuke Zhu
NeurIPS3
2022 RARE: Renewable Energy Aware Resource Management in Datacenters
Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi
JSSPP2
2022 ST-MAML : A stochastic-task based method for task-heterogeneous meta-learning
abstract
Optimization-based meta-learning typically assumes tasks are sampled from a single distribution - an assumption that oversimplifies and limits the diversity of tasks that meta-learning can model. Handling tasks from multiple distributions is challenging for meta-learning because it adds ambiguity to task identities. This paper proposes a novel method, ST-MAML, that empowers model-agnostic meta-learning (MAML) to learn from multiple task distributions. ST-MAML encodes tasks using a stochastic neural network module, that summarizes every task with a stochastic representation. The proposed Stochastic Task (ST) strategy learns a distribution of solutions for an ambiguous task and allows a meta-model to self-adapt to the current task. ST-MAML also propagates the task representation to enhance input variable encodings. Empirically, we demonstrate that ST-MAML outperforms the state-of-the-art on two few-shot image classification tasks, one curve regression benchmark, one image completion problem, and a real-world temperature prediction application.
Zhe Wang 0025, Jake Grigsby, Arshdeep Sekhon, Yanjun Qi
UAI2