EDBT 2026 Demo / reviewers in the wild / expert
Jake Grigsby
dblp:276/6109
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 71% Deep learning architectures and training · 20% Transfer learning and domain adaptation · 9% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning |
1.5 | 2 | 2024 | AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024 AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024 |
Machine learning › Reinforcement learning
meta-reinforcement learning |
1.5 | 2 | 2024 | AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024 AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024 |
Machine learning › Reinforcement learning
multi-task reinforcement learning |
1.4 | 2 | 2024 | AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024 Cross-Episodic Curriculum for Transformer Agents · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
transformer |
1.0 | 2 | 2024 | AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers · NeurIPS 2024 AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.7 | 1 | 2023 | PGrad: Learning Principal Gradients For Domain Generalization · ICLR 2023 |
Machine learning › Reinforcement learning
imitation learning |
0.7 | 1 | 2023 | Cross-Episodic Curriculum for Transformer Agents · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.2 | 1 | 2024 | AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.4off-policy reinforcement learning · 0.8hindsight relabeling · 0.8classification-based actor-critic · 0.8principal gradient · 0.7meta-learning · 0.7cross-episodic curriculum · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AMAGO: Scalable In-Context Reinforcement Learning for Adaptive AgentsabstractWe introduce AMAGO, an in-context Reinforcement Learning (RL) agent that uses sequence models to tackle the challenges of generalization, long-term memory, and meta-learning. Recent works have shown that off-policy learning can make in-context RL with recurrent policies viable. Nonetheless, these approaches require extensive tuning and limit scalability by creating key bottlenecks in agents' memory capacity, planning horizon, and model size. AMAGO revisits and redesigns the off-policy in-context approach to successfully train long-sequence Transformers over entire rollouts in parallel with end-to-end RL. Our agent is scalable and applicable to a wide range of problems, and we demonstrate its strong performance empirically in meta-RL and long-term memory domains. AMAGO's focus on sparse rewards and off-policy data also allows in-context learning to extend to goal-conditioned problems with challenging exploration. When combined with a multi-goal hindsight relabeling scheme, AMAGO can solve a previously difficult category of open-world domains, where agents complete many possible instructions in procedurally generated environments. Jake Grigsby, Linxi Fan, Yuke Zhu |
ICLR | 1 |
| 2024 | Launchpad: Learning to Schedule Using Offline and Online RL Methods
Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi |
JSSPP | 2 |
| 2024 | AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with TransformersabstractLanguage models trained on diverse datasets unlock generalization by in-context learning. Reinforcement Learning (RL) policies can achieve a similar effect by meta-learning within the memory of a sequence model. However, meta-RL research primarily focuses on adapting to minor variations of a single task. It is difficult to scale towards more general behavior without confronting challenges in multi-task optimization, and few solutions are compatible with meta-RL's goal of learning from large training sets of unlabeled tasks. To address this challenge, we revisit the idea that multi-task RL is bottlenecked by imbalanced training losses created by uneven return scales across different tasks. We build upon recent advancements in Transformer-based (in-context) meta-RL and evaluate a simple yet scalable solution where both an agent's actor and critic objectives are converted to classification terms that decouple optimization from the current scale of returns. Large-scale comparisons in Meta-World ML45, Multi-Game Procgen, Multi-Task POPGym, Multi-Game Atari, and BabyAI find that this design unlocks significant progress in online multi-task adaptation and memory problems without explicit task labels. Jake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi, Amy Zhang 0001, Yuke Zhu |
NeurIPS | 1 |
| 2023 | PGrad: Learning Principal Gradients For Domain Generalization
Zhe Wang 0025, Jake Grigsby, Yanjun Qi |
ICLR | 2 |
| 2023 | Cross-Episodic Curriculum for Transformer AgentsabstractWe present a new algorithm, Cross-Episodic Curriculum (CEC), to boost the learning efficiency and generalization of Transformer agents. Central to CEC is the placement of cross-episodic experiences into a Transformer’s context, which forms the basis of a curriculum. By sequentially structuring online learning trials and mixed-quality demonstrations, CEC constructs curricula that encapsulate learning progression and proficiency increase across episodes. Such synergy combined with the potent pattern recognition capabilities of Transformer models delivers a powerful cross-episodic attention mechanism. The effectiveness of CEC is demonstrated under two representative scenarios: one involving multi-task reinforcement learning with discrete control, such as in DeepMind Lab, where the curriculum captures the learning progression in both individual and progressively complex settings; and the other involving imitation learning with mixed-quality data for continuous control, as seen in RoboMimic, where the curriculum captures the improvement in demonstrators' expertise. In all instances, policies resulting from CEC exhibit superior performance and strong generalization. Code is open-sourced on the project website https://cec-agent.github.io/ to facilitate research on Transformer agent learning. Lucy Xiaoyang Shi, Yunfan Jiang 0001, Jake Grigsby, Linxi Fan, Yuke Zhu |
NeurIPS | 3 |
| 2022 | RARE: Renewable Energy Aware Resource Management in Datacenters
Vanamala Venkataswamy, Jake Grigsby, Andrew Grimshaw, Yanjun Qi |
JSSPP | 2 |
| 2022 | ST-MAML : A stochastic-task based method for task-heterogeneous meta-learningabstractOptimization-based meta-learning typically assumes tasks are sampled from a single distribution - an assumption that oversimplifies and limits the diversity of tasks that meta-learning can model. Handling tasks from multiple distributions is challenging for meta-learning because it adds ambiguity to task identities. This paper proposes a novel method, ST-MAML, that empowers model-agnostic meta-learning (MAML) to learn from multiple task distributions. ST-MAML encodes tasks using a stochastic neural network module, that summarizes every task with a stochastic representation. The proposed Stochastic Task (ST) strategy learns a distribution of solutions for an ambiguous task and allows a meta-model to self-adapt to the current task. ST-MAML also propagates the task representation to enhance input variable encodings. Empirically, we demonstrate that ST-MAML outperforms the state-of-the-art on two few-shot image classification tasks, one curve regression benchmark, one image completion problem, and a real-world temperature prediction application. Zhe Wang 0025, Jake Grigsby, Arshdeep Sekhon, Yanjun Qi |
UAI | 2 |