Esmaeil Seraj

dblp:169/3595 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0002-0147-1037ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Multi-agent systems · 61% Reinforcement learning · 32% Robot manipulation · 7%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
multi-robot coordination
2.032024
Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination · IEEE Trans. Robotics 2024
Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot Teams · NeurIPS 2023
A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot Teams · IEEE Trans. Robotics 2022
Knowledge, reasoning and agents › Multi-agent systems › multi-robot coordination
heterogeneous robot team coordination
1.322024
Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination · IEEE Trans. Robotics 2024
A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot Teams · IEEE Trans. Robotics 2022
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.132024
Embodied, Intelligent Communication for Multi-Agent Cooperation · AAAI 2023
Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination · IEEE Trans. Robotics 2024
A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot Teams · IEEE Trans. Robotics 2022
Knowledge, reasoning and agents › Multi-agent systems › emergent communication
learned communication
0.812024
Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination · IEEE Trans. Robotics 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication
0.812024
Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination · IEEE Trans. Robotics 2024
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
0.712023
Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot Teams · NeurIPS 2023
Machine learning › Reinforcement learning
imitation learning
0.712023
Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot Teams · NeurIPS 2023
Robotics › Robot manipulation
learning from demonstration
0.712023
Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot Teams · NeurIPS 2023
Knowledge, reasoning and agents › Multi-agent systems
multi-agent collaboration
0.712023
Embodied, Intelligent Communication for Multi-Agent Cooperation · AAAI 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
0.612022
Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming · ICLR 2022
Knowledge, reasoning and agents › Multi-agent systems
multi-robot systems
0.212023
Embodied, Intelligent Communication for Multi-Agent Cooperation · AAAI 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
decentralized partially observable markov decision process
0.212022
A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot Teams · IEEE Trans. Robotics 2022

Methods — techniques the papers use, named apart from their topics

heterogeneous graph attention network · 0.8binarized messaging · 0.8reverse model · 0.7planning · 0.7mutual information maximization · 0.7multi-agent reinforcement learning · 0.7model-based control · 0.7mutual information · 0.6iterated reasoning · 0.6hierarchical planning · 0.6
YearPublicationVenuePosition
2024 Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination
abstract
High-performing human–human teams learn intelligent and efficient communication and coordination strategies to maximize their joint utility. These teams implicitly understand the different roles of heterogeneous team members and adapt their communication protocols accordingly. Multiagent reinforcement learning (MARL) has attempted to develop computational methods for synthesizing such joint coordination–communication strategies, but emulating heterogeneous communication patterns across agents with different state, action, and observation spaces has remained a challenge. Without properly modeling agent heterogeneity, as in prior MARL work that leverages homogeneous graph networks, communication becomes less helpful and can even deteriorate the team's performance. In the past, we proposed heterogeneous policy networks (HetNet) to learn efficient and diverse communication models for coordinating cooperative heterogeneous teams. In this extended work, we extend HetNet to support scaling heterogeneous robot teams. Building on heterogeneous graph-attention networks, we show that HetNet not only facilitates learning heterogeneous collaborative policies, but also enables end-to-end training for learning highly efficient binarized messaging. Our empirical evaluation shows that HetNet sets a new state-of-the-art in learning coordination and communication strategies for heterogeneous multiagent teams by achieving an 5.84% to 707.65% performance improvement over the next-best baseline across multiple domains while simultaneously achieving a 200× reduction in the required communication bandwidth.
Esmaeil Seraj, Rohan R. Paleja, Luis Pimentel, Kin Man Lee, Zheyuan Wang, Matthew Sklar, John Z. Zhang, Zahi M. Kakish, Matthew C. Gombolay
IEEE Trans. Robotics1
2023 Embodied, Intelligent Communication for Multi-Agent Cooperation
abstract
High-performing human teams leverage intelligent and efficient communication and coordination strategies to collaboratively maximize their joint utility. Inspired by teaming behaviors among humans, I seek to develop computational methods for synthesizing intelligent communication and coordination strategies for collaborative multi-robot systems. I leverage both classical model-based control and planning approaches as well as data-driven methods such as Multi-Agent Reinforcement Learning (MARL) to provide several contributions towards enabling emergent cooperative teaming behavior across both homogeneous and heterogeneous (including agents with different capabilities) robot teams.
Esmaeil Seraj
AAAI1
2023 Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot Teams
abstract
Extending recent advances in Learning from Demonstration (LfD) frameworks to multi-robot settings poses critical challenges such as environment non-stationarity due to partial observability which is detrimental to the applicability of existing methods. Although prior work has shown that enabling communication among agents of a robot team can alleviate such issues, creating inter-agent communication under existing Multi-Agent LfD (MA-LfD) frameworks requires the human expert to provide demonstrations for both environment actions and communication actions, which necessitates an efficient communication strategy on a known message spaces. To address this problem, we propose Mixed-Initiative Multi-Agent Apprenticeship Learning (MixTURE). MixTURE enables robot teams to learn from a human expert-generated data a preferred policy to accomplish a collaborative task, while simultaneously learning emergent inter-agent communication to enhance team coordination. The key ingredient to MixTURE's success is automatically learning a communication policy, enhanced by a mutual-information maximizing reverse model that rationalizes the underlying expert demonstrations without the need for human generated data or an auxiliary reward function. MixTURE outperforms a variety of relevant baselines on diverse data generated by human experts in complex heterogeneous domains. MixTURE is the first MA-LfD framework to enable learning multi-robot collaborative policies directly from real human data, resulting in ~44% less human workload, and ~46% higher usability score.
Esmaeil Seraj, Jerry Xiong, Mariah Schrum, Matthew C. Gombolay
NeurIPS1
2022 Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming
Sachin Konan, Esmaeil Seraj, Matthew C. Gombolay
ICLR2
2022 Multi-UAV planning for cooperative wildfire coverage and tracking with quality-of-service guarantees
Esmaeil Seraj, Andrew Silva, Matthew C. Gombolay
Auton. Agents Multi Agent Syst.1
2022 A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot Teams
abstract
We propose a collaborative planning and control algorithm to enhance cooperation for composite teams of autonomous robots in dynamic environments. Composite robot teams are groups of agents that perform different tasks according to their respective capabilities in order to accomplish an overarching mission. Examples of such teams include groups of perception agents (can only sense) and action agents (can only manipulate) working together to perform disaster response tasks. Coordinating robots in a composite team is a challenging problem due to the heterogeneity in the robots’ characteristics and their tasks. Here, we propose a coordination framework for composite robot teams. The proposed framework consists of two hierarchical modules: First, A multiagent state-action-reward-time-state-action algorithm in multiagent partially observable semi-Markov decision process as the high-level decision-making module to enable perception agents to learn to surveil in an environment with an unknown number of dynamic targets and second, a low-level coordinated control and planning module that ensures probabilistically guaranteed support for action agents. Simulation and physical robot implementations of our algorithms on a multiagent robot testbed demonstrated the efficacy and feasibility of our coordination framework by reducing the overall operation times in a benchmark wildfire-fighting case study.
Esmaeil Seraj, Letian Chen, Matthew C. Gombolay
IEEE Trans. Robotics1