Somdeb Majumdar

dblp:63/8320 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2024
0009-0005-4873-4729ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 51% Video understanding and tracking · 28% Autonomous driving · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › population-based learning › evolutionary learning › population-based reinforcement learning
evolutionary reinforcement learning
0.822020
Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020
Collaborative Evolutionary Reinforcement Learning · ICML 2019
Computer vision › Video understanding and tracking › multimodal video understanding › audio-visual video understanding
active speaker detection
0.612022
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection · ECCV (35) 2022
Computer vision › Video understanding and tracking › human motion prediction
hand motion prediction
0.612022
Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos · CVPR 2022
Robotics › Autonomous driving
trajectory prediction
0.612022
Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos · CVPR 2022
Machine learning › Reinforcement learning › relational reinforcement learning
graph reinforcement learning
0.512021
Optimizing Memory Placement using Evolutionary Graph Reinforcement Learning · ICLR 2021
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412020
Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020
Machine learning › Reinforcement learning › reward design
reward shaping
0.412020
Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020
Machine learning › Reinforcement learning
sparse reward
0.412020
Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020
Machine learning › Reinforcement learning
deep reinforcement learning
0.412019
Collaborative Evolutionary Reinforcement Learning · ICML 2019
Machine learning › Reinforcement learning
exploration
0.412019
Collaborative Evolutionary Reinforcement Learning · ICML 2019
Computer vision › Video understanding and tracking
egocentric video understanding
0.212022
Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos · CVPR 2022
Machine learning › Graph learning › graph neural network
spatio-temporal graph
0.212022
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection · ECCV (35) 2022
Machine learning › Deep learning architectures and training › neural network training
neuroevolution
0.112020
Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination · ICML 2020
Machine learning › Reinforcement learning
continuous control
0.112019
Collaborative Evolutionary Reinforcement Learning · ICML 2019

Methods — techniques the papers use, named apart from their topics

evolutionary graph reinforcement learning · 1.0neuroevolution · 0.8transformer · 0.6self-attention · 0.6probabilistic sampling · 0.6graph neural network · 0.6gradient-based optimization · 0.4evolutionary algorithm · 0.4replay buffer · 0.4TD3 · 0.4
YearPublicationVenuePosition
2024 FloorSet - a VLSI Floorplanning Dataset with Design Constraints of Real-World SOCs
abstract
Floorplanning for systems-on-a-chip (SoCs) and its sub-systems is a crucial and non-trivial step of the physical design flow. It represents a difficult combinatorial optimization problem. A typical large scale SoC with 120 partitions generates a search-space of ~ 10250. As novel machine learning (ML) approaches emerge to tackle such problems, there is a growing need for a modern benchmark that comprises a large training dataset and performance metrics that better reflect real-world constraints and objectives compared to existing benchmarks. To address this need, we present FloorSet - two comprehensive datasets of synthetic fixed-outline floorplan layouts that reflect the distribution of real SoCs. Each dataset has 1M training samples and 100 test samples where each sample is a synthetic floor-plan. FloorSet-Prime comprises fully-abutted rectilinear partitions and near-optimal wire-length. A simplified dataset that reflects early design phases, FloorSet-Lite comprises rectangular partitions, with < 5% white-space and near-optimal wire-length. Both datasets define hard constraints seen in modern design flows such as shape constraints, edge-affinity, grouping constraints, and pre-placement constraints. FloorSet is intended to spur fundamental research on large-scale constrained optimization problems. Crucially, FloorSet alleviates the core issue of reproducibility in modern ML driven solutions to such problems. FloorSet is available as an open-source repository for the research community1.
Uday Mallappa, Hesham Mostafa, Michael Galkin, Mariano Phielipp, Somdeb Majumdar
ICCAD5
2023 Exploiting Long-Term Dependencies for Generating Dynamic Scene Graphs
abstract
Dynamic scene graph generation from a video is challenging due to the temporal dynamics of the scene and the inherent temporal fluctuations of predictions. We hypothesize that capturing long-term temporal dependencies is the key to effective generation of dynamic scene graphs. We propose to learn the long-term dependencies in a video by capturing the object-level consistency and inter-object relationship dynamics over object-level long-term tracklets using transformers. Experimental results demonstrate that our Dynamic Scene Graph Detection Transformer (DSG- DETR) outperforms state-of-the-art methods by a significant margin on the benchmark dataset Action Genome. Our ablation studies validate the effectiveness of each component of the proposed approach. The source code is available at https://github.com/Shengyu-Feng/DS G-DETR.
Shengyu Feng, Hesham Mostafa, Marcel Nassar, Somdeb Majumdar, Subarna Tripathi
WACV4
2022 Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos
abstract
We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e., interaction hotspots). This relatively low-dimensional representation provides a con-crete description of future interactions. To tackle this task, we first provide an automatic way to collect trajectory and hotspots labels on large-scale data. We then use this data to train an Object-Centric Transformer (OCT) model for prediction. Our model performs hand and object interaction reasoning via the self-attention mechanism in Transformers. OCT also provides a probabilistic framework to sample the future trajectory and hotspots to handle uncertainty in prediction. We perform experi-ments on the Epic-Kitchens-55, Epic-Kitchens-100 and EGTEA Gaze+ datasets, and show that OCT significantly outperforms state-of the-art approaches by a large margin. Project page is available at https://stevenlsw.github.io/hoi-forecast.
Subarna Tripathi, Somdeb Majumdar, Xiaolong Wang 0004
CVPR3
2022 Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
Kyle Min 0001, Sourya Roy, Subarna Tripathi, Tanaya Guha, Somdeb Majumdar
ECCV (35)5
2022 Neuroevolution-enhanced multi-objective optimization for mixed-precision quantization
abstract
Mixed-precision quantization is a powerful tool to enable memory and compute savings of neural network workloads by deploying different sets of bit-width precisions on separate compute operations. In this work, we present a flexible and scalable framework for automated mixed-precision quantization that concurrently optimizes task performance, memory compression, and compute savings through multi-objective evolutionary computing. Our framework centers on Neuroevolution-Enhanced Multi-Objective Optimization (NEMO), a novel search method, which combines established search methods with the representational power of neural networks. Within NEMO, the population is divided into structurally distinct sub-populations, or species, which jointly create the Pareto frontier of solutions for the multi-objective problem. At each generation, species perform separate mutation and crossover operations, and are re-sized in proportion to the goodness of their contribution to the Pareto frontier. In our experiments, we define a graph-based representation to describe the underlying workload, enabling us to deploy graph neural networks trained by NEMO via neuroevolution, to find Pareto optimal configurations for MobileNet-V2, ResNet50 and ResNeXt-101-32×8d. Compared to the state-of-the-art, we achieve competitive results on memory compression and superior results for compute compression. Further analysis reveals that the graph representation and the species-based approach employed by NEMO are critical to finding optimal solutions.
Santiago Miret, Vui Seng Chua, Mattias Marder, Mariano Phielipp, Nilesh Jain, Somdeb Majumdar
GECCO6
2022 Learning Intrinsic Symbolic Rewards in Reinforcement Learning
abstract
Learning effective policies for sparse objectives is a key challenge in Deep Reinforcement Learning (RL). A common approach is to design task-related dense rewards to improve task learnability. While such rewards are easily interpreted, they rely on heuristics and domain expertise. Alternate approaches that train neural networks to discover dense surrogate rewards avoid heuristics, but are high-dimensional, black-box solutions offering little interpretability. In this paper, we present a method that discovers dense rewards in the form of low-dimensional symbolic trees - thus making them more tractable for analysis. The trees use simple functional operators to map an agent's observations to a scalar reward, which then supervises the policy gradient learning of a neural network policy. We test our method on continuous action spaces in Mujoco and discrete action spaces in Atari and Pygame environments. We show that the discovered dense rewards are an effective signal for an RL policy to solve the benchmark tasks. Notably, we significantly outperform a widely used, contemporary neural-network based reward-discovery algorithm in all environments considered.
Hassam Ullah Sheikh, Shauharda Khadka, Santiago Miret, Somdeb Majumdar, Mariano Phielipp
IJCNN4
2021 MAEDyS: multiagent evolution via dynamic skill selection
abstract
Evolving effective coordination strategies in tightly coupled multi-agent settings with sparse team fitness evaluations is challenging. It relies on multiple agents simultaneously stumbling upon the goal state to generate a learnable feedback signal. In such settings, estimating an agent's contribution to the overall team performance is extremely difficult, leading to a well-known structural credit assignment problem. This problem is further exacerbated when agents must complete sub-tasks with added spatial and temporal constraints, and different sub-tasks may require different local skills. We introduce MAEDyS, Multiagent Evolution via Dynamic Skill Selection, a hybrid bi-level optimization framework that augments evolutionary methods with policy gradient methods to generate effective coordination policies. MAEDyS learns to dynamically switch between multiple local skills towards optimizing the team fitness. It adopts fast policy gradients to learn several local skills using dense local rewards. It utilizes an evolutionary process to optimize the delayed team fitness by recruiting the most optimal skill at any given time. The ability to switch between various local skills during an episode eliminates the need for designing heuristic mixing functions. We evaluate MAEDyS in complex multiagent coordination environments with spatial and temporal constraints and show that it outperforms prior methods.
Enna Sachdeva, Shauharda Khadka, Somdeb Majumdar, Kagan Tumer
GECCO3
2021 Optimizing Memory Placement using Evolutionary Graph Reinforcement Learning
Shauharda Khadka, Estelle Aflalo, Mattias Marder, Avrech Ben-David, Santiago Miret, Shie Mannor, Tamir Hazan, Somdeb Majumdar
ICLR9
2020 Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination
abstract
Many cooperative multiagent reinforcement learning environments provide agents with a sparse team-based reward, as well as a dense agent-specific reward that incentivizes learning basic skills. Training policies solely on the team-based reward is often difficult due to its sparsity. Also, relying solely on the agent-specific reward is sub-optimal because it usually does not capture the team coordination objective. A common approach is to use reward shaping to construct a proxy reward by combining the individual rewards. However, this requires manual tuning for each environment. We introduce Multiagent Evolutionary Reinforcement Learning (MERL), a split-level training platform that handles the two objectives separately through two optimization processes. An evolutionary algorithm maximizes the sparse team-based objective through neuroevolution on a population of teams. Concurrently, a gradient-based optimizer trains policies to only maximize the dense agent-specific rewards. The gradient-based policies are periodically added to the evolutionary population as a way of information transfer between the two optimization processes. This enables the evolutionary algorithm to use skills learned via the agent-specific rewards toward optimizing the global objective. Results demonstrate that MERL significantly outperforms state-of-the-art methods, such as MADDPG, on a number of difficult coordination benchmarks.
Somdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer, Kagan Tumer
ICML1
2019 Collaborative Evolutionary Reinforcement Learning
abstract
Deep reinforcement learning algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically struggle with achieving effective exploration and are extremely sensitive to the choice of hyperparameters. One reason is that most approaches use a noisy version of their operating policy to explore - thereby limiting the range of exploration. In this paper, we introduce Collaborative Evolutionary Reinforcement Learning (CERL), a scalable framework that comprises a portfolio of policies that simultaneously explore and exploit diverse regions of the solution space. A collection of learners - typically proven algorithms like TD3 - optimize over varying time-horizons leading to this diverse portfolio. All learners contribute to and use a shared replay buffer to achieve greater sample efficiency. Computational resources are dynamically distributed to favor the best learners as a form of online algorithm selection. Neuroevolution binds this entire process to generate a single emergent learner that exceeds the capabilities of any individual learner. Experiments in a range of continuous control benchmarks demonstrate that the emergent learner significantly outperforms its composite learners while remaining overall more sample-efficient - notably solving the Mujoco Humanoid benchmark where all of its composite learners (TD3) fail entirely in isolation.
Shauharda Khadka, Somdeb Majumdar, Tarek Nassar, Zach Dwiel, Evren Tumer, Santiago Miret, Yinyin Liu, Kagan Tumer
ICML2
2012 A closed-loop system for artifact mitigation in ambulatory electrocardiogram monitoring
abstract
Motion artifacts interfere with electrocardiogram (ECG) detection and information processing. In this paper, we present an independent component analysis based technique to mitigate these signal artifacts. We propose a new statistical measure to enable an automatic identification and removal of independent components, which correspond to the sources of noise. For the first time, we also present a signal-dependent closed-loop system for the quality assessment of the denoised ECG. In one experiment, noisy data is obtained by the addition of calibrated amounts of noise from the MIT-BIH NST database to the AHA ECG database. Arrhythmia classification based on a state-of-the-art algorithm with the direct use of noisy data thus obtained shows sensitivity and positive predictivity values of 87.7% and 90.0%, respectively, at an input signal SNR of -9 dB. Detection with the use of ECG data denoised by the proposed approach exhibits significant improvement in the performance of the classifier with the corresponding results being 96.5% and 99.1%, respectively. In a related lab trial, we demonstrate a reduction in RMS error of instantaneous heart rate estimates from 47.2% to 7.0% with the use of 56 minutes of denoised ECG from four physically active subjects. To validate our experiments, we develop a closed-loop, ambulatory ECG monitoring platform, which consumes 2.17 mW of power and delivers a data rate of 33 kbps over a dedicated UWB link.
Mohammed Shoaib, Gene Marsh, Harinath Garudadri, Somdeb Majumdar
DATE4