Dipanjan Das 0003

dblp:232/3022 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0003-2325-0646ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot manipulation · 29% Planning, search and constraint satisfaction · 25% Reinforcement learning · 16%
Computer graphics and multimedia
1 paper
Computer animation and physical simulation · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
1.522024
Task Planning for Visual Room Rearrangement under Partial Observability · ICLR 2024
Task Planning for Object Rearrangement in Multi-Room Environments · AAAI 2024
Robotics › Robot manipulation
object rearrangement
1.022024
Task Planning for Object Rearrangement in Multi-Room Environments · AAAI 2024
Task Planning for Visual Room Rearrangement under Partial Observability · ICLR 2024
Robotics › Robot navigation and mapping
object search
0.812024
Task Planning for Visual Room Rearrangement under Partial Observability · ICLR 2024
Machine learning › Reinforcement learning
partial observability
0.812024
Task Planning for Visual Room Rearrangement under Partial Observability · ICLR 2024
Robotics › Robot manipulation › object rearrangement
visual room rearrangement
0.812024
Task Planning for Visual Room Rearrangement under Partial Observability · ICLR 2024
Machine learning › Generative modeling
generative adversarial network
0.412020
Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture · ECCV (30) 2020
Natural language and speech › Speech recognition and synthesis
speech-driven animation
0.412020
Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture · ECCV (30) 2020
Computer animation and physical simulation
facial animation
0.412020
Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture · ECCV (30) 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.212024
Task Planning for Object Rearrangement in Multi-Room Environments · AAAI 2024
Machine learning › Reinforcement learning
deep reinforcement learning
0.212024
Task Planning for Object Rearrangement in Multi-Room Environments · AAAI 2024

Methods — techniques the papers use, named apart from their topics

large language model commonsense knowledge · 1.5deep reinforcement learning · 1.5generative adversarial network · 0.9cascaded GANs · 0.9graph-based state representation · 0.8directed spatial graph · 0.8cross-entropy method · 0.8cluster-biased sampling · 0.8
YearPublicationVenuePosition
2024 Task Planning for Object Rearrangement in Multi-Room Environments
abstract
Object rearrangement in a multi-room setup should produce a reasonable plan that reduces the agent's overall travel and the number of steps. Recent state-of-the-art methods fail to produce such plans because they rely on explicit exploration for discovering unseen objects due to partial observability and a heuristic planner to sequence the actions for rearrangement. This paper proposes a novel task planner to efficiently plan a sequence of actions to discover unseen objects and rearrange misplaced objects within an untidy house to achieve a desired tidy state. The proposed method introduces several innovative techniques, including (i) a method for discovering unseen objects using commonsense knowledge from large language models, (ii) a collision resolution and buffer prediction method based on Cross-Entropy Method to handle blocked goal and swap cases, (iii) a directed spatial graph-based state space for scalability, and (iv) deep reinforcement learning (RL) for producing an efficient plan to simultaneously discover unseen objects and rearrange the visible misplaced ones to minimize the overall traversal. The paper also presents new metrics and a benchmark dataset called MoPOR to evaluate the effectiveness of the rearrangement planning in a multi-room setting. The experimental results demonstrate that the proposed method effectively addresses the multi-room rearrangement problem.
Karan Mirakhor, Dipanjan Das 0003, Brojeshwar Bhowmick
AAAI3
2024 Task Planning for Visual Room Rearrangement under Partial Observability
abstract
This paper presents a novel hierarchical task planner under partial observability that empowers an embodied agent to use visual input to efficiently plan a sequence of actions for simultaneous object search and rearrangement in an untidy room, to achieve a desired tidy state. The paper introduces (i) a novel Search Network that utilizes commonsense knowledge from large language models to find unseen objects, (ii) a Deep RL network trained with proxy reward, along with (iii) a novel graph-based state representation to produce a scalable and effective planner that interleaves object search and rearrangement to minimize the number of steps taken and overall traversal of the agent, as well as to resolve blocked goal and swap cases, and (iv) a sample-efficient cluster-biased sampling for simultaneous training of the proxy reward network along with the Deep RL network. Furthermore, the paper presents new metrics and a benchmark dataset - RoPOR, to measure the effectiveness of rearrangement planning. Experimental results show that our method significantly outperforms the state-of-the-art rearrangement methods Weihs et al. (2021a); Gadre et al. (2022); Sarch et al. (2022); Ghosh et al. (2022).
Karan Mirakhor, Dipanjan Das 0003, Brojeshwar Bhowmick
ICLR3
2022 Planning Large-scale Object Rearrangement Using Deep Reinforcement Learning
abstract
Object rearrangement is about moving a set of objects from an initial state to a goal state through task and motion planning. Existing methods either show poor scalability in number of objects they can handle, or do not generalize well across situations, or need explicit running buffers to avoid collisions during placements. In this paper, we propose a deep-RL based task planning method to solve large-scale object rearrangement problems. Given the source and target state of objects in the form of images, our method determines a collision-free object movement plan. Our method produces a feasible plan in discrete-continuous action space where picking the selected objects are discrete actions followed by a set of continuous actions to place the object. We propose a novel hierarchical dense reward structure to train our deep-RL network to make our method more sample efficient using the AI2Thor simulator. We show that our method works well on unseen publicly available datasets and on a publicly available simulation environment such as Pybullet thereby demonstrating the superiority of our method in terms of generalizability. To the best of our knowledge, our method is the first one that demonstrates the rearrangement across different scenarios from 2D surfaces such as tabletops to 3D rooms with a large number of objects and without any explicit need of buffer space.
Dipanjan Das 0003, Marichi Agarwal, Brojeshwar Bhowmick
IJCNN2
2020 Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture
Dipanjan Das 0003, Sandika Biswas, Sanjana Sinha, Brojeshwar Bhowmick
ECCV (30)1
2020 Variational Clustering: Leveraging Variational Autoencoders for Image Clustering
abstract
Recent advances in deep learning have shown their ability to learn strong feature representations for images. The task of image clustering naturally requires good feature representations to capture the distribution of the data and subsequently differentiate data points from one another. Often these two aspects are dealt with independently and thus traditional feature learning alone does not suffice in partitioning the data meaningfully. Variational Autoencoders (VAEs) naturally lend themselves to learning data distributions in a latent space. Since we wish to efficiently discriminate between different clusters in the data, we propose a method based on VAEs where we use a Gaussian Mixture prior to help cluster the images accurately. We jointly learn the parameters of both the prior and the posterior distributions. Our method represents a true Gaussian Mixture VAE. This way, our method simultaneously learns a prior that captures the latent distribution of the images and a posterior to help discriminate well between data points. We also propose a novel reparametrization of the latent space consisting of a mixture of discrete and continuous variables. One key takeaway is that our method generalizes better across different datasets without using any pre-training or learnt models, unlike existing methods, allowing it to be trained from scratch in an end-to-end manner. We verify our efficacy and generalizability experimentally by achieving state-of-the-art results among unsupervised methods on a variety of datasets. To the best of our knowledge, we are the first to pursue image clustering using VAEs in a purely unsupervised manner on real image datasets.
Vignesh Prasad, Dipanjan Das 0003, Brojeshwar Bhowmick
IJCNN2
2019 Deep Representation Learning Characterized by Inter-Class Separation for Image Clustering
abstract
Despite significant advances in clustering methods in recent years, the outcome of clustering of a natural image dataset is still unsatisfactory due to two important drawbacks. Firstly, clustering of images needs a good feature representation of an image and secondly, we need a robust method which can discriminate these features for making them belonging to different clusters such that intra-class variance is less and inter-class variance is high. Often these two aspects are dealt with independently and thus the features are not sufficient enough to partition the data meaningfully. In this paper, we propose a method where we discover these features required for the separation of the images using deep autoencoder. Our method learns the image representation features automatically for the purpose of clustering and also select a coherent image and an incoherent image simultaneously for a given image so that the feature representation learning can learn better discriminative features for grouping the similar images in a cluster and at the same time separating the dissimilar images across clusters. Experiment results show that our method produces significantly better result than the state-of-the-art methods and we also show that our method is more generalized across different dataset without using any pre-trained model like other existing methods.
Dipanjan Das 0003, Ratul Ghosh, Brojeshwar Bhowmick
WACV1