Brojeshwar Bhowmick

dblp:88/7529 · DBLP profile ↗
← Back
41ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0001-9291-5889ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 8 since 2021Systems, architecture and hardware · 9 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Sketch3R: Rapid and Realistic 3D VR Sketch Creation to Shape Retrieval
abstract
Large 3D shape repositories are rapidly expanding, driven by advances in generative modeling, making efficient shape retrieval increasingly important for authoring tools. While text queries capture high-level semantics, they often fail to convey precise geometric details. 3D sketches provide a more expressive means of representing shape geometry, and recent AR/VR developments have made sketch-based retrieval practical. However, existing 3D sketch datasets face three major limitations: (1) reliance on quad meshes or voxel hulls, which often fail on complex or non-manifold shapes; (2) use of fixed-size point clouds that discard stroke connectivity and limit geometric fidelity; and (3) dependence on expensive curve-based or multi-view rendering pipelines, which hinder large-scale data generation. Limited point cloud representations also fail to capture sketch connectivity and topology when used to train retrieval models. To address these challenges, we propose Sketch3R, a scalable framework that converts arbitrary 3D meshes into human-like VR sketches using a graph-based representation that preserves stroke connectivity and adapts to sketch complexity. Leveraging this representation, Sketch3R employs a lightweight graph-attention Siamese network for efficient and accurate sketch-to-shape retrieval. Experiments demonstrate that our method outperforms prior approaches in both accuracy and speed, while robustly handling 3D shapes across diverse topologies.
Mritunjoy Halder, Shivam Ashok Shukla, Lokender Tiwari, Raghav Mittal, Brojeshwar Bhowmick
WACV5
2026 ObjectMeshDeform : Towards recovering precise 3D geometry of real objects via image-guided mesh deformation of 3D generative priors
abstract
3D Generative Models that synthesize high fidelity 3D assets from single view or multi-view images cannot recover precise 3D geometry and dimensions of real-world objects, which is often required by practical applications. On the other hand, multi-view 3D reconstruction methods based on structure from motion, implicit surfaces, gaussian splatting, fail to recover high fidelity object meshes with dense geometry, shape regularity, smooth surfaces present in real world objects. In this paper we propose a novel approach that leverages a 3D mesh prior synthesized by generative models pre-trained on large scale 3D synthetic datasets. Our method automatically refines the 3D geometry of the generated meshes for improved geometric precision using only sparse multi-view input images, while maintaining the geometric fidelity of the generated prior meshes. Our method can automatically reconstruct meshes from images of real world objects, without requiring any additional large-scale training data or manual inputs for shape deformation.
Siddharth Katageri, Sanjana Sinha, Soumyadip Maity, Brojeshwar Bhowmick
WACV5
2025 SketchTo3DGen : GenAI Powered Articulation Ready 3D Asset Ideation using 3D Sketches and Audio Descriptions
abstract
We present SketchTo3DGen, a novel system for rapid 3D content ideation on a VR headset. SketchTo3DGen combines freehand 3D sketching and audio descriptions to generate photo-realistic 3D assets on-the-fly. Running on a Meta Quest headset with a Unity application, our system leverages remote GPU-accelerated services for AI-driven content creation using intuitive inputs. The user can draw a 3D sketch in mid-air and describe the intended asset verbally; our pipeline transcribes and normalizes the speech into a text prompt, selects informative viewpoints of the 3D VR sketch, generates corresponding images via a state-of-the-art text-to-image model, and finally reconstructs a 3D mesh using an image-to-3D generator. The entire workflow is experienced in VR with minimal interface elements. We describe the design motivations, technical pipeline, and user interaction details of SketchTo3DGen. This VR in-headset pipeline, using intuitive inputs in the form of hand-drawn 3D VR sketch and speech, streamlines 3D modeling, accelerating the generation of articulation ready 3D assets.
Shivam Ashok Shukla, Raghav Mittal, Lokender Tiwari, Brojeshwar Bhowmick
VRST4
2025 DisFlowEm : One-Shot Emotional Talking Head Generation Using Disentangled Pose and Expression Flow-Guidance
abstract
Generating realistic one-shot emotional talking head animation on arbitrary faces is a challenging problem, as it requires realistic emotions, head movements, identity preser-vation, and accurate lip sync. Existing emotional talking face generation methods either fail to retain the identity information of arbitrary subjects owing to the limited variability of existing emotional datasets, or they fail to capture emotions accurately even if they preserve identity of arbitrary faces. Moreover, most of the methods rely on additional input videos for driving poses and/or or expressions on the generated video. For practical applications, it is in-feasible to obtain driving videos of the same or different subject with variations in head pose, expressions etc. In this paper, we propose a novel approach for Audio-driven Emotional Talking Head generation from a single image, with emotion-controllable head pose generation. Unlike existing methods, our method does not require a driving video either for pose or emotions, and can generate different emotions and diverse head pose variations from input speech and a single image of an arbitrary subject in neutral emotion. Our method overcomes the limitations of existing emotional audio-visual datasets by learning a disentangled approach for optical flow computation approach for pose and expression. Using our proposed method of independently computing pose-driven and expression-driven optical flow, our image generation network can be pretrained on a large dataset with greater pose variability but lacking emotion annotations. The expression flow generation branch is fine-tuned on a smaller emotional dataset to accurately capture different emotions not present in the original dataset, while retaining the pose variability from the original dataset. We present extensive experiments to demonstrate the superior-ity of our proposed method in generating talking head animation with accurate emotions, diverse head movements, and generalization to arbitrary faces.
Sanjana Sinha, Brojeshwar Bhowmick, Lokender Tiwari, Sushovan Chanda
WACV2
2024 Task Planning for Object Rearrangement in Multi-Room Environments
abstract
Object rearrangement in a multi-room setup should produce a reasonable plan that reduces the agent's overall travel and the number of steps. Recent state-of-the-art methods fail to produce such plans because they rely on explicit exploration for discovering unseen objects due to partial observability and a heuristic planner to sequence the actions for rearrangement. This paper proposes a novel task planner to efficiently plan a sequence of actions to discover unseen objects and rearrange misplaced objects within an untidy house to achieve a desired tidy state. The proposed method introduces several innovative techniques, including (i) a method for discovering unseen objects using commonsense knowledge from large language models, (ii) a collision resolution and buffer prediction method based on Cross-Entropy Method to handle blocked goal and swap cases, (iii) a directed spatial graph-based state space for scalability, and (iv) deep reinforcement learning (RL) for producing an efficient plan to simultaneously discover unseen objects and rearrange the visible misplaced ones to minimize the overall traversal. The paper also presents new metrics and a benchmark dataset called MoPOR to evaluate the effectiveness of the rearrangement planning in a multi-room setting. The experimental results demonstrate that the proposed method effectively addresses the multi-room rearrangement problem.
Karan Mirakhor, Dipanjan Das 0003, Brojeshwar Bhowmick
AAAI4
2024 Task Planning for Visual Room Rearrangement under Partial Observability
abstract
This paper presents a novel hierarchical task planner under partial observability that empowers an embodied agent to use visual input to efficiently plan a sequence of actions for simultaneous object search and rearrangement in an untidy room, to achieve a desired tidy state. The paper introduces (i) a novel Search Network that utilizes commonsense knowledge from large language models to find unseen objects, (ii) a Deep RL network trained with proxy reward, along with (iii) a novel graph-based state representation to produce a scalable and effective planner that interleaves object search and rearrangement to minimize the number of steps taken and overall traversal of the agent, as well as to resolve blocked goal and swap cases, and (iv) a sample-efficient cluster-biased sampling for simultaneous training of the proxy reward network along with the Deep RL network. Furthermore, the paper presents new metrics and a benchmark dataset - RoPOR, to measure the effectiveness of rearrangement planning. Experimental results show that our method significantly outperforms the state-of-the-art rearrangement methods Weihs et al. (2021a); Gadre et al. (2022); Sarch et al. (2022); Ghosh et al. (2022).
Karan Mirakhor, Dipanjan Das 0003, Brojeshwar Bhowmick
ICLR4
2024 Anticipate & Act: Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments†
abstract
Assistive agents performing household tasks such as making the bed or cooking breakfast often compute and execute actions that accomplish one task at a time. However, efficiency can be improved by anticipating upcoming tasks and computing an action sequence that jointly achieves these tasks. State-of-the-art methods for task anticipation use data-driven deep networks and Large Language Models (LLMs), but they do so at the level of high-level tasks and/or require many training examples. Our framework leverages the generic knowledge of LLMs through a small number of prompts to perform high-level task anticipation, using the anticipated tasks as goals in a classical planning system to compute a sequence of finer-granularity actions that jointly achieve these goals. We ground and evaluate our framework’s abilities in realistic scenarios in the VirtualHome environment and demonstrate a 31% reduction in execution time compared with a system that does not consider upcoming tasks.
Raghav Arora, Shivam Singh, Karthik Swaminathan, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
ICRA6
2023 Sequence-Agnostic Multi-Object Navigation
abstract
The Multi-Object Navigation (MultiON) task requires a robot to localize an instance (each) of multiple object classes. It is a fundamental task for an assistive robot in a home or a factory. Existing methods for MultiON have viewed this as a direct extension of Object Navigation (ON), the task of localising an instance of one object class, and are pre-sequenced, i.e., the sequence in which the object classes are to be explored is provided in advance. This is a strong limitation in practical applications characterized by dynamic changes. This paper describes a deep reinforcement learning framework for sequence-agnostic MultiON based on an actor-critic architecture and a suitable reward specification. Our framework leverages past experiences and seeks to reward progress toward individual as well as multiple target object classes. We use photo-realistic scenes from the Gibson benchmark dataset in the AI Habitat 3D simulation environment to experimentally show that our method performs better than a pre-sequenced approach and a state of the art ON method extended to MultiON.
Nandiraju Gireesh, Ahana Datta, Snehasis Banerjee, Mohan Sridharan, Brojeshwar Bhowmick, K. Madhava Krishna
ICRA6
2023 SCARP: 3D Shape Completion in ARbitrary Poses for Improved Grasping
abstract
Recovering full 3D shapes from partial observations is a challenging task that has been extensively addressed in the computer vision community. Many deep learning methods tackle this problem by training 3D shape generation networks to learn a prior over the full 3D shapes. In this training regime, the methods expect the inputs to be in a fixed canonical form, without which they fail to learn a valid prior over the 3D shapes. We propose SCARP, a model that performs Shape C ompletion in ARbitrary Poses. Given a partial pointcloud of an object, SCARP learns a disentangled feature representation of pose and shape by relying on rotationally equivariant pose features and geometric shape features trained using a multi-tasking objective. Unlike existing methods that depend on an external canonicalization method, SCARP performs canonicalization, pose estimation, and shape completion in a single network, improving the performance by 45% over the existing baselines. In this work, we use SCARP for improving grasp proposals on tabletop objects. By completing partial tabletop objects directly in their observed poses, SCARP enables a SOTA grasp proposal network improve their proposals by 71.2% on partial shapes. Project page: https://bipashasen.github.io/scarp
Bipasha Sen, Aditya Agarwal, Gaurav Singh 0012, Brojeshwar Bhowmick, Srinath Sridhar 0002, K. Madhava Krishna
ICRA4
2023 Learning Arc-Length Value Function for Fast Time-Optimal Pick and Place Sequence Planning and Execution
abstract
This paper presents a real-time algorithm for computing the optimal sequence and motion plans for a fixed-base manipulator to pick and place a set of given objects. The optimality is defined in terms of the total execution time of the sequence or its proxy, the arc-length in the joint-space. The fundamental complexity stems from the fact that the optimality metric depends on the joint motion, but the task specification is in the end-effector space. Moreover, mapping between a pair of end-effector positions to the shortest arc-length joint trajectory is not analytic; instead, it entails solving a complex trajectory optimization problem. Existing works ignore this complex mapping and use the Euclidean distance in the end-effector space to compute the sequence. In this paper, we overcome the reliance on the Euclidean distance heuristic by introducing a novel data-driven technique to estimate the optimal arc-length cost in joint space (a.k.a the value function) between two given end-effector positions. We parametrize the value function as a Neural Network and motivate a niche choice for its architecture, inspired by the works on metric learning. The learned value function is then used as an edge cost in a capacitated vehicle routing problem (CVRP) set-up to compute the optimal visitation sequence. Finally, we optimize over the input space of the learnt value function network to propose a novel Inverse Kinematics (IK) algorithm that produces substantially shorter joint arc-length trajectories than existing approaches while executing the computed optimal sequence. We show that our sequence planner, in combination with our proposed IK, offers a substantial improvement in joint arc-length over existing state-of-the-art while maintaining scalability to a large number of objects.
Prajwal Thakur, M. Nomaan Qureshi, Arun Kumar Singh 0001, Y. V. S. Harish, Pushkal Katara, Houman Masnavi, K. Madhava Krishna, Brojeshwar Bhowmick
IJCNN8
2023 Exploring Social Motion Latent Space and Human Awareness for Effective Robot Navigation in Crowded Environments
abstract
This work proposes a novel approach to social robot navigation by learning to generate robot controls from a social motion latent space. By leveraging this social motion latent space, the proposed method achieves significant improvements in social navigation metrics such as success rate, navigation time, and trajectory length while producing smoother (less jerk and angular deviations) and more anticipatory trajectories. The superiority of the proposed method is demonstrated through comparison with baseline models in various scenarios. Additionally, the concept of humans' awareness towards the robot is introduced into the social robot navigation framework, showing that incorporating human awareness leads to shorter and smoother trajectories owing to humans' ability to positively interact with the robot.
Junaid Ahmed Ansari, Satyajit Tourani, Gourav Kumar, Brojeshwar Bhowmick
IROS4
2023 CLIPGraphs: Multimodal Graph Networks to Infer Object-Room Affinities
abstract
This paper introduces a novel method for determining the best room to place an object in, for embodied scene rearrangement. While state-of-the-art approaches rely on large language models (LLMs) or reinforcement learned (RL) policies for this task, our approach, CLIPGraphs, efficiently combines commonsense domain knowledge, data-driven methods, and recent advances in multimodal learning. Specifically, it (a) encodes a knowledge graph of prior human preferences about the room location of different objects in home environments, (b) incorporates vision-language features to support multimodal queries based on images or text, and (c) uses a graph network to learn object-room affinities based on embeddings of the prior knowledge and the vision-language features. We demonstrate that our approach provides better estimates of the most appropriate location of objects from a benchmark set of object categories in comparison with state-of-the-art baselines.11Supplementary material and code: https://clipgraphs.github.io
Raghav Arora, Ahana Datta, Snehasis Banerjee, Brojeshwar Bhowmick, Krishna Murthy Jatavallabhula, Mohan Sridharan, K. Madhava Krishna
RO-MAN5
2023 GarSim: Particle Based Neural Garment Simulator
abstract
We present a particle-based neural garment simulator (dubbed as GarSim) that can simulate template garments on the target arbitrary body poses. Existing learning-based methods majorly work for specific garment type (e.g. top, skirt, etc) or garment topology, and needs retraining for a new type of garment. Similarly, some methods focus on a particular fabric, body shape, and pose. To circumvent these limitations, our method fundamentally learns the physical dynamics of the garment vertices conditioned on underlying body shape, motion, and fabric properties to generalize across garment types, topology, and fabric along with different body shape and pose. In particular, we represent the garment as a graph, where the nodes represent the physical state of the garment vertices, and the edges represent the relation between the two nodes. The nodes and edges of the garment graph encode various properties of garments and the human body to compute the dynamics of the vertices through a learned message-passing. Learning of such dynamics of the garment vertices conditioned on underlying body motion and fabric properties enables our method to be trained simultaneously for multiple types of garments (e.g., tops, skirts, etc) with arbitrary mesh resolutions, varying topologies, and fabric properties. Our experimental results show that GarSim with less amount of training data not only outperforms the SOTA methods on challenging CLOTH3D dataset both qualitatively and quantitatively, but also works reliably well on the unseen poses obtained from YouTube videos, and give satisfactory results on unseen cloth types which were not present during the training.
Lokender Tiwari, Brojeshwar Bhowmick
WACV2
2023 Robo-vision! 3D mesh generation of a scene for a robot for planar and non-planar complex objects
Swapna Agarwal, Soumyadip Maity, Hrishav Bakul Barua, Brojeshwar Bhowmick
Multim. Tools Appl.4
2022 Emotion-Controllable Generalized Talking Face Generation
abstract
Despite the significant progress in recent years, very few of the AI-based talking face generation methods attempt to render natural emotions. Moreover, the scope of the methods is majorly limited to the characteristics of the training dataset, hence they fail to generalize to arbitrary unseen faces. In this paper, we propose a one-shot facial geometry-aware emotional talking face generation method that can generalize to arbitrary faces. We propose a graph convolutional neural network that uses speech content feature, along with an independent emotion input to generate emotion and speech-induced motion on facial geometry-aware landmark representation. This representation is further used in our optical flow-guided texture generation network for producing the texture. We propose a two-branch texture generation network, with motion and texture branches designed to consider the motion and texture content independently. Compared to the previous emotion talking face methods, our method can adapt to arbitrary faces captured in-the-wild by fine-tuning with only a single image of the target identity in neutral emotion.
Sanjana Sinha, Sandika Biswas, Ravindra Yadav, Brojeshwar Bhowmick
IJCAI4
2022 Planning Large-scale Object Rearrangement Using Deep Reinforcement Learning
abstract
Object rearrangement is about moving a set of objects from an initial state to a goal state through task and motion planning. Existing methods either show poor scalability in number of objects they can handle, or do not generalize well across situations, or need explicit running buffers to avoid collisions during placements. In this paper, we propose a deep-RL based task planning method to solve large-scale object rearrangement problems. Given the source and target state of objects in the form of images, our method determines a collision-free object movement plan. Our method produces a feasible plan in discrete-continuous action space where picking the selected objects are discrete actions followed by a set of continuous actions to place the object. We propose a novel hierarchical dense reward structure to train our deep-RL network to make our method more sample efficient using the AI2Thor simulator. We show that our method works well on unseen publicly available datasets and on a publicly available simulation environment such as Pybullet thereby demonstrating the superiority of our method in terms of generalizability. To the best of our knowledge, our method is the first one that demonstrates the rearrangement across different scenarios from 2D surfaces such as tabletops to 3D rooms with a large number of objects and without any explicit need of buffer space.
Dipanjan Das 0003, Marichi Agarwal, Brojeshwar Bhowmick
IJCNN5
2022 IndoLayout: Leveraging Attention for Extended Indoor Layout Estimation from an RGB Image
abstract
In this work, we propose IndoLayout, a novel real-time approach for generating high-quality occupancy maps from an RGB image for indoor scenes. Such occupancy maps are often crucial for path-planning and mapping in indoor environments but are often built using only information contained in the ego view. In contrast, our approach also predicts occupancy values beyond immediately visible regions from just a monocular image, leveraging learnt priors from indoor scenes. Hence, our proposed network can produce a hallucinated, amodal scene layout that includes areas occluded in the RGB image, such as a navigable floor behind a desk. Specifically, we propose a novel architecture that uses self-attention and adversarial learning to vastly improve the quality of the predicted layout. We evaluate our model on several photorealistic indoor datasets and outperform previous relevant work on all metrics that measure layout quality, including newly adopted ones. Finally, we demonstrate the effectiveness of our method by showing significant improvements on the PointNav task over similar approaches using IndoLayout. For more details, please refer to the project page: https://indolayout.github.io/.
Shantanu Singh, Jaidev Shriram, Shaantanu Kulkarni, Brojeshwar Bhowmick, K. Madhava Krishna
IROS4
2021 RTVS: A Lightweight Differentiable MPC Framework for Real-Time Visual Servoing
abstract
Recent data-driven approaches to visual servoing have shown improved performances over classical methods due to precise feature matching and depth estimation. Some recent servoing approaches use a model predictive control (MPC) framework which generalise well to novel environments and are capable of incorporating dynamic constraints, but are computationally intractable in real-time, making it difficult to deploy in real-world scenarios. On the contrary, single-step methods optimise greedily and achieve high servoing rates, but lack the benefits of the MPC multi-step ahead formulation. In this paper, we make the best of both worlds and propose a lightweight visual servoing MPC framework which generates optimal control near real-time at a frequency of 10.52 Hz. This work utilises the differential cross-entropy sampling method for quick and effective control generation along with a lightweight neural network, significantly improving the servoing frequency. We also propose a flow depth normalisation layer which ameliorates the issue of inferior predictions of two view depth from the flow network. We conduct extensive experimentation on the Habitat simulator and show a notable decrease in servoing time in comparison with other approaches that optimise over a time horizon. We achieve the right balance between time and performance for visual servoing in six degrees of freedom (6DoF), while retaining the advantageous MPC formulation. Our code and dataset are publicly available†.
M. Nomaan Qureshi, Pushkal Katara, Harit Pandya, Y. V. S. Harish, AadilMehdi J. Sanchawala, Gourav Kumar, Brojeshwar Bhowmick, K. Madhava Krishna
IROS8
2021 GCExp: Goal-Conditioned Exploration for Object Goal Navigation
abstract
In this paper, we address the highly challenging problem of object goal navigation. The agent, in an unseen environment, has to perceive its surroundings to identify and navigate towards potential regions where the specified goal category can occur. Rather than developing goal driven exploration policies, we aim to adapt the existing exploration policies that maximize scene coverage to be goal-conditioned. Thus, we propose a standalone scene understanding module to identify potential regions where the goal occurs. We also propose Goal-Conditioned Exploration (GCExp), an algorithm that entails the integration of our novel scene understanding module with any existing exploration policy. We test our solution in photo-realistic simulation environments using state-of-the-art exploration policy, Active Neural Slam [1], and show improved performance over the same on every evaluation metric.
Gulshan Kumar, Narasimhan Sai Shankar, Himansu Didwania, Ruddra Dev Roychoudhury, Brojeshwar Bhowmick, K. Madhava Krishna
RO-MAN5
2020 Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture
Dipanjan Das 0003, Sandika Biswas, Sanjana Sinha, Brojeshwar Bhowmick
ECCV (30)4
2020 Variational Clustering: Leveraging Variational Autoencoders for Image Clustering
abstract
Recent advances in deep learning have shown their ability to learn strong feature representations for images. The task of image clustering naturally requires good feature representations to capture the distribution of the data and subsequently differentiate data points from one another. Often these two aspects are dealt with independently and thus traditional feature learning alone does not suffice in partitioning the data meaningfully. Variational Autoencoders (VAEs) naturally lend themselves to learning data distributions in a latent space. Since we wish to efficiently discriminate between different clusters in the data, we propose a method based on VAEs where we use a Gaussian Mixture prior to help cluster the images accurately. We jointly learn the parameters of both the prior and the posterior distributions. Our method represents a true Gaussian Mixture VAE. This way, our method simultaneously learns a prior that captures the latent distribution of the images and a posterior to help discriminate well between data points. We also propose a novel reparametrization of the latent space consisting of a mixture of discrete and continuous variables. One key takeaway is that our method generalizes better across different datasets without using any pre-training or learnt models, unlike existing methods, allowing it to be trained from scratch in an end-to-end manner. We verify our efficacy and generalizability experimentally by achieving state-of-the-art results among unsupervised methods on a variety of datasets. To the best of our knowledge, we are the first to pursue image clustering using VAEs in a purely unsupervised manner on real image datasets.
Vignesh Prasad, Dipanjan Das 0003, Brojeshwar Bhowmick
IJCNN3
2020 Identity-Preserving Realistic Talking Face Generation
abstract
Speech-driven facial animation is useful for a variety of applications such as telepresence, chatbots, etc. The necessary attributes of having a realistic face animation are 1) audiovisual synchronization (2) identity preservation of the target individual (3) plausible mouth movements (4) presence of natural eye blinks. The existing methods mostly address the audiovisual lip synchronization, and few recent works have addressed synthesis of natural eye blinks for overall video realism. In this paper, we propose a method for identity-preserving realistic facial animation from speech. We first generate person-independent facial landmarks from audio using DeepSpeech features for invariance to different voices, accents, etc. To add realism, we impose eye blinks on facial landmarks using unsupervised learning and retarget the person-independent landmarks to person-specific landmarks to preserve the identity-related facial structure which helps in generation of plausible mouth shapes of the target identity. Finally, we use LSGAN to generate the facial texture from person-specific facial landmarks, using an attention mechanism that helps to preserve identity-related texture. An extensive comparison of our proposed method with the current state-of-the-art methods demonstrate a significant improvement in terms of lip synchronization accuracy, image reconstruction quality, sharpness, and identity-preservation. A user study also reveals improved realism of our animation results over the state-of-the-art methods. To the best of our knowledge, this is the first work in speech-driven 2D facial animation that simultaneously addresses all the above-mentioned attributes of a realistic speech driven face animation.
Sanjana Sinha, Sandika Biswas, Brojeshwar Bhowmick
IJCNN3
2020 Simple means Faster: Real-Time Human Motion Forecasting in Monocular First Person Videos on CPU
abstract
We present a simple, fast, and light-weight RNN based framework for forecasting future locations of humans in first person monocular videos. The primary motivation for this work was to design a network which could accurately predict future trajectories at a very high rate on a CPU. Typical applications of such a system would be a social robot or a visual assistance system "for all", as both cannot afford to have high compute power to avoid getting heavier, less power efficient, and costlier. In contrast to many previous methods which rely on multiple type of cues such as camera ego-motion or 2D pose of the human, we show that a carefully designed network model which relies solely on bounding boxes can not only perform better but also predicts trajectories at a very high rate while being quite low in size of approximately 17 MB. Specifically, we demonstrate that having an auto-encoder in the encoding phase of the past information and a regularizing layer in the end boosts the accuracy of predictions with negligible overhead. We experiment with three first person video datasets: CityWalks, FPL and JAAD. Our simple method trained on CityWalks surpasses the prediction accuracy of state-of-the-art method (STED) while being 9.6x faster on a CPU (STED runs on a GPU). We also demonstrate that our model can transfer zero-shot or after just 15% fine-tuning to other similar datasets and perform on par with the state-of-the-art methods on such datasets (FPL and DTP). To the best of our knowledge, we are the first to accurately forecast trajectories at a very high prediction rate of 78 trajectories per second on CPU.
Junaid Ahmed Ansari, Brojeshwar Bhowmick
IROS2
2019 Multi-modal Image Stitching with Nonlinear Optimization
abstract
Despite significant advances in recent years, the problem of image stitching still lacks a robust solution. Most of the feature based image stitching algorithms perform image alignment based on either homography-based transformation or content-preserving warping. Pairwise homography-based approach miserably fails to handle parallax whereas content-preserving warping approach does not preserve the structural property of the images. In this paper, we propose a nonlinear optimization to find out the global homographies using pairwise homography estimates and point correspondences. We further compute local warping based alignment to mitigate the aberration caused by noises in the global homography estimation. To this end, we incorporate geometric as well as photometric constraints to design our cost function which is minimized to obtain better alignment after the global registration, thus producing accurate image stitching. Experimental results on various open datasets demonstrate that our proposed method outperforms state-of-the-art image stitching algorithms.
Arindam Saha, Soumyadip Maity, Brojeshwar Bhowmick
ICASSP3
2019 Lifting 2d Human Pose to 3d : A Weakly Supervised Approach
abstract
Estimating 3d human pose from monocular images is a challenging problem due to the variety and complexity of human poses and the inherent ambiguity in recovering depth from the single view. Recent deep learning based methods show promising results by using supervised learning on 3d pose annotated datasets. However, the lack of large-scale 3d annotated training data captured under in-the-wild settings makes the 3d pose estimation difficult for in-the-wild poses. Few approaches have utilized training images from both 3d and 2d pose datasets in a weakly-supervised manner for learning 3d poses in unconstrained settings. In this paper, we propose a method which can effectively predict 3d human pose from 2d pose using a deep neural network trained in a weakly-supervised manner on a combination of ground-truth 3d pose and ground-truth 2d pose. Our method uses re-projection error minimization as a constraint to predict the 3d locations of body joints, and this is crucial for training on data where the 3d ground-truth is not present. Since minimizing re-projection error alone may not guarantee an accurate 3d pose, we also use additional geometric constraints on skeleton pose to regularize the pose in 3d. We demonstrate the superior generalization ability of our method by cross-dataset validation on a challenging 3d benchmark dataset MPI-INF-3DHP containing in the wild 3d poses.
Sandika Biswas, Sanjana Sinha, Kavya Gupta, Brojeshwar Bhowmick
IJCNN4
2019 Talk to the Vehicle: Language Conditioned Autonomous Navigation of Self Driving Cars
abstract
We propose a novel pipeline that blends encodings from natural language and 3D semantic maps obtained from visual imagery to generate local trajectories that are executed by a low-level controller. The pipeline precludes the need for a prior registered map through a local waypoint generator neural network. The waypoint generator network (WGN) maps semantics and natural language encodings (NLE) to local waypoints. A local planner then generates a trajectory from the ego location of the vehicle (an outdoor car in this case) to these locally generated waypoints while a low-level controller executes these plans faithfully. The efficacy of the pipeline is verified in the CARLA simulator environment as well as on local semantic maps built from real-world KITTI dataset. In both these environments (simulated and real-world) we show the ability of the WGN to generate waypoints accurately by mapping NLE of varying sequence lengths and levels of complexity. We compare with baseline approaches and show significant performance gain over them. And finally, we show real implementations on our electric car verifying that the pipeline lends itself to practical and tangible realizations in uncontrolled outdoor settings. In loop execution of the proposed pipeline that involves repetitive invocations of the network is critical for any such language-based navigation framework. This effort successfully accomplishes this thereby bypassing the need for prior metric maps or strategies for metric level localization during traversal.
Sriram N. N., Tirth Maniar, Jayaganesh Kalyanasundaram, Vineet Gandhi, Brojeshwar Bhowmick, K. Madhava Krishna
IROS5
2019 A Hierarchical Network for Diverse Trajectory Proposals
abstract
Autonomous explorative robots frequently encounter scenarios where multiple future trajectories can be pursued. Often these are cases with multiple paths around an obstacle or trajectory options towards various frontiers. Humans in such situations can inherently perceive and reason about the surrounding environment to identify several possibilities of either manoeuvring around the obstacles or moving towards various frontiers. In this work, we propose a 2 stage Convolutional Neural Network architecture which mimics such an ability to map the perceived surroundings to multiple trajectories that a robot can choose to traverse. The first stage is a Trajectory Proposal Network which suggests diverse regions in the environment which can be occupied in the future. The second stage is a Trajectory Sampling network which provides a finegrained trajectory over the regions proposed by Trajectory Proposal Network. We evaluate our framework in diverse and complicated real life settings. For the outdoor case, we use the KITTI dataset and our own outdoor driving dataset. In the indoor setting, we use an autonomous drone to navigate various scenarios and also a ground robot which can explore the environment using the trajectories proposed by our framework. Our experiments suggest that the framework is able to develop a semantic understanding of the obstacles, open regions and identify diverse trajectories that a robot can traverse. Our comparisons portray the performance gain of the proposed architecture over a diverse set of methods against which it is compared.
Sriram N. N., Gourav Kumar, Abhay Singh, M. Siva Karthik, Saket Saurav, Brojeshwar Bhowmick, K. Madhava Krishna
IV6
2019 Deep Representation Learning Characterized by Inter-Class Separation for Image Clustering
abstract
Despite significant advances in clustering methods in recent years, the outcome of clustering of a natural image dataset is still unsatisfactory due to two important drawbacks. Firstly, clustering of images needs a good feature representation of an image and secondly, we need a robust method which can discriminate these features for making them belonging to different clusters such that intra-class variance is less and inter-class variance is high. Often these two aspects are dealt with independently and thus the features are not sufficient enough to partition the data meaningfully. In this paper, we propose a method where we discover these features required for the separation of the images using deep autoencoder. Our method learns the image representation features automatically for the purpose of clustering and also select a coherent image and an incoherent image simultaneously for a given image so that the feature representation learning can learn better discriminative features for grouping the similar images in a cluster and at the same time separating the dissimilar images across clusters. Experiment results show that our method produces significantly better result than the state-of-the-art methods and we also show that our method is more generalized across different dataset without using any pre-trained model like other existing methods.
Dipanjan Das 0003, Ratul Ghosh, Brojeshwar Bhowmick
WACV3
2019 SfMLearner++: Learning Monocular Depth & Ego-Motion Using Meaningful Geometric Constraints
abstract
Most geometric approaches to monocular Visual Odometry (VO) provide robust pose estimates, but sparse or semi-dense depth estimates. Off late, deep methods have shown good performance in generating dense depths and VO from monocular images by optimizing the photometric consistency between images. Despite being intuitive, a naive photometric loss does not ensure proper pixel correspondences between two views, which is the key factor for accurate depth and relative pose estimations. It is a well known fact that simply minimizing such an error is prone to failures. We propose a method using Epipolar constraints to make the learning more geometrically sound. We use the Essential matrix, obtained using Nistér's Five Point Algorithm, for enforcing meaningful geometric constraints on the loss, rather than using it as labels for training. Our method, although simplistic but more geometrically meaningful, uses lesser number of parameters to give a comparable performance to state-of-the-art methods which use complex losses and large networks showing the effectiveness of using epipolar constraints. Such a geometrically constrained learning method performs successfully even in cases where simply minimizing the photometric error would fail.
Vignesh Prasad, Brojeshwar Bhowmick
WACV2
2018 Indoor Dense Depth Map at Drone Hovering
abstract
Autonomous Micro Aerial Vehicles (MAVs) gained tremendous attention in recent years. Autonomous flight in indoor requires a dense depth map for navigable space detection which is the fundamental component for autonomous navigation. In this paper, we address the problem of reconstructing dense depth while a drone is hovering (small camera motion) in indoor scenes using already estimated cameras and sparse point cloud obtained from a vSLAM. We start by segmenting the scene based on sudden depth variation using sparse 3D points and introduce a patch-based local plane fitting via energy minimization which combines photometric consistency and co-planarity with neighbouring patches. The method also combines a plane sweep technique for image segments having almost no sparse point for initialization. Experiments show, the proposed method produces better depth for indoor in artificial lighting condition, low-textured environment compared to earlier literature in small motion.
Arindam Saha, Soumyadip Maity, Brojeshwar Bhowmick
ICIP3
2018 Constructing Category-Specific Models for Monocular Object-SLAM
abstract
We present a new paradigm for real-time object-oriented SLAM with a monocular camera. Contrary to previous approaches, that rely on object-level models, we construct category-level models from CAD collections which are now widely available. To alleviate the need for huge amounts of labeled data, we develop a rendering pipeline that enables synthesis of large datasets from a limited amount of manually labeled data. Using data thus synthesized, we learn category-level models for object deformations in 3D, as well as discriminative object features in 2D. These category models are instance-independent and aid in the design of object landmark observations that can be incorporated into a generic monocular SLAM framework. Where typical object-SLAM approaches usually solve only for object and camera poses, we also estimate object shape on-the-fty, allowing for a wide range of objects from the category to be present in the scene. Moreover, since our 2D object features are learned discriminatively, the proposed object-SLAM system succeeds in several scenarios where sparse feature-based monocular SLAM fails due to insufficient features or parallax. Also, the proposed category-models help in object instance retrieval, useful for Augmented Reality (AR) applications. We evaluate the proposed framework on multiple challenging real-world scenes and show - to the best of our knowledge - first results of an instance-independent monocular object-SLAM system and the benefits it enjoys over feature-based SLAM methods.
Parv Parkhiya, Rishabh Khawad, Krishna Murthy Jatavallabhula, Brojeshwar Bhowmick, K. Madhava Krishna
ICRA4
2018 Coupled Analysis Dictionary Learning to inductively learn inversion: Application to real-time reconstruction of Biomedical signals
abstract
This work addresses the problem of reconstructing biomedical signals from their lower dimensional projections. Traditionally Compressed Sensing (CS) based techniques have been employed for this task. These are transductive inversion processes; the problem with these approaches is that the inversion is time-consuming and hence not suitable for real-time applications. With the recent advent of deep learning, Stacked Sparse Denoising Autoencoder (SSDAE) has been used for learning inversion in an inductive setup. The training period for inductive learning is large but is very fast during application - capable of real-time speed. This work proposes a new approach for inductive learning of the inversion process. It is based on Coupled Analysis Dictionary Learning. Results on Biomedical signal reconstruction show that our proposed approach is very fast and yields result far better than CS and SSDAE.
Kavya Gupta, Brojeshwar Bhowmick, Angshul Majumdar
IJCNN2
2018 Robust Adaptive Heart-Rate Monitoring Using Face Videos
abstract
Heart rate (HR) monitoring is indispensable for several real-world scenarios, especially when acquired in a non-contact manner. It can be accomplished using face videos acquired from ubiquitous cameras in an inexpensive, non-invasive and unobtrusive manner. But the HR monitoring can be erroneous when the video contains facial expressions, out-of-plane movements, change in camera parameters (like focus) and variations in environmental factors (like illumination). The proposed system mitigates these problems for improving the HR monitoring. For this, it defines an adaptive temporal signal selection mechanism which identifies and removes the facial areas affected by facial expressions. Moreover, it introduces a novel post-processing mechanism which perform HR monitoring by utilizing face reconstruction and quality. The post-processing is used when the face video contains facial movements. Experimental results reveal that incorporation of adaptive temporal signal selection and post-processing mechanisms can significantly improve the HR monitoring. It depicts that the Pearson correlation between actual and estimated HR is 0.95 while the average absolute error is 1.63 beats per minute, which indicates that the proposed system provides good HR monitoring.
Puneet Gupta 0002, Brojeshwar Bhowmick, Arpan Pal 0001
WACV2
2017 3D point cloud registration with shape constraint
abstract
In this paper, a shape-constrained iterative algorithm is proposed to register a rigid template point-cloud to a given reference point-cloud. The algorithm embeds a shape-based similarity constraint into the principle of gravitation. The shape-constrained gravitation, as induced by the reference, controls the movement of the template such that at each iteration, the template better aligns with the reference in terms of shape. This constraint enables the alignment in difficult conditions introduced by change (presence of outliers and/or missing parts), translation, rotation and scaling. We discuss efficient implementation techniques with least manual intervention. The registration is shown to be important for change detection in the 3D point-cloud. The algorithm is compared with three state-of-the-art registration approaches. The experiments are done on both synthetic and real-world data. The proposed algorithm is shown to perform better in the presence of big rotation, structured and unstructured outliers and missing data.
Swapna Agarwal, Brojeshwar Bhowmick
ICIP2
2017 Motion blur removal via coupled autoencoder
abstract
In this paper we propose a joint optimization technique for coupled autoencoder which learns the autoencoder weights and coupling map (between source and target) simultaneously. The technique is applicable to any transfer learning problem. In this work, we propose a new formulation that recasts deblurring as a transfer learning problem; it is solved using the proposed coupled autoencoder. The proposed technique can operate on-the-fly; since it does not require solving any costly inverse problem. Experiments have been carried out on state-of-the-art techniques; our method yields better quality images in shorter operating times.
Kavya Gupta, Brojeshwar Bhowmick, Angshul Majumdar
ICIP2
2017 Accurate heart-rate estimation from face videos using quality-based fusion
abstract
Estimating heart rate (HR) accurately using face videos acquired from a low cost camera in contactless manner is of paramount importance for many real-world applications. Such existing systems perform spuriously due to change in camera parameters, respiration, facial expressions and environmental factors. This paper mitigates the issues for accurate HR estimation. The face video consisting of frontal, profile or multiple faces is divided into multiple overlapping fragments to determine HR estimates. The HR estimates are fused using quality-based fusion which aims to minimize illumination and face deformations. Experimental results demonstrate that the proposed system exhibit better performance than the state of the art systems and establishes the efficacy of the quality-based fusion in HR estimation.
Puneet Gupta 0002, Brojeshwar Bhowmick, Arpan Pal 0001
ICIP2
2017 Divide and conquer: A hierarchical approach to large-scale structure-from-motion
Brojeshwar Bhowmick, Suvam Patra, Avishek Chatterjee, Venu Madhav Govindu, Subhashis Banerjee
Comput. Vis. Image Underst.1
2016 An efficient and robust method of virtual augmentation of eye-glass for easy shopping
abstract
We present a novel pipeline for augmenting a 3D eye-glass mesh into a person's face. While doing so, we take care about the proper fitment of the glass in terms of pupilary distance computed automatically, which is user-friendly in compare to standard marker based approaches. Our method also performs rigid eye-glass temple correction during augmentation followed by tracking to present realistic rendering to user. Moreover, our rendering system also allows user to change the material properties of the mesh so that eye-glass can be rendered in different color and shading. With these unique features, our system is robust as well as attractive for e-shopping. Experimental results show the efficacy of our system in terms of pupilary distance, realistic rendering with dynamic shading controllable by user.
Apurbaa Mallik, Brojeshwar Bhowmick
ICARCV2
2016 Quantification of balance in single limb stance using kinect
abstract
This paper presents a novel single limb body balance analysis system which will aid medical practitioners to analyze crucial factor for fall risk minimization, injury prevention, fitness and rehabilitation programs. We use skeleton data obtained from Microsoft Kinect which captures full human body as well as ensures user's privacy. A new eigen vector based curvature analysis algorithm is developed to compute single limb stance (SLS) duration on the skeleton data. Two parameters vibration-jitter and force per unit mass (FPUM) are derived for each body part to assess postural stability during SLS. Experimental results show the efficacy of our system to apply it in medical domain.
Kingshuk Chakravarty, Suraj Suman, Brojeshwar Bhowmick, Aniruddha Sinha, Abhijit Das 0003
ICASSP3
2014 Divide and Conquer: Efficient Large-Scale Structure from Motion Using Graph Partitioning
Brojeshwar Bhowmick, Suvam Patra, Avishek Chatterjee, Venu Madhav Govindu, Subhashis Banerjee
ACCV (2)1
2008 A multi-stage neural network aided system for detection of microcalcifications in digitized mammograms
Nikhil R. Pal, Brojeshwar Bhowmick, Sanjaya K. Patel, Srimanta Pal
Neurocomputing2