Sandika Biswas

dblp:165/0874 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-8264-3346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 54% Generative modeling · 18% Speech recognition and synthesis · 18%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 66% Geometric modeling and processing · 34%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
speech-driven animation
1.022022
Emotion-Controllable Generalized Talking Face Generation · IJCAI 2022
Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture · ECCV (30) 2020
Computer vision › 3D vision
3d reconstruction
0.812024
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene · NeurIPS 2024
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.812024
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene · NeurIPS 2024
Computer vision › 3D vision
neural radiance field
0.812024
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene · NeurIPS 2024
Computer vision › 3D vision › 3d reconstruction › object reconstruction
template-free reconstruction
0.812024
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene · NeurIPS 2024
Computer vision › Face, body and person analysis
facial animation
0.612022
Emotion-Controllable Generalized Talking Face Generation · IJCAI 2022
Machine learning › Generative modeling › face synthesis
talking face generation
0.612022
Emotion-Controllable Generalized Talking Face Generation · IJCAI 2022
Machine learning › Generative modeling
generative adversarial network
0.412020
Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture · ECCV (30) 2020
Computer animation and physical simulation
facial animation
0.412020
Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture · ECCV (30) 2020
Geometric modeling and processing › deformable models
deformable object modeling
0.212024
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

linear blend skinning · 1.5invertible neural network · 1.5generative adversarial network · 0.9cascaded GANs · 0.9optical flow · 0.6graph convolutional neural network · 0.6
YearPublicationVenuePosition
2024 Shape-prior Free Space-time Neural Radiance Field for 4D Semantic Reconstruction of Dynamic Scene from Sparse-View RGB Videos
abstract
Many applications in Augmented/Virtual Reality or robotics require precise geometry modeling of individual elements in a dynamic scene under a sparse-view camera setup, without any prior information about their semantic labels or shapes. In our research, we introduce a 3D shape prior-free Neural Radiance Field-based technique for detailed geometry reconstruction under human-object interactions, offering an explicit surface reconstruction with semantic labels for reconstructed geometry. Our approach harnesses the capabilities of an Invertible Neural Network to learn a deformation function that effectively connects local (current-frame input) and canonical spaces for each of the components under motion. The deformation process is guided by temporal constraints from multi-frame, facilitating the precise reconstruction of the complex interactions between humans and objects. Our experimental evaluations highlight the effectiveness of our framework, demonstrating its ability to accurately represent both the comprehensive object-compositional scene and individual components over state-of-the-art methods, under complex interactions between the scene entities. This research is deemed to mark a significant stride in semantic 3D geometry modeling within dynamic interactive environments, relying solely on sparse multi-view RGB data.
Sandika Biswas, Biplab Banerjee, Seyed Hamid Rezatofighi
IROS1
2024 TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene
abstract
Despite advancements in Neural Implicit models for 3D surface reconstruction, handling dynamic environments with interactions between arbitrary rigid, non-rigid, or deformable entities remains challenging. The generic reconstruction methods adaptable to such dynamic scenes often require additional inputs like depth or optical flow or rely on pre-trained image features for reasonable outcomes. These methods typically use latent codes to capture frame-by-frame deformations. Another set of dynamic scene reconstruction methods, are entity-specific, mostly focusing on humans, and relies on template models. In contrast, some template-free methods bypass these requirements and adopt traditional LBS (Linear Blend Skinning) weights for a detailed representation of deformable object motions, although they involve complex optimizations leading to lengthy training times. To this end, as a remedy, this paper introduces TFS-NeRF, a template-free 3D semantic NeRF for dynamic scenes captured from sparse or single-view RGB videos, featuring interactions among two entities and more time-efficient than other LBS-based approaches. Our framework uses an Invertible Neural Network (INN) for LBS prediction, simplifying the training process. By disentangling the motions of interacting entities and optimizing per-entity skinning weights, our method efficiently generates accurate, semantically separable geometries. Extensive experiments demonstrate that our approach produces high-quality reconstructions of both deformable and non-deformable objects in complex interactions, with improved training efficiency compared to existing methods. The code and models will be available on our github page.
Sandika Biswas, Qianyi Wu, Biplab Banerjee, Seyed Hamid Rezatofighi
NeurIPS1
2022 Emotion-Controllable Generalized Talking Face Generation
abstract
Despite the significant progress in recent years, very few of the AI-based talking face generation methods attempt to render natural emotions. Moreover, the scope of the methods is majorly limited to the characteristics of the training dataset, hence they fail to generalize to arbitrary unseen faces. In this paper, we propose a one-shot facial geometry-aware emotional talking face generation method that can generalize to arbitrary faces. We propose a graph convolutional neural network that uses speech content feature, along with an independent emotion input to generate emotion and speech-induced motion on facial geometry-aware landmark representation. This representation is further used in our optical flow-guided texture generation network for producing the texture. We propose a two-branch texture generation network, with motion and texture branches designed to consider the motion and texture content independently. Compared to the previous emotion talking face methods, our method can adapt to arbitrary faces captured in-the-wild by fine-tuning with only a single image of the target identity in neutral emotion.
Sanjana Sinha, Sandika Biswas, Ravindra Yadav, Brojeshwar Bhowmick
IJCAI2
2020 Speech-Driven Facial Animation Using Cascaded GANs for Learning of Motion and Texture
Dipanjan Das 0003, Sandika Biswas, Sanjana Sinha, Brojeshwar Bhowmick
ECCV (30)2
2020 Identity-Preserving Realistic Talking Face Generation
abstract
Speech-driven facial animation is useful for a variety of applications such as telepresence, chatbots, etc. The necessary attributes of having a realistic face animation are 1) audiovisual synchronization (2) identity preservation of the target individual (3) plausible mouth movements (4) presence of natural eye blinks. The existing methods mostly address the audiovisual lip synchronization, and few recent works have addressed synthesis of natural eye blinks for overall video realism. In this paper, we propose a method for identity-preserving realistic facial animation from speech. We first generate person-independent facial landmarks from audio using DeepSpeech features for invariance to different voices, accents, etc. To add realism, we impose eye blinks on facial landmarks using unsupervised learning and retarget the person-independent landmarks to person-specific landmarks to preserve the identity-related facial structure which helps in generation of plausible mouth shapes of the target identity. Finally, we use LSGAN to generate the facial texture from person-specific facial landmarks, using an attention mechanism that helps to preserve identity-related texture. An extensive comparison of our proposed method with the current state-of-the-art methods demonstrate a significant improvement in terms of lip synchronization accuracy, image reconstruction quality, sharpness, and identity-preservation. A user study also reveals improved realism of our animation results over the state-of-the-art methods. To the best of our knowledge, this is the first work in speech-driven 2D facial animation that simultaneously addresses all the above-mentioned attributes of a realistic speech driven face animation.
Sanjana Sinha, Sandika Biswas, Brojeshwar Bhowmick
IJCNN2
2019 Lifting 2d Human Pose to 3d : A Weakly Supervised Approach
abstract
Estimating 3d human pose from monocular images is a challenging problem due to the variety and complexity of human poses and the inherent ambiguity in recovering depth from the single view. Recent deep learning based methods show promising results by using supervised learning on 3d pose annotated datasets. However, the lack of large-scale 3d annotated training data captured under in-the-wild settings makes the 3d pose estimation difficult for in-the-wild poses. Few approaches have utilized training images from both 3d and 2d pose datasets in a weakly-supervised manner for learning 3d poses in unconstrained settings. In this paper, we propose a method which can effectively predict 3d human pose from 2d pose using a deep neural network trained in a weakly-supervised manner on a combination of ground-truth 3d pose and ground-truth 2d pose. Our method uses re-projection error minimization as a constraint to predict the 3d locations of body joints, and this is crucial for training on data where the 3d ground-truth is not present. Since minimizing re-projection error alone may not guarantee an accurate 3d pose, we also use additional geometric constraints on skeleton pose to regularize the pose in 3d. We demonstrate the superior generalization ability of our method by cross-dataset validation on a challenging 3d benchmark dataset MPI-INF-3DHP containing in the wild 3d poses.
Sandika Biswas, Sanjana Sinha, Kavya Gupta, Brojeshwar Bhowmick
IJCNN1