EDBT 2026 Demo / reviewers in the wild / expert
Siddharth Srivastava 0004
dblp:64/3431-4
· DBLP profile ↗
19ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise Aware Audio-Visual Speech DenoisingabstractWe address the problem of speech denoising where the goal is to extract clean speech signal from a noisy signal. Traditionally, the task of denoising has been performed using audio modality only. However, human speech perception is inherently multimodal where cues from visual modality are used to understand the speech better in a noisy environment. Similar observation has been made with computational denoising methods, i.e., performance of audio only model improves after adding visual modality. Inspired by these findings we propose a novel audio-visual network for adaptively combining both modalities for the task of speech audio denoising. We show that extracting noise from mixed audio and using it as a conditioning signal, improves speech denoising performance. To estimate the noise, we use both audio and visual modalities, i.e., lip region of the speaker, to extract the non-speech/silent regions from it. The silent regions enable us to estimate better noise profile to eliminate from the signal. Our proposed network uses self and cross attention framework between audio and video features, along the temporal dimension, to model correlations between the two modalities. We evaluate the proposed approach on a large scale audio-visual dataset VoxCeleb2 and obtain state-of-the-art results. We also demonstrate generalization to unseen speakers at test time. Kranti Kumar Parida, Siddharth Srivastava 0004, Gaurav Sharma 0004 |
IEEE Trans. Multim. | 2 |
| 2025 | Preserve Anything: Controllable Image Synthesis with Object PreservationabstractWe introduce \textit{Preserve Anything}, a novel method for controlled image synthesis that addresses key limitations in object preservation and semantic consistency in text-to-image (T2I) generation. Existing approaches often fail (i) to preserve multiple objects with fidelity, (ii) maintain semantic alignment with prompts, or (iii) provide explicit control over scene composition. To overcome these challenges, the proposed method employs an N-channel ControlNet that integrates (i) object preservation with size and placement agnosticism, color and detail retention, and artifact elimination, (ii) high-resolution, semantically consistent backgrounds with accurate shadows, lighting, and prompt adherence, and (iii) explicit user control over background layouts and lighting conditions. Key components of our framework include object preservation and background guidance modules, enforcing lighting consistency and a high-frequency overlay module to retain fine details while mitigating unwanted artifacts. We introduce a benchmark dataset consisting of 240K natural images filtered for aesthetic quality and 18K 3D-rendered synthetic images with metadata such as lighting, camera angles, and object relationships. This dataset addresses the deficiencies of existing benchmarks and allows a complete evaluation. Empirical results demonstrate that our method achieves state-of-the-art performance, significantly improving feature-space fidelity (FID 15.26) and semantic alignment (CLIP-S 32.85) while maintaining competitive aesthetic quality. We also conducted a user study to demonstrate the efficacy of the proposed work on unseen benchmark and observed a remarkable improvement of $\sim25\%$, $\sim19\%$, $\sim13\%$, and $\sim14\%$ in terms of prompt alignment, photorealism, the presence of AI artifacts, and natural aesthetics over existing works. Prasen Kumar Sharma, Neeraj Matiyali, Siddharth Srivastava 0004, Gaurav Sharma 0004 |
ICCV | 3 |
| 2024 | OmniVec2 - A Novel Transformer Based Network for Large Scale Multimodal and Multitask LearningabstractWe present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image, video, audio, text, depth, point cloud, time series, tabular, graph, X-ray, infrared, IMU, and hyper-spectral. The proposed approach utilizes modality specialized tokenizers, a shared transformer architecture, and cross-attention mechanisms to project the data from different modalities into a unified embedding space. It addresses multimodal and multitask scenarios by incorpo-rating modality-specific task heads for different tasks in respective modalities. We propose a novel pretraining strategy with iterative modality switching to initialize the network, and a training algorithm which trades off fully joint training over all modalities, with training on pairs of modalities at a time. We provide comprehensive evaluation across 25 datasets from 12 modalities and show state of the art performances, demonstrating the effectiveness of the proposed architecture, pretraining strategy and adapted multitask training. Siddharth Srivastava 0004, Gaurav Sharma 0004 |
CVPR | 1 |
| 2024 | ICPR 2024 Leaf Inspect Competition: Leaf Instance Segmentation and Counting
Swati Bhugra, Prerana Mukherjee, Vinay Kaushik, Siddharth Srivastava 0004, Viswanathan Chinnusamy, Brejesh Lall, Santanu Chaudhary |
ICPR (34) | 4 |
| 2024 | Talk2BEV: Language-enhanced Bird's-eye View Maps for Autonomous DrivingabstractThis work introduces Talk2BEV, a large vision-language model (LVLM)1interface for bird’s-eye view (BEV) maps commonly used in autonomous driving. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set of object categories and driving scenarios, Talk2BEV eliminates the need for BEV-specific training, relying instead on well-performing pre-trained LVLMs. This enables a single system to cater to a variety of autonomous driving tasks encompassing visual and spatial reasoning, predicting the intents of traffic actors, and decision-making based on visual cues. We extensively evaluate Talk2BEV on a large number of scene understanding tasks that rely on both the ability to interpret freeform natural language queries, and in grounding these queries to the visual context embedded into the language-enhanced BEV map. To enable further research in LVLMs for autonomous driving scenarios, we develop and release Talk2BEV-Bench, a benchmark encompassing 1000 human-annotated BEV scenarios, with more than 20,000 questions and ground-truth responses from the NuScenes dataset. We encourage the reader to view the demos on our project page: https://llmbev.github.io/talk2bev/ Tushar Choudhary, Vikrant Dewangan, Shivam Chandhok, Shubham Priyadarshan, Anushka Jain, Arun Kumar Singh 0001, Siddharth Srivastava 0004, Krishna Murthy Jatavallabhula, K. Madhava Krishna |
ICRA | 7 |
| 2024 | OmniVec: Learning robust representations with cross modal sharingabstractMajority of research in learning based methods has been towards designing and training networks for specific tasks. However, many of the learning based tasks, across modalities, share commonalities and could be potentially tackled in a joint framework. We present an approach in such direction, to learn multiple tasks, in multiple modalities, with a unified architecture. The proposed network is composed of task specific encoders, a common trunk in the middle, followed by task specific prediction heads. We first pre-train it by self-supervised masked training, followed by sequential training for the different tasks. We train the network on all major modalities, e.g. visual, audio, text and 3D, and report results on 22 diverse and challenging public benchmarks. We demonstrate empirically that, using a joint network to train across modalities leads to meaningful information sharing and this allows us to achieve state-of-the-art results on most of the benchmarks. We also show generalization of the trained network on cross-modal tasks as well as unseen datasets and tasks. Siddharth Srivastava 0004, Gaurav Sharma 0004 |
WACV | 1 |
| 2023 | Hierarchical Multi-task Learning via Task Affinity GroupingsabstractMulti-task learning (MTL) permits joint task learning based on a shared deep learning architecture and multiple loss functions. Despite the recent advances in MTL, one loss often dominates the learning optimization in multiple unrelated tasks. This often results in poor performance compared to the corresponding single task learning. To overcome the aforementioned "negative transfer", we propose a novel hierarchical framework that leverages task relations via inter-task affinity to supervise multi-task learning. Specifically, the inter-task affinity generated task sets, with low-level task set and complex task set at the bottom and top layers respectively, enables iterative multi-task information sharing. In addition, it also alleviates simultaneous image annotations for multiple tasks. The proposed framework achieves state-of-the-art results on classification, detection, semantic segmentation and depth estimation across three standard benchmarks. Furthermore, with state of the results on two benchmarks for image retrieval task, we also demonstrate that the embeddings learned using such a framework provide good generalization and robust representation learning. Siddharth Srivastava 0004, Swati Bhugra, Vinay Kaushik, Brejesh Lall |
ICIP | 1 |
| 2022 | Beyond Mono to Binaural: Generating Binaural Audio from Mono Audio with Depth and Cross Modal AttentionabstractBinaural audio gives the listener an immersive experience and can enhance augmented and virtual reality. However, recording binaural audio requires specialized setup with a dummy human head having microphones in left and right ears. Such a recording setup is difficult to build and setup, therefore mono audio has become the preferred choice in common devices. To obtain the same impact as binaural audio, recent efforts have been directed towards lifting mono audio to binaural audio conditioned on the visual input from the scene. Such approaches have not used an important cue for the task: the distance of different sound producing objects from the microphones. In this work, we argue that depth map of the scene can act as a proxy for inducing distance information of different objects in the scene, for the task of audio binauralization. We propose a novel encoder-decoder architecture with a hierarchical attention mechanism to encode image, depth and audio feature jointly. We design the network on top of state-of-the-art transformer networks for image and depth representation. We show empirically that the proposed method outperforms state-of-the-art methods comfortably for two challenging public datasets FAIR-Play and MUSIC-Sereo. We also demonstrate with qualitative results that the method is able to focus on the right information required for the task. The qualitative results are available at our project page https://krantiparida.github.io/projects/bmonobinaural.html Kranti Kumar Parida, Siddharth Srivastava 0004, Gaurav Sharma 0004 |
WACV | 2 |
| 2021 | Beyond Image to Depth: Improving Depth Prediction Using EchoesabstractWe address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth estimation. We propose an end-to-end deep learning based pipeline utilizing RGB images, binaural echoes and estimated material properties of various objects within a scene. We argue that the relation between image, echoes and depth, for different scene elements, is greatly influenced by the properties of those elements, and a method designed to leverage this information can lead to significantly improved depth estimation from audio visual inputs. We propose a novel multi modal fusion technique, which incorporates the material properties explicitly while combining audio (echoes) and visual modalities to predict the scene depth. We show empirically, with experiments on Replica dataset, that the proposed method obtains 28% improvement in RMSE compared to the state-of-the-art audio-visual depth prediction method. To demonstrate the effectiveness of our method on larger dataset, we report competitive performance on Matterport3D, proposing to use it as a multimodal depth prediction benchmark with echoes for the first time. We also analyse the proposed method with exhaustive ablation experiments and qualitative results. The code and models are available at https://krantiparida.github.io/projects/bimgdepth.html Kranti Kumar Parida, Siddharth Srivastava 0004, Gaurav Sharma 0004 |
CVPR | 2 |
| 2021 | Exploiting Local Geometry for Feature and Graph Construction for Better 3D Point Cloud Processing with Graph Neural NetworksabstractWe propose simple yet effective improvements in point representations and local neighborhood graph construction within the general framework of graph neural networks (GNNs) for 3D point cloud processing. As a first contribution, we propose to augment the vertex representations with important local geometric information of the points, followed by nonlinear projection using a MLP. As a second contribution, we propose to improve the graph construction for GNNs for 3D point clouds. The existing methods work with a k-NN based approach for constructing the local neighborhood graph. We argue that it might lead to reduction in coverage in case of dense sampling by sensors in some regions of the scene. The proposed methods aims to counter such problems and improve coverage in such cases. As the traditional GNNs were designed to work with general graphs, where vertices may have no geometric interpretations, we see both our proposals as augmenting the general graphs to incorporate the geometric nature of 3D point clouds. While being simple, we demonstrate with multiple challenging benchmarks, with relatively clean CAD models, as well as with real world noisy scans, that the proposed method achieves state of the art results on benchmarks for 3D classification (ModelNet40) , part segmentation (ShapeNet) and semantic segmentation (Stanford 3D Indoor Scenes Dataset). We also show that the proposed network achieves faster training convergence, i.e. ∼ 40% less epochs for classification. The project details are available at https://siddharthsrivastava.github.io/publication/geomgcnn/ Siddharth Srivastava 0004, Gaurav Sharma 0004 |
ICRA | 1 |
| 2021 | Self Attention Guided Depth Completion using RGB and Sparse LiDAR Point CloudsabstractWe address the problem of completing per pixel dense depth map using a single RGB image and the sparse point cloud of the scene. Depth prediction from RGB image is a hard problem and while dense point clouds obtained from LiDAR sensors can be used in addition to RGB image, the cost of such sensors is a significant barrier. Having LiDAR sensors which capture sparse point clouds is a reasonable middle ground. We propose a novel architecture which incorporates geometric primitives and self attention mechanisms, to improve the prediction. The motivation of self attention is to capture the correlations between scene and object elements, e.g. between the right and left window of car, early on in the network. While that for using geometric primitives is to have a high level clustering cue to enable the network to exploit similar correlations. In addition, we enforce complimentarity in the predictions made with RGB and sparse LiDAR respectively, this forces the two corresponding branches to focus on hard areas which are not already well predicted by the other branch. With exhaustive experiments on KITTI depth completion benchmark, NYU v2 and Matterport3D we show that the proposed method provides state-of-the-art results. Siddharth Srivastava 0004, Gaurav Sharma 0004 |
IROS | 1 |
| 2020 | An Online Learning Approach for Dengue Fever ClassificationabstractThis paper introduces a novel approach for dengue fever classification based on online learning paradigms. The proposed approach is suitable for practical implementation as it enables learning using only a few training samples. With time, the proposed approach is capable of learning incrementally from the data collected without need for retraining the model or redeployment of the prediction engine. Additionally, we also provide a comprehensive evaluation of machine learning methods for prediction of dengue fever. The input to the proposed pipeline comprises of recorded patient symptoms and diagnostic investigations. Offline classifier models have been employed to obtain baseline scores to establish that the feature set is optimal for classification of dengue. The primary benefit of the online detection model presented in the paper is that it has been established to effectively identify patients with high likelihood of dengue disease, and experiments on scalability in terms of number of training and test samples validate the use of the proposed model. Siddharth Srivastava 0004, Sumit Soman, Astha Rai, Amarjeet Singh Cheema |
CBMS | 1 |
| 2020 | Using Scene Graphs for Detecting Visual RelationshipsabstractIn this paper we solve the problem of detecting relationships between pairs of objects in an image. We develop spatially aware word embeddings using scene graphs and use joint feature representations containing visual, spatial and semantic embeddings from the input images to train a deep network on the task of relationship detection. Further, we propose to utilize context aligned scene graph embeddings from the train set, without requiring explicit availability of scene graphs at test time. We show that the proposed method outperforms the state-of-the-art methods for predicate detection and provides competing results on relationship detection. We also show the generalization ability of the proposed method by performing predictions under zero shot settings. Further, we also provide an exhaustive empirical evaluation on each component of the proposed network. Anurag Tripathi, Siddharth Srivastava 0004, Brejesh Lall, Santanu Chaudhury |
ICPR | 2 |
| 2020 | Analysing Risk of Coronary Heart Disease through Discriminative Neural NetworksabstractThe application of data mining, machine learning and artificial intelligence techniques in the field of diagnostics is not a new concept, and these techniques have been very successfully applied in a variety of applications, especially in dermatology and cancer research. But, in the case of medical problems that involve tests resulting in true or false (binary classification), the data generally has a class imbalance with samples majorly belonging to one class (ex: a patient undergoes a regular test and the results are false). Such disparity in data causes problems when trying to model predictive systems on the data. In critical applications like diagnostics, this class imbalance cannot be overlooked and must be given extra attention. In our research, we depict how we can handle this class imbalance through neural networks using a discriminative model and contrastive loss using a Siamese neural network structure. Such a model does not work on a probability-based approach to classify samples into labels. Instead it uses a distance-based approach to differentiate between samples classified under different labels. The code is available at this https URL Ayush Khaneja, Siddharth Srivastava 0004, Astha Rai, Amarjeet Singh Cheema, Praveen K. Srivastava |
ICPRAM | 2 |
| 2019 | Learning 2D to 3D Lifting for Object Detection in 3D for Autonomous VehiclesabstractWe address the problem of 3D object detection from 2D monocular images in autonomous driving scenarios. We propose to lift the 2D images to 3D representations using learned neural networks and leverage existing networks working directly on 3D data to perform 3D object detection and localization. We show that, with carefully designed training mechanism and automatically selected minimally noisy data, such a method is not only feasible, but gives higher results than many methods working on actual 3D inputs acquired from physical sensors. On the challenging KITTI benchmark, we show that our 2D to 3D lifted method outperforms many recent competitive 3D networks while significantly outperforming previous state-of-the-art for 3D detection from monocular images. We also show that a late fusion of the output of the network trained on generated 3D images, with that trained on real 3D images, improves performance. We find the results very interesting and argue that such a method could serve as a highly reliable backup in case of malfunction of expensive 3D sensors, if not potentially making them redundant, at least in the case of low human injury risk autonomous navigation scenarios like warehouse automation. Siddharth Srivastava 0004, Frédéric Jurie, Gaurav Sharma 0004 |
IROS | 1 |
| 2019 | Guided Compositional Generative Adversarial NetworksabstractIn this paper, we propose to synthesize natural images from a set of input objects. The proposed technique generates a scene which has high correlation with the provided set of input objects while also maintaining the natural placement of objects within the scene. The technique constitutes of a generative adversarial network trained on a large corpus of objects and natural scenes. This is in contrast with earlier works where the objective was to generate a natural scene from a noise vector or conditioning the network over a variable. However, such methods have limitations in their ability to control the objects within the generated images. On the contrary, we show that by training a Generative Adversarial Network with raw image pixels as input, we can generate scenes which constitute the objects as well as generate the surrounding environment suitable for the combination of the input objects. We provide qualitative and quantitative results on challenging MS-COCO dataset to show the effectiveness of the proposed technique. Anurag Tripathi, Siddharth Srivastava 0004, Brejesh Lall, Santanu Chaudhury |
SMC | 2 |
| 2019 | DeepPoint3D: Learning discriminative local descriptors using deep metric learning on 3D point clouds
Siddharth Srivastava 0004, Brejesh Lall |
Pattern Recognit. Lett. | 1 |
| 2018 | Joint Image Classification and Annotation Prediction Using Iterative Learning on Local NeighbourhoodabstractImage annotation (tag) and classification play a critical role in many computer vision applications, such as image retrieval, scene understanding, scene description etc. While, databases such as ImageNet have high quality labels for images, in real world, a large number of images have missing labels or tags that completely describe the contents of an image. To solve this problem, in this paper, we work on the hypothesis that class and tag information are correlated and propose a joint optimization for image classification and annotation. We construct a unified cost function to learn the class scoring vectors as well as tag scoring vectors. The proposed approach achieves state-of-the-art results on benchmark datasets for joint tag prediction and classification. Anurag Tripathi, Siddharth Srivastava 0004, Santanu Chaudhury, Brejesh Lall |
SMC | 2 |
| 2018 | Large Scale Novel Object Discovery in 3DabstractWe present a method for discovering never-seen-before objects in 3D point clouds obtained from sensors like Microsoft Kinect. We generate supervoxels directly from the point cloud data and use them with a Siamese network, built on a recently proposed 3D convolutional neural network architecture. We use known objects to train a non-linear embedding of supervoxels, by optimizing the criteria that supervoxels which fall on the same object should be closer than those which fall on different objects, in the embedding space. We test on unknown objects, which were not seen during training, and perform clustering in the learned embedding space of supervoxels to effectively perform novel object discovery. We validate the method with extensive experiments, quantitatively showing that it can discover numerous unseen objects while being trained on only a few dense 3D models. We also show very good qualitative results of object discovery in point cloud data when the test objects, either specific instances or even categories, were never seen during training. Siddharth Srivastava 0004, Gaurav Sharma 0004, Brejesh Lall |
WACV | 1 |