Shivam Pande

dblp:249/1920 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
12since 2021 · last 2024
0000-0003-2020-9306ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Plant Detection from Ultra High Resolution Remote Sensing Images: A Semantic Segmentation Approach Based on Fuzzy Loss
abstract
In this study, we tackle the challenge of identifying plant species from ultra high resolution (UHR) remote sensing images. Our approach involves introducing an RGB remote sensing dataset, characterized by millimeter-level spatial resolution, meticulously curated through several field expeditions across a mountainous region in France covering various landscapes. The task of plant species identification is framed as a semantic segmentation problem for its practical and efficient implementation across vast geographical areas. However, when dealing with segmentation masks, we confront instances where distinguishing boundaries between plant species and their background is challenging. We tackle this issue by introducing a fuzzy loss within the segmentation model. Instead of utilizing one-hot encoded ground truth (GT), our model incorporates Gaussian filter refined GT, introducing stochasticity during training. First experimental results obtained on both our UHR dataset and a public dataset are presented, showing the relevance of the proposed methodology, as well as the need for future improvement.
Shivam Pande, Baki Uzun, Florent Guiotte, Minh-Tan Pham, Thomas Corpetti, Florian Delerue, Sébastien Lefèvre
IGARSS1
2024 Mapping Earth Mounds from Space
abstract
Regular patterns of vegetation are considered widespread landscapes, although their global extent has never been estimated. Among them, spotted landscapes are of particular interest in the context of climate change. Indeed, regularly spaced vegetation spots in semi-arid shrublands result from extreme resource depletion and prefigure catastrophic shift of the ecosystem to a homogeneous desert, while termite mounds also producing spotted landscapes were shown to increase robustness to climate change. Yet, their identification at large scale calls for automatic methods, for instance using the popular deep learning framework, able to cope with a vast amount of remote sensing data, e.g., optical satellite imagery. In this paper, we tackle this problem and benchmark some state-of-the-art deep networks on several landscapes and geographical areas. Despite the promising results we obtained, we found that more research is needed to be able to map automatically these earth mounds from space.
Baki Uzun, Shivam Pande, Gwendal Cachin-Bernard, Minh-Tan Pham, Sébastien Lefèvre, Rumsais Blatrix, Doyle McKey
IGARSS2
2024 Domain Adaptive 3D Shape Retrieval from Monocular Images
abstract
In this work, we address the novel and challenging problem of domain adaptive 3D shape retrieval from single 2D images (DA-IBSR). While the existing image-based 3D shape retrieval (IBSR) problem focuses on modality alignment for retrieving a matchable 3D shape from a shape repository given a 2D image query, it does not consider any distribution shift between the training and testing image-shape pairs, making the performance of off-the-shelves IBSR methods subpar. In contrast, the proposed DA-IBSR addresses the non-trivial problem of modality shift as well distribution shift across training and test sets. To address these issues, we propose an end-to-end trainable model called DAIS-NET. Our objective is to align the images and shapes separately from both domains while simultaneously learn a shared embedding space for the 2D and 3D modalities. The former problem is addressed by separately employing maximum mean discrepancy loss across the 2D images and 3D shapes of the two domains. To address the modality alignment, we incorporate the notion of negative sample mining and employ triplet loss to bridge the gap between positive 2D-3D pairs (of same class) and increase the separation between negative 2D-3D pairs (of different class). Additionally, we employ an entropy minimization strategy to align the unlabeled target domain data in the semantic space. To evaluate our proposed approach, we define the experimental setting of DA-IBSR on the following benchmarks: SHREC’14 ↔ Pix3D and ShapeNet ↔ SHREC’14. Considering the novelty of the problem statement, we have demonstrated that the issue of domain gap is prevalent by comparing our method with the existing literature. Additionally, through extensive evaluations, we demonstrate the capability of DAIS-NET to successfully mitigate this domain gap in image based 3D shape retrieval.
Harsh Pal, Ritwik Khandelwal, Shivam Pande, Biplab Banerjee, Srikrishna Karanam
WACV3
2023 Semi-Supervised Learning for Hyperspectral Images by Non Parametrically Predicting View AssignmentCRediT
abstract
Hyperspectral image (HSI) classification is gaining a lot of momentum in present time because of high inherent spectral information within the images. However, these images suffer from the problem of curse of dimensionality and usually require a large number samples for tasks such as classification, especially in supervised setting. Recently, to effectively train the deep learning models with minimal labelled samples, the unlabeled samples are also being leveraged in self-supervised and semi-supervised setting. In this work, we leverage the idea of semi-supervised learning to assist the discriminative self-supervised pretraining of the models. The proposed method takes different augmented views of the unlabeled samples as input and assigns them the same pseudo-label corresponding to the labelled sample from the downstream task. We train our model on two HSI datasets, anemly Houston dataset (from data fusion contest, 2013) and Pavia university dataset, and show that the proposed approach performs better than self-supervised approach and supervised training.
Shivam Pande, Nassim Ait Ali Braham, Yi Wang 0072, Conrad M. Albrecht, Biplab Banerjee, Xiao Xiang Zhu 0001
IGARSS1
2023 Visual Question Answering in Remote Sensing with Cross-Attention and Multimodal Information Bottleneck
abstract
In this research, we deal with the problem of visual question answering (VQA) in remote sensing. While remotely sensed images contain information significant for the task of identification and object detection, they pose a great challenge in their processing because of high dimensionality, volume and redundancy. Furthermore, processing image information jointly with language features adds additional constraints, such as mapping the corresponding image and language features. To handle this problem, we propose a cross attention based approach combined with information maximization. The CNN-LSTM based cross-attention highlights the information in the image and language modalities and establishes a connection between the two, while information maximization learns a low dimensional bottleneck layer, that has all the relevant information required to carry out the VQA task. We evaluate our method on two VQA remote sensing datasets of different resolutions. For the high resolution dataset, we achieve an overall accuracy of 79.11% and 73.87% for the two test sets while for the low resolution dataset, we achieve an overall accuracy of 85.98%.
Jayesh Songara, Shivam Pande, Shabnam Choudhury, Biplab Banerjee, Rajbabu Velmurugan
IGARSS2
2023 Self-supervision assisted multimodal remote sensing image classification with coupled self-looping convolution networks
Shivam Pande, Biplab Banerjee
Neural Networks1
2022 RSINet: Inpainting Remotely Sensed Images Using Triple GAN Framework
abstract
We tackle the problem of image inpainting in the remote sensing domain. Remote sensing images possess high resolution and geographical variations, that render the conventional inpainting methods less effective. This further entails the requirement of models with high complexity to sufficiently capture the spectral, spatial and textural nuances within an image, emerging from its high spatial variability. To this end, we propose a novel inpainting method that individually focuses on each aspect of an image such as edges, colour and texture using a task specific GAN. Moreover, each individual GAN also incorporates the attention mechanism that explicitly extracts the spectral and spatial features. To ensure consistent gradient flow, the model uses residual learning paradigm, thus simultaneously working with high and low level features. We evaluate our model, alongwith previous state of the art models, on the two well known remote sensing datasets, Open Cities AI and Earth on Canvas, and achieve competitive performance. The code can be referred here: https://github.com/advaitkumar3107/RSINet.
Advait Kumar, Dipesh Tamboli, Shivam Pande, Biplab Banerjee
IGARSS3
2022 Feedback Convolution Based Autoencoder for Dimensionality Reduction in Hyperspectral Images
abstract
Hyperspectral images (HSI) possess a very high spectral res-olution (due to innumerous bands), which makes them invalu-able in the remote sensing community for landuse/land cover classification. However, the multitude of bands forces the algorithms to consume more data for better performance. To tackle this, techniques from deep learning are often explored, most prominently convolutional neural networks (CNN) based autoencoders. However, one of the main limitations of conventional CNNs is that they only have forward connections. This prevents them to generate robust representations since the information from later layers is not used to refine the earlier layers. Therefore, we introduce a 1D-convolutional autoencoder based on feedback connections for hyperspec-tral dimensionality reduction. Feedback connections create self-updating loops within the network, which enable it to use future information to refine past layers. Hence, the low dimensional code has more refined information for efficient classification. The performance of our method is evaluated on Indian pines 2010 and Indian pines 1992 HSI datasets, where it surpasses the existing approaches.
Shivam Pande, Biplab Banerjee
IGARSS1
2021 Two Headed Dragons: Multimodal Fusion And Cross Modal Transactions
abstract
As the field of remote sensing is evolving, we witness the accumulation of information from several modalities, such as multispectral (MS), hyperspectral (HSI), LiDAR etc. Each of these modalities possess its own distinct characteristics and when combined synergistically, perform very well in the recognition and classification tasks. However, fusing multiple modalities in remote sensing is cumbersome due to highly disparate domains. Furthermore, the existing methods do not facilitate cross-modal interactions. To this end, we propose a novel transformer based fusion method for HSI and LiDAR modalities. The model is composed of stacked auto encoders that harness the cross key-value pairs for HSI and LiDAR, thus establishing a communication between the two modalities, while simultaneously using the CNNs to extract the spectral and spatial information from HSI and LiDAR. We test our model on Houston (Data Fusion Contest – 2013) and MUUFL Gulfport datasets and achieve competitive results.
Rupak Bose, Shivam Pande, Biplab Banerjee
ICIP2
2021 Attention Based Convolution Autoencoder for Dimensionality Reduction in Hyperspectral Images
abstract
Hyperspectral images (HSIs) are being actively used for land use/land cover classification owing to their high spectral resolution. However, this leads to the problem of high dimensionality, making the algorithms data hungry. To resolve these issues, deep learning techniques, such as convolution neural networks (CNNs) based autoencoders, are used. However, traditional CNNs tend to focus on all the features irrespective of their importance, leading to weaker representations. To overcome this, we incorporate attention modules in our autoencoder architecture. These attention modules explicitly focus on more important wavelengths, leading to better transformation of the features in the low dimension. In the proposed method, the attention driven encoder transforms high dimension features to low dimensions, considering their relative importance, while the CNN based decoder reconstructs the original features. We evaluate our method on Indian pines 2010 and Indian pines 1992 hyperspectral datasets, where it surpasses the previous approaches.
Shivam Pande, Biplab Banerjee
IGARSS1
2021 Bidirectional GRU Based Autoencoder for Dimensionality Reduction in Hyperspectral Images
abstract
Hyperspectral images (HSI) are being extensively used in land use/land cover classification because they possess high spectral resolution. Although, this leads to better reflectance distinguishability, the problem of high dimensionality also occurs, making the algorithms data greedy. To counter it, deep learning models, such as autoencoders, are being used. To exploit the contiguous nature of HSIs, sequential models like recurrent neural network (RNNs) are adopted. However, for longer sequences, RNNs exhibit vanishing gradients. Also, they fail to incorporate the future information, limiting their scope. Hence, we propose a Bidirectional Gated Recurrent Unit based autoencoder (BiGRUAE), to project the high dimensional features to a low dimensional space. The bidirectional nature captures the information, both from past and future states, while the gating mechanism of GRU prevents the vanishing gradient. We evaluate our method on two hyperspectral datasets, namely, Indian pines 2010 and Salinas, where our method surpasses the benchmark methods.
Shivam Pande, Biplab Banerjee
IGARSS1
2021 Adaptive hybrid attention network for hyperspectral image classification
Shivam Pande, Biplab Banerjee
Pattern Recognit. Lett.1
2020 MT-UNET: A Novel U-Net Based Multi-Task Architecture For Visual Scene Understanding
abstract
We tackle the problem of deep end-to-end multi-task learning (MTL) for jointly performing image segmentation and depth estimation from monocular images. It is proven already that learning several related tasks together helps in attaining improved performance per task than training them autonomously. To this end, we follow the typical U-Net based encoder-decoder architecture (MT-UNet) where the densely connected deep convolutional neural network (CNN) based feature encoder is shared among the tasks while the soft attention based task-specific decoder modules produce the desired outputs. Additionally, we encourage cross-talk (CT) between the tasks by introducing cross-task skip connections at the decoder end with adaptive weight learning for the task-specific loss functions in the final cost measure. We validate the proposed framework on the challenging CityScapes and NYUv2 datasets, where our method sharply outperforms the current state-of-the-art.
Ankit Jha, Awanish Kumar, Shivam Pande, Biplab Banerjee, Subhasis Chaudhuri
ICIP3
2020 Dimensionality Reduction Using 3D Residual Autoencoder for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) are actively used for land-use/land-cover classification. However, HSIs suffer from the problem of high dimensionality and high spectral-spatial variability leading to requirement of large number of training samples. Deep learning offers several approaches to handle the aforementioned problems but is limited with its own problem of vanishing gradient that creeps in deeper networks (especially CNNs). In this paper, we propose an autoencoder (3D ResAE) that uses 3D convolutions and residual blocks to project the high-dimensional HSI features to a low-dimensional space. 3D convolutions effectively handle the spectral-spatial characteristics whereas, residual block adds an identity mapping, thereby tackling the issue of vanishing gradient. Furthermore, 3D deconvolutions are used to reconstruct the original features, while the network is trained in a semi-supervised manner. Our proposed method is tested on Indian pines and Salinas hyperspectral datasets and the results clearly demonstrate its effectiveness in classification.
Shivam Pande, Biplab Banerjee
IGARSS1