Chalavadi Vishnu

dblp:302/7458 · also Vishnu Chalavadi · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
20since 2021 · last 2025
0000-0001-9184-3545ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 TransRefine: Transformer-augmented feature refinement for zero-shot scene classification in remote sensing images
Damalla Rambabu, Pratik Abhijeet Bendre, Gayathri C., Rajeshreddy Datla, Chalavadi Vishnu
Pattern Recognit.5
2024 A Hybrid Embedding for Generalized Zero-Shot Scene Classification in Remote Sensing Images
abstract
Generalized zero-shot learning (GZSL) is a prominent approach for implementing zero-shot learning involving unseen and seen classes in the classification stage. Many existing GZSL methods in remote sensing images use word vectors for semantic exploration that inadequately describe unseen scene classes. This paper proposes a novel embedding approach (WDV-ZRS) that combines word2vec and data2vec embedding techniques to enhance the classification accuracy of unseen classes in remote sensing images. Word2vec generates a vector representation of a word based on its context usage, capturing semantic relationships between words. Data2vec, derived from self-supervised learning, generates a continuous and contextualized latent representation, leveraging the strengths of the standard transformer architecture. The proposed WDV-ZRS leverages the semantic features of word2vec and data2vec to construct a discriminative semantic space for characterizing remote sensing scene classes. Experimental results and analysis on three benchmark datasets for scene classification in remote sensing images demonstrate the effectiveness of WDV-ZRS, surpassing existing GZSL methods.
Damalla Rambabu, Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
AVSS3
2024 RSZero-CSAT: Zero-Shot Scene Classification in Remote Sensing Imagery using a Cross Semantic Attribute-guided Transformer
abstract
Zero-shot learning (ZSL) based scene classification aims to recognize unseen classes by transferring semantic information from seen classes. The applicability of ZSL for scene classification in remote sensing images becomes challenging due to the complexity of scenes. Earlier attention-based methods are ineffective for extracting discriminative region-based features within a single image. This limitation hinders their ability to achieve transferability and accurately localize object attributes, which is essential for extracting discriminative region-based features. Hence, we propose a method for zero-shot scene classification in remote sensing images using a cross-semantic attribute-guided Transformer named RSZero-CSAT. Firstly, the semantic information is acquired using shared remote sensing semantic attributes to localize object attributes that characterize discriminative region features. Then, a Transformer in RSZero-CSAT is employed to localize object attributes within visual features accurately, enhancing the effectiveness of semantic information transfer in ZSL. Specifically, the RSZero-CSAT employs a semantic attribute → visual Transformer (SAVT) and a visual → semantic attribute Transformer (VSAT) components to extract visual features guided by semantic attributes and semantic attribute features guided by visual features, respectively. Further, SAVT and VSAT mutually learn and collaborate to obtain semantically enriched visual representations, leveraging prediction-level and feature-level semantic collaborative losses for capturing crucial semantic information. Finally, the semantically enriched visual representations obtained from SAVT and VSAT are combined to facilitate visual-semantic interactions in collaboration with class semantic vectors to classify ZSL. Our experimental results demonstrate the impact of RSZero-CSAT in improving the performance of unseen classes on four scene classification benchmark datasets in remote sensing images. The code is available at https://github.com/rs-scn-cls/rszero-csat.
Damalla Rambabu, G. Swetha, Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
IJCNN4
2024 Adaptive temporal aggregation for table tennis shot recognition
Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan
Neurocomputing2
2024 TSANet: Forecasting traffic congestion patterns from aerial videos using graphs and transformers
K. Naveen Kumar, Debaditya Roy, Thakur Ashutosh Suman, Chalavadi Vishnu, C. Krishna Mohan
Pattern Recognit.4
2024 Memory Guided Transformer With Spatio-Semantic Visual Extractor for Medical Report Generation
abstract
Medicalimaging-based report writing for effective diagnosis in radiology is time-consuming and can be error-prone by inexperienced radiologists. Automatic reporting helps radiologists avoid missed diagnoses and saves valuable time. Recently, transformer-based medical report generation has become prominent in capturing long-term dependencies of sequential data with its attention mechanism. Nevertheless, input features obtained from traditional visual extractor of conventional transformers do not capture spatial and semantic information of an image. So, the transformer is unable to capture fine-grained details and may not produce detailed descriptive reports of radiology images. Therefore, we propose a spatio-semantic visual extractor (SSVE) to capture multi-scale spatial and semantic information from radiology images. Here, we incorporate two types of networks in ResNet 101 backbone architecture, i.e. (i) deformable network at the intermediate layer of ResNet 101 that utilizes deformable convolutions in order to obtain spatially invariant features, and (ii) semantic network at the final layer of backbone architecture which uses dilated convolutions to extract rich multi-scale semantic information. Further, these network representations are fused to encode fine-grained details of radiology images. The performance of our proposed model outperforms existing works on two radiology report datasets, i.e., IU X-ray and MIMIC-CXR.
Peketi Divya, Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics3
2024 Towards a Transitional Weather Scene Recognition Approach for Autonomous Vehicles
abstract
Driving in adverse weather conditions is a key challenge for autonomous vehicles (AV). Typical scene perception models perform poorly in rainy, foggy, snowy, and cloudy conditions. In addition, we observe transition states between extremes (cloudy to rainy, rainy to sunny, etc.) in nature with variations in adversity. It is crucial to define and understand these transition states in order to develop robust AV perception models. Existing research works on classification focused on identifying extreme weather conditions. However, there is a lack of emphasis on the transition between these extreme weather scenes. Hence, this paper proposes an approach to define and understand six intermediate weather transition states: sunny to rainy, rainy to sunny, and others. Firstly, we propose a way to interpolate the intermediate weather transition data using a variational autoencoder and extract its spatial features using VGG. Further, we model the temporal distribution of these spatial features using a gated recurrent unit to classify the corresponding transition state. Also, we introduce a large-scale dataset called the AIWD6: Adverse Intermediate Weather Driving dataset, generated for three different time intervals. Experimental results on the AIWD6 dataset demonstrate that our model efficiently generates weather transition conditions for AV technology. Also, the spatio-temporal deep neural network can effectively classify the adverse weather transition states for different time intervals.
Madhavi Kondapally, K. Naveen Kumar, Chalavadi Vishnu, C. Krishna Mohan
IEEE Trans. Intell. Transp. Syst.3
2023 Multi-class object classification using deep learning models in automotive object detection scenarios
abstract
This paper presents two deep learning models using a multi-perspective convolutional neural network (CNN) for classifying objects in the context of intelligent transportation systems (ITS). The proposed model categorizes objects accurately, enabling them to make well-informed decisions in multi-object (such as Persons, Trucks, Motorbikes, Cars, and Cyclists.) detection in complex scenarios for automotive applications. The custom backbone model is designed based on experimentation with the VGG backbone network based on the VGG backbone network, incorporating a multilayer prediction head and custom feature extraction blocks for classifying multiple objects in complex scenes. The model is to extract abstract features and features at multiple scales with a custom-designed feature extraction backbone with multiple blocks. The proposed models are lightweight and require fewer computational resources for high classification performance. The automotive publicly available dataset with 19800 images and labels has been used. Results show that when we experimented with the VGG backbone CNN model, the classification accuracy of 99.64% is achieved, and on the other hand, the classification accuracy of custom backbone CNN is 99.46%. The performance of the proposed custom model is also compared to those of pre-trained benchmark models. The experimental findings presented in this paper show that the proposed models achieve higher accuracy than the pre-trained models.
Soumya Abbu, Linga Reddy Cenkeramaddi, Chalavadi Vishnu, C. Krishna Mohan
ICMV3
2023 FLWGAN: Federated Learning with Wasserstein Generative Adversarial Network for Brain Tumor Segmentation
abstract
Recently, the potential of deep learning in identifying complex patterns is gaining research interest in medical applications specifically for brain tumor diagnosis. To segment tumors accurately in brain MRIs, there is a need for a large amount of data for training deep learning models. Also, hospitals cannot share patient data for centralization on the server since health records are prone to privacy and ownership challenges. To deal with these challenges, we set up an efficient federated learning (FL) pipeline with Wasserstein generative adversarial networks (FLWGAN) to ensure data privacy and data sufficiency. FL preserves the data privacy of clients by sharing only the trained model parameters to a centralized server instead of raw data. A modified 3D Wasserstein generative adversarial network with gradient penalty (WGAN-GP) and is incorporated at the client side to generate image-segmentation pairs for efficient training segmentation models. Here, 3D-UNet with an attention module is used for the brain MRI segmentation. The attention module is integrated into a 3D-UNet encoder network for effective brain tumor segmentation. Our approach aims to allow each client to benefit from locally available real data and synthetic data. This process enhances the learning performance while respecting data privacy. The efficacy of our proposed pipeline is demonstrated on the brain tumor task of the medical segmentation decathlon (MSD) dataset. We designed FLWGAN frameworks for predicting four segmentation tasks, i.e., whole tumor (WT), enhanced tumor (ET), tumor core (TC), and multiclass. Our proposed approach achieves state of the art performance in terms of various segmentation metrics.
Peketi Divya, Chalavadi Vishnu, C. Krishna Mohan, Yen-Wei Chen 0001
IJCNN2
2023 EVAA - Exchange Vanishing Adversarial Attack on LiDAR Point Clouds in Autonomous Vehicles
abstract
In addition to RGB camera sensors, LiDAR (Light Detection and Ranging) plays an important role in autonomous vehicles (AVs) to perceive their surroundings. Deep neural networks (DNNs) are able to achieve cutting-edge 3D object detection and segmentation performance using LiDAR point clouds. LiDAR-enabled autonomous vehicles provide human perception by segmenting LiDAR point clouds into meaningful regions and providing semantic context to the AV user. However, the generation of point clouds to provide semantic segmentation in AVs is not reliable and secure, which may result in traffic accidents. We propose a novel adversarial attack against LiDAR point clouds in autonomous vehicles in this paper. We devised an exchange vanishing adversarial attack (EVAA) to deceive LiDAR point clouds by introducing targeted noise on specific objects (e.g., vehicles and driveways). On two autonomous driving datasets with 3D object annotations, NuScenes and PandaSet, we evaluate the performance of our proposed attack framework. We achieve an attack success rate (ASR) of ≈63% and ASR of ≈29% on both NuScenes and PandaSet datasets, respectively.
Chalavadi Vishnu, Jayesh Khandelwal, C. Krishna Mohan, Linga Reddy Cenkeramaddi
IEEE Trans. Geosci. Remote. Sens.1
2022 Adaptive spatial and temporal aggregation for table tennis shot recognition
abstract
Action recognition is one of the challenging video understanding tasks in computer vision. Although there has been extensive research in the task of classifying coarse-grained actions, existing methods are still limited in differentiating actions with low inter-class and high intra-class variation. Particularly, the table tennis sport that involves shots of high inter-class similarity, subtle variations, occlusion, and view-point variations. While a few datasets have been available for event spotting and shot recognition, these benchmarks are mostly recorded in a constrained environment with a clear view/perception of shots executed by players. In this paper, we introduce a Table tennis shots 1.0 dataset consisting of 9000 videos of 6 fine-grained actions collected in an unconstrained manner to analyze the performance of both players. To effectively recognise these different types of table tennis shots, we propose an adaptive spatial and temporal aggregation method that can handle the spatial and temporal interactions concerning the subtle variations among shots and low inter-class variations. Our method consists of three components, namely, (i) feature extraction module, (ii) spatial aggregation network, and (iii) temporal aggregation network. The feature extraction module is a 3D convolutional neural network (3D-CNN) that captures the spatial and temporal characteristics of table tennis shots. In order to capture the interaction among the elements of the extracted 3D-CNN feature maps efficiently, we employ spatial aggregation network to obtain the compact spatial representation. Later, we propose to replace the final global average pooling layer (GAP) with the temporal aggregation network to overcome the loss of motion information due to averaging of temporal features. This temporal aggregation network utilizes the attention mechanism of bidirectional encoder representations from Transformers (BERT) to model the significant temporal interactions among the shots effectively. We demonstrate that our proposed approach improves the performance of existing 3D-CNN methods by ~10% on the Table tennis shots 1.0 dataset.We also show the performance of our approach on other action recognition datasets, namely, UCF-101 and HMDB-51.
Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan
ICMV2
2022 STIP-GCN: Space-time interest points graph convolutional network for action recognition
abstract
Action recognition requires modelling the interactions between either human & human or human & objects. Re-cently, graph convolutional neural networks (GCNs) are exploited to effectively capture the structure of action by modelling the relationship among entities present in a video. However, most of the approaches depend on the effectiveness of object detection frameworks to detect the entities. In this paper, we propose a graph-based framework for action recognition to model the spatio-temporal interactions among the entities in a video without any object-level supervision. First, we obtain the salient space-time interest points (STIP) that contain rich information about the significant local variations in space and time by using the Harris 3D detector. In order to incorporate the local appearance and motion information of the entities, either low-level or deep features are extracted around these STIPs. Next, we build a graph by considering the extracted STIPs as nodes and are connected by spatial edges and temporal edges. These edges are determined based on a membership function that measures the similarity of entities associated with the STIPs. Finally, GCN is employed on the given graph to provide reasoning among different entities present in a video. We evaluate our method on three widely used datasets, namely, UCF-101, HMDB-51, SSV2 to demonstrate the efficacy of the proposed approach.
Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan
IJCNN2
2022 M-FFN: multi-scale feature fusion network for image captioning
Jeripothula Prudviraj, Chalavadi Vishnu, C. Krishna Mohan
Appl. Intell.2
2022 mSODANet: A network for multi-scale object detection in aerial images using hierarchical dilated convolutions
Chalavadi Vishnu, Jeripothula Prudviraj, Rajeshreddy Datla, Sobhan Babu Chintapalli, C. Krishna Mohan
Pattern Recognit.1
2022 Fine-grained action recognition using dynamic kernels
Sravani Yenduri, Nazil Perveen, Chalavadi Vishnu, C. Krishna Mohan
Pattern Recognit.3
2022 AAP-MIT: Attentive Atrous Pyramid Network and Memory Incorporated Transformer for Multisentence Video Description
abstract
Generating multi-sentence descriptions for video is considered to be the most complex task in computer vision and natural language understanding due to the intricate nature of video-text data. With the recent advances in deep learning approaches, the multi-sentence video description has achieved an impressive progress. However, learning rich temporal context representation of visual sequences and modelling long-term dependencies of natural language descriptions is still a challenging problem. Towards this goal, we propose an Attentive Atrous Pyramid network and Memory Incorporated Transformer (AAP-MIT) for multi-sentence video description. The proposed AAP-MIT incorporates the effective representation of visual scene by distilling the most informative and discriminative spatio-temporal features of video data at multiple granularities and further generates the highly summarized descriptions. Profoundly, we construct AAP-MIT with three major components: i) a temporal pyramid network, which builds the temporal feature hierarchy at multiple scales by convolving the local features at temporal space, ii) a temporal correlation attention to learn the relations among various temporal video segments, and iii) the memory incorporated transformer, which augments the new memory block in language transformer to generate highly descriptive natural language sentences. Finally, the extensive experiments on ActivityNet Captions and YouCookII datasets demonstrate the substantial superiority of AAP-MIT over the existing approaches.
Jeripothula Prudviraj, Malipatel Indrakaran Reddy, Chalavadi Vishnu, C. Krishna Mohan
IEEE Trans. Image Process.3
2021 A multimodal semantic segmentation for airport runway delineation in panchromatic remote sensing images
abstract
Monitoring airport runways in panchromatic remote sensing images is helpful for both civil and strategic communities in effective utilization of the large-area acquisitions. This paper proposes a novel multimodal semantic segmentation approach for effective delineation of the runways in panchromatic remote sensing images. The proposed approach aims to learn complementary information from two modalities, namely, panchromatic image and digital elevation model (DEM) to obtain discriminative features of the runway. The fusion of image features and the corresponding terrain information is performed by stacking the image and DEM by leveraging the merits of both Transformers and U-Net architecture. We perform the experiments on Cartosat-1 panchromatic satellite images with the corresponding Cartosat-1 DEM scenes. The experimental results demonstrate a significant contribution of terrain information to the segmentation process in achieving the contours of airport runways effectively.
Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
ICMV2
2021 A framework to derive geospatial attributes for aircraft type recognition in large-scale remote sensing images
abstract
Aircraft type recognition remains challenging, due to their tiny sizes and geometric distortions in large-scale panchromatic satellite images. This paper proposes a framework for aircraft type recognition by focusing on shape preservation, spatial transformations, and geospatial attributes derivation. First, we construct an aircraft segmentation model to obtain masks representing the shape of aircrafts by employing a learnable shape-preserved and deformable network in the mask RCNN architecture. Then, the orientation of the segmented aircrafts is determined by estimating the symmetrical axes using their gradient information. Besides template matching, we derive the length and width of aircrafts using the geotagged information of images to further categorize the types of aircrafts. Also, we present an effective inferencing mechanism to overcome the issue of partial detection or missing aircrafts in large-scale images. The efficacy of the proposed framework is demonstrated on large-scale panchromatic images with ground sampling distances of 0.65m (C2S).
Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
ICMV2
2021 Scene Classification in Remote Sensing Images using Dynamic Kernels
abstract
Classification of scenes across multi-sensor remote sensing images with different spatial, spectral, temporal resolutions involves identification of variable length spatial patterns of objects in a scene. So, it necessitates the use of local representations from different regions of a scene in order to comprehend the scene formation. In this paper, we propose a dynamic kernel based representation to handle the patterns of variable lengths in the scenes of remote sensing images. These kernels help to assimilate spatial variability captured using convolutional features in a Gaussian mixture model. The statistics of GMM facilitate the dynamic kernels in preserving the local spatial similarities while handling the changes in spatial content globally within the same scene. The efficacy of the proposed method using two variants of the dynamic kernels is demonstrated on three benchmark scene classification datasets, namely, UCM Land Use (21 classes), Aerial image dataset (30 classes), and NWPU-RESISC45 (45 classes). Our experiments show that the mean interval kernel is better discriminative as it makes use of first and second-order statistics of GMM.
Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
IJCNN2
2021 Attentive Contextual Network for Image Captioning
abstract
Existing image captioning approaches fail to generate fine-grained captions due to the lack of rich encoding representation of an image. In this paper, we present an attentive contextual network (ACN) to learn the spatially transformed image features and dense multi-scale contextual information of an image to generate semantically meaningful captions. At first, we construct deformable network on intermediate layers of convolutional neural network (CNN) to cultivate spatial invariant features. And the multi-scale contextual features are produced by employing contextual network on top of last layers of CNN. Then, we exploit attention mechanism on contextual network to extract dense contextual features. Further, the extracted spatial and contextual features are combined to encode the holistic representation of an image. Finally, a multi-stage caption decoder with visual attention module is incorporated to generate fine-grained captions. The performance of the proposed approach is demonstrated on COCO dataset, the largest dataset for image captioning.
Jeripothula Prudviraj, Chalavadi Vishnu, C. Krishna Mohan
IJCNN2
2017 Detection of motorcyclists without helmet in videos using convolutional neural network
abstract
In order to ensure the safety measures, the detection of traffic rule violators is a highly desirable but challenging task due to various difficulties such as occlusion, illumination, poor quality of surveillance video, varying whether conditions, etc. In this paper, we present a framework for automatic detection of motorcyclists driving without helmets in surveillance videos. In the proposed approach, first we use adaptive background subtraction on video frames to get moving objects. Later convolutional neural network (CNN) is used to select motorcyclists among the moving objects. Again, we apply CNN on upper one fourth part for further recognition of motorcyclists driving without a helmet. The performance of the proposed approach is evaluated on two datasets, IITH_Helmet_1 contains sparse traffic and IITH_Helmet_2 contains dense traffic, respectively. The experiments on real videos successfully detect 92.87% violators with a low false alarm rate of 0.5% on an average and thus shows the efficacy of the proposed approach.
Chalavadi Vishnu, Dinesh Singh 0001, C. Krishna Mohan, Sobhan Babu Chintapalli
IJCNN1
2016 Visual Big Data Analytics for Traffic Monitoring in Smart City
abstract
The application such as video surveillance for traffic control in smart cities needs to analyze the large amount (hours/days) of video footage in order to locate the people who are violating the traffic rules. The traditional computer vision techniques are unable to analyze such a huge amount of visual data generated in real-time. So, there is a need for visual big data analytics which involves processing and analyzing large scale visual data such as images or videos to find semantic patterns that are useful for interpretation. In this paper, we propose a framework for visual big data analytics for automatic detection of bike-riders without helmet in city traffic. We also discuss challenges involved in visual big data analytics for traffic control in a city scale surveillance data and explore opportunities for future research.
Dinesh Singh 0001, Chalavadi Vishnu, C. Krishna Mohan
ICMLA2