C. Krishna Mohan

dblp:30/4639 · also Chalavadi Krishna Mohan, Krishna Mohan C · DBLP profile ↗
← Back
95ranked-venue papers
1as first author
57since 2021 · last 2026
0000-0002-7316-0836ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 23 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 9 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Task-Free Online Replay with Contrastive Learning and Dynamic Herding
Rahul Biswas, C. Krishna Mohan, Subhrajit Nag, Sandipan Dandapat
ICPR (14)2
2026 Reducing Spurious Detections in Video-Based Aerial Object Recognition: A Multi-Stage Contextual, Temporal, and Multi-Scale Attention Architecture
Shubham Kumar Dubey, J. V. Satyanarayana, C. Krishna Mohan
ICPRAM3
2026 A Unified Framework Combining Clustering Algorithms, SMT-Based Reasoning, and Automated Semantic Analysis to Reduce False Positives in Video Analysis
Shubham Kumar Dubey, J. V. Satyanarayana, C. Krishna Mohan
ICPRAM3
2026 Bridging the Domain Gap in Small Multimodal Models: A Dual-level Alignment Perspective
Aveen Dayal, Peketi Divya, Nidhi Tiwari, Linga Reddy Cenkeramaddi, C. Krishna Mohan, Abhinav Kumar 0001
WACV5
2026 CaRS: A Causal Intervention Segmentation Framework and Benchmark Dataset for Autonomous Driving under Transitional Weather Conditions
abstract
Autonomous vehicles must excel in safety-critical perception tasks, especially in adverse weather conditions. In addition, transitional weather shifts in nature, such as sunny to rainy, rainy to cloudy, etc., pose abrupt illumination changes that can distort object boundaries and degrade segmentation performance. Existing research focuses mainly on segmentation in clear and discrete weather conditions, leaving a gap in addressing the issues of transitional weather scenarios. Hence, we propose a novel method called causal road and rest segmentation (CaRS) that utilizes causal intervention to mitigate the confounding bias due to transitional weather changes. We use dual complementary attention modules, one for causal and another for confounding feature extraction. These modules complement each other and are fine-tuned via an adversarial min-max approach to reduce confounding bias and enhance segmentation performance. Also, our CaRS method concurrently performs road semantic segmentation and instance segmentation of vehicles and pedestrians. Further, we introduce a transitional weather-driving dataset for segmentation (TWDS16) using a spurious correlation generator that leverages data interpolation to produce 16 weather transitions. We evaluate the performance of CaRS on TWDS16, along with three other benchmark datasets, namely, Foggy Cityscapes, RainCityscapes, and BDD100K. The experimental results validate the efficacy of the proposed method in mitigating confounding influences, leading to improved mIoU for semantic segmentation and mAP for instance segmentation across diverse datasets.
Madhavi Kondapally, K. Naveen Kumar, C. Krishna Mohan, Sobhan Babu
WACV3
2026 3DA-net: a dual-attention-based network integrating global and local context for enhanced 3D object detection
abstract
Accurate 3D object detection from LiDAR data is vital for enhancing road safety, enabling efficient traffic management, and supporting reliable path planning in autonomous navigation systems. However, LiDAR point clouds suffer from inherent challenges such as sparsity, occlusion, and variations in point density, which can significantly impact detection accuracy. To address these challenges, we introduce 3DA-Net, a dual-attention-based network that integrates global and local context for enhanced 3D object detection. We begin by converting raw LiDAR point clouds into structured voxel representations, which are then processed through a hybrid dual-attention encoder. In this encoder, global attention modules capture high-level semantic dependencies across the entire scene, while local attention focuses on fine-grained geometric structures within neighborhoods. This dual-attention mechanism is further strengthened with point-wise and channel-wise attention, which enhances the model’s ability to capture both spatial and contextual information, which is essential for 3D perception. Our design incorporates a custom backbone for robust feature extraction from voxel-based pseudo-image representations, coupled with a feature pyramid network for efficient multi-scale feature learning. Evaluations on the KITTI dataset show that 3DA-Net achieves AP40 scores of 95.91% (easy), 94.78% (moderate), and 91.98% (hard) for cars in bird’s-eye view detection, and 95.83%, 94.58%, and 90.03% in 3D detection, outperforming strong LiDAR-based detection baselines. Significant improvements in pedestrian and cyclist detection further demonstrate the robustness and generalizability of our method in complex driving environments.
Soumya Abbu, Linga Reddy Cenkeramaddi, C. Krishna Mohan
Appl. Intell.3
2026 Hybrid federated continual graph contrastive learning for evolving money laundering threats
Zarka Bashir, Mridula Verma, C. Krishna Mohan
Data Min. Knowl. Discov.3
2026 An improved multi-instance learning model with clinical-guided cross-attention for postoperative early recurrence prediction of hepatocellular carcinoma using histopathological images
Gan Zhan, Fang Wang 0030, Yinhao Li 0002, Rahul Kumar Jain 0001, Qingqing Chen 0001, Lanfen Lin, Hongjie Hu, C. Krishna Mohan, Yen-Wei Chen 0001
Neurocomputing10
2026 Optimal Transport Barycentric Aggregation for Byzantine-Resilient Federated Learning
abstract
Federated learning (FL) has emerged as a promising solution to enable distributed learning without sharing sensitive data. However, FL is vulnerable to data poisoning attacks, where malicious clients inject malicious data during training to compromise the global model. Existing FL defenses suffer from the assumptions of independent and identically distributed (IID) model updates, asymptotic optimal error rate bounds, and strong convexity in the optimization problem. Hence, we propose a novel framework called Federated Learning Optimal Transport (FLOT) that leverages the Wasserstein barycentric technique to obtain a global model from a set of locally trained non-IID models on client devices. In addition, we introduce a loss function-based rejection (LFR) mechanism to suppress malicious updates and a dynamic weighting scheme to optimize the Wasserstein barycentric aggregation function. We provide the theoretical proof of the Byzantine resilience and convergence of FLOT to highlight its efficacy. We evaluate FLOT on four benchmark datasets: GTSRB, KBTS, CIFAR10, and EMNIST. The experimental results underscore the practical significance of FLOT as an effective defense mechanism against data poisoning attacks in FL while maintaining high accuracy and scalability. Also, we observe that FLOT serves as a robust client selection technique under no attack, which demonstrates its effectiveness.
K. Naveen Kumar, Srinivasa Rao Chalamala, Ajeet Kumar Singh, C. Krishna Mohan
IEEE Trans. Big Data4
2025 CE-KD: Class-Wise Expert-Based Knowledge Distillation for Facial Expression Recognition
abstract
Knowledge distillation (KD) is a model compression technique that transfers knowledge from a complex and well-trained teacher model to a compact student model, thereby enabling the student to mimic the performance and behavior of the teacher. However, traditional KD methods struggle with long-tailed facial expression recognition (FER), as FER datasets often exhibit severe class imbalance. For instance, certain expressions (e.g., happiness) are overrepresented, while others (e.g., fear) have significantly fewer samples. This innate class-imbalanced property of FER leads to suboptimal knowledge transfer for the underrepresented expressions (i.e. biased learning and poor generalization for underrepresented classes). To address this issue, this paper introduces the CE-KD, a Class-wise Expert-based Knowledge Distillation framework that enables the student model to effectively learn fine-grained expression details and high-level emotional concepts from specialized emotion experts. This improves the generalizability of the FER models. Extensive experiments on benchmark datasets — FERPlus and RAF-DB — demonstrate that our CE-KD framework provides a practical solution to implement efficient FER systems in real-world applications while maintaining robust performance across different emotion expressions.
Khin Cho Win, Zahid Akhtar, C. Krishna Mohan
AVSS3
2025 Fortifying Federated Learning Towards Trustworthiness via Auditable Data Valuation and Verifiable Client Contribution
abstract
Ensuring auditability and verifiability in Federated Learning (FL) is both challenging and essential to guarantee that local data remains untampered and client updates are trustworthy. Recent FL frameworks assess client contributions through a trusted central server using various client selection and aggregation techniques. However, reliance on a central server can create a single point of failure, making it vulnerable to privacy-centric attacks and limiting its ability to audit and verify client-side data contributions due to restricted access. In addition, data quality and fairness evaluations are often inadequate, failing to distinguish between high-impact contributions and those from low-quality or poisoned data. To address these challenges, we propose Federated Auditable and Verifiable Data valuation (FAVD), a privacy-preserving method that ensures auditability and verifiability of client contributions through data valuation, independent of any central authority or predefined training algorithm. FAVD utilizes shared local data density functions to construct a global density function, aligning data contributions and facilitating effective valuation prior to local model training. This proactive approach improves transparency in data valuation and ensures that only benign updates are generated, even in the presence of malicious data. Further, to mitigate privacy risks associated with sharing data density functions, we add Gaussian noise to each client’s local density function before sharing it with the server. We theoretically demonstrate the convergence, auditability, and verifiability of FAVD, along with its resilience against data poisoning threats. Our experiments on five diverse benchmarks, including three medical datasets, show that FAVD achieves significant performance gains, accurate data valuation, and fair client contributions under threat, highlighting its reliability as a trustworthy FL approach.
K. Naveen Kumar, Ranjeet Ranjan Jha, C. Krishna Mohan, Ravindra Babu Tallamraju
CVPR3
2025 Fed-SMTDA: A Novel Framework for Federated Source-Free Multi-Target Domain Adaptation Using Feature Clustering and Adaptive Aggregation
abstract
Federated Learning (FL) enables collaborative model training across decentralized devices while maintaining data privacy. However, in many real-world scenarios, accessing both source and labeled target data on these devices is not feasible due to privacy concerns, data heterogeneity, and resource constraints. This paper introduces a novel approach to address this challenge, termed Federated Source-Free Multi-Target Domain Adaptation (Fed-SMTDA). In this framework, client devices possess only unlabeled target data, while the server has a pre-trained model derived from source data. Our method leverages the intrinsic structure of the target domain at each client by clustering similar features, facilitating more effective feature extraction and assignment. To improve model performance, we incorporate an entropy regularization term to minimize class confusion, ensuring cleaner decision boundaries. Additionally, we introduce a dynamic aggregation strategy called Weight Adjustment (WA), where the server adjusts the weights assigned to client models during aggregation based on the observed generalization gap across clients. This adaptive approach improves the overall robustness and generalization of the federated model, enabling it to perform effectively across diverse and unlabeled target domains. Fed-SMTDA is evaluated on the OfficeHome and PACS datasets. Its performance is compared against two centralized baselines – one that has full access to labels across domains and is termed Oracle, and another where the model is trained only with the source domain labels, which we call Source-Only. Fed-SMTDA delivers consistently better performance than the Source-Only model, and its performance is upper-bounded by the Oracle model.
Challapalli Phanindra Revanth, Sumohana S. Channappayya, C. Krishna Mohan
IJCNN3
2025 Semantic-assisted report generation with memory enhanced transformer using context-aware visual extractor
Peketi Divya, Partha Sarathi Chakraborty 0003, C. Krishna Mohan, Yen-Wei Chen 0001
Appl. Intell.3
2025 Minimal data poisoning attack in federated learning for medical image classification: An attacker perspective
K. Naveen Kumar, C. Krishna Mohan, Linga Reddy Cenkeramaddi, Navchetan Awasthi
Artif. Intell. Medicine2
2025 CPL-PL: Contrapositive Learning-Based Pseudo-Labeling for Semi-Supervised Scene Classification in Remote Sensing Images
abstract
Scene classification in remote sensing images is a challenging task due to the limited availability of labeled data and the high intra-class variability in complex landscapes. Semi-supervised learning (SSL) has emerged as an effective approach to leverage the limited labeled data in utilizing a large amount of unlabeled data for improved classification. Pseudo-labeling, a widely used SSL technique, determines suitable labels to unlabeled data based on high-confidence model predictions. However, traditional pseudo-labeling methods suffer from confirmation bias, where incorrect labels reinforce errors, degrading model performance. To address this, we propose Contrapositive Learning-based Pseudo-Labeling (CPL-PL), a novel method designed specifically for remote sensing scene classification. CPL-PL introduces a Contrapositive Loss that enforces feature consistency for similar scenes while ensuring representation separation for dissimilar ones, leading to more reliable pseudo-label assignments. Our approach mitigates pseudo-label noise, enhances feature discrimination, and improves classification robustness. Experimental results on benchmark remote sensing datasets demonstrate that CPL-PL significantly outperforms conventional pseudo-labeling strategies, especially in low-label regimes. The proposed method provides a promising direction for advancing semi-supervised scene classification in remote sensing images.
G. Swetha, Rajeshreddy Datla, Sobhan Babu Chintapalli, C. Krishna Mohan
IEEE Geosci. Remote. Sens. Lett.4
2025 Federated Learning Minimal Model Replacement Attack Using Optimal Transport: An Attacker Perspective
abstract
Federated learning (FL) has emerged as a powerful collaborative learning approach that enables client devices to train a joint machine learning model without sharing private data. However, the decentralized nature of FL makes it highly vulnerable to adversarial attacks from multiple sources. There are diverse FL data poisoning and model poisoning attack methods in the literature. Nevertheless, most of them focus only on the attack’s impact and do not consider the attack budget and attack visibility. These factors are essential to effectively comprehend the adversary’s rationale in designing an attack. Hence, our work highlights the significance of considering these factors by providing an attacker perspective in designing an attack with a low budget, low visibility, and high impact. Also, existing attacks that use total neuron replacement and randomly selected neuron replacement approaches only cater to these factors partially. Therefore, we propose a novel federated learning minimal model replacement attack (FL-MMR) that uses optimal transport (OT) for minimal neural alignment between a surrogate poisoned model and the benign model. Later, we optimize the attack budget in a three-fold adaptive fashion by considering critical learning periods and introducing the replacement map. In addition, we comprehensively evaluate our attack under three threat scenarios using three large-scale datasets: GTSRB, CIFAR10, and EMNIST. We observed that our FL-MMR attack drops global accuracy to$\approx 35\%$less with merely 0.54% total attack budget and lower attack visibility than other attacks. The results confirm that our method aligns closely with the attacker’s viewpoint compared to other methods.
K. Naveen Kumar, C. Krishna Mohan, Linga Reddy Cenkeramaddi
IEEE Trans. Inf. Forensics Secur.2
2025 Leveraging Mixture Alignment for Multi-Source Domain Adaptation
abstract
In a conventional Domain Adaptation (DA) setting, we only have one source and target domain, whereas, in many real-world applications, data is often collected from several related sources in different conditions. This has led to a more practical and challenging knowledge transfer problem called Multi-source Domain Adaptation (MDA). Several methodologies, such as prototype matching, explicit distance discrepancy, adversarial learning, etc., have been considered to tackle the MDA problem in recent years. Among them, the adversarial-based learning framework is a popular methodology for transferring knowledge from multiple sources to target domains using a minmax optimization strategy. Despite the advances in adversarial-based methods, several limitations exist, such as the need for a classifier-aware discrepancy metric to align the domains and the need to consider target samples' consistency and semantic information while aligning the domains. To mitigate these issues, in this work, we propose a novel adversarial learning MDA algorithm, MDAMA, which aligns the target domain with a mixture distribution that consists of source domains. MDAMA uses margin-based discrepancy and augmented intermediate distributions to align the domains effectively. We also propose consistency of target samples by confidence thresholding and transfer of semantic information from multiple source domains to the augmented target domain to further improve the performance of the target domain. We extensively experiment with the MDAMA algorithm on popular real-world MDA datasets such as OfficeHome, Office31, PACS, Office-Caltech, and DomainNet. We evaluate the MDAMA model on these benchmark datasets and demonstrate top performance in all of them.
Aveen Dayal, Shrusti S., Linga Reddy Cenkeramaddi, C. Krishna Mohan, Abhinav Kumar 0001
IEEE Trans. Image Process.4
2024 A Hybrid Embedding for Generalized Zero-Shot Scene Classification in Remote Sensing Images
abstract
Generalized zero-shot learning (GZSL) is a prominent approach for implementing zero-shot learning involving unseen and seen classes in the classification stage. Many existing GZSL methods in remote sensing images use word vectors for semantic exploration that inadequately describe unseen scene classes. This paper proposes a novel embedding approach (WDV-ZRS) that combines word2vec and data2vec embedding techniques to enhance the classification accuracy of unseen classes in remote sensing images. Word2vec generates a vector representation of a word based on its context usage, capturing semantic relationships between words. Data2vec, derived from self-supervised learning, generates a continuous and contextualized latent representation, leveraging the strengths of the standard transformer architecture. The proposed WDV-ZRS leverages the semantic features of word2vec and data2vec to construct a discriminative semantic space for characterizing remote sensing scene classes. Experimental results and analysis on three benchmark datasets for scene classification in remote sensing images demonstrate the effectiveness of WDV-ZRS, surpassing existing GZSL methods.
Damalla Rambabu, Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
AVSS4
2024 Precision Guided Approach to Mitigate Data Poisoning Attacks in Federated Learning
abstract
Federated Learning (FL) is a collaborative learning paradigm enabling participants to collectively train a shared machine learning model while preserving the privacy of their sensitive data. Nevertheless, the inherent decentralized and data-opaque characteristics of FL render its susceptibility to data poisoning attacks. These attacks introduce malformed or malicious inputs during local model training, subsequently influencing the global model and resulting in erroneous predictions. Current FL defense strategies against data poisoning attacks either involve a trade-off between accuracy and robustness or necessitate the presence of a uniformly distributed root dataset at the server. To overcome these limitations, we present FedZZ, which harnesses a zone-based deviating update (ZBDU) mechanism to effectively counter data poisoning attacks in FL. The ZBDU approach identifies the clusters of benign clients whose collective updates exhibit notable deviations from those of malicious clients engaged in data poisoning attack. Further, we introduce a precision-guided methodology that actively characterizes these client clusters (zones), which in turn aids in recognizing and discarding malicious updates at the server. Our evaluation of FedZZ across two widely recognized datasets: CIFAR10 and EMNIST, demonstrate its efficacy in mitigating data poisoning attacks, surpassing the performance of prevailing state-of-the-art methodologies in both single and multi-client attack scenarios and varying attack volumes. Notably, FedZZ also functions as a robust client selection strategy, even in highly non-IID and attack-free scenarios. Moreover, in the face of escalating poisoning rates, the model accuracy attained by FedZZ displays superior resilience compared to existing techniques. For instance, when confronted with a 50% presence of malicious clients, FedZZ sustains an accuracy of 67.43%, while the accuracy of the second-best solution, FL-Defender, diminishes to 43.36%.
K. Naveen Kumar, C. Krishna Mohan, Aravind Machiry
CODASPY2
2024 Revamping Federated Learning Security from a Defender's Perspective: A Unified Defense with Homomorphic Encrypted Data Space
abstract
Federated Learning (FL) facilitates clients to collaborate on training a shared machine learning model without exposing individual private data. Nonetheless, FL remains susceptible to utility and privacy attacks, notably evasion data poisoning and model inversion attacks, compromising the system's efficiency and data privacy. Existing FL defenses are often specialized to a particular single attack, lacking generality and a comprehensive defender's perspective. To address these challenges, we introduce Federated Cryptography Defense (FCD), a unified single framework aligning with the defender's perspective. FCD employs row-wise transposition cipher based data encryption with a secret key to counter both evasion black-box data poisoning and model inversion attacks. The crux of FCD lies in transferring the entire learning process into an encrypted data space and using a novel distillation loss guided by the Kullback-Leibler (KL) divergence. This measure compares the probability distributions of the local pretrained teacher model's predictions on normal data and the local student model's predictions on the same data in FCD's encrypted form. By working within this encrypted space, FCD eliminates the need for decryption at the server, resulting in reduced computational complexity. We demonstrate the practical feasibility of FCD and apply it to defend against evasion utility attack on benchmark datasets (GTSRB, KBTS, CIFAR10, and EMNIST). We further extend FCD for defending against model inversion attack in split FL on the CIFAR100 dataset. Our experiments across the diverse attack and FL settings demonstrate practical feasibility and robustness against utility evasion (impact > 30) and privacy attacks (MSE > 73) compared to the second best method.
K. Naveen Kumar, Reshmi Mitra, C. Krishna Mohan
CVPR3
2024 Improving Unsupervised Domain Adaptation: A Pseudo-candidate Set Approach
Aveen Dayal, Rishabh Lalla, Linga Reddy Cenkeramaddi, C. Krishna Mohan, Abhinav Kumar 0001, Vineeth N. Balasubramanian
ECCV (32)4
2024 FL-PSeC: Federated Learning-Pseudo Labeled Medical Image Segmentation with Personalized Class Balancing Semi-supervised Approach
Ishu Priya, C. Krishna Mohan
ICPR (10)2
2024 Object Detection in Transitional Weather Conditions for Autonomous Vehicles
abstract
Navigating safely and dependably through challenging weather conditions poses a significant hurdle for autonomous vehicles (AVs). While state-of-the-art object detection models have demonstrated superior performance on standard benchmark datasets, their accuracy is compromised by visual variations introduced by adverse weather conditions. In addition, we naturally observe the continuous shifts between discrete weather conditions (cloudy to rainy, rainy to sunny, etc.), with variation in different levels of adversity. The current object detection research predominantly concentrates on identifying objects in discrete weather conditions (extremely cloudy, rainy, etc.). However, there is a lack of emphasis on continuous shifts between these stationary weather conditions. In response to this challenge, we introduce a pioneering solution, the Multi-Scale Adaptive Transformer (mSAT). This innovative approach amalgamates a Domain Adaptive Network (DAN), adept at identifying continuous weather-invariant features across various scales, with a transformer network tailored for object detection. Our method is evaluated on the AIWD6 dataset, showcasing its efficacy in addressing the impact of adverse weather conditions on object detection. Our approach effectively mitigates the domain discrepancy, enabling adaptation to various continuous weather shifts. Later, we introduce three novel metrics for evaluating object detection performance on continuous weather data along with standard metrics. Our proposed mSAT, designed to operate on various intensity levels of weather with unlabeled target data, achieves 74.6 mAP on the AIWD6 dataset. Experimental results demonstrate that our model adapts to continuous weather shifts and effectively performs object detection.
Madhavi Kondapally, K. Naveen Kumar, C. Krishna Mohan
IJCNN3
2024 RSZero-CSAT: Zero-Shot Scene Classification in Remote Sensing Imagery using a Cross Semantic Attribute-guided Transformer
abstract
Zero-shot learning (ZSL) based scene classification aims to recognize unseen classes by transferring semantic information from seen classes. The applicability of ZSL for scene classification in remote sensing images becomes challenging due to the complexity of scenes. Earlier attention-based methods are ineffective for extracting discriminative region-based features within a single image. This limitation hinders their ability to achieve transferability and accurately localize object attributes, which is essential for extracting discriminative region-based features. Hence, we propose a method for zero-shot scene classification in remote sensing images using a cross-semantic attribute-guided Transformer named RSZero-CSAT. Firstly, the semantic information is acquired using shared remote sensing semantic attributes to localize object attributes that characterize discriminative region features. Then, a Transformer in RSZero-CSAT is employed to localize object attributes within visual features accurately, enhancing the effectiveness of semantic information transfer in ZSL. Specifically, the RSZero-CSAT employs a semantic attribute → visual Transformer (SAVT) and a visual → semantic attribute Transformer (VSAT) components to extract visual features guided by semantic attributes and semantic attribute features guided by visual features, respectively. Further, SAVT and VSAT mutually learn and collaborate to obtain semantically enriched visual representations, leveraging prediction-level and feature-level semantic collaborative losses for capturing crucial semantic information. Finally, the semantically enriched visual representations obtained from SAVT and VSAT are combined to facilitate visual-semantic interactions in collaboration with class semantic vectors to classify ZSL. Our experimental results demonstrate the impact of RSZero-CSAT in improving the performance of unseen classes on four scene classification benchmark datasets in remote sensing images. The code is available at https://github.com/rs-scn-cls/rszero-csat.
Damalla Rambabu, G. Swetha, Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
IJCNN5
2024 PointBi-FPN: An Extention to Pointpillars for LiDAR 3D Object Detection in Autonomous Vehicles Using Bi-Directional Feature Pyramid Network
abstract
Autonomous vehicles increasingly rely on accurate three-dimensional (3D) object detection for safe navigation. While two-dimensional (2D) methods offer computational efficiency, the shift to 3D detection enhances precision in understanding environments. Point-based and voxel-based approaches are accurate but computationally intensive for onboard deployment. Pillar-based methods like PointPillars provide efficiency but may lack detection accuracy compared to voxel-based approaches. In this paper, we focus on enhancing the performance of PointPillars, a popular pillar-based detector, using LiDAR data. The proposed approach (PointBi-FPN) employs a bi-directional feature pyramid network (Bi-FPN) as a backbone, that aggregates multiscale features in the input data, providing a holistic view of the environment, which is crucial for detecting objects of varying sizes and distances accurately. Bi-FPN improves the model's understanding of complex scenes, making it more robust to occlusions. Through extensive experimentation on the KITTI dataset, the PointBi-FPN approach demonstrates better performance across all three detection benchmarks: BEV (Bird's Eye View), 3D (Three-Dimensional), and AOS (Average Orientation Similarity). Notably, the proposed approach exhibits significant improvements, particularly in accurately detecting objects labeled as “hard” difficulty.
Soumya Abbu, C. Krishna Mohan, Linga Reddy Cenkeramaddi, Sobhan Babu
TENCON2
2024 Robust Facial Emotion Recognition System via De-Pooling Feature Enhancement and Weighted Exponential Moving Average
Khin Cho Win, Zahid Akhtar, C. Krishna Mohan
TENCON3
2024 Learning scene-vectors for remote sensing image scene classification
Rajeshreddy Datla, Nazil Perveen, C. Krishna Mohan
Neurocomputing3
2024 Adaptive temporal aggregation for table tennis shot recognition
Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan
Neurocomputing3
2024 Semantic segmentation of breast cancer images using DenseNet with proposed PSPNet
Suresh Samudrala, C. Krishna Mohan
Multim. Tools Appl.2
2024 An integrated approach for prediction of magnitude using deep learning techniques
Anushka Joshi, Balasubramanian Raman, C. Krishna Mohan
Neural Comput. Appl.3
2024 The Impact of Adversarial Attacks on Federated Learning: A Survey
abstract
Federated learning (FL) has emerged as a powerful machine learning technique that enables the development of models from decentralized data sources. However, the decentralized nature of FL makes it vulnerable to adversarial attacks. In this survey, we provide a comprehensive overview of the impact of malicious attacks on FL by covering various aspects such as attack budget, visibility, and generalizability, among others. Previous surveys have primarily focused on the multiple types of attacks and defenses but failed to consider the impact of these attacks in terms of their budget, visibility, and generalizability. This survey aims to fill this gap by providing a comprehensive understanding of the attacks' effect by identifying FL attacks with low budgets, low visibility, and high impact. Additionally, we address the recent advancements in the field of adversarial defenses in FL and highlight the challenges in securing FL. The contribution of this survey is threefold: first, it provides a comprehensive and up-to-date overview of the current state of FL attacks and defenses. Second, it highlights the critical importance of considering the impact, budget, and visibility of FL attacks. Finally, we provide ten case studies and potential future directions towards improving the security and privacy of FL systems.
K. Naveen Kumar, C. Krishna Mohan, Linga Reddy Cenkeramaddi
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 TSANet: Forecasting traffic congestion patterns from aerial videos using graphs and transformers
K. Naveen Kumar, Debaditya Roy, Thakur Ashutosh Suman, Chalavadi Vishnu, C. Krishna Mohan
Pattern Recognit.5
2024 Monte Carlo DropBlock for modeling uncertainty in object detection
Sai Harsha Yelleni, Deepshikha Kumari, P. K. Srijith, C. Krishna Mohan
Pattern Recognit.4
2024 Memory Guided Transformer With Spatio-Semantic Visual Extractor for Medical Report Generation
abstract
Medicalimaging-based report writing for effective diagnosis in radiology is time-consuming and can be error-prone by inexperienced radiologists. Automatic reporting helps radiologists avoid missed diagnoses and saves valuable time. Recently, transformer-based medical report generation has become prominent in capturing long-term dependencies of sequential data with its attention mechanism. Nevertheless, input features obtained from traditional visual extractor of conventional transformers do not capture spatial and semantic information of an image. So, the transformer is unable to capture fine-grained details and may not produce detailed descriptive reports of radiology images. Therefore, we propose a spatio-semantic visual extractor (SSVE) to capture multi-scale spatial and semantic information from radiology images. Here, we incorporate two types of networks in ResNet 101 backbone architecture, i.e. (i) deformable network at the intermediate layer of ResNet 101 that utilizes deformable convolutions in order to obtain spatially invariant features, and (ii) semantic network at the final layer of backbone architecture which uses dilated convolutions to extract rich multi-scale semantic information. Further, these network representations are fused to encode fine-grained details of radiology images. The performance of our proposed model outperforms existing works on two radiology report datasets, i.e., IU X-ray and MIMIC-CXR.
Peketi Divya, Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan, Yen-Wei Chen 0001
IEEE J. Biomed. Health Informatics4
2024 Towards a Transitional Weather Scene Recognition Approach for Autonomous Vehicles
abstract
Driving in adverse weather conditions is a key challenge for autonomous vehicles (AV). Typical scene perception models perform poorly in rainy, foggy, snowy, and cloudy conditions. In addition, we observe transition states between extremes (cloudy to rainy, rainy to sunny, etc.) in nature with variations in adversity. It is crucial to define and understand these transition states in order to develop robust AV perception models. Existing research works on classification focused on identifying extreme weather conditions. However, there is a lack of emphasis on the transition between these extreme weather scenes. Hence, this paper proposes an approach to define and understand six intermediate weather transition states: sunny to rainy, rainy to sunny, and others. Firstly, we propose a way to interpolate the intermediate weather transition data using a variational autoencoder and extract its spatial features using VGG. Further, we model the temporal distribution of these spatial features using a gated recurrent unit to classify the corresponding transition state. Also, we introduce a large-scale dataset called the AIWD6: Adverse Intermediate Weather Driving dataset, generated for three different time intervals. Experimental results on the AIWD6 dataset demonstrate that our model efficiently generates weather transition conditions for AV technology. Also, the spatio-temporal deep neural network can effectively classify the adverse weather transition states for different time intervals.
Madhavi Kondapally, K. Naveen Kumar, Chalavadi Vishnu, C. Krishna Mohan
IEEE Trans. Intell. Transp. Syst.4
2023 Multi-class object classification using deep learning models in automotive object detection scenarios
abstract
This paper presents two deep learning models using a multi-perspective convolutional neural network (CNN) for classifying objects in the context of intelligent transportation systems (ITS). The proposed model categorizes objects accurately, enabling them to make well-informed decisions in multi-object (such as Persons, Trucks, Motorbikes, Cars, and Cyclists.) detection in complex scenarios for automotive applications. The custom backbone model is designed based on experimentation with the VGG backbone network based on the VGG backbone network, incorporating a multilayer prediction head and custom feature extraction blocks for classifying multiple objects in complex scenes. The model is to extract abstract features and features at multiple scales with a custom-designed feature extraction backbone with multiple blocks. The proposed models are lightweight and require fewer computational resources for high classification performance. The automotive publicly available dataset with 19800 images and labels has been used. Results show that when we experimented with the VGG backbone CNN model, the classification accuracy of 99.64% is achieved, and on the other hand, the classification accuracy of custom backbone CNN is 99.46%. The performance of the proposed custom model is also compared to those of pre-trained benchmark models. The experimental findings presented in this paper show that the proposed models achieve higher accuracy than the pre-trained models.
Soumya Abbu, Linga Reddy Cenkeramaddi, Chalavadi Vishnu, C. Krishna Mohan
ICMV4
2023 False positive elimination in object detection methods for videos
abstract
A robust object detection algorithm is essential while detecting objects in videos and real time scenarios, where false positives might result in unwanted outcomes. Our goal here is to observe how Simple Online and Real-time Tracking with a Deep association metric (Deep SORT) algorithm for Multi-Object Tracking (MOT) can be used to minimize false positives, from a state of the art detection algorithm like You Only Look Once (YOLO), by using the Kalman filter approach. An auto encoder based feature extractor has been used, instead of the standard CNN networks like ResNet-50 to further improve speed of the detector. There have been other MOT algorithms in the recent times which give good results, but are not as real time efficient as the simple yet efficient Deep SORT method. Experimental analysis has shown how Autoencoder based Deep SORT performs in contrast to native Deep SORT and YOLO, in eliminating false positive detections.
Shubham Kumar Dubey, J. V. Satyanarayana, C. Krishna Mohan
ICMV3
2023 A data parallel approach for distributed neural networks to achieve faster convergence
abstract
The availability of large datasets has significantly contributed to recent advancements in deep Convolutional Neural Network (CNN) models. However, training a large CNN model using such datasets is a time-consuming task. This issue has been addressed by the parallelization and distribution of data/model during the training process. There are two ways to implement distributed deep learning processes: data parallelism and model parallelism. Data parallelism involves distributing the dataset across multiple workers, allowing them to process different portions simultaneously. While increasing the number of workers can reduce computation time, it also introduces additional communication time. In some cases, the increased communication time can outweigh the benefits gained from reduced computation time. In this paper, our focus is on reducing the overall computation time of data parallel approach by employing two strategies. First, we emphasize the preservation of dataset distribution across all workers, ensuring that each worker has access to representative data. Second, we explore the localization of parameters and the quantization of gradients to three levels: {-1, 0, 1} to reduce communication delays between the server and workers, as well as between workers themselves. By adopting these two strategies, we aim to enhance the performance of data parallel approach in the distributed deep learning processes. As a result of preserving the distribution of the data while sampling the entire data, each partition retains a similar mean and variance (capturing important first and second-order statistics). This approach guarantees that all worker machines train their local models on uniformly distributed data instead of random distribution. Additionally, localizing parameters limits the communication between the server and workers to gradients only. Furthermore, by quantizing gradients to 2-bits, we successfully achieve our objective of reducing computation time by enabling faster convergence without compromising test or validation accuracy. The experimental results demonstrate that employing these strategies in distributed deep learning effectively reduces communication overhead and leads to faster convergence when compared to methods that utilize random data sampling. These improvements were observed across multiple datasets such as MNIST, CIFAR-10, and Tiny ImageNet.
C. Nagaraju, Yenda Ramesh, C. Krishna Mohan
ICMV3
2023 FLWGAN: Federated Learning with Wasserstein Generative Adversarial Network for Brain Tumor Segmentation
abstract
Recently, the potential of deep learning in identifying complex patterns is gaining research interest in medical applications specifically for brain tumor diagnosis. To segment tumors accurately in brain MRIs, there is a need for a large amount of data for training deep learning models. Also, hospitals cannot share patient data for centralization on the server since health records are prone to privacy and ownership challenges. To deal with these challenges, we set up an efficient federated learning (FL) pipeline with Wasserstein generative adversarial networks (FLWGAN) to ensure data privacy and data sufficiency. FL preserves the data privacy of clients by sharing only the trained model parameters to a centralized server instead of raw data. A modified 3D Wasserstein generative adversarial network with gradient penalty (WGAN-GP) and is incorporated at the client side to generate image-segmentation pairs for efficient training segmentation models. Here, 3D-UNet with an attention module is used for the brain MRI segmentation. The attention module is integrated into a 3D-UNet encoder network for effective brain tumor segmentation. Our approach aims to allow each client to benefit from locally available real data and synthetic data. This process enhances the learning performance while respecting data privacy. The efficacy of our proposed pipeline is demonstrated on the brain tumor task of the medical segmentation decathlon (MSD) dataset. We designed FLWGAN frameworks for predicting four segmentation tasks, i.e., whole tumor (WT), enhanced tumor (ET), tumor core (TC), and multiclass. Our proposed approach achieves state of the art performance in terms of various segmentation metrics.
Peketi Divya, Chalavadi Vishnu, C. Krishna Mohan, Yen-Wei Chen 0001
IJCNN3
2023 MADG: Margin-based Adversarial Learning for Domain Generalization
abstract
Domain Generalization (DG) techniques have emerged as a popular approach to address the challenges of domain shift in Deep Learning (DL), with the goal of generalizing well to the target domain unseen during the training. In recent years, numerous methods have been proposed to address the DG setting, among which one popular approach is the adversarial learning-based methodology. The main idea behind adversarial DG methods is to learn domain-invariant features by minimizing a discrepancy metric. However, most adversarial DG methods use 0-1 loss based $\mathcal{H}\Delta\mathcal{H}$ divergence metric. In contrast, the margin loss-based discrepancy metric has the following advantages: more informative, tighter, practical, and efficiently optimizable. To mitigate this gap, this work proposes a novel adversarial learning DG algorithm, $\textbf{MADG}$, motivated by a margin loss-based discrepancy metric. The proposed $\textbf{MADG}$ model learns domain-invariant features across all source domains and uses adversarial training to generalize well to the unseen target domain. We also provide a theoretical analysis of the proposed $\textbf{MADG}$ model based on the unseen target error bound. Specifically, we construct the link between the source and unseen domains in the real-valued hypothesis space and derive the generalization bound using margin loss and Rademacher complexity. We extensively experiment with the $\textbf{MADG}$ model on popular real-world DG datasets, VLCS, PACS, OfficeHome, DomainNet, and TerraIncognita. We evaluate the proposed algorithm on DomainBed's benchmark and observe consistent performance across all the datasets.
Aveen Dayal, Vimal KB, Linga Reddy Cenkeramaddi, C. Krishna Mohan, Abhinav Kumar 0001, Vineeth N. Balasubramanian
NeurIPS4
2023 FONN: Federated Optimization with Nys-Newton
abstract
Federated optimization or federated learning (FL) involves optimization of the global model or the server model by minimizing the global loss function which is weighted average of all the local loss functions. The optimization of the global model requires faster convergence to reduce the number of communication rounds or global iterations which is one of the major challenge in federated optimization. This paper propose FONN which handles this communication overhead in federated optimization by utilizing Nys-Newton, while updating local models. As compared to existing state-of-the-art FL algorithms, SCAFFOLD, GIANT and DONE, utilization of Nys-Newton leads to better convergence and reduction in communication rounds or global iterations while achieving a desired performance from the global model which may be observed from the experimental results on various heterogeneously partitioned datasets.
C. Nagaraju, Mrinmay Sen, C. Krishna Mohan
TENCON3
2023 Incorporating attentive multi-scale context information for image captioning
Jeripothula Prudviraj, Sravani Yenduri, C. Krishna Mohan
Multim. Tools Appl.3
2023 EVAA - Exchange Vanishing Adversarial Attack on LiDAR Point Clouds in Autonomous Vehicles
abstract
In addition to RGB camera sensors, LiDAR (Light Detection and Ranging) plays an important role in autonomous vehicles (AVs) to perceive their surroundings. Deep neural networks (DNNs) are able to achieve cutting-edge 3D object detection and segmentation performance using LiDAR point clouds. LiDAR-enabled autonomous vehicles provide human perception by segmenting LiDAR point clouds into meaningful regions and providing semantic context to the AV user. However, the generation of point clouds to provide semantic segmentation in AVs is not reliable and secure, which may result in traffic accidents. We propose a novel adversarial attack against LiDAR point clouds in autonomous vehicles in this paper. We devised an exchange vanishing adversarial attack (EVAA) to deceive LiDAR point clouds by introducing targeted noise on specific objects (e.g., vehicles and driveways). On two autonomous driving datasets with 3D object annotations, NuScenes and PandaSet, we evaluate the performance of our proposed attack framework. We achieve an attack success rate (ASR) of ≈63% and ASR of ≈29% on both NuScenes and PandaSet datasets, respectively.
Chalavadi Vishnu, Jayesh Khandelwal, C. Krishna Mohan, Linga Reddy Cenkeramaddi
IEEE Trans. Geosci. Remote. Sens.3
2022 Structural representative network for remote sensing image captioning
abstract
Current encoder-decoder methods for remote sensing image captioning (RSIC) avoids fine-grained structural representation of objects due to the lack of prominent encoding frameworks. This paper proposes a novel structural representative network (SRN) for acquiring fine-grained structures of remote sensing images (RSI) for generating semantically meaningful captions. Initially, we employ SRN on top of the final layers of the convolutional neural network (CNN) for attaining the spatially transformed RSI features. A multi-stage decoder is incorporated into the extracted features of SRN to produce fine-grained meaningful captions. The efficacy of our proposed methodology is exhibited on two RSIC datasets, i.e Sydney-Captions dataset, and the UCM-Captions dataset.
Jaya Sharma, Peketi Divya, Sravani Yenduri, B. H. Shekar, C. Krishna Mohan
ICMV5
2022 Performance analysis of deep neural networks for COVID-19 detection from chest radiographs
abstract
Contrary to the World Health Organization’s (WHO) and the medical community’s projections, Covid-19, which started in Wuhan, China, in December 2019, still doesn’t show any signs of progressing to the endemic stage or slowing down any time soon. It continues to wreak havoc on the lives and livelihood of thousands of people every day. There is general agreement that the best way to contain this dangerous virus is through testing and isolation. Therefore, in these epidemic times, developing an automated Covid-19 detection method is of utmost importance. This study uses three different Machine Learning classifiers, such as Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression (LR), along with five Transfer Learning models such as DenseNet121, DenseNet169, ResNet50, ResNet152V2, and Xception as feature extraction methods for identifying Covid-19. Five different datasets are used to assess the models’ performance to generalize. There are encouraging findings, with the best one being the combination of DenseNet121 and DenseNet169 together with SVM and LR.
B. H. Shekar, Shazia Mannan, Habtu Hailu, C. Krishna Mohan, Linga Reddy Cenkeramaddi
ICMV4
2022 Adaptive spatial and temporal aggregation for table tennis shot recognition
abstract
Action recognition is one of the challenging video understanding tasks in computer vision. Although there has been extensive research in the task of classifying coarse-grained actions, existing methods are still limited in differentiating actions with low inter-class and high intra-class variation. Particularly, the table tennis sport that involves shots of high inter-class similarity, subtle variations, occlusion, and view-point variations. While a few datasets have been available for event spotting and shot recognition, these benchmarks are mostly recorded in a constrained environment with a clear view/perception of shots executed by players. In this paper, we introduce a Table tennis shots 1.0 dataset consisting of 9000 videos of 6 fine-grained actions collected in an unconstrained manner to analyze the performance of both players. To effectively recognise these different types of table tennis shots, we propose an adaptive spatial and temporal aggregation method that can handle the spatial and temporal interactions concerning the subtle variations among shots and low inter-class variations. Our method consists of three components, namely, (i) feature extraction module, (ii) spatial aggregation network, and (iii) temporal aggregation network. The feature extraction module is a 3D convolutional neural network (3D-CNN) that captures the spatial and temporal characteristics of table tennis shots. In order to capture the interaction among the elements of the extracted 3D-CNN feature maps efficiently, we employ spatial aggregation network to obtain the compact spatial representation. Later, we propose to replace the final global average pooling layer (GAP) with the temporal aggregation network to overcome the loss of motion information due to averaging of temporal features. This temporal aggregation network utilizes the attention mechanism of bidirectional encoder representations from Transformers (BERT) to model the significant temporal interactions among the shots effectively. We demonstrate that our proposed approach improves the performance of existing 3D-CNN methods by ~10% on the Table tennis shots 1.0 dataset.We also show the performance of our approach on other action recognition datasets, namely, UCF-101 and HMDB-51.
Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan
ICMV3
2022 STIP-GCN: Space-time interest points graph convolutional network for action recognition
abstract
Action recognition requires modelling the interactions between either human & human or human & objects. Re-cently, graph convolutional neural networks (GCNs) are exploited to effectively capture the structure of action by modelling the relationship among entities present in a video. However, most of the approaches depend on the effectiveness of object detection frameworks to detect the entities. In this paper, we propose a graph-based framework for action recognition to model the spatio-temporal interactions among the entities in a video without any object-level supervision. First, we obtain the salient space-time interest points (STIP) that contain rich information about the significant local variations in space and time by using the Harris 3D detector. In order to incorporate the local appearance and motion information of the entities, either low-level or deep features are extracted around these STIPs. Next, we build a graph by considering the extracted STIPs as nodes and are connected by spatial edges and temporal edges. These edges are determined based on a membership function that measures the similarity of entities associated with the STIPs. Finally, GCN is employed on the given graph to provide reasoning among different entities present in a video. We evaluate our method on three widely used datasets, namely, UCF-101, HMDB-51, SSV2 to demonstrate the efficacy of the proposed approach.
Sravani Yenduri, Chalavadi Vishnu, C. Krishna Mohan
IJCNN3
2022 M-FFN: multi-scale feature fusion network for image captioning
Jeripothula Prudviraj, Chalavadi Vishnu, C. Krishna Mohan
Appl. Intell.3
2022 ClarifyNet: A high-pass and low-pass filtering based CNN for single image dehazing
Onkar Susladkar, Gayatri Deshmukh, Subhrajit Nag, Ananya Mantravadi, Dhruv Makwana, Sujitha Ravichandran, R. Sai Chandra Teja, Gajanan H. Chavhan, C. Krishna Mohan, Sparsh Mittal
J. Syst. Archit.9
2022 mSODANet: A network for multi-scale object detection in aerial images using hierarchical dilated convolutions
Chalavadi Vishnu, Jeripothula Prudviraj, Rajeshreddy Datla, Sobhan Babu Chintapalli, C. Krishna Mohan
Pattern Recognit.5
2022 Fine-grained action recognition using dynamic kernels
Sravani Yenduri, Nazil Perveen, Chalavadi Vishnu, C. Krishna Mohan
Pattern Recognit.4
2022 AAP-MIT: Attentive Atrous Pyramid Network and Memory Incorporated Transformer for Multisentence Video Description
abstract
Generating multi-sentence descriptions for video is considered to be the most complex task in computer vision and natural language understanding due to the intricate nature of video-text data. With the recent advances in deep learning approaches, the multi-sentence video description has achieved an impressive progress. However, learning rich temporal context representation of visual sequences and modelling long-term dependencies of natural language descriptions is still a challenging problem. Towards this goal, we propose an Attentive Atrous Pyramid network and Memory Incorporated Transformer (AAP-MIT) for multi-sentence video description. The proposed AAP-MIT incorporates the effective representation of visual scene by distilling the most informative and discriminative spatio-temporal features of video data at multiple granularities and further generates the highly summarized descriptions. Profoundly, we construct AAP-MIT with three major components: i) a temporal pyramid network, which builds the temporal feature hierarchy at multiple scales by convolving the local features at temporal space, ii) a temporal correlation attention to learn the relations among various temporal video segments, and iii) the memory incorporated transformer, which augments the new memory block in language transformer to generate highly descriptive natural language sentences. Finally, the extensive experiments on ActivityNet Captions and YouCookII datasets demonstrate the substantial superiority of AAP-MIT over the existing approaches.
Jeripothula Prudviraj, Malipatel Indrakaran Reddy, Chalavadi Vishnu, C. Krishna Mohan
IEEE Trans. Image Process.4
2022 Detection of Collision-Prone Vehicle Behavior at Intersections Using Siamese Interaction LSTM
abstract
As a large proportion of road accidents occur at intersections, monitoring traffic safety of intersections is important. Existing approaches are designed to investigate accidents in lane-based traffic. However, such approaches are not suitable in a lane-less mixed-traffic environment where vehicles often ply very close to each other. Hence, we propose an approach called Siamese Interaction Long Short-Term Memory network (SILSTM) to detect collision prone vehicle behavior. The SILSTM network learns the interaction trajectory of a vehicle that describes the interactions of a vehicle with its neighbors at an intersection. Among the hundreds of interactions for every vehicle, there maybe only some interactions that may be unsafe, and hence, a temporal attention layer is used in the SILSTM network. Furthermore, the comparison of interaction trajectories requires labeling the trajectories as either unsafe or safe, but such a distinction is highly subjective, especially in lane-less traffic. Hence, in this work, we compute the characteristics of interaction trajectories involved in accidents using the collision energy model. The interaction trajectories that match accident characteristics are labeled as unsafe while the rest are considered safe. Finally, there is no existing dataset that allows us to monitor a particular intersection for a long duration. Therefore, we introduce the SkyEye dataset that contains 1 hour of continuous aerial footage from each of the 4 chosen intersections in the city of Ahmedabad in India. A detailed evaluation of SILSTM on the SkyEye dataset shows that unsafe (collision-prone) interaction trajectories can be effectively detected at different intersections.
Debaditya Roy, Tetsuhiro Ishizaka, C. Krishna Mohan, Atsushi Fukuda
IEEE Trans. Intell. Transp. Syst.3
2021 A multimodal semantic segmentation for airport runway delineation in panchromatic remote sensing images
abstract
Monitoring airport runways in panchromatic remote sensing images is helpful for both civil and strategic communities in effective utilization of the large-area acquisitions. This paper proposes a novel multimodal semantic segmentation approach for effective delineation of the runways in panchromatic remote sensing images. The proposed approach aims to learn complementary information from two modalities, namely, panchromatic image and digital elevation model (DEM) to obtain discriminative features of the runway. The fusion of image features and the corresponding terrain information is performed by stacking the image and DEM by leveraging the merits of both Transformers and U-Net architecture. We perform the experiments on Cartosat-1 panchromatic satellite images with the corresponding Cartosat-1 DEM scenes. The experimental results demonstrate a significant contribution of terrain information to the segmentation process in achieving the contours of airport runways effectively.
Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
ICMV3
2021 A framework to derive geospatial attributes for aircraft type recognition in large-scale remote sensing images
abstract
Aircraft type recognition remains challenging, due to their tiny sizes and geometric distortions in large-scale panchromatic satellite images. This paper proposes a framework for aircraft type recognition by focusing on shape preservation, spatial transformations, and geospatial attributes derivation. First, we construct an aircraft segmentation model to obtain masks representing the shape of aircrafts by employing a learnable shape-preserved and deformable network in the mask RCNN architecture. Then, the orientation of the segmented aircrafts is determined by estimating the symmetrical axes using their gradient information. Besides template matching, we derive the length and width of aircrafts using the geotagged information of images to further categorize the types of aircrafts. Also, we present an effective inferencing mechanism to overcome the issue of partial detection or missing aircrafts in large-scale images. The efficacy of the proposed framework is demonstrated on large-scale panchromatic images with ground sampling distances of 0.65m (C2S).
Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
ICMV3
2021 Scene Classification in Remote Sensing Images using Dynamic Kernels
abstract
Classification of scenes across multi-sensor remote sensing images with different spatial, spectral, temporal resolutions involves identification of variable length spatial patterns of objects in a scene. So, it necessitates the use of local representations from different regions of a scene in order to comprehend the scene formation. In this paper, we propose a dynamic kernel based representation to handle the patterns of variable lengths in the scenes of remote sensing images. These kernels help to assimilate spatial variability captured using convolutional features in a Gaussian mixture model. The statistics of GMM facilitate the dynamic kernels in preserving the local spatial similarities while handling the changes in spatial content globally within the same scene. The efficacy of the proposed method using two variants of the dynamic kernels is demonstrated on three benchmark scene classification datasets, namely, UCM Land Use (21 classes), Aerial image dataset (30 classes), and NWPU-RESISC45 (45 classes). Our experiments show that the mean interval kernel is better discriminative as it makes use of first and second-order statistics of GMM.
Rajeshreddy Datla, Chalavadi Vishnu, C. Krishna Mohan
IJCNN3
2021 Attentive Contextual Network for Image Captioning
abstract
Existing image captioning approaches fail to generate fine-grained captions due to the lack of rich encoding representation of an image. In this paper, we present an attentive contextual network (ACN) to learn the spatially transformed image features and dense multi-scale contextual information of an image to generate semantically meaningful captions. At first, we construct deformable network on intermediate layers of convolutional neural network (CNN) to cultivate spatial invariant features. And the multi-scale contextual features are produced by employing contextual network on top of last layers of CNN. Then, we exploit attention mechanism on contextual network to extract dense contextual features. Further, the extracted spatial and contextual features are combined to encode the holistic representation of an image. Finally, a multi-stage caption decoder with visual attention module is incorporated to generate fine-grained captions. The performance of the proposed approach is demonstrated on COCO dataset, the largest dataset for image captioning.
Jeripothula Prudviraj, Chalavadi Vishnu, C. Krishna Mohan
IJCNN3
2020 ULSAM: Ultra-Lightweight Subspace Attention Module for Compact Convolutional Neural Networks
abstract
The capability of the self-attention mechanism to model the long-range dependencies has catapulted its deployment in vision models. Unlike convolution operators, self-attention offers infinite receptive field and enables compute- efficient modeling of global dependencies. However, the existing state-of-the-art attention mechanisms incur high compute and/or parameter overheads, and hence unfit for compact convolutional neural networks (CNNs). In this work, we propose a simple yet effective "Ultra-Lightweight Subspace Attention Mechanism" (ULSAM), which infers different attention maps for each feature map subspace. We argue that leaning separate attention maps for each feature subspace enables multi-scale and multi-frequency feature representation, which is more desirable for fine-grained image classification. Our method of subspace attention is orthogonal and complementary to the existing state-of-the- arts attention mechanisms used in vision models. ULSAM is end-to-end trainable and can be deployed as a plug-and- play module in the pre-existing compact CNNs. Notably, our work is the first attempt that uses a subspace attention mechanism to increase the efficiency of compact CNNs. To show the efficacy of ULSAM, we perform experiments with MobileNet-V1 and MobileNet-V2 as backbone architectures on ImageNet-1K and three fine-grained image classification datasets. We achieve ≈13% and ≈25% reduction in both the FLOPs and parameter counts of MobileNet-V2 with a 0.27% and more than 1% improvement in top-1 accuracy on the ImageNet-1K and fine-grained image classification datasets (respectively). Code and trained models are available at https://github.com/Nandan91/ULSAM.
Rajat Saini, Nandan Kumar Jha, Bedanta Das, Sparsh Mittal, C. Krishna Mohan
WACV5
2020 Facial Expression Recognition in Videos Using Dynamic Kernels
abstract
Recognition of facial expressions across various actors, contexts, and recording conditions in real-world videos involves identifying local facial movements. Hence, it is important to discover the formation of expressions from local representations captured from different parts of the face. So in this paper, we propose a dynamic kernel-based representation for facial expressions that assimilates facial movements captured using local spatio-temporal representations in a large universal Gaussian mixture model (uGMM). These dynamic kernels are used to preserve local similarities while handling global context changes for the same expression by utilizing the statistics of uGMM. We demonstrate the efficacy of dynamic kernel representation using three different dynamic kernels, namely, explicit mapping based, probability-based, and matching-based, on three standard facial expression datasets, namely, MMI, AFEW, and BP4D. Our evaluations show that probability-based kernels are the most discriminative among the dynamic kernels. However, in terms of computational complexity, intermediate matching kernels are more efficient as compared to the other two representations.
Nazil Perveen, Debaditya Roy, C. Krishna Mohan
IEEE Trans. Image Process.3
2020 Echocardiogram Analysis Using Motion Profile Modeling
abstract
Echocardiography is a widely used and cost-effective medical imaging procedure that is used to diagnose cardiac irregularities. To capture the various chambers of the heart, echocardiography videos are captured from different angles called views to generate standard images/videos. Automatic classification of these views allows for faster diagnosis and analysis. In this work, we propose a representation for echo videos which encapsulates the motion profile of various chambers and valves that helps effective view classification. This variety of motion profiles is captured in a large Gaussian mixture model called universal motion profile model (UMPM). In order to extract only the relevant motion profiles for each view, a factor analysis based decomposition is applied to the means of the UMPM. This results in a low-dimensional representation called motion profile vector (MPV) which captures the distinctive motion signature for a particular view. To evaluate MPVs, a dataset called ECHO 1.0 is introduced which contains around 637 video clips of the four major views: a) parasternal long-axis view (PLAX), b) parasternal short-axis (PSAX), c) apical four-chamber view (A4C), and d) apical two-chamber view (A2C). We demonstrate the efficacy of motion profile-vectors over other spatio-temporal representations. Further, motion profile-vectors can classify even poorly captured videos with high accuracy which shows the robustness of the proposed representation.
Inayathullah Ghori, Debaditya Roy, Renu John, C. Krishna Mohan
IEEE Trans. Medical Imaging4
2019 Deep Spatio-Temporal Representation for Detection of Road Accidents Using Stacked Autoencoder
abstract
Vision-based detection of road accidents using traffic surveillance video is a highly desirable but challenging task. In this paper, we propose a novel framework for automatic detection of road accidents in surveillance videos. The proposed framework automatically learns feature representation from the spatiotemporal volumes of raw pixel intensity instead of traditional hand-crafted features. We consider the accident of the vehicles as an unusual incident. The proposed framework extracts deep representation using denoising autoencoders trained over the normal traffic videos. The possibility of an accident is determined based on the reconstruction error and the likelihood of the deep representation. For the likelihood of the deep representation, an unsupervised model is trained using one class support vector machine. Also, the intersection points of the vehicle's trajectories are used to reduce the false alarm rate and increase the reliability of the overall system. We evaluated out proposed approach on real accident videos collected from the CCTV surveillance network of Hyderabad City in India. The experiments on these real accident videos demonstrate the efficacy of the proposed approach.
Dinesh Singh 0001, C. Krishna Mohan
IEEE Trans. Intell. Transp. Syst.2
2019 Unsupervised Universal Attribute Modeling for Action Recognition
abstract
A fixed dimensional representation for action clips of varying lengths has been proposed in the literature using aggregation models like bag-of-words and Fisher vector. These representations are high dimensional and require classification techniques for action recognition. In this paper, we propose a framework for unsupervised extraction of a discriminative low-dimensional representation called action-vector. To start with, local spatio-temporal features are utilized to capture the action attributes implicitly in a large Gaussian mixture model called the universal attribute model (UAM). To enhance the contribution of the significant attributes in each action clip, a maximum aposteriori adaptation of the UAM means is performed for each clip. This results in a concatenated mean vector called super action vector (SAV) for each action clip. However, the SAV is still high dimensional because of the presence of redundant attributes. Hence, we employ factor analysis to represent every SAV only in terms of the few important attributes contributing to the action clip. This leads to a low-dimensional representation called action-vector. This entire procedure requires no class labels and produces action-vectors that are distinct representations of each action irrespective of the inter-actor variability encountered in unconstrained videos. An evaluation on trimmed action datasets UCF101 and HMDB51 demonstrates the efficacy of action-vectors for action classification over state-of-the-art techniques. Moreover, we also show that action-vectors can adequately represent untrimmed videos from the THUMOS14 dataset and produce classification results comparable to existing techniques.
Debaditya Roy, K. Sri Rama Murty, C. Krishna Mohan
IEEE Trans. Multim.3
2018 Projection-SVM: Distributed Kernel Support Vector Machine for Big Data using Subspace Partitioning
abstract
The training of kernel support vector machine (SVM) is a computationally complex task for large datasets where the number of samples ranges in millions. This is because kernel matrix (in general not sparse) is both computation expensive and memory intensive. Existing methods hardly achieve a linear scale and suffer from high approximation loss. We propose Projection-SVM, a distributed implementation of kernel support vector machine for large datasets using subspace partitioning. In subspace partitioning, a decision tree is constructed on projection of data along the direction of maximum variance (i.e., dominant eigenvector) to obtain smaller partitions (i.e., subspaces) of the dataset. On each of these partitions, a kernel SVM is trained independently over a cluster thereby reducing the overall training time. Also, it results in reducing the prediction time significantly. We demonstrate the efficacy of the proposed approach on eight standard large datasets from various application domains, namely, mnist8m, kddcup99, webspam, etc. where Projection-SVM is on an average 150 times faster than sequential SVM while maintaining the classification accuracy. The experimental results also show the superiority of the Projection-SVM over the state-of-the-art approaches for distributed kernel SVMs, such as DCSVM, CASVM, and DTSVM.
Dinesh Singh 0001, C. Krishna Mohan
IEEE BigData2
2018 Fast-BoW: Scaling Bag-of-Visual-Words Generation
Dinesh Singh 0001, Abhijeet Bhure, Sumit Mamtani, C. Krishna Mohan
BMVC4
2018 Action Recognition Based on Discriminative Embedding of Actions Using Siamese Networks
abstract
Actions can be recognized effectively when the various atomic attributes forming the action are identified and combined in the form of a representation. In this paper, a low-dimensional representation is extracted from a pool of attributes learned in a universal Gaussian mixture model using factor analysis. However, such a representation cannot adequately discriminate between actions with similar attributes. Hence, we propose to classify such actions by leveraging the corresponding class labels. We train a Siamese deep neural network with a contrastive loss on the low-dimensional representation. We show that Siamese networks allow effective discrimination even between similar actions. The efficacy of the proposed approach is demonstrated on two benchmark action datasets, HMDB51 and MPII Cooking Activities. On both the datasets, the proposed method improves the state-of-the-art performance considerably.
Debaditya Roy, C. Krishna Mohan, K. Sri Rama Murty
ICIP2
2018 Snatch theft detection in unconstrained surveillance videos using action attribute modelling
Debaditya Roy, C. Krishna Mohan
Pattern Recognit. Lett.2
2018 Spontaneous Expression Recognition Using Universal Attribute Model
abstract
Spontaneous expression recognition refers to recognizing non-posed human expressions. In literature, most of the existing approaches for expression recognition mainly rely on manual annotations by experts, which is both time-consuming and difficult to obtain. Hence, we propose an unsupervised framework for spontaneous expression recognition that preserves discriminative information for the videos of each expression without using annotations. Initially, a large Gaussian mixture model called universal attribute model (UAM) is trained to learn the attributes of various expressions implicitly. Attributes are the movements of various facial muscles that are combined to form a particular facial expression. Then a concatenated mean vector called the super expression-vector (SEV) is formed by using a maximum a posteriori adaptation of the UAM means for each expression clip. This SEV contains attributes from all the expressions resulting in a high dimensional representation. To retain only the attributes of that particular expression clip, the SEV is decomposed using factor analysis to produce a low-dimensional expression-vector. This procedure does not require any class labels and produces expression-vectors that are distinct for each expression irrespective of high inter-actor variability present in spontaneous expressions. On spontaneous expression datasets like BP4D and AFEW, we demonstrate that expression-vector achieves better performance than state-of-the-art techniques. Further, we also show that UAM trained on a constrained dataset can be effectively used to recognize expressions in unconstrained expression videos.
Nazil Perveen, Debaditya Roy, C. Krishna Mohan
IEEE Trans. Image Process.3
2018 An Information Bottleneck Approach to Optimize the Dictionary of Visual Data
abstract
In this paper, we propose a novel information theoretic approach to obtain compact and discriminative dictionary of visual data. This approach squeezes discriminative information from the dictionary for efficient representation using information bottleneck. The dictionary is optimized from the initial sparse dictionary, which is learned from action data. In this, a constraint information optimization problem is formulated in which mutual information between the initial and optimized dictionary is minimized while maximizing mutual information between optimized dictionary and class labels. We use an effective similarity measure, Jensen-Shannon divergence with adaptive weightages, for class distributions of each dictionary atom. These adaptive weightages are obtained based on the usage of the dictionary atom among different classes. The resultant dictionary becomes discriminative and compact, while retaining maximum information with fewer atoms. Using simple reconstruction error, we test computational efficiency of the proposed method without compromising classification accuracy on popular benchmark datasets. It is further demonstrated how efficiently discriminative information is retained by comparing the classification performance of the dictionary before and after the removal of redundant dictionary atoms.
Shyju Wilson, C. Krishna Mohan
IEEE Trans. Multim.2
2017 Action-vectors: Unsupervised movement modeling for action recognition
abstract
Representation and modelling of movements play a significant role in recognising actions in unconstrained videos. However, explicit segmentation and labelling of movements are non-trivial because of the variability associated with actors, camera viewpoints, duration etc. Therefore, we propose to train a GMM with a large number of components termed as a universal movement model (UMM). This UMM is trained using motion boundary histograms (MBH) which capture the motion trajectories associated with the movements across all possible actions. For a particular action video, the MAP adapted mean vectors of the UMM are concatenated to form a fixed dimensional representation referred to as “super movement vector” (SMV). However, SMV is still high dimensional and hence, Baum-Welch statistics extracted from the UMM are used to arrive at a compact representation for each action video, which we refer to as an “action-vector”. It is shown that even without the use of class labels, action-vectors provide a more discriminatory representation of action classes translating to a 8 % relative improvement in classification accuracy for action-vectors based on MBH features over naïve MBH features on the UCF101 dataset. Furthermore, action-vectors projected with LDA achieve 93% accuracy on the UCF101 dataset which rivals state-of-the-art deep learning techniques.
Debaditya Roy, K. Sri Rama Murty, C. Krishna Mohan
ICASSP3
2017 Detection of motorcyclists without helmet in videos using convolutional neural network
abstract
In order to ensure the safety measures, the detection of traffic rule violators is a highly desirable but challenging task due to various difficulties such as occlusion, illumination, poor quality of surveillance video, varying whether conditions, etc. In this paper, we present a framework for automatic detection of motorcyclists driving without helmets in surveillance videos. In the proposed approach, first we use adaptive background subtraction on video frames to get moving objects. Later convolutional neural network (CNN) is used to select motorcyclists among the moving objects. Again, we apply CNN on upper one fourth part for further recognition of motorcyclists driving without a helmet. The performance of the proposed approach is evaluated on two datasets, IITH_Helmet_1 contains sparse traffic and IITH_Helmet_2 contains dense traffic, respectively. The experiments on real videos successfully detect 92.87% violators with a low false alarm rate of 0.5% on an average and thus shows the efficacy of the proposed approach.
Chalavadi Vishnu, Dinesh Singh 0001, C. Krishna Mohan, Sobhan Babu Chintapalli
IJCNN3
2017 Human action recognition in RGB-D videos using motion sequence information and deep learning
Earnest Paul Ijjina, C. Krishna Mohan
Pattern Recognit.2
2017 Graph formulation of video activities for abnormal activity recognition
Dinesh Singh 0001, C. Krishna Mohan
Pattern Recognit.2
2017 Coherent and Noncoherent Dictionaries for Action Recognition
abstract
In this letter, we propose sparsity-based coherent and noncoherent dictionaries for action recognition. First, the input data are divided into different clusters and the number of clusters depends on the number of action categories. Within each cluster, we seek data items of each action category. If the number of data items exceeds threshold in any action category, these items are labeled as coherent. In a similar way, all coherent data items from different clusters form a coherent group of each action category, and data that are not part of the coherent group belong to noncoherent group of each action category. These coherent and noncoherent groups are learned using K-singular value decomposition dictionary learning. Since the coherent group has more similarity among data, only few atoms need to be learned. In the noncoherent group, there is a high variability among the data items. So, we propose an orthogonal-projection-based selection to get optimal dictionary in order to retain maximum variance in the data. Finally, the obtained dictionary atoms of both groups in each action category are combined and then updated using the limited Broyden-Fletcher-Goldfarb-Shanno optimization algorithm. The experiments are conducted on challenging datasets HMDB51 and UCF50 with action bank features and achieve comparable result using this state-of-the-art feature.
Shyju Wilson, C. Krishna Mohan
IEEE Signal Process. Lett.2
2017 DiP-SVM : Distribution Preserving Kernel Support Vector Machine for Big Data
abstract
In literature, the task of learning a support vector machine for large datasets has been performed by splitting the dataset into manageable sized “partitions” and training a sequential support vector machine on each of these partitions separately to obtain local support vectors. However, this process invariably leads to the loss in classification accuracy as global support vectors may not have been chosen as local support vectors in their respective partitions. We hypothesize that retaining the original distribution of the dataset in each of the partitions can help solve this issue. Hence, we present DiP-SVM, a distribution preserving kernel support vector machine where the first and second order statistics of the entire dataset are retained in each of the partitions. This helps in obtaining local decision boundaries which are in agreement with the global decision boundary, thereby reducing the chance of missing important global support vectors. We show that DiP-SVM achieves a minimal loss in classification accuracy among other distributed support vector machine techniques on several benchmark datasets. We further demonstrate that our approach reduces communication overhead between partitions leading to faster execution on large datasets and making it suitable for implementation in cloud environments.
Dinesh Singh 0001, Debaditya Roy, C. Krishna Mohan
IEEE Trans. Big Data3
2016 Distributed quadratic programming solver for kernel SVM using genetic algorithm
abstract
Support vector machine (SVM) is a powerful tool for classification and regression problems, however, its time and space complexities make it unsuitable for large datasets. In this paper, we present GeneticSVM, an evolutionary computing based distributed approach to find optimal solution of quadratic programming (QP) for kernel support vector machine. In Ge-neticSVM, novel encoding method and crossover operation help in obtaining the better solution. In order to train a SVM from large datasets, we distribute the training task over the graphics processing units (GPUs) enabled cluster. It leverages the benefit of the GPUs for large matrix multiplication. The experiments show better performance in terms of classification accuracy as well as computational time on standard datasets like GISETTE, ADULT, etc.
Dinesh Singh 0001, C. Krishna Mohan
CEC2
2016 Classification of medical images using edge-based features and sparse representation
abstract
In this paper, an approach for classification of medical images using edge-based features is proposed. We demonstrate that the edge information extracted from an image by dividing the image into patches and each patch into concentric circular regions provide discriminative information useful for classification of medical images by considering 18 categories of radiological medical images namely, skull, hand, breast, cranium, hip, cervical spin, pelvis, radiocarpaljoint, elbow etc.,. The ability of On-line Dictionary Learning (ODL) to achieve sparse representation of an image is exploited to develop dictionaries for each class using edge-based feature. A low rate of misclassification error for these test images validates the effectiveness of edge-based features and On-line Dictionary Learning models for classification of medical images.
C. Krishna Mohan
ICASSP2
2016 Discriminative feature extraction from X-ray images using deep convolutional neural networks
abstract
Feature extraction is one of the most important phases of medical image classification which requires extensive domain knowledge. Convolutional Neural Networks (CNN) have been successfully used for feature extraction in images from different domains involving a lot of classes. In this paper, CNNs are exploited to extract a hierarchical and discriminative representation of X-ray images. This representation is then used for classification of the X-ray images as various parts of the body. Visualization of the feature maps in the hidden layers show that features learnt by the CNN resemble the essential features which help discern the discrimination among different body parts. A comparison on the standard IRMA X-ray image dataset demonstrates that the CNNs easily outperform classifiers with hand-engineered features.
Debaditya Roy, C. Krishna Mohan
ICASSP3
2016 Spontaneous Facial Expression Recognition: A Part Based Approach
abstract
A part-based approach for spontaneous expression recognition using audio-visual feature and deep convolution neural network (DCNN) is proposed. The ability of convolution neural network to handle variations in translation and scale is exploited for extracting visual features. The sub-regions, namely, eye and mouth parts extracted from the video faces are given as an input to the deep CNN (DCNN) inorder to extract convnet features. The audio features, namely, voice-report, voice intensity, and other prosodic features are used to obtain complementary information useful for classification. The confidence scores of the classifier trained on different facial parts and audio information are combined using different fusion rules for recognizing expressions. The effectiveness of the proposed approach is demonstrated on acted facial expression in wild (AFEW) dataset.
Nazil Perveen, Dinesh Singh 0001, C. Krishna Mohan
ICMLA3
2016 Visual Big Data Analytics for Traffic Monitoring in Smart City
abstract
The application such as video surveillance for traffic control in smart cities needs to analyze the large amount (hours/days) of video footage in order to locate the people who are violating the traffic rules. The traditional computer vision techniques are unable to analyze such a huge amount of visual data generated in real-time. So, there is a need for visual big data analytics which involves processing and analyzing large scale visual data such as images or videos to find semantic patterns that are useful for interpretation. In this paper, we propose a framework for visual big data analytics for automatic detection of bike-riders without helmet in city traffic. We also discuss challenges involved in visual big data analytics for traffic control in a city scale surveillance data and explore opportunities for future research.
Dinesh Singh 0001, Chalavadi Vishnu, C. Krishna Mohan
ICMLA3
2016 Automatic detection of bike-riders without helmet using surveillance videos in real-time
abstract
In this paper, we propose an approach for automatic detection of bike-riders without helmet using surveillance videos in real time. The proposed approach first detects bike riders from surveillance video using background subtraction and object segmentation. Then it determines whether bike-rider is using a helmet or not using visual features and binary classifier. Also, we present a consolidation approach for violation reporting which helps in improving reliability of the proposed approach. In order to evaluate our approach, we have provided a performance comparison of three widely used feature representations namely histogram of oriented gradients (HOG), scale-invariant feature transform (SIFT), and local binary patterns (LBP) for classification. The experimental results show detection accuracy of 93.80% on the real world surveillance data. It has also been shown that proposed approach is computationally less expensive and performs in real-time with a processing time of 11.58 ms per frame.
Kunal Dahiya, Dinesh Singh 0001, C. Krishna Mohan
IJCNN3
2016 Human action recognition using genetic algorithms and convolutional neural networks
Earnest Paul Ijjina, C. Krishna Mohan
Pattern Recognit.2
2016 Sparsity-inducing dictionaries for effective action classification
Debaditya Roy, C. Krishna Mohan
Pattern Recognit.3
2016 Classification of human actions using pose-based features and stacked auto encoder
Earnest Paul Ijjina, C. Krishna Mohan
Pattern Recognit. Lett.2
2015 Multi-level classification: A generic classification method for medical datasets
abstract
Classification of medical data is one of the most challenging pattern recognition problems. As stated in literature a single classifier is unable to solve all medical image classification problems due to high sensitivity to noise and other imperfections like data imbalance. So, several individual classifiers have been studied to solve the different types of classification problems arising in medical datasets but all have proven to be useful on some specific datasets. Hence, in this paper, we propose a generic multi-level classification approach for medical datasets using sparsity based dictionary learning and support vector machine approaches. The proposed technique demonstrates the following advantages: 1) gives better performance of classification accuracy over all datasets 2) solves imbalanced data problems 3) needs no fusion and ensemble methods in multi-level classification. The results presented on the 5 standard UCI medical datasets demonstrate that the efficacy of the proposed multi-level classification technique.
R. Bharath, Pachamuthu Rajalakshmi, C. Krishna Mohan
HealthCom4
2015 Nearest Neighbor Minutia Quadruplets Based Fingerprint Matching with Reduced Time and Space Complexity
abstract
The fingerprint biometric is often used as the primary source of person authentication in a large population person identity system because fingerprints have unique properties like distinctiveness and persistence. However, the large volumes of fingerprint data may lead to the scalability issues which are to be addressed in the context of memory and computational complexity. In this paper, an attempt is made to develop an efficient fingerprint matching algorithm using nearest neighbor minutia quadruplets (NNMQ). These minutia quadruplets are both rotation and translation invariant. Experimental results demonstrate that the proposed fingerprint matching algorithm achieves the reduced space and time complexities with the publicly available standard fingerprint benchmark databases FVC ongoing, FVC2000 and FVC2004.
A. Tirupathi Rao, Nalla Pattabhi Ramaiah, V. Raghavendra Reddy, C. Krishna Mohan
ICMLA4
2015 Feature selection using Deep Neural Networks
abstract
Feature descriptors involved in video processing are generally high dimensional in nature. Even though the extracted features are high dimensional, many a times the task at hand depends only on a small subset of these features. For example, if two actions like running and walking have to be identified, extracting features related to the leg movement of the person is enough. Since, this subset is not known apriori, we tend to use all the features, irrespective of the complexity of the task at hand. Selecting task-aware features may not only improve the efficiency but also the accuracy of the system. In this work, we propose a supervised approach for task-aware selection of features using Deep Neural Networks (DNN) in the context of action recognition. The activation potentials contributed by each of the individual input dimensions at the first hidden layer are used for selecting the most appropriate features. The selected features are found to give better classification performance than the original high-dimensional features. It is also shown that the classification performance of the proposed feature selection technique is superior to the low-dimensional representation obtained by principal component analysis (PCA).
Debaditya Roy, K. Sri Rama Murty, C. Krishna Mohan
IJCNN3
2015 Content based medical image retrieval using dictionary learning
R. Ramu Naidu, C. S. Sastry 0001, C. Krishna Mohan
Neurocomputing4
2014 Human Action Recognition Based on MOCAP Information Using Convolution Neural Networks
abstract
Human action recognition is an important component in semantic analysis of human behavior. In this paper, we propose an approach for human action recognition based on motion capture (MOCAP) information using convolutional neural networks (CNN). Distance based metrics computed from MOCAP information of only three human joints are used in the computation of features. The range and temporal variation of these distance metrics are considered in the design of features which are discriminative for action recognition. A convolutional neural network capable of recognizing local patterns is used to identify human actions from the temporal variation of these features, which are distorted due to the inconsistency in the execution of actions across observations and subjects. Experiments conducted on Berkeley MHAD dataset demonstrate the effectiveness of the proposed approach.
Earnest Paul Ijjina, C. Krishna Mohan
ICMLA2
2014 Human Action Recognition Based on Recognition of Linear Patterns in Action Bank Features Using Convolutional Neural Networks
abstract
In this paper, we proposed a deep convolutional network architecture for recognizing human actions in videos using action bank features. Action bank features computed against of a predefined set of videos known as an action bank, contain linear patterns representing the similarity of the video against the action bank videos. Due to the independence of the patterns across action bank features, a convolutional neural network with linear masks is considered to capture the local patterns associated with each action. The knowledge gained through training is used to assign an action label to videos during testing. Experiments conducted on UCF50 dataset demonstrates the effectiveness of the proposed approach in capturing and recognizing these linear local patterns.
Earnest Paul Ijjina, C. Krishna Mohan
ICMLA2
2014 One-Shot Periodic Activity Recognition Using Convolutional Neural Networks
abstract
Activities capture vital facts for the semantic analysis of human behavior. In this paper, we propose a method for recognizing human activities based on periodic actions from a single instance using convolutional neural networks (CNN). The height of the foot above the ground is considered as features to discriminate human locomotion activities. The periodic nature of actions in these activities is exploited to generate the training cases from a single instance using a sliding window. Also, the capability of a convolutional neural network to learn local visual patterns is exploited for human activity recognition. Experiments on Carnegie Mellon University (CMU) Mocap dataset demonstrate the effectiveness of the proposed approach.
Earnest Paul Ijjina, C. Krishna Mohan
ICMLA2
2014 Facial Expression Recognition Using Kinect Depth Sensor and Convolutional Neural Networks
abstract
Facial expression recognition is an active area of research with applications in the design of Human Computer Interaction (HCI) systems. In this paper, we propose an approach for facial expression recognition using deep convolutional neural networks (CNN) based on features generated from depth information only. The Gradient direction information of depth data is used to represent facial information, due its invariance to distance from the sensor. The ability of a convolutional neural networks (CNN) to learn local discriminative patterns from data is used to recognize facial expressions from the representation of unregistered facial images. Experiments conducted on EURECOM kinect face dataset demonstrate the effectiveness of the proposed approach.
Earnest Paul Ijjina, C. Krishna Mohan
ICMLA2
2014 Music genre classification using On-line Dictionary Learning
abstract
In this paper, an approach for music genre classification based on sparse representation using MARSYAS features is proposed. The MARSYAS feature descriptor consisting of timbral texture, pitch and beat related features is used for the classification of music genre. On-line Dictionary Learning (ODL) is used to achieve sparse representation of the features for developing dictionaries for each musical genre. We demonstrate the efficacy of the proposed framework on the Latin Music Database (LMD) consisting of over 3000 tracks spanning 10 genres namely Axé, Bachata, Bolero, Forró, Gaúcha, Merengue, Pagode, Salsa, Sertaneja and Tango.
Debaditya Roy, C. Krishna Mohan
IJCNN3
2010 Efficient clustering approach using incremental and hierarchical clustering methods
abstract
There are many clustering methods available and each of them may give a different grouping of datasets. It is proven that hybrid clustering algorithms give efficient results over the other algorithms. In this paper, we propose an efficient hybrid clustering algorithm by combining the features of leader's method which is an incremental clustering method and complete linkage algorithm which is a hierarchical clustering procedure. It is most common to find the dissimilarity between two clusters as the distance between their centorids or the distance between two closest (or farthest) data points. However, these measures may not give efficient clustering results in all cases. So, we propose a new similarity measure, known as cohesion to find the intercluster distance. By using this measure of cohesion, a two level clustering algorithm is proposed, which runs in linear time to the size of input data set. We demonstrate the effectiveness of the clustering procedure by using the leader's algorithm and cohesion similarity measure. The proposed method works in two steps: In the first step, the features of incremental and hierarchical clustering methods are combined to partition the input data set into several smaller subclusters. In the second step, subclusters are merged continuously based on cohesion similarity measure. We demonstrate the effectiveness of this framework for the web mining applications.
C. Krishna Mohan
IJCNN2
2008 Video Shot Segmentation Using Late Fusion Technique
abstract
In this paper, a new method for detecting shot boundaries in video sequences using a late fusion technique is proposed. The method uses color histogram as the feature, and processes each bin separately for detecting shot boundaries. The decisions from individual bins are combined later for hypothesizing the presence of shot boundaries. The method provides a certain degree of robustness against illumination and camera/object motion, as it ignores small changes in the bins. While the early fusion techniques rely on the extent of change in color information, the proposed technique relies on the number of significant changes. Experimental results successfully validate the new method and show that it can effectively detect both abrupt and gradual transitions.
C. Krishna Mohan, Dhananjaya Gowda, Bayya Yegnanarayana
ICMLA1
2004 Content-Based Video Classification Using Support Vector Machines
Vakkalanka Suresh, C. Krishna Mohan, R. Kumaraswamy 0001, Bayya Yegnanarayana
ICONIP2