Kaushik Roy 0003

dblp:r/KaushikRoy3 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0002-9026-5322ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MARVEL: Multi-modal Analysis and Reasoning for Violence Explanation with Large Language Models
abstract
Violence detection in surveillance footage remains a critical challenge for public safety; however, existing approaches suffer from a fundamental limitation: their black-box nature prevents security personnel from understanding the rationale behind automated alerts. This paper introduces MARVEL (Multimodal Analysis and Reasoning for Violence Explanation with Large Language Models), a novel framework that addresses this interpretability gap by combining advanced violence detection with comprehensive explainability. Our approach integrates three complementary modalities, pose dynamics, motion patterns, and scene context, through an attention-based fusion mechanism, while employing hierarchical explainability through SHAP and LIME to identify violence-indicative regions at both feature and frame levels. The distinguishing contribution of MARVEL lies in its ability to generate human-readable security narratives through Large Language Model integration, transforming technical detection outputs into actionable intelligence reports. Extensive evaluation on the RWF-2000 and RLVS datasets demonstrates robust performance, achieving 92. 4% overall precision with 92. 8% F1 score on the combined dataset. In particular, the system exhibits strong generalization capabilities, attaining 86.7% accuracy on RWF-2000 surveillance footage and an exceptional 98. 9% accuracy in various scenarios from RLVS. Beyond quantitative metrics, MARVEL generates detailed incident reports that articulate specific behavioral indicators, temporal patterns, and risk assessments. This research work represents a novel comprehensive solution to explainable violence detection, bridging the gap between automated surveillance technology and practical security operations through interpretable multimodal analysis coupled with natural language explanation generation.
Kushal Badal, Subhram Dasgupta, Md Tasnim Alam, Xiaohong Yuan, Kaushik Roy 0003
ICMLA5
2025 Efficient Visual Cyberbullying Detection via Knowledge Distillation from EfficientNet to Lightweight CNNs
abstract
The rise of social media platforms has resulted in a significant rise in cyberbullying, especially through visual content that can inflict considerable psychological damage on victims. Although current methods have shown effectiveness in identifying cyberbullying via textual analysis, the detection of image-based cyberbullying continues to be a significantly underexplored and important issue. This study introduces a framework that uses knowledge distillation to develop an efficient and interpretable system to detect cyberbullying in images. This study employs EfficientNet-B5, which incorporates dual attention mechanisms, as a teacher model to facilitate training of a lightweight enhanced sequential CNN student model. The knowledge distillation process allows the student model to achieve competitive performance while ensuring appropriate computational efficiency for real-time deployment. To improve transparency and trust in automated decision making, we incorporate various explainability techniques such as GradCAM, GradCAM++, and LayerCAM to offer detailed insight into model predictions. The experimental results on 19,300 images indicate that the EfficientNet-B5 teacher model achieves an accuracy of 98.34%. In contrast, the enhanced CNN student model maintains competitive performance with 96.59% accuracy, utilizing significantly fewer computational resources. This reflects a mere 1.75% decrease in performance while making considerable efficiency improvements. The explainability analysis demonstrates that our model effectively identifies pertinent visual cues, including threatening gestures, offensive symbols, and contextual elements associated with cyberbullying behavior. This study addresses the increasing demand for efficient and interpretable AI systems that improve online safety and ensure transparency in automated content management processes.
Pal Dave, Subhram Dasgupta, Kushal Badal, Xiaohong Yuan, Kaushik Roy 0003
ICMLA5
2025 Genomic privacy and security in the era of artificial intelligence and quantum computing
abstract
The rapid advancements in sequencing technologies have greatly increased access to genomic data stored in public databases. This has raised significant privacy and security concerns. This review emphasizes the importance of protecting genomic data by analyzing vulnerabilities in current storage and sharing practices. It examines the risks genetic databases face from cyber-attacks and internal breaches, focusing especially on advanced AI-driven threats and quantum computing vulnerabilities. The review explores machine learning methods designed to secure data. It highlights algorithms that prioritize privacy while maintaining data confidentiality, such as differential privacy, federated learning, and synthetic data generation using Generative Adversarial Networks (GANs). Findings demonstrate progress in mitigating common privacy breaches like re-identification and inference attacks. However, persistent vulnerabilities remain, particularly to emerging threats such as model inversion and membership inference attacks. The review advocates an integrated approach combining robust legislative frameworks with advanced technology to address genomic privacy challenges. It calls for intensified research efforts to safeguard genomic information. In particular, there is an urgent need to adopt quantum-resistant cryptographic methods, including lattice-based encryption and blockchain-integrated security frameworks. The paper emphasizes the necessity for genomics researchers to prioritize data privacy and security. This ensures responsible handling of genomic information in research.
Richard Annan, Justin Noland, Kamaria Perkins, Xiaohong Yuan, Kaushik Roy 0003, Letu Qingge
Discov. Comput.5
2024 Deep Fake Video Classification with Sequential Input Frames Using Hybrid Deep Learning Model and Bayesian Optimization
abstract
With the advancement of Artificial Intelligence, creating synthetic content has become effortless. Manipulated content, such as synthesized images and videos known as Deep-fakes, is now prevalent on the internet and social media platforms. These deepfakes are generated using deep learning techniques, such as Generative Adversarial Networks (GANs) or Autoen-coders (AEs). They have the potential to spread false information and are particularly concerning in cases involving identity theft, especially when involving well-known celebrities. Distinguishing these fake videos with the human eye is challenging, as they closely resemble real ones. This research utilizes popular deep learning (DL) models, such as CNN (VGG19) combined with stacked LSTM, for the classification of real and fake videos. Sequences of input frames are stacked and fed into the hybrid model, which is developed using the TensorFlow framework for this study. Our approach applies video-level classification rather than frame-level classification, and it utilizes Bayesian Optimization (BO) for hyperparameter tuning, distinguishing it from other recent methods. The VGG19-LSTM model is trained and tested against FaceForensics dataset(FF++) to determine the best parameters and achieved the highest test accuracy of 95.6% for deepfakes dataset.
Swetha Chittam, Kaushik Roy 0003, Xiaohong Yuan
ICMLA3
2024 Interpretable Deep Learning Model for Multiclass Brain Tumor Classification
abstract
While deep learning techniques like convolutional neural networks (CNNs) have shown promise in automating brain tumor detection from magnetic resonance imaging (MRI) images, a critical gap remains in model accuracy and efficiency for multi-class model classification and model explainability. This lack of accuracy and interpretability hinders trust and adoption in clinical settings. This study introduces an innovative approach by enhancing the traditional ResNet50 architecture with advanced regularization techniques and additional convolutional layers to address these challenges. Our model improves training efficiency and robustness, achieving an impressive 98% accuracy on multi-class MRI image classification. Furthermore, integrating Gradient-weighted Class Activation Mapping(Grad-CAM) with modified ResNet50 architecture can improve patient outcomes by explaining automated brain tumor detection.
Raihana Tasnim, Kaushik Roy 0003, Madhuri Siddula
ICMLA2
2023 Machine learning based fileless malware traffic classification using image visualization
abstract
Abstract In today’s interconnected world, network traffic is replete with adversarial attacks. As technology evolves, these attacks are also becoming increasingly sophisticated, making them even harder to detect. Fortunately, artificial intelligence (AI) and, specifically machine learning (ML), have shown great success in fast and accurate detection, classification, and even analysis of such threats. Accordingly, there is a growing body of literature addressing how subfields of AI/ML (e.g., natural language processing (NLP)) are getting leveraged to accurately detect evasive malicious patterns in network traffic. In this paper, we delve into the current advancements in ML-based network traffic classification using image visualization. Through a rigorous experimental methodology, we first explore the process of network traffic to image conversion. Subsequently, we investigate how machine learning techniques can effectively leverage image visualization to accurately classify evasive malicious traces within network traffic. Through the utilization of production-level tools and utilities in realistic experiments, our proposed solution achieves an impressive accuracy rate of 99.48% in detecting fileless malware, which is widely regarded as one of the most elusive classes of malicious software.
Fikirte Ayalke Demmese, Ajaya Neupane, Sajad Khorsandroo, May Wang, Kaushik Roy 0003
Cybersecur.5
2023 Blind Image Quality Assessment via Multiperspective Consistency
abstract
Blind image quality assessment (BIQA) has made significant progress, but it remains a challenging problem due to the wide variation in image content and the diverse nature of distortions. To address these challenges and improve the adaptability of BIQA algorithms to different image contents and distortions, we propose a novel model that incorporates multiperspective consistency. Our approach introduces a multiperspective strategy to extract features from various viewpoints, enabling us to capture more beneficial cues from the image content. To map the extracted features to a scalar score, we employ a content‐aware hypernetwork architecture. Additionally, we integrate all perspectives by introducing a consistency supervision strategy, which leverages cues from each perspective and enforces a learning consistency constraint between them. To evaluate the effectiveness of our proposed approach, we conducted extensive experiments on five representative datasets. The results demonstrate that our method outperforms state‐of‐the‐art techniques on both authentic and synthetic distortion image databases. Furthermore, our approach exhibits excellent generalization ability. The source code is publicly available at https://github.com/gn-share/multi-perspective .
Letu Qingge, Yuanchen Huang, Kaushik Roy 0003, Yanggui Li, Pei Yang 0004
Int. J. Intell. Syst.4
2022 Performance Benchmark of Machine Learning-Based Methodology for Swahili News Article Categorization
abstract
As data increases at unprecedented rates, so does the need to classify this data, including news article data. Unfortunately, most news article categorization research utilizes global languages such as English or Spanish, and not much research considers low-resource languages like Swahili. Testing multiple classifiers and preprocessing methods, we show that the SVM model with tokenization and stop word removal has the highest accuracy (85.13%) scores for Swahili news article categorization. These results from the first publicly available peer-reviewed Swahili news article dataset provide benchmark performance for Swahili news article categorization and contribute to lean Swahili text classification research.
Shaun Anthony Little, Kaushik Roy 0003, Ahmed Al Hamoud
ICMLA2
2022 Machine Learning Techniques to Predict Real Time Thermal Comfort, Preference, Acceptability, and Sensation for Automation of HVAC Temperature
Yaa Takyiwaa Acquaah, Balakrishna Gokaraju, Raymond C. Tesiero, Kaushik Roy 0003
IEA/AIE4
2022 Intrusion-Based Attack Detection Using Machine Learning Techniques for Connected Autonomous Vehicle
Mansi Bhavsar, Kaushik Roy 0003, John C. Kelly, Balakrishna Gokaraju
IEA/AIE2
2022 Deepfake Detection Using CNN Trained on Eye Region
Tony Gwyn, Letu Qingge, Kaushik Roy 0003
IEA/AIE4
2022 Face Authentication from Masked Face Images Using Deep Learning on Periocular Biometrics
Jeffrey J. Hernandez V., Rodney Dejournett, Udayasri Nannuri, Tony Gwyn, Xiaohong Yuan, Kaushik Roy 0003
IEA/AIE6
2022 Correction: DeepSuccinylSite: a deep learning based approach for protein succinylation site prediction
abstract
Results: Using an independent test set of experimentally identified succinylation sites, our method achieved efficiency scores of 79%, 68.7% and 0.27 for sensitivity, specificity and MCC respectively, with an area under the receiver operator characteristic (ROC) curve of 0.8.In side-by-side comparisons with previously described succinylation site predictors, DeepSuccinylSite produces similar or better results compared to the other state-of-the-art predictors.On page 7, Last paragraph on right should be changed from Consequently, DeepSuccinylSite achieved a significantly higher performance as measured by MCC.Indeed, DeepSuccinylSite exhibited an ~ 62% increase in MCC when compared to the next highest method, GPSuc.to: Consequently, DeepSuccinylSite achieved an MCC score (at decision boundary of 0.5) on par with the top performingmethod, GPSuc.On page 2, in Table 1, the negative data of Independent Test should be 2977 rather than 254.On page 8, in Table 6, the MCC data of DeepSuccinylSite should be 0.27 rather than 0.48.
Niraj Thapa, Meenal Chaudhari, Sean McManus, Kaushik Roy 0003, Robert H. Newman, Hiroto Saigo, Dukka B. KC
BMC Bioinform.4
2022 Non-volume preserving-based fusion to group-level emotion recognition on crowd videos
Kha Gia Quach, T. Hoang Ngan Le, Chi Nhan Duong, Ibsa Jalata, Kaushik Roy 0003, Khoa Luu
Pattern Recognit.5
2020 Vec2Face: Unveil Human Faces From Their Blackbox Features in Face Recognition
abstract
Unveiling face images of a subject given his/her high-level representations extracted from a blackbox Face Recognition engine is extremely challenging. It is because the limitations of accessible information from that engine including its structure and uninterpretable extracted features. This paper presents a novel generative structure with Bijective Metric Learning, namely Bijective Generative Adversarial Networks in a Distillation framework (DiBiGAN), for synthesizing faces of an identity given that person's features. In order to effectively address this problem, this work firstly introduces a bijective metric so that the distance measurement and metric learning process can be directly adopted in image domain for an image reconstruction task. Secondly, a distillation process is introduced to maximize the information exploited from the blackbox face recognition engine. Then a Feature-Conditional Generator Structure with Exponential Weighting Strategy is presented for a more robust generator that can synthesize realistic faces with ID preservation. Results on several benchmarking datasets including CelebA, LFW, AgeDB, CFP-FP against matching engines have demonstrated the effectiveness of DiBiGAN on both image realism and ID preservation properties.
Chi Nhan Duong, Thanh-Dat Truong, Khoa Luu, Kha Gia Quach, Kaushik Roy 0003
CVPR6
2020 Evaluation of Local Binary Pattern Algorithm for User Authentication with Face Biometric
abstract
In the ever-changing world of computer security and user authentication, the username/password standard is becoming increasingly outdated. Using the same username and password across multiple accounts and websites leaves a user open to vulnerabilities, and the need to remember multiple usernames and passwords feels very unnecessary in the current digital age. Authentication methods of the future need to be reliable and fast, while maintaining the ability to provide secure access. Augmenting traditional username-password standard with face biometric is proposed in the literature to enhance the user authentication. However, this technique still needs an extensive evaluation study to show how reliable and effective it will be under different settings. Local Binary Pattern (LBP) is a discrete yet powerful texture classification scheme, which works particularly well with image classification for facial recognition. The system proposed here strives to examine and test various LBP configurations to determine their image classification accuracy. The most favorable configurations of LBP should be examined as a potential way to augment the current username and password standard by increasing their security with facial biometrics.
Tony Gwyn, Mustafa Atay, Kaushik Roy 0003, Albert C. Esterline
ICMLA3
2020 DeepSuccinylSite: a deep learning based approach for protein succinylation site prediction
abstract
BACKGROUND: Protein succinylation has recently emerged as an important and common post-translation modification (PTM) that occurs on lysine residues. Succinylation is notable both in its size (e.g., at 100 Da, it is one of the larger chemical PTMs) and in its ability to modify the net charge of the modified lysine residue from + 1 to - 1 at physiological pH. The gross local changes that occur in proteins upon succinylation have been shown to correspond with changes in gene activity and to be perturbed by defects in the citric acid cycle. These observations, together with the fact that succinate is generated as a metabolic intermediate during cellular respiration, have led to suggestions that protein succinylation may play a role in the interaction between cellular metabolism and important cellular functions. For instance, succinylation likely represents an important aspect of genomic regulation and repair and may have important consequences in the etiology of a number of disease states. In this study, we developed DeepSuccinylSite, a novel prediction tool that uses deep learning methodology along with embedding to identify succinylation sites in proteins based on their primary structure. RESULTS: Using an independent test set of experimentally identified succinylation sites, our method achieved efficiency scores of 79%, 68.7% and 0.48 for sensitivity, specificity and MCC respectively, with an area under the receiver operator characteristic (ROC) curve of 0.8. In side-by-side comparisons with previously described succinylation predictors, DeepSuccinylSite represents a significant improvement in overall accuracy for prediction of succinylation sites. CONCLUSION: Together, these results suggest that our method represents a robust and complementary technique for advanced exploration of protein succinylation.
Niraj Thapa, Meenal Chaudhari, Sean McManus, Kaushik Roy 0003, Robert H. Newman, Hiroto Saigo, Dukka B. KC
BMC Bioinform.4
2019 Comparing Learning-Based Methods for Identifying Disaster-Related Tweets
abstract
During natural disaster events, social media users generate a significant amount of data, some of which are valuable for relief efforts and emergency management. In this paper, we study the nature of social media content generated during two natural disasters. We provide a system for labeling Twitter data and classifying disaster-related tweets. The focus is on Twitter data generated before, during, and after Hurricane Florence and Hurricane Michael, both of which occurred in 2018. Then we apply various machine learning models on the labeled data and provide a performance comparison between the models using features generated from TfidfVectorizer and CountVectorizer Approaches.
Nasser Assery, Xiaohong Yuan, Sultan Almalki, Kaushik Roy 0003, Xiuli Qu
ICMLA4
2019 Touch-Based Active Cloud Authentication Using Traditional Machine Learning and LSTM on a Distributed Tensorflow Framework
abstract
In this modern world, mobile devices have been paired with the cloud environment to scale the voluminous amount of generated data. The implementation comes at the cost of privacy as proprietary data can be stolen in transit to the cloud, or victims’ phones can be seized along with synced data from cloud. The attacker can gain access to the phone through shoulder surfing, or even spoofing attacks. Our approach is to mitigate this issue by proposing an active cloud authentication framework using touch biometric pattern. To the best of our knowledge, active cloud authentication using touch dynamics for mobile cloud computing has not been explored in the literature. This research creates a proof of concept that will lead into a simulated cloud framework for active authentication. Given the amount of data captured by the mobile device from user activity, it can be a computationally intensive process for the mobile device to handle with such limited resources. To solve this, we simulated a post-transmission process of data to the cloud so that we could implement the authentication process within the cloud. We evaluated the touch data using traditional machine learning algorithms, such as Random Forest (RF), Support Vector Machine (SVM), and also using a deep learning classifier, the Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) algorithms. The novelty of this work is two-fold. First, we develop a distributed tensorflow framework for cloud authentication using touch biometric pattern. This framework helps alleviate the drawback of the computationally intensive recognition of the substantial amount of raw data from the user. Second, we apply the RF, SVM, and a deep learning classifier, the LSTM-RNN, on the touch data to evaluate the performance of the proposed authentication scheme. The proposed approach shows a promising performance with an accuracy of 99.0361% using RF on the distributed tensorflow framework.
Dylan J. Gunn, Rushit Dave, Xiaohong Yuan, Kaushik Roy 0003
Int. J. Comput. Intell. Appl.5
2019 An Empirical Evaluation of User Movement Data on Smartphones
abstract
Movement data can be collected and used to add new security features and functionality to users’ mobile devices. Measuring a user’s movement using mobile devices allows for the use of behavioral biometrics. This assessment could introduce a shift in our current methods for securing mobile devices: instead of physical attributes like fingerprints or our face, the use of behavioral attributes like the way we walk or perform some personal activity. In this paper, an empirical evaluation of different classification techniques is conducted on user movement data. The datasets used in this empirical evaluation contain accelerometer data that were collected during various experiments from several mobile devices, including smartphones, smart watches, and other accelerometer sensors. We aggregated the user movement data and provided them as input into five traditional machine learning algorithms. The classification performances of the data were compared with a deep learning technique, the Long Short-Term Memory-Recurrent Neural Network (LSTM-RNN). The LSTM-RNN achieved its highest accuracy at 89% compared to 97% from a traditional machine learning algorithm, specifically the k-Nearest Neighbor (k-NN) algorithm on wrist-worn accelerometer data, thus showing the LSTM to be a less viable option.
Christopher Kelley, Janelle C. Mason, Albert C. Esterline, Kaushik Roy 0003
Int. J. Comput. Intell. Appl.4
2018 Comparison of Pre-Trained Word Vectors for Arabic Text Classification Using Deep Learning Approach
abstract
Artificial Intelligence (AI) has been used widely to extract people's opinions from social media websites. However, most of the existing works focus on eliciting the features from English text. In this paper, we describe an Arabic text sentiment analysis approach using a Deep Neural network, namely Long Short-Term Memory Recurrent Neural Network (LSTM-RNN). In this research, we investigate how the different pre-trained Word Embedding (WE) models affect our model's accuracy. The dataset includes Arabic corpus collected from Twitter. The results show significant improvement in Arabic text classification.
Ali Alwehaibi, Kaushik Roy 0003
ICMLA2
2018 Anti-spoofing Approach Using Deep Convolutional Neural Network
Prosenjit Chatterjee, Kaushik Roy 0003
IEA/AIE2
2018 Classifying Political Tweets Using Naïve Bayes and Support Vector Machines
Ahmed Al Hamoud, Ali Alwehaibi, Kaushik Roy 0003, Marwan Bikdash
IEA/AIE3
2018 An Evaluation of User Movement Data
Janelle C. Mason, Christopher Kelley, Bisoye Olaleye, Albert C. Esterline, Kaushik Roy 0003
IEA/AIE5
2014 Multispectral iris recognition utilizing hough transform and modified LBP
abstract
This paper presents a multispectral iris recognition scheme using Circular Hough Transform (CHT) and a modified Local Binary Pattern (mLBP) feature extraction technique. The CHT is used to localize the iris regions from the multispectral iris images. We also apply the binary thresholding and edge detection techniques in an effort to reduce the effects of over and under segmentation in multispectral iris images in which iris and pupil boundaries are not clearly separable. Furthermore, we apply mLBP in an attempt to elicit the iris feature elements. The mLBP technique combines both the sign and magnitude features for the improvement of iris texture classification performance. The identification and verification performance of the proposed scheme is validated using a multispectral iris dataset of 3120 images.
Khary Popplewell, Kaushik Roy 0003, Foysal Ahmad, Joseph Shelton
SMC2
2014 Iris recognition using Level Set and hGEFE
abstract
In this paper, we deploy a Fuzzy C-Means Clustering with a Level Set (FCMLS) method in an effort to localize the nonideal iris images accurately. We apply Genetic and Evolutionary Feature Extraction (GEFE), a method that evolves Local Binary Pattern (LBP) based feature extractors in order to elicit the most discriminating biometric features. In addition, a hybrid Genetic and Evolutionary Feature Weighting/Selection (GEFeWS) method is applied to select and weight the most important features. GEFeWS uses a genetic and evolutionary computation (GEC) to evolve a population of real-coded feature masks (FMs). We apply GEFeWS on features extracted by GEFE, and we refer to this technique as hybrid GEFE/GEFeWS, or hGEFE. Results show that hGEFE provides a significant increase in recognition accuracy while reducing the number of features being used when compared to just using GEFE alone.
Joseph Shelton, Kaushik Roy 0003, Foysal Ahmad, Brian O'Connor
SMC2