VLDB 2026 Research / reviewers in the wild / expert
Mst. Shapna Akter
dblp:331/5591
· DBLP profile ↗
16ranked-venue papers in the field
7as first author
16since 2021 · last 2026
0000-0002-9859-6265ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 13 (6 first)Database Systems & Data Management · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing m6Am Site Prediction Through Deep Learning on Diverse mRNA Sequence Landscapes: The m6Am-DLcat Approach
Alfredo Cuzzocrea, Tasmin Karim, Md. Shazzad Hossain Shaon, Md. Fahim Sultan, Mst. Shapna Akter |
DATA (1) | 5 |
| 2025 | A Benchmark Dataset for Code-Level Vulnerability Detection and Analysis
Tasmin Karim, Mst. Shapna Akter, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2025 | ResVul-LLM: A Neurosymbolic Framework Combining Large Language Models and Symbolic Reasoning for C/C++ Vulnerability Analysis
Md. Shazzad Hossain Shaon, Mst. Shapna Akter, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2025 | P3R: Parallel Plugin-Based Parameter Efficient Fine-Tuning for Code Understanding Through Hierarchical Representation Refinement
Md. Fahim Sultan, Mst. Shapna Akter, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2025 | CodeVul+: A Structure-Aware Framework for Cross-Repository Vulnerability Detection
Md. Fahim Sultan, Mst. Shapna Akter, Alfredo Cuzzocrea |
IEEE Big Data | 2 |
| 2025 | Neuro-Symbolic Methods in Natural Language Processing: A Review
Mst. Shapna Akter, Md. Fahim Sultan, Alfredo Cuzzocrea |
DATA | 1 |
| 2025 | Improving Software Security Through a LLM-Based Vulnerability Detection Model
Syeda Sadia Alam, Mst. Shapna Akter, Alfredo Cuzzocrea |
DEXA (1) | 2 |
| 2024 | EnStack: An Ensemble Stacking Framework of Large Language Models for Enhanced Vulnerability Detection in Source CodeabstractAutomated detection of software vulnerabilities is critical for enhancing security, yet existing methods often struggle with the complexity and diversity of modern codebases. In this paper, we propose a novel ensemble stacking approach that synergizes multiple pre-trained large language models (LLMs)—CodeBERT, GraphCodeBERT, and UniXcoder—to improve vulnerability detection in source code. Our method uniquely combines the semantic understanding of CodeBERT, the structural code representations of GraphCodeBERT, and the cross-modal capabilities of UniXcoder. By fine-tuning these models on the Draper VDISC dataset and integrating their predictions using meta-classifiers such as Logistic Regression, Support Vector Machines (SVM), Random Forest, and XGBoost, we effectively capture complex code patterns that individual models may miss. The meta-classifiers aggregate the strengths of each model, enhancing overall predictive performance. Our ensemble demonstrates significant performance gains over existing methods, with notable improvements in accuracy, precision, recall, F1-score, and AUC-score. This advancement addresses the challenge of detecting subtle and complex vulnerabilities in diverse programming contexts. The results suggest that our ensemble stacking approach offers a more robust and comprehensive solution for automated vulnerability detection, potentially influencing future AI-driven security practices. Shahriyar Zaman Ridoy, Md. Shazzad Hossain Shaon, Alfredo Cuzzocrea, Mst. Shapna Akter |
IEEE Big Data | 4 |
| 2024 | NeuroBooster: A Robust Classifier for the Discovery of Neuropeptide Sequences based on Meta-learning ApproachabstractNeuropeptides (NPs) are fragile proteins that serve as essential signaling molecules in the neurological system, playing a key role in modulating various physiological processes. Identifying particular neuropeptide sequences relevant to specific disorders would be beneficial for accelerating the development of diagnostic tools. The study proposed another approach to detecting NPs with multi-layer perception (MLP) and a bagging classifier-based meta-learning method called NeuroBooster. This investigation initially focused on five feature extractions based on composition, such as AAC, PAAC, physicochemical properties, QSO, and transfer-learning, such as Bert, and F2V strategies. Subsequently, we used the XGB feature selection method in the Bert and F2V methods to obtain the most 100D crucial features. The predicted probabilistic outcomes of NPs from the 8 preliminary models merged and derived a two-stage dataset with 40 dimensions of features and transmitted them into three classic models and two meta-models, through rigorous criteria for evaluation. Compared with the existing predictor, our proposed model NeuroBooster achieved a higher accuracy of 91.91% in the independent test method. Consequently, we discovered important features in these five models underscoring that physicochemical properties are potential targets for identification, thereby revealing new avenues for therapies. Md. Shazzad Hossain Shaon, Md. Fahim Sultan, Tasmin Karim, Md. Shoaib Hossain Alshan, Alfredo Cuzzocrea, Mst. Shapna Akter |
IEEE Big Data | 6 |
| 2024 | An Advanced Liver Disease Detection Tool with a Stacking-Ensemble-based Machine Learning ApproachabstractLiver diseases (LD) encompass a variety of disorders associated with the liver, including infectious hepatitis, obesity, cirrhosis, and malignancy, which constitute significant health issues in the world. Due to the minimal symptoms, comprehending the disease becomes extremely difficult until its severe stages; earlier detection is advantageous for appropriate action, which could save lives. This study developed a StackLD framework based on the stacking-ensemble-based machine learning approach. In the preliminary phase, we collected the Indian Liver Patient Dataset (ILPD) dataset, which contains 11 features, and the dataset highlighted a significant discrepancy. To overcome this, we used the SMOTE to rebalance the dataset, facilitating the development of robust machine learning models. We applied 7 different models such as XGB, LGBM, DT, KNN, RF, KNN, and stacking approaches with several evaluation metrics on independent test methods The analysis presented that the stacking technique executed superior in accuracy, sensitivity, specificity, and area under the curve, with values of 0.8622, 0.8933, 0.8369, and 0.9275, respectively. These outcomes indicate that our approach effectively differentiates between positive and negative classes. This study illustrates that the Alkphos, Sgot, and Sgpt elements have a significant role in determining liver disease (LD) from various features. As a result, a web server was built using these attributes, demonstrating that our model accurately predicts liver disorders at an early stage which is useful insight into the medical field with the potential to improve diagnostic procedures and patient outcomes. Md. Shazzad Hossain Shaon, Md. Fahim Sultan, Tasmin Karim, Alfredo Cuzzocrea, Mst. Shapna Akter |
IEEE Big Data | 5 |
| 2023 | Quantum Cryptography for Enhanced Network Security: A Comprehensive Survey of Research, Developments, and Future DirectionsabstractWith the ever-growing concern for internet security, the field of quantum cryptography emerges as a promising solution for enhancing the security of networking systems. In this paper, 20 notable papers from leading conferences and journals are reviewed and categorized based on their focus on various aspects of quantum cryptography, including key distribution, quantum bit commitment, post-quantum cryptography, and counterfactual quantum key distribution. The paper explores the motivations and challenges of employing quantum cryptography, addressing security and privacy concerns along with existing solutions. Secure key distribution, a critical component in ensuring the confidentiality and integrity of transmitted information over a network, is emphasized in the discussion. The survey examines the potential of quantum cryptography to enable secure key exchange between parties, even when faced with eavesdropping, and other applications of quantum cryptography. Additionally, the paper analyzes the methodologies, findings, and limitations of each reviewed study, pinpointing trends such as the increasing focus on practical implementation of quantum cryptography protocols and the growing interest in post-quantum cryptography research. Furthermore, the survey identifies challenges and open research questions, including the need for more efficient quantum repeater networks, improved security proofs for continuous variable quantum key distribution, and the development of quantum-resistant cryptographic algorithms, showing future directions for the field of quantum cryptography. Mst. Shapna Akter, Juanjose Rodriguez-Cardenas, Hossain Shahriar, Alfredo Cuzzocrea, Fan Wu 0013 |
IEEE Big Data | 1 |
| 2023 | A Trustable LSTM-Autoencoder Network for Cyberbullying Detection on Social Media Using Synthetic DataabstractSocial media cyberbullying has a detrimental effect on human life. As online social networking grows daily, the amount of hate speech also increases. Such terrible content can cause depression and actions related to suicide. This paper proposes a trustable LSTM-Autoencoder Network for cyberbullying detection on social media using synthetic data. We have demonstrated a cutting-edge method to address data availability difficulties by producing machine-translated data. However, several languages such as Hindi and Bangla still lack adequate investigations due to a lack of datasets. We carried out experimental identification of aggressive comments on Hindi, Bangla, and English datasets using the proposed model and traditional models, including Long Short-Term Memory (LSTM), Bidirectional Long ShortTerm Memory (BiLSTM), LSTM-Autoencoder, Word2vec, Bidirectional Encoder Representations from Transformers (BERT), and Generative Pre-trained Transformer 2 (GPT-2) models. We employed evaluation metrics such as f1-score, accuracy, precision, and recall to assess the models’ performance. Our proposed model outperformed all the models on all datasets, achieving the highest accuracy of 95%. Our model achieves state-of-the-art results among all the previous works on the dataset we used in this paper. Mst. Shapna Akter, Hossain Shahriar, Alfredo Cuzzocrea, Fan Wu 0013, Juanjose Rodriguez-Cardenas |
IEEE Big Data | 1 |
| 2022 | Software Supply Chain Vulnerabilities Detection in Source Code: Performance Comparison between Traditional and Quantum Machine Learning AlgorithmsabstractThe software supply chain (SSC) attack has become one of the crucial issues that are being increased rapidly with the advancement of the software development domain. In general, SSC attacks execute during the software development processes lead to vulnerabilities in software products targeting downstream customers and even involved stakeholders. Machine Learning approaches are proven in detecting and preventing software security vulnerabilities. Besides, emerging quantum machine learning can be promising in addressing SSC attacks. Considering the distinction between traditional and quantum machine learning, performance could be varies based on the proportions of the experimenting dataset. In this paper, we conduct a comparative analysis between quantum neural networks (QNN) and conventional neural networks (NN) with a software supply chain attack dataset known as ClaMP. Our goal is to distinguish the performance between QNN and NN and to conduct the experiment, we develop two different models for QNN and NN by utilizing Pennylane for quantum and TensorFlow and Keras for traditional respectively. We evaluated the performance of both models with different proportions of the ClaMP dataset to identify the f1 score, recall, precision, and accuracy. We also measure the execution time to check the efficiency of both models. The demonstration result indicates that execution time for QNN is slower than NN with a higher percentage of datasets. Due to recent advancements in QNN, a large level of experiments shall be carried out to understand both models accurately in our future research. Mst. Shapna Akter, Md. Jobair Hossain Faruk, Nafisa Anjum, Mohammad Masum, Hossain Shahriar, Akond Ashfaque Ur Rahman, Fan Wu 0013, Alfredo Cuzzocrea |
IEEE Big Data | 1 |
| 2022 | Deep Learning Approach for Classifying the Aggressive Comments on Social Media: Machine Translated Data Vs Real Life DataabstractAggressive comments on social media negatively impact human life. Such offensive contents are responsible for depression and suicidal-related activities. Since online social networking is increasing day by day, the hate content is also increasing. Several investigations have been done on the domain of cyberbullying, cyberaggression, hate speech, etc. The majority of the inquiry has been done in the English language. Some languages (Hindi and Bangla) still lack proper investigations due to the lack of a dataset. This paper particularly worked on the Hindi, Bangla, and English datasets to detect aggressive comments and have shown a novel way of generating machine-translated data to resolve data unavailability issues. A fully machine-translated English dataset has been analyzed with the models such as the Long Short term memory model (LSTM), Bidirectional Long-short term memory model (BiLSTM), LSTM-Autoencoder, word2vec, Bidirectional Encoder Representations from Transformers (BERT), and generative pre-trained transformer (GPT-2) to make an observation on how the models perform on a machine-translated noisy dataset. We have compared the performance of using the noisy data with two more datasets such as raw data, which does not contain any noises, and semi-noisy data, which contains a certain amount of noisy data. We have classified both the raw and semi-noisy data using the aforementioned models. To evaluate the performance of the models, we have used evaluation metrics such as F1-score, accuracy, precision, and recall. We have achieved the highest accuracy on raw data using the gpt2 model, semi-noisy data using the BERT model, and fully machine-translated data using the BERT model. Since many languages do not have proper data availability, our approach will help researchers create machine-translated datasets for several analysis purposes. Mst. Shapna Akter, Hossain Shahriar, Nova Ahmed, Alfredo Cuzzocrea |
IEEE Big Data | 1 |
| 2022 | Handwritten Word Recognition using Deep Learning Approach: A Novel Way of Generating Handwritten WordsabstractA handwritten word recognition system comes with issues such as-lack of large and diverse datasets. It is necessary to resolve such issues since millions of official documents can be digitized by training deep learning models using a large and diverse dataset. Due to the lack of data availability, the trained model does not give the expected result. Thus, it has a high chance of showing poor results. This paper proposes a novel way of generating diverse handwritten word images using handwritten characters. The idea of our project is to train the BiLSTM-CTC architecture with generated synthetic handwritten words. The whole approach shows the process of generating two types of large and diverse handwritten word datasets: overlapped and non-overlapped. Since handwritten words also have issues like overlapping between two characters, we have tried to put it into our experimental part. We have also demonstrated the process of recognizing handwritten documents using the deep learning model. For the experiments, we have targeted the Bangla language, which lacks the handwritten word dataset, and can be followed for any language. Our approach is less complex and less costly than traditional GAN models. Finally, we have evaluated our model using Word Error Rate (WER), accuracy, f1-score, precision, and recall metrics. The model gives 39% WER score, 92% percent accuracy, and 92% percent f1 scores using non-overlapped data and 63% percent WER score, 83% percent accuracy, and 85% percent f1 scores using overlapped data. Mst. Shapna Akter, Hossain Shahriar, Alfredo Cuzzocrea, Nova Ahmed, Carson K. Leung |
IEEE Big Data | 1 |
| 2022 | Multi-class Skin Cancer Classification Architecture Based on Deep Convolutional Neural NetworkabstractSkin cancer is a deadly disease. Melanoma is a type of skin cancer responsible for the high mortality rate. Early detection of skin cancer can enable patients to treat the disease and minimize the death rate. Skin cancer detection is challenging since different types of skin lesions share high similarities. This paper proposes a computer-based deep learning approach that will accurately identify different kinds of skin lesions. Deep learning approaches can detect skin cancer very accurately since the models learn each pixel of an image. Sometimes humans can get confused by the similarities of the skin lesions, which we can minimize by involving the machine. However, not all deep learning approaches can give better predictions. Some deep learning models have limitations, leading the model to a false-positive result. We have introduced several deep learning models to classify skin lesions to distinguish skin cancer from different types of skin lesions. Before classifying the skin lesions, data preprocessing and data augmentation methods are used. Finally, a Convolutional Neural Network (CNN) model and six transfer learning models such as Resnet-50, VGG-16, Densenet, Mobilenet, Inceptionv3, and Xception are applied to the publically available benchmark HAM10000 dataset to classify seven classes of skin lesions and to conduct a comparative analysis. The models will detect skin cancer by differentiating the cancerous cell from the non-cancerous ones. The models’ performance is measured using performance metrics such as precision, recall, f1 score, and accuracy. We receive accuracy of 90, 88, 88, 87, 82, and 77 percent for inceptionv3, Xception, Densenet, Mobilenet, Resnet, CNN, and VGG16, respectively. Furthermore, we develop five different stacking models such as inceptionv3-inceptionv3, Densenet-mobilenet, inceptionv3-Xception, Resnet50-Vgg16, and stack-six for classifying the skin lesions and found that the stacking models perform poorly. We achieve the highest accuracy of 78 percent among all the stacking models. Mst. Shapna Akter, Hossain Shahriar, Sweta Sneha, Alfredo Cuzzocrea |
IEEE Big Data | 1 |