VLDB 2026 Research / reviewers in the wild / expert
Md. Abdul Awal
dblp:189/7484
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A metamorphic testing perspective on knowledge distillation for language models of code: Does the student deeply mimic the teacher?abstractTransformer-based language models of code have achieved state-of-the-art performance across a wide range of software analytics tasks, but their practical deployment remains limited due to high computational costs, slow inference speeds, and significant environmental impact. To address these challenges, recent research has increasingly explored knowledge distillation as a method for compressing a large language model of code (the teacher) into a smaller model (the student) while maintaining performance. However, the degree to which a student model deeply mimics the predictive behavior and internal representations of its teacher remains largely unexplored, as current accuracy-based evaluation provides only a surface-level view of model quality and often fails to capture more profound discrepancies in behavioral fidelity between the teacher and student models. To address this gap, we empirically show that the student model often fails to deeply mimic the teacher model, resulting in up to 285% greater performance drop under adversarial attacks, which is not captured by traditional accuracy-based evaluation. In addition, to capture discrepancies in behavioral fidelity, we propose MetaCompress , a metamorphic testing framework that systematically evaluates behavioral fidelity by comparing the outputs of teacher and student models under a set of behavior-preserving metamorphic relations. We evaluate MetaCompress on two widely studied tasks, clone detection and vulnerability prediction, using compressed versions of popular language models of code, CodeBERT and GraphCodeBERT , obtained via three different knowledge distillation techniques: Compressor , AVATAR , and MORPH . The results show that MetaCompress identifies up to 62% behavioral discrepancies in student models, underscoring the need for behavioral fidelity evaluation within the knowledge distillation pipeline. Furthermore, the ablation study indicates that MetaCompress is robust and effectively detects behavioral fidelity divergence, even when both the teacher and student models are assessed using transformed inputs. These results position MetaCompress as a practical framework for evaluating the behavioral fidelity of compressed language models of code derived through knowledge distillation. Md. Abdul Awal, Mrigank Rochan, Chanchal Kumar Roy |
J. Syst. Softw. | 1 |
| 2025 | Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?abstractLarge Language Model (LLM)-generated data is increasingly used in software analytics, but it is unclear how this data compares to human-written data, particularly when models are exposed to adversarial scenarios. Adversarial attacks can compromise the reliability and security of software systems, so understanding how LLM-generated data performs under these conditions, compared to human-written data, which serves as the benchmark for model performance, can provide valuable insights into whether LLM-generated data offers similar robustness and effectiveness. To address this gap, we systematically evaluate and compare the quality of human-written and LLM-generated data for fine-tuning robust pre-trained models (PTMs) in the context of adversarial attacks. We evaluate the robustness of six widely used PTMs, fine-tuned on human-written and LLM-generated data, before and after adversarial attacks. This evaluation employs nine state-of-the-art (SOTA) adversarial attack techniques across three popular software analytics tasks: clone detection, code summarization, and sentiment analysis in code review discussions. Additionally, we analyze the quality of the generated adversarial examples using eleven similarity metrics. Our findings reveal that while PTMs fine-tuned on LLM-generated data perform competitively with those fine-tuned on human-written data, they exhibit less robustness against adversarial attacks in software analytics tasks. Our study underscores the need for further exploration into enhancing the quality of LLM-generated training data to develop models that are both high-performing and capable of withstanding adversarial attacks in software analytics. Md. Abdul Awal, Mrigank Rochan, Chanchal Kumar Roy |
EASE | 1 |
| 2025 | Fully Quanvolutional Networks for Time Series ClassificationabstractDespite the advancements in quantum convolution or quanvolution, challenges persist in making quanvolution scalable, efficient, and applicable to multi-dimensional data. Existing quanvolutional networks heavily rely on classical layers, with minimal quantum involvement due to inherent limitations in current quanvolution algorithms. Moreover, the application of quanvolution in the domain of 1D data remains largely unexplored. To address these limitations, we propose a new quanvolution algorithm-Quanv1D-capable of processing arbitrary-channel 1D data, handling variable kernel sizes, and generating a customizable number of feature maps, along with a classification network-fully quanvolutional network (FQN)-built solely using Quanv1D layers. Quanv1D is inspired by the classical Conv1D and stands out from the quanvolution literature by being fully trainable, modular, and freely scalable with a self-regularizing feature. To evaluate FQN, we tested it on 20 UEA and UCR time series datasets, both univariate and multivariate, and benchmarked its performance against state-of-the-art convolutional models (both quantum and classical). We found FQN to outperform all compared models in terms of average accuracy while using significantly fewer parameters. Additionally, to assess the viability of FQN on real hardware, we conducted a shot-based analysis across all the datasets to simulate statistical quantum noise and found our model robust and equally efficient. Nabil Anan Orka, Ehtashamul Haque, Md. Abdul Awal, Mohammad Ali Moni |
KDD (2) | 3 |
| 2025 | Investigating adversarial attacks in software analytics via machine learning explainability
Md. Abdul Awal, Mrigank Rochan, Chanchal Kumar Roy |
Softw. Qual. J. | 1 |
| 2025 | HepNet: Deep Neural Network for Classification of Early-Stage Hepatic Steatosis Using Microwave SignalsabstractHepatic steatosis, a key factor in chronic liver diseases, is difficult to diagnose early. This study introduces a classifier for hepatic steatosis using microwave technology, validated through clinical trials. Our method uses microwave signals and deep learning to improve detection to reliable results. It includes a pipeline with simulation data, a new deep-learning model called HepNet, and transfer learning. The simulation data, created with 3D electromagnetic tools, is used for training and evaluating the model. HepNet uses skip connections in convolutional layers and two fully connected layers for better feature extraction and generalization. Calibration and uncertainty assessments ensure the model's robustness. Our simulation achieved an F1-score of 0.91 and a confidence level of 0.97 for classifications with entropy ≤0.1, outperforming traditional models like LeNet (0.81) and ResNet (0.87). We also use transfer learning to adapt HepNet to clinical data with limited patient samples. Using1H-MRS as the standard for two microwave liver scanners, HepNet achieved high F1-scores of 0.95 and 0.88 for 94 and 158 patient samples, respectively, showing its clinical potential. Sazid Hasan, Aida Brankovic, Md. Abdul Awal, Sasan Ahdi Rezaeieh, Shelley E. Keating, Amin M. Abbosh |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | EvaluateXAI: A framework to evaluate the reliability and consistency of rule-based XAI techniques for software analytics tasks
Md. Abdul Awal, Chanchal Kumar Roy |
J. Syst. Softw. | 1 |
| 2021 | PreDTIs: prediction of drug-target interactions based on multiple feature information using gradient boosting framework with data balancing and feature selection techniquesabstractDiscovering drug-target (protein) interactions (DTIs) is of great significance for researching and developing novel drugs, having a tremendous advantage to pharmaceutical industries and patients. However, the prediction of DTIs using wet-lab experimental methods is generally expensive and time-consuming. Therefore, different machine learning-based methods have been developed for this purpose, but there are still substantial unknown interactions needed to discover. Furthermore, data imbalance and feature dimensionality problems are a critical challenge in drug-target datasets, which can decrease the classifier performances that have not been significantly addressed yet. This paper proposed a novel drug-target interaction prediction method called PreDTIs. First, the feature vectors of the protein sequence are extracted by the pseudo-position-specific scoring matrix (PsePSSM), dipeptide composition (DC) and pseudo amino acid composition (PseAAC); and the drug is encoded with MACCS substructure fingerings. Besides, we propose a FastUS algorithm to handle the class imbalance problem and also develop a MoIFS algorithm to remove the irrelevant and redundant features for getting the best optimal features. Finally, balanced and optimal features are provided to the LightGBM Classifier to identify DTIs, and the 5-fold CV validation test method was applied to evaluate the prediction ability of the proposed method. Prediction results indicate that the proposed model PreDTIs is significantly superior to other existing methods in predicting DTIs, and our model could be used to discover new drugs for unknown disorders or infections, such as for the coronavirus disease 2019 using existing drugs compounds and severe acute respiratory syndrome coronavirus 2 protein sequences. S. M. Hasan Mahmud, Md. Abdul Awal, Kawsar Ahmed, Mohammad Ali Moni |
Briefings Bioinform. | 4 |
| 2020 | Design of sEMG-based clench force estimator in FPGA using artificial neural networks
Sheikh Shanawaz Mostafa, Md. Abdul Awal, Mohiuddin Ahmad, Fernando Morgado Dias |
Neural Comput. Appl. | 2 |
| 2017 | An automatic fast optimization of Quadratic Time-frequency Distribution using the hybrid genetic algorithm
Md. Abdul Awal, Boualem Boashash |
Signal Process. | 1 |