Roozbeh Razavi-Far

dblp:41/7578 · DBLP profile ↗
← Back
8ranked-venue papers in the field
2as first author
7since 2021 · last 2026
0000-0002-4330-3656ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 3Database Systems & Data Management · 2 (1 first)Data Mining & Knowledge Discovery · 2 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Divided We Fall: Defending against adversarial attacks via soft-gated fractional mixture-of-experts with randomized adversarial training
abstract
Machine learning is a powerful tool enabling full automation of a huge number of tasks without explicit programming. Despite recent progress of machine learning in different domains, these models have shown vulnerabilities when they are exposed to adversarial threats. Adversarial threats aim to hinder the machine learning models from satisfying their objectives. They can create adversarial perturbations, which are imperceptible to humans’ eyes but have the ability to cause misclassification during inference. In this paper, we propose a defense system, which devises an adversarial training module within mixture-of-experts architecture to enhance its robustness against white-box evasion attacks. In our proposed defense system, we use nine pre-trained classifiers (experts) with ResNet-18 as their backbone. During end-to-end training, the parameters of all experts and the gating mechanism are jointly updated allowing further optimization of the experts. Our proposed defense system outperforms prior MoE-based defenses under strong white-box FGSM and PGD evaluation on CIFAR-10 and SVHN. The use of multiple experts increases training time and compute relative to single-network baselines; however, inference scales approximately linearly with the number of experts and is substantially cheaper than training.
Mohammad Meymani, Roozbeh Razavi-Far
Inf. Sci.2
2025 SBAN: A Framework & Multi-Dimensional Dataset for Large Language Model Pre-Training and Software Code Mining
abstract
This paper introduces SBAN (Source code, Binary, Assembly, and Natural Language Description), a large-scale, multi-dimensional dataset designed to advance the pre-training and evaluation of large language models (LLMs) for software code analysis. SBAN comprises more than 3 million samples, including 2.9 million benign and 672,000 malware respectively, each represented across four complementary layers: binary code, assembly instructions, natural language descriptions, and source code. This unique multimodal structure enables research on cross-representation learning, semantic understanding of software, and automated malware detection. Beyond security applications, SBAN supports broader tasks such as code translation, code explanation, and other software mining tasks involving heterogeneous data. It is particularly suited for scalable training of deep models, including transformers and other LLM architectures. By bridging low-level machine representations and high-level human semantics, SBAN provides a robust foundation for building intelligent systems that reason about code. We believe that this dataset opens new opportunities for mining software behavior, improving security analytics, and enhancing LLM capabilities in pre-training and fine-tuning tasks for software code mining.
Hamed Jelodar, Mohammad Meymani, Samita Bai, Roozbeh Razavi-Far, Ali A. Ghorbani 0001
ICDM4
2025 On the consistency of GNN explanations for malware detection
abstract
Control Flow Graphs (CFGs) are critical for analyzing program execution and characterizing malware behavior. With the growing adoption of Graph Neural Networks (GNNs), CFG-based representations have proven highly effective for malware detection. This study proposes a novel framework that dynamically constructs CFGs and embeds node features using a hybrid approach combining rule-based encoding and autoencoder-based embedding. A GNN-based classifier is then constructed to detect malicious behavior from the resulting graph representations. To improve model interpretability, we apply state-of-the-art explainability techniques, including GNNExplainer, PGExplainer, and CaptumExplainer, the latter is utilized three attribution methods: Integrated Gradients, Guided Backpropagation, and Saliency. In addition, we introduce a novel aggregation method, called RankFusion, that integrates the outputs of the top-performing explainers to enhance the explanation quality. We also evaluate explanations using two subgraph extraction strategies, including the proposed Greedy Edge-wise Composition (GEC) method for improved structural coherence. A comprehensive evaluation using accuracy, fidelity, and consistency metrics demonstrates the effectiveness of the proposed framework in terms of accurate identification of malware samples and generating reliable and interpretable explanations.
Hossein Shokouhi-Nejad, Griffin Higgins, Roozbeh Razavi-Far, Hesamodin Mohammadian, Ali A. Ghorbani 0001
Inf. Sci.3
2023 Constrained Generative Adversarial Learning for Dimensionality Reduction
abstract
Emerging data-driven technologies and big data analytics generate and deal with high-dimensional data. Transformation of such data into a low-dimensional feature space brings about numerous benefits, such as a more discriminant feature space, performance enhancement, less computational burden, and facilitating data visualization. This paper proposes a novel dimensionality reduction algorithm based on generative adversarial networks to tackle the issues related to high-dimensional data and common challenges in dimensionality reduction. To this aim, two constraints are defined to preserve the characteristics of the original data while rectifying the data distribution upon transformation. Formulating the transformation as sequential projections, the proposed Constrained Adversarial Dimensionality Reduction (CADR) method finds a set of sequential projection vectors that lead to a feature space in which between-class separability and within-class integrity are satisfied. This is while the transformed data perfectly comply with the pairwise affinity correlation in the original feature space. To evaluate the proposed method, nine advanced dimensionality reduction techniques are employed to enable a comparative study. The experiments are performed on several real-world benchmark datasets in terms of classification accuracy, F-measure, and G-mean. The obtained results show that the CADR could yield classification performance at a satisfactory level and outperforms the other competitors.
Ehsan Hallaji, Maryam Farajzadeh-Zanjani, Roozbeh Razavi-Far, Vasile Palade, Mehrdad Saif
IEEE Trans. Knowl. Data Eng.3
2022 Consensus-Based Decision Support Model and Fusion Architecture for Dynamic Decision Making
Hossein Hassani 0003, Roozbeh Razavi-Far, Mehrdad Saif, Enrique Herrera-Viedma
Inf. Sci.2
2022 An integrated framework for diagnosing process faults with incomplete features
Roozbeh Razavi-Far, Mehrdad Saif, Vasile Palade, Shiladitya Chakrabarti
Knowl. Inf. Syst.1
2021 Imputation-Based Ensemble Techniques for Class Imbalance Learning
abstract
Correct classification of rare samples is a vital data mining task and of paramount importance in many research domains. This article mainly focuses on the development of the novel class-imbalance learning techniques, which make use of oversampling methods integrated with bagging and boosting ensembles. Two novel oversampling strategies based on the single and the multiple imputation methods are proposed. The proposed techniques aim to create useful synthetic minority class samples, similar to the original minority class samples, by estimation of missing values that are already induced in the minority class samples. The re-balanced datasets are then used to train base-learners of the ensemble algorithms. In addition, the proposed techniques are compared with the commonly used class imbalance learning methods in terms of three performance metrics including AUC, F-measure, and G-mean over several synthetic binary class datasets. The empirical results show that the proposed multiple imputation-based oversampling combined with bagging significantly outperforms other competitors.
Roozbeh Razavi-Far, Maryam Farajzadeh-Zanjani, Boyu Wang 0004, Mehrdad Saif, Shiladitya Chakrabarti
IEEE Trans. Knowl. Data Eng.1
2020 A neuro-wavelet based approach for diagnosing bearing defects
Niloofar Gharesi, Mohammad Mahdi Arefi, Roozbeh Razavi-Far, Jafar Zarei, Shen Yin
Adv. Eng. Informatics3