Swarup Roy

dblp:120/3692 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-0011-3633ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Analyzing shallow and deep classifiers in detecting unseen attacks on internet of things network
abstract
Intrusion Detection Systems (IDSs) for Internet of Things (IoT) networks face increasing difficulty when encountering attack behaviors that differ from those observed during training. While most existing IDS solutions assume closed set conditions, real deployments often experience distribution shifts caused by evolving or previously unobserved attack patterns. This study evaluates the robustness of widely used machine learning and deep learning classifiers when trained on a subset of attack classes and tested on withheld attack categories within IoT traffic datasets. We conduct a comprehensive experimental study using two publicly available IoT traffic datasets, N_BaIoT and BoT-IoT, evaluating thirteen shallow and deep learning classifiers under both standard multiclass (seen-attack) settings and held-out(unseen) attack-class scenarios. The results shows significant miss rates across most classifiers when evaluated on unseen attack classes. This underlines the limitations of closed-set learning for this task. Among the evaluated models, XGBoost consistently achieves higher overall accuracy and lower false negative rates compared to other classifiers, outperforming the least effective DenseNet-based model by up to 37% in terms of mean accuracy across multiple experimental setups. However, despite achieving high classification accuracy under distribution shifts, these models continue to assign unseen attack samples to known classes, underscoring their inability to explicitly identify traffic as novel. The findings emphasize that strong closed-set performance does not necessarily equate to true unseen attack detection, motivating the need for open-set and novelty-aware intrusion detection approaches in realistic IoT security scenarios.
Lekhika Chettri, Justin Leo, Swarup Roy, Jugal K. Kalita
Peer Peer Netw. Appl.3
2025 GT-GRN : a graph transformer framework for enhanced gene regulatory network inference via multimodal embedding of expression data and existing network knowledge
abstract
The inference of gene regulatory networks (GRNs) is critical for understanding the regulatory mechanisms underlying cellular development, functional specialization, and disease progression. Predicting regulatory gene interactions-often framed as a link prediction task-is a foundational step toward modeling cellular behavior. However, GRN inference from gene coexpression data alone is limited by noise, low interpretability, and difficulty in capturing indirect regulatory signals. Additionally, challenges such as data sparsity, nonlinearity, and complex gene interactions hinder accurate network reconstruction. To address these issues, we propose, a novel graph transformer (GT) based framework (GT-GRN) that enhances GRN inference by integrating multimodal gene embeddings. Our method combines three complementary sources of information: (i) autoencoder-based embeddings, which capture high-dimensional gene expression patterns while preserving biological signals; (ii) structural embeddings, derived from previously inferred GRNs and encoded via random walks and a Bidirectional Encoder Representations from Transformers (BERT) based language model to learn global gene representations; (iii) positional encodings, capturing each gene's role within the network topology . These heterogeneous features are fused and processed using a GT, allowing the joint modeling of both local and global regulatory structures. Experimental results on benchmark datasets show that GT-GRN outperforms existing GRN inference methods in predictive accuracy and robustness. Furthermore, it reconstructs cell-type-specific GRNs with high fidelity and produces gene embeddings that generalize to other tasks such as cell-type annotation.
Binon Teji, Swarup Roy, Dinabandhu Bhandari, Jugal K. Kalita
Briefings Bioinform.2
2025 Network-based analysis of Alzheimer's Disease genes using multi-omics network integration with graph diffusion
abstract
Alzheimer's Disease (AD) is a complex neurodegenerative disorder affecting millions worldwide. Despite extensive research, the mechanisms behind AD remain elusive. Many studies suggest that disease-responsible genes often act as hub genes in biological networks. However, this assumption requires further investigation in the context of AD. To examine the network characteristics of known AD genes, it is crucial to construct a highly confident network, which is challenging to achieve using a single data source. This work integrates multi-omics networks inferred from microarray, single-cell RNA sequencing, and single-nuclei RNA sequencing expression data, weighted with protein interaction and gene ontology information. We generate a high-quality integrated network by utilizing various inference methods and combining them through a graph diffusion-based integration approach. This network is then analyzed to investigate the properties of known AD-specific genes. Our findings reveal that AD genes are not always high-degree or central hub nodes in the network. Instead, these genes are distributed across different quartiles of degree centrality while maintaining significant interconnections for effective regulation. Furthermore, our study highlights that peripheral genes, often overlooked, also play crucial roles by connecting to relevant disease nodes and hub genes. These findings challenge the conventional understanding that AD-responsible genes are primarily the hub genes in the network, offering new insights into the complex regulatory mechanisms of AD and suggesting novel directions for future research.
Softya Sebastian, Swarup Roy, Jugal K. Kalita
J. Biomed. Informatics2
2024 Application of Generative Graph Models in Biological Network Regeneration: A Selective Review and Qualitative Analysis
abstract
Biological networks are essential for understanding the complex cellular mechanisms of living organisms. Simulating biological networks allows researchers to model and understand complex cellular processes without the need for extensive and costly experiments. The ability to generate synthetic graphs that closely resemble these real-world, complex biological processes is a vital research area. Graph generation in the context of biological networks is particularly significant because it can lead to insights into cellular functions, disease mechanisms, and therapeutic targets. Recent advances in deep learning, particularly in graph generative models, have opened new avenues for applications in the biological domain. These advancements have the potential to revolutionize our understanding of biological systems. However, despite the development of numerous effective graph generation models, there has been limited work assessing the qualitative aspects of the generated graphs within the context of biological networks. Addressing this gap is crucial for ensuring that synthetic graphs are not only structurally accurate but also biologically meaningful.Although various graph generation models are available, their application to biological network recreation has been limited. In this paper, we focus on graph generation, specifically edge and node-independent models for biological networks. We assess four candidate models across four gene expression networks. Our systematic assessment examines the models’ qualitative aspects, including graph structural properties, generation diversity, and computational efficiency. Our findings highlight the strengths and limitations of current models, offering insights to guide the development of more robust graph generation techniques that accurately replicate biological network characteristics.
Binon Teji, Swarup Roy, Pietro H. Guzzi, Dinabandhu Bhandari
BIBM2
2023 A generic parallel framework for inferring large-scale gene regulatory networks from expression profiles: application to Alzheimer's disease network
abstract
The inference of large-scale gene regulatory networks is essential for understanding comprehensive interactions among genes. Most existing methods are limited to reconstructing networks with a few hundred nodes. Therefore, parallel computing paradigms must be leveraged to construct large networks. We propose a generic parallel framework that enables any existing method, without re-engineering, to infer large networks in parallel, guaranteeing quality output. The framework is tested on 15 inference methods (not limited to) employing in silico benchmarks and real-world large expression matrices, followed by qualitative and speedup assessment. The framework does not compromise the quality of the base serial inference method. We rank the candidate methods and use the top-performing method to infer an Alzheimer's Disease (AD) affected network from large expression profiles of a triple transgenic mouse model consisting of 45,101 genes. The resultant network is further explored to obtain hub genes that emerge functionally related to the disease. We partition the network into 41 modules and conduct pathway enrichment analysis, revealing that a good number of participating genes are collectively responsible for several brain disorders, including AD. Finally, we extract the interactions of a few known AD genes and observe that they are periphery genes connected to the network's hub genes. Availability: The R implementation of the framework is downloadable from https://github.com/Netralab/GenericParallelFramework.
Softya Sebastian, Swarup Roy, Jugal K. Kalita
Briefings Bioinform.2
2023 Enhanced U-Net segmentation with ensemble convolutional neural network for automated skin disease classification
Dasari Anantha Reddy, Swarup Roy, Rakesh Tripathi
Knowl. Inf. Syst.2
2022 CARES: CAuse Recognition for Emotion in Suicide Notes
Soumitra Ghosh, Swarup Roy, Asif Ekbal, Pushpak Bhattacharyya
ECIR (2)2
2021 Data science in unveiling COVID-19 pathogenesis and diagnosis: evolutionary origin to drug repurposing
abstract
MOTIVATION: The outbreak of novel severe acute respiratory syndrome coronavirus (SARS-CoV-2, also known as COVID-19) in Wuhan has attracted worldwide attention. SARS-CoV-2 causes severe inflammation, which can be fatal. Consequently, there has been a massive and rapid growth in research aimed at throwing light on the mechanisms of infection and the progression of the disease. With regard to this data science is playing a pivotal role in in silico analysis to gain insights into SARS-CoV-2 and the outbreak of COVID-19 in order to forecast, diagnose and come up with a drug to tackle the virus. The availability of large multiomics, radiological, bio-molecular and medical datasets requires the development of novel exploratory and predictive models, or the customisation of existing ones in order to fit the current problem. The high number of approaches generates the need for surveys to guide data scientists and medical practitioners in selecting the right tools to manage their clinical data. RESULTS: Focusing on data science methodologies, we conduct a detailed study on the state-of-the-art of works tackling the current pandemic scenario. We consider various current COVID-19 data analytic domains such as phylogenetic analysis, SARS-CoV-2 genome identification, protein structure prediction, host-viral protein interactomics, clinical imaging, epidemiological research and drug discovery. We highlight data types and instances, their generation pipelines and the data science models currently in use. The current study should give a detailed sketch of the road map towards handling COVID-19 like situations by leveraging data science experts in choosing the right tools. We also summarise our review focusing on prime challenges and possible future research directions. CONTACT: [email protected], [email protected].
Jayanta Kumar Das, Giuseppe Tradigo, Pierangelo Veltri, Pietro H. Guzzi, Swarup Roy
Briefings Bioinform.5
2021 A scheme for inferring viral-host associations based on codon usage patterns identifies the most affected signaling pathways during COVID-19
Jayanta Kumar Das, Subhadip Chakraborty, Swarup Roy
J. Biomed. Informatics3
2020 Assessing the Effectiveness of Causality Inference Methods for Gene Regulatory Networks
abstract
Causality inference is the use of computational techniques to predict possible causal relationships for a set of variables, thereby forming a directed network. Causality inference in Gene Regulatory Networks (GRNs) is an important, yet challenging task due to the limits of available data and lack of efficiency in existing causality inference techniques. A number of techniques have been proposed and applied to infer causal relationships in various domains, although they are not specific to regulatory network inference. In this paper, we assess the effectiveness of methods for inferring causal GRNs. We introduce seven different inference methods and apply them to infer directed edges in GRNs. We use time-series expression data from the DREAM challenges to assess the methods in terms of quality of inference and rank them based on performance. The best method is applied to Breast Cancer data to infer a causal network. Experimental results show that Causation Entropy is best, however, highly time-consuming and not feasible to use in a relatively large network. We infer Breast Cancer GRN with the second-best method, Transfer Entropy. The topological analysis of the network reveals that top out-degree genes such as SLC39A5 which are considered central genes, play important role in cancer progression.
Syed Sazzad Ahmed, Swarup Roy, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Chemical Characterization of Interacting Genes in Few Subnetworks of Alzheimer's Disease
Antara Sengupta, Pabitra Pal Choudhury, Hazel N. Manners, Pietro H. Guzzi, Swarup Roy
BIBM5
2018 A detection framework for semantic code clones and obfuscated code
Abdullah Sheneamer, Swarup Roy, Jugal K. Kalita
Expert Syst. Appl.2
2017 Network based algorithms for module extraction from RNASeq data: A quantitative assessment
abstract
Genes participating in a common module may cause clinically similar diseases and shares the common genetic origin of their associated disease phenotypes. Identifying such modules may be helpful in system level understanding of biological and cellular processes or their disruption caused in associated diseases. The choose dofthe appropriate method for gene selection is a difficult task. In this work we discuss and compare selective module finding methods.
Monica Jha, Pierangelo Veltri, Pietro H. Guzzi, Swarup Roy
BIBM4
2017 Performing local network alignment by ensembling global aligners
abstract
Interactions among proteins are important mechanisms in living cells. The whole set of interactions is often referred to as a protein-protein interaction network (PIN). Comparison among such networks may discover conserved (or disrupted) patterns of interactions among species. Such comparison is performed using network alignment algorithms. They help analyse PPI networks for a better understanding of biological processes such as finding conserved regions between species, giving us insight into their evolution. However, there is no best aligner or standard evaluation measure to assess the quality of alignments. In this work, we use several aligners to produce an ensembled result, which can further improve individual aligners' alignment quality. Two basic ensemble approaches are used: One by finding majority node mappings from aligners and another by combining their results into one final alignment. These alignments are then evaluated based on three scoring schemes: Gene Ontology Consistency (GOC), Node Coverage (NCV) and Generalised Symmetric Substructure Score (GS3) using IsoBase PPI networks. Results show that the majority voting based ensemble scheme performs well in GS3while the ensemble by the union of the decision by different aligners produces satisfactory outcomes in comparison in GOC and NCV scores.
Hazel N. Manners, Ahed Elmsallati, Pietro H. Guzzi, Swarup Roy, Jugal K. Kalita
BIBM4
2017 Schemes for Labeling Semantic Code Clones using Machine Learning
abstract
Machine learning approaches built to identify code clones fail to perform well due to insufficient training samples and have been restricted only up to Type-III clones. A majority of the publicly available code clone corpora are incomplete in nature and lack labeled samples for semantic or Type-IV clones. We present here two schemes for labeling all types of clones including Type-IV clones. We restrict our study to Java code only. First, we use an unsupervised approach to label Type-IV clones and validate them using expert Java programmers. Next, we present a supervised scheme for labeling (or classifying) unknown samples based on labeled samples derived from our first scheme. We evaluate the performance of our schemes using six well-known Java code clone corpora and report on the quality of produced clones in terms of kappa agreement, mean error and accuracy scores. Results show that both schemes produce high quality code clones facilitating future use of machine learning in detecting clones of Type-IV.
Abdullah Sheneamer, Hanan Hazazi, Swarup Roy, Jugal K. Kalita
ICMLA3
2016 Incremental Approach for Detecting Arbitrary and Embedded Cluster Structures
Keshab Nath, Swarup Roy, Sukumar Nandi
MEDI2
2015 MODULA: A network module based local protein interaction network alignment method
abstract
Biological networks are usually used to model interactions among biological macromolecules in a cells. For instance protein-protein interaction networks (PIN) are used to model and analyse the set of interactions among proteins. The comparison of networks may result in the identification of conserved patterns of interactions corresponding to biological relevant entities such as protein complexes and pathways. Several algorithms, known as network alignment algorithms, have been proposed to unravel relations between different species at the interactome level. Algorithms may be categorized in two main classes: merge and mine and mine and merge. Algorithms belonging to the first class initially merge input network into a single integrated and then mine such networks. Conversely algorithms belonging to the second class initially analyze separately two input networks then integrate such results. In this paper we present MODULA (Network Module based PPI Aligner), a novel approach for local network alignment that belong to the second class. The algorithm at first identifies compact modules from input networks. Modules of both networks are then matched using functional knowledge. Then it uses high scoring pairs of modules as seeds to build a bigger alignment. In order to asses MODULA we compared it to the state of the art local alignment algorithms over a rather extensive and updated dataset.
Pietro H. Guzzi, Pierangelo Veltri, Swarup Roy, Jugal K. Kalita
BIBM3
2014 Reconstruction of gene co-expression network from microarray data using local expression patterns
abstract
BACKGROUND: Biological networks connect genes, gene products to one another. A network of co-regulated genes may form gene clusters that can encode proteins and take part in common biological processes. A gene co-expression network describes inter-relationships among genes. Existing techniques generally depend on proximity measures based on global similarity to draw the relationship between genes. It has been observed that expression profiles are sharing local similarity rather than global similarity. We propose an expression pattern based method called GeCON to extract Gene CO-expression Network from microarray data. Pair-wise supports are computed for each pair of genes based on changing tendencies and regulation patterns of the gene expression. Gene pairs showing negative or positive co-regulation under a given number of conditions are used to construct such gene co-expression network. We construct co-expression network with signed edges to reflect up- and down-regulation between pairs of genes. Most existing techniques do not emphasize computational efficiency. We exploit a fast correlogram matrix based technique for capturing the support of each gene pair to construct the network. RESULTS: We apply GeCON to both real and synthetic gene expression data. We compare our results using the DREAM (Dialogue for Reverse Engineering Assessments and Methods) Challenge data with three well known algorithms, viz., ARACNE, CLR and MRNET. Our method outperforms other algorithms based on in silico regulatory network reconstruction. Experimental results show that GeCON can extract functionally enriched network modules from real expression data. CONCLUSIONS: In view of the results over several in-silico and real expression datasets, the proposed GeCON shows satisfactory performance in predicting co-expression network in a computationally inexpensive way. We further establish that a simple expression pattern matching is helpful in finding biologically relevant gene network. In future, we aim to introduce an enhanced GeCON to identify Protein-Protein interaction network complexes by incorporating variable density concept.
Swarup Roy, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
BMC Bioinform.1
2013 CoBi: Pattern Based Co-Regulated Biclustering of Gene Expression Data
Swarup Roy, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Pattern Recognit. Lett.1
2012 Deterministic Approach for Biclustering of Co-Regulated Genes from Gene Expression Data
abstract
This paper presents an expression pattern based biclustering technique for grouping both positively and negatively regulated genes together as co-regulated genes from microarray expression data. Most interesting variants of this problem are NP-complete requiring either large computational effort or the use of lossy heuristics to short circuit the calculation. Our approach deterministically finds all biclusters using a non-greedy approach in polynomial time. Various real datasets have been used for experiments and results are excellent.
Swarup Roy, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
KES1