Jugal K. Kalita

dblp:78/5662 · also Jugal Kalita, Jugal Kumar Kalita · DBLP profile ↗
← Back
97ranked-venue papers
5as first author
36since 2021 · last 2026
0000-0002-8765-7018ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 4 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 4 since 2021Computer networks · 6 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Security and privacy · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Toxicity in online platforms and AI systems: A survey of needs, challenges, mitigations, and future directions
Smita Khapre, Melkamu Mersha, Hassan Shakil, Jonali Baruah, Jugal K. Kalita
Expert Syst. Appl.5
2026 Structured grounding is critical for reliable intelligence report generation with LLMs
Hassan Shakil, William Armstrong, Zachary Pugh, Dylan Tokita, Jugal K. Kalita
Expert Syst. Appl.5
2026 Agentic AI systems: A systematic survey of multi-agent architectures, cognitive foundations, interaction, explainability, security, and performance evaluation
Smita Khapre, Melkamu Mersha, Zainab Olalekan, Hassan Shakil, Samer Iskander, Asser Moustafa, Jugal K. Kalita
Neurocomputing7
2026 Explainable AI: Context-aware layer-wise integrated gradients for explaining transformer models
Melkamu Mersha, Jugal K. Kalita
Neurocomputing2
2026 Video background music generation using hybrid shared mixture-of-experts multimodal Transformer
Khang Nhut Lam, Thuan Trung Phan, Khang Minh Duong, Thuy Phuong Nguyen, Tai Tran-Anh Pham, Van Lam Le, Jugal K. Kalita
Neural Comput. Appl.7
2026 Analyzing shallow and deep classifiers in detecting unseen attacks on internet of things network
abstract
Intrusion Detection Systems (IDSs) for Internet of Things (IoT) networks face increasing difficulty when encountering attack behaviors that differ from those observed during training. While most existing IDS solutions assume closed set conditions, real deployments often experience distribution shifts caused by evolving or previously unobserved attack patterns. This study evaluates the robustness of widely used machine learning and deep learning classifiers when trained on a subset of attack classes and tested on withheld attack categories within IoT traffic datasets. We conduct a comprehensive experimental study using two publicly available IoT traffic datasets, N_BaIoT and BoT-IoT, evaluating thirteen shallow and deep learning classifiers under both standard multiclass (seen-attack) settings and held-out(unseen) attack-class scenarios. The results shows significant miss rates across most classifiers when evaluated on unseen attack classes. This underlines the limitations of closed-set learning for this task. Among the evaluated models, XGBoost consistently achieves higher overall accuracy and lower false negative rates compared to other classifiers, outperforming the least effective DenseNet-based model by up to 37% in terms of mean accuracy across multiple experimental setups. However, despite achieving high classification accuracy under distribution shifts, these models continue to assign unseen attack samples to known classes, underscoring their inability to explicitly identify traffic as novel. The findings emphasize that strong closed-set performance does not necessarily equate to true unseen attack detection, motivating the need for open-set and novelty-aware intrusion detection approaches in realistic IoT security scenarios.
Lekhika Chettri, Justin Leo, Swarup Roy, Jugal K. Kalita
Peer Peer Netw. Appl.4
2025 Automatic Summarization of Long Documents (Student Abstract)
abstract
A vast amount of textual data is added to the internet daily, making utilization and interpretation of textual data difficult and cumbersome. As a result, automatic text summarization is crucial for extracting relevant information, saving precious time. Although many transformer models excel in summarization, they are constrained by their input size, preventing them from processing texts longer than their context size. This study introduces several novel algorithms that allow any LLM to efficiently overcome its input size limitation, effectively utilizing its full potential without any architectural modifications. We test our algorithms on texts with more than 70,000 words, and our experiments show a significant increase in BERTScore with competitive ROUGE scores.
Naman Chhibbar, Jugal K. Kalita
AAAI2
2025 Linear Decoding of Morphology Relations in Language Models (Student Abstract)
abstract
The recent success of transformer language models owes much to their conversational fluency, which includes linguistic and morphological proficiency. An affine Taylor approximation has been found to be a good approximation for transformer computations over certain factual and encyclopedic relations. We show that the truly linear approximation W s, where s is a early layer representation of the base form and W is a local model derivative, is necessary and sufficient to approximate morphological derivation, achieving above 80% top-1 accuracy across most morphological tasks in the Bigger Analogy Test Set. We argue that many morphological forms in transformer models are likely linearly encoded.
Eric Xia, Jugal K. Kalita
AAAI2
2025 Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development
abstract
Agentic AI and Multi-Agent Systems are poised to dominate industry and society imminently. Powered by goal-driven autonomy, they represent a powerful form of generative AI, marking a transition from reactive content generation into proactive multitasking capabilities. As an exemplar, we propose an architecture of a multi-agent system for the implementation phase of the software engineering process. We also present a comprehensive threat model for the proposed system. We demonstrate that while such systems can generate code quite accurately, they are vulnerable to attacks, including code injection. Due to their autonomous design and lack of humans in the loop, these systems cannot identify and respond to attacks by themselves. This paper analyzes the vulnerability of multi-agent systems and concludes that the coder-reviewer-tester architecture is more resilient than both the coder and coder-tester architectures, but is less efficient at writing code. We find that by adding a security analysis agent, we mitigate the loss in efficiency while achieving even better resiliency. We conclude by demonstrating that the security analysis agent is vulnerable to advanced code injection attacks, showing that embedding poisonous few-shot examples in the injected code can increase the attack success rate from 0% to 71.95%.
Brian Bowers, Smita Khapre, Jugal K. Kalita
ICMLA3
2025 Drug Repurposing Using Deep Embedded Clustering and Graph Neural Networks
abstract
Drug repurposing has historically been an economically infeasible process for identifying novel uses for abandoned drugs. Modern machine learning has enabled the identification of complex biochemical intricacies in candidate drugs; however, many studies rely on simplified datasets with known drug-disease similarities. We propose a machine learning pipeline that uses unsupervised deep embedded clustering, combined with supervised graph neural network link prediction to identify new drug-disease links from multi-omic data. Unsupervised autoencoder and cluster training reduced the dimensionality of omic data into a compressed latent embedding. A total of 9,022 unique drugs were partitioned into 35 clusters with a mean silhouette score of 0.8550. Graph neural networks achieved strong statistical performance, with a prediction accuracy of 0.901, receiver operating characteristic area under the curve of 0.960, and F1-Score of 0.901. A ranked list comprised of 477 per-cluster link probabilities exceeding 99 percent was generated. This study could provide new drug-disease link prospects across unrelated disease domains, while advancing the understanding of machine learning in drug repurposing studies.
Luke Delzer, Robert Kroleski, Ali K. AlShami, Jugal K. Kalita
ICMLA4
2025 Solving Math Word Problems Using Estimation Verification and Equation Generation
abstract
Large Language Models (LLMs) excel at various tasks, including problem-solving and question-answering. However, LLMs often find Math Word Problems (MWPs) challenging because solving them requires a range of reasoning and mathematical abilities with which LLMs seem to struggle. Recent efforts have helped LLMs solve more complex MWPs with improved prompts. This study proposes a novel method that initially prompts an LLM to create equations from a decomposition of the question, followed by using an external symbolic equation solver to produce an answer. To ensure the accuracy of the obtained answer, inspired by an established recommendation of math teachers, the LLM is instructed to solve the MWP a second time, but this time with the objective of estimating the correct answer instead of solving it exactly. The estimation is then compared to the generated answer to verify. If verification fails, an iterative rectification process is employed to ensure the correct answer is eventually found. This approach achieves new state-of-the-art results on datasets used by prior published research on numeric and algebraic MWPs, improving the previous best results by nearly two percent on average. In addition, the approach obtains satisfactory results on trigonometric MWPs, a task not previously attempted to the authors’ best knowledge. This study also introduces two new datasets, SVAMPClean and Trig300, to further advance the testing of LLMs’ reasoning abilities.1
Mitchell Piehl, Dillon Wilson, Ananya Kalita, Jugal K. Kalita
ICMLA4
2025 Optimized detection of cyber-attacks on IoT networks via hybrid deep learning models
Ahmed Bensaoud, Jugal K. Kalita
Ad Hoc Networks2
2025 GT-GRN : a graph transformer framework for enhanced gene regulatory network inference via multimodal embedding of expression data and existing network knowledge
abstract
The inference of gene regulatory networks (GRNs) is critical for understanding the regulatory mechanisms underlying cellular development, functional specialization, and disease progression. Predicting regulatory gene interactions-often framed as a link prediction task-is a foundational step toward modeling cellular behavior. However, GRN inference from gene coexpression data alone is limited by noise, low interpretability, and difficulty in capturing indirect regulatory signals. Additionally, challenges such as data sparsity, nonlinearity, and complex gene interactions hinder accurate network reconstruction. To address these issues, we propose, a novel graph transformer (GT) based framework (GT-GRN) that enhances GRN inference by integrating multimodal gene embeddings. Our method combines three complementary sources of information: (i) autoencoder-based embeddings, which capture high-dimensional gene expression patterns while preserving biological signals; (ii) structural embeddings, derived from previously inferred GRNs and encoded via random walks and a Bidirectional Encoder Representations from Transformers (BERT) based language model to learn global gene representations; (iii) positional encodings, capturing each gene's role within the network topology . These heterogeneous features are fused and processed using a GT, allowing the joint modeling of both local and global regulatory structures. Experimental results on benchmark datasets show that GT-GRN outperforms existing GRN inference methods in predictive accuracy and robustness. Furthermore, it reconstructs cell-type-specific GRNs with high fidelity and produces gene embeddings that generalize to other tasks such as cell-type annotation.
Binon Teji, Swarup Roy, Dinabandhu Bhandari, Jugal K. Kalita
Briefings Bioinform.4
2025 Explainable AI: XAI-guided context-aware data augmentation
Melkamu Mersha, Mesay Gemeda Yigezu, Atnafu Lambebo Tonja, Hassan Shakil, Samer Iskander, Olga Kolesnikova, Jugal K. Kalita
Expert Syst. Appl.7
2025 A novel active learning approach to label one million unknown malware variants
Ahmed Bensaoud, Jugal K. Kalita
Int. J. Approx. Reason.2
2025 Network-based analysis of Alzheimer's Disease genes using multi-omics network integration with graph diffusion
abstract
Alzheimer's Disease (AD) is a complex neurodegenerative disorder affecting millions worldwide. Despite extensive research, the mechanisms behind AD remain elusive. Many studies suggest that disease-responsible genes often act as hub genes in biological networks. However, this assumption requires further investigation in the context of AD. To examine the network characteristics of known AD genes, it is crucial to construct a highly confident network, which is challenging to achieve using a single data source. This work integrates multi-omics networks inferred from microarray, single-cell RNA sequencing, and single-nuclei RNA sequencing expression data, weighted with protein interaction and gene ontology information. We generate a high-quality integrated network by utilizing various inference methods and combining them through a graph diffusion-based integration approach. This network is then analyzed to investigate the properties of known AD-specific genes. Our findings reveal that AD genes are not always high-degree or central hub nodes in the network. Instead, these genes are distributed across different quartiles of degree centrality while maintaining significant interconnections for effective regulation. Furthermore, our study highlights that peripheral genes, often overlooked, also play crucial roles by connecting to relevant disease nodes and hub genes. These findings challenge the conventional understanding that AD-responsible genes are primarily the hub genes in the network, offering new insights into the complex regulatory mechanisms of AD and suggesting novel directions for future research.
Softya Sebastian, Swarup Roy, Jugal K. Kalita
J. Biomed. Informatics3
2025 Neural network translations for building SentiWordNets
Khang Nhut Lam, Trung Phuong Le, Khanh Cong Ngu, Kien Trung Le, Phuc Minh Le, Huy Hoang-Dang Nguyen, Jugal K. Kalita
J. Intell. Inf. Syst.7
2025 Evaluating the effectiveness of XAI techniques for encoder-based language models
Melkamu Mersha, Mesay Gemeda Yigezu, Jugal K. Kalita
Knowl. Based Syst.3
2025 SMART-vision: survey of modern action recognition techniques in vision
Ali K. AlShami, Ryan Rabinowitz, Khang Nhut Lam, Yousra Shleibik, Melkamu Mersha, Terrance E. Boult, Jugal K. Kalita
Multim. Tools Appl.7
2024 Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT
abstract
This research examines the effectiveness of OpenAI's GPT models as independent evaluators of text summaries generated by six transformer-based models from Hugging Face: DistilBART, BERT, ProphetNet, T5, BART, and PEGASUS. We evaluated these summaries based on the essential properties of a high-quality summary: concision, relevance, coherence, and readability using traditional metrics such as ROUGE and Latent Semantic Analysis (LSA). Uniquely, we employ GPT not as a summarizer, but as an evaluator, allowing it to independently assess summary quality without predefined metrics. Our analysis revealed significant correlations between GPT evaluations and traditional metrics, particularly in assessing relevance and coherence. The results demonstrate the potential of GPT as a robust tool for evaluating text summaries, providing insights that complement established metrics while providing a basis for comparative analysis of transformer-based models in natural language processing tasks.
Hassan Shakil, Atqiya Munawara Mahi, Phuoc Nguyen, Zeydy Ortiz, Jugal K. Kalita, Mamoun T. Mardini
ICMLA5
2024 MaskPure: Improving Defense Against Text Adversaries with Stochastic Purification
Harrison Gietz, Jugal K. Kalita
NLDB (1)2
2024 Survey of continuous deep learning methods and techniques used for incremental learning
Justin Leo, Jugal K. Kalita
Neurocomputing2
2024 Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction
Melkamu Mersha, Khang Nhut Lam, Joseph Wood, Ali K. AlShami, Jugal K. Kalita
Neurocomputing5
2024 Abstractive text summarization: State of the art, challenges, and improvements
Hassan Shakil, Ahmad Farooq, Jugal K. Kalita
Neurocomputing3
2024 CNN-LSTM and transfer learning models for malware classification based on opcodes and API calls
Ahmed Bensaoud, Jugal K. Kalita
Knowl. Based Syst.2
2023 Training-free Neural Architecture Search for RNNs and Transformers
abstract
Neural architecture search (NAS) has allowed for the automatic creation of new and effective neural network architectures, offering an alternative to the laborious process of manually designing complex architectures.However, traditional NAS algorithms are slow and require immense amounts of computing power.Recent research has investigated training-free NAS metrics for image classification architectures, drastically speeding up search algorithms.In this paper, we investigate trainingfree NAS metrics for recurrent neural network (RNN) and BERT-based transformer architectures, targeted towards language modeling tasks.First, we develop a new trainingfree metric, named hidden covariance, that predicts the trained performance of an RNN architecture and significantly outperforms existing training-free metrics.We experimentally evaluate the effectiveness of the hidden covariance metric on the NAS-Bench-NLP benchmark.Second, we find that the current search space paradigm for transformer architectures is not optimized for training-free neural architecture search.Instead, a simple qualitative analysis can effectively shrink the search space to the best performing architectures.This conclusion is based on our investigation of existing training-free metrics and new metrics developed from recent transformer pruning literature, evaluated on our own benchmark of trained BERT architectures.Ultimately, our analysis shows that the architecture search space and the training-free metric must be developed together in order to achieve effective results.Our source code is available at https://github. com/aaronserianni/training-free-nas.
Aaron Serianni, Jugal K. Kalita
ACL (1)2
2023 A generic parallel framework for inferring large-scale gene regulatory networks from expression profiles: application to Alzheimer's disease network
abstract
The inference of large-scale gene regulatory networks is essential for understanding comprehensive interactions among genes. Most existing methods are limited to reconstructing networks with a few hundred nodes. Therefore, parallel computing paradigms must be leveraged to construct large networks. We propose a generic parallel framework that enables any existing method, without re-engineering, to infer large networks in parallel, guaranteeing quality output. The framework is tested on 15 inference methods (not limited to) employing in silico benchmarks and real-world large expression matrices, followed by qualitative and speedup assessment. The framework does not compromise the quality of the base serial inference method. We rank the candidate methods and use the top-performing method to infer an Alzheimer's Disease (AD) affected network from large expression profiles of a triple transgenic mouse model consisting of 45,101 genes. The resultant network is further explored to obtain hub genes that emerge functionally related to the disease. We partition the network into 41 modules and conduct pathway enrichment analysis, revealing that a good number of participating genes are collectively responsible for several brain disorders, including AD. Finally, we extract the interactions of a few known AD genes and observe that they are periphery genes connected to the network's hub genes. Availability: The R implementation of the framework is downloadable from https://github.com/Netralab/GenericParallelFramework.
Softya Sebastian, Swarup Roy, Jugal K. Kalita
Briefings Bioinform.3
2023 Pose2Trajectory: Using transformers on body pose to predict tennis player's trajectory
Ali K. AlShami, Terrance E. Boult, Jugal K. Kalita
J. Vis. Commun. Image Represent.3
2022 DEGnext: classification of differentially expressed genes from RNA-seq data using a convolutional neural network with transfer learning
abstract
BACKGROUND: A limitation of traditional differential expression analysis on small datasets involves the possibility of false positives and false negatives due to sample variation. Considering the recent advances in deep learning (DL) based models, we wanted to expand the state-of-the-art in disease biomarker prediction from RNA-seq data using DL. However, application of DL to RNA-seq data is challenging due to absence of appropriate labels and smaller sample size as compared to number of genes. Deep learning coupled with transfer learning can improve prediction performance on novel data by incorporating patterns learned from other related data. With the emergence of new disease datasets, biomarker prediction would be facilitated by having a generalized model that can transfer the knowledge of trained feature maps to the new dataset. To the best of our knowledge, there is no Convolutional Neural Network (CNN)-based model coupled with transfer learning to predict the significant upregulating (UR) and downregulating (DR) genes from both trained and untrained datasets. RESULTS: We implemented a CNN model, DEGnext, to predict UR and DR genes from gene expression data obtained from The Cancer Genome Atlas database. DEGnext uses biologically validated data along with logarithmic fold change values to classify differentially expressed genes (DEGs) as UR and DR genes. We applied transfer learning to our model to leverage the knowledge of trained feature maps to untrained cancer datasets. DEGnext's results were competitive (ROC scores between 88 and 99[Formula: see text]) with those of five traditional machine learning methods: Decision Tree, K-Nearest Neighbors, Random Forest, Support Vector Machine, and XGBoost. DEGnext was robust and effective in terms of transferring learned feature maps to facilitate classification of unseen datasets. Additionally, we validated that the predicted DEGs from DEGnext were mapped to significant Gene Ontology terms and pathways related to cancer. CONCLUSIONS: DEGnext can classify DEGs into UR and DR genes from RNA-seq cancer datasets with high performance. This type of analysis, using biologically relevant fine-tuning data, may aid in the exploration of potential biomarkers and can be adapted for other disease datasets.
Tulika Kakati, Dhruba Kumar Bhattacharyya, Jugal K. Kalita, Trina M. Norden-Krichmar
BMC Bioinform.3
2022 Deep multi-task learning for malware image classification
Ahmed Bensaoud, Jugal K. Kalita
J. Inf. Secur. Appl.2
2022 UIPBC: An effective clustering for scRNA-seq data analysis without user input
Hussain Ahmed Chowdhury, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Knowl. Based Syst.3
2022 Incremental Deep Neural Network Learning Using Classification Confidence Thresholding
abstract
Most modern neural networks for classification fail to take into account the concept of the unknown. Trained neural networks are usually tested in an unrealistic scenario with only examples from a closed set of known classes. In an attempt to develop a more realistic model, the concept of working in an open set environment has been introduced. This in turn leads to the concept of incremental learning where a model with its own architecture and initial trained set of data can identify unknown classes during the testing phase and autonomously update itself if evidence of a new class is detected. Some problems that arise in incremental learning are inefficient use of resources to retrain the classifier repeatedly and the decrease of classification accuracy as multiple classes are added over time. This process of instantiating new classes is repeated as many times as necessary, accruing errors. To address these problems, this article proposes the classification confidence threshold (CT) approach to prime neural networks for incremental learning to keep accuracies high by limiting forgetting. A lean method is also used to reduce resources used in the retraining of the neural network. The proposed method is based on the idea that a network is able to incrementally learn a new class even when exposed to a limited number samples associated with the new class. This method can be applied to most existing neural networks with minimal changes to network architecture.
Justin Leo, Jugal K. Kalita
IEEE Trans. Neural Networks Learn. Syst.2
2021 DeepSplicer: An Improved Method of Splice Sites Prediction using Deep Learning
abstract
Post-transcriptional splicing of ribonucleic acid (mRNA) entails removing regions of RNA sequences (Introns) that do not include information for protein synthesis. Thus, accurate splicing site detection is integral for understanding gene structure and, as a result, protein synthesis for biological and medicinal applications. However, the necessity to develop an advanced computational algorithm arises because existing splice site (SS) prediction methods are either computationally inefficient or expensive. Considering this, we present DeepSplicer-a deep learning-based Convolutional Neural Network (CNN) model for locating splice sites. In this work, we compared the ability of the existing SS prediction algorithms model to identify SS in organisms-Homo sapiens, Oryza sativa japonica, Arabidopsis thaliana, DrosophUa melanogaster, and Caenorhabditis elegans-to ours. Using a 5-fold cross-validation test, DeepSplicer achieves an accuracy of 96.65% for acceptor homo sapiens dataset and 94.75% for donor homo sapiens dataset. The datasets used and models generated are available at our GitHub repository here: https://github.com/OluwadareLab/DeeoSolicer.
Victor Akpokiro, Oluwatosin Oluwadare, Jugal K. Kalita
ICMLA3
2021 Character-level Adversarial Examples in Arabic
abstract
Several adversarial attacks have been pro-posed in the domains of computer vision and natural language processing (NLP). However, most attacks in the NLP domain have been applied to evaluate deep neural networks (DNNs) that were trained on English corpora. This paper proposes the first set of character-level adversarial attacks designed for models trained on Arabic. We present an efficient method to generate character-level adversarial examples against neural classifiers. Our method relies on flip operations that were designed based on the most common spelling mistakes that non-native Arabic learners make. We find that only a few manipulations are needed to mislead powerful and popular DNN-based classifiers trained on Arabic corpora.
Basemah Alshemali, Jugal K. Kalita
ICMLA2
2021 UIFDBC: Effective density based clustering to find clusters of arbitrary shapes without user input
Hussain Ahmed Chowdhury, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Expert Syst. Appl.3
2021 A Survey of the Usages of Deep Learning for Natural Language Processing
abstract
Over the last several years, the field of natural language processing has been propelled forward by an explosion in the use of deep learning models. This article provides a brief introduction to the field and a quick overview of deep learning architectures and methods. It then sifts through the plethora of recent studies and summarizes a large assortment of relevant contributions. Analyzed research areas include several core linguistic processing issues in addition to many applications of computational linguistics. A discussion of the current state of the art is then provided along with recommendations for future research in the field.
Daniel W. Otter, Julian R. Medina, Jugal K. Kalita
IEEE Trans. Neural Networks Learn. Syst.3
2020 Improving the Reliability of Deep Neural Networks in NLP: A Review
Basemah Alshemali, Jugal K. Kalita
Knowl. Based Syst.2
2020 Multi-task learning for natural language processing in the 2020s: Where are we going?
Joseph Worsham, Jugal K. Kalita
Pattern Recognit. Lett.2
2020 Assessing the Effectiveness of Causality Inference Methods for Gene Regulatory Networks
abstract
Causality inference is the use of computational techniques to predict possible causal relationships for a set of variables, thereby forming a directed network. Causality inference in Gene Regulatory Networks (GRNs) is an important, yet challenging task due to the limits of available data and lack of efficiency in existing causality inference techniques. A number of techniques have been proposed and applied to infer causal relationships in various domains, although they are not specific to regulatory network inference. In this paper, we assess the effectiveness of methods for inferring causal GRNs. We introduce seven different inference methods and apply them to infer directed edges in GRNs. We use time-series expression data from the DREAM challenges to assess the methods in terms of quality of inference and rank them based on performance. The best method is applied to Breast Cancer data to infer a causal network. Experimental results show that Causation Entropy is best, however, highly time-consuming and not feasible to use in a relatively large network. We infer Breast Cancer GRN with the second-best method, Transfer Entropy. The topological analysis of the network reveals that top out-degree genes such as SLC39A5 which are considered central genes, play important role in cancer progression.
Syed Sazzad Ahmed, Swarup Roy, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Differential Expression Analysis of RNA-seq Reads: Overview, Taxonomy, and Tools
abstract
Analysis of RNA-sequence (RNA-seq) data is widely used in transcriptomic studies and it has many applications. We review RNA-seq data analysis from RNA-seq reads to the results of differential expression analysis. In addition, we perform a descriptive comparison of tools used in each step of RNA-seq data analysis along with a discussion of important characteristics of these tools. A taxonomy of tools is also provided. A discussion of issues in quality control and visualization of RNA-seq data is also included along with useful tools. Finally, we provide some guidelines for the RNA-seq data analyst, along with research issues and challenges which should be addressed.
Hussain Ahmed Chowdhury, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 (Differential) Co-Expression Analysis of Gene Expression: A Survey of Best Practices
abstract
Analysis of gene expression data is widely used in transcriptomic studies to understand functions of molecules inside a cell and interactions among molecules. Differential co-expression analysis studies diseases and phenotypic variations by finding modules of genes whose co-expression patterns vary across conditions. We review the best practices in gene expression data analysis in terms of analysis of (differential) co-expression, co-expression network, differential networking, and differential connectivity considering both microarray and RNA-seq data along with comparisons. We highlight hurdles in RNA-seq data analysis using methods developed for microarrays. We include discussion of necessary tools for gene expression analysis throughout the paper. In addition, we shed light on scRNA-seq data analysis by including preprocessing and scRNA-seq in co-expression analysis along with useful tools specific to scRNA-seq. To get insights, biological interpretation and functional profiling is included. Finally, we provide guidelines for the analyst, along with research issues and challenges which should be addressed.
Hussain Ahmed Chowdhury, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.3
2019 Active learning to detect DDoS attack using ranked features
Rup Kumar Deka, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Comput. Commun.3
2018 Genre Identification and the Compositional Effect of Genre in Literature
abstract
Recent advances in Natural Language Processing are finding ways to place an emphasis on the hierarchical nature of text instead of representing language as a flat sequence or unordered collection of words or letters. A human reader must capture multiple levels of abstraction and meaning in order to formulate an understanding of a document. In this paper, we address the problem of developing approaches which are capable of working with extremely large and complex literary documents to perform Genre Identification. The task is to assign the literary classification to a full-length book belonging to a corpus of literature, where the works on average are well over 200,000 words long and genre is an abstract thematic concept. We introduce the Gutenberg Dataset for Genre Identification. Additionally, we present a study on how current deep learning models compare to traditional methods for this task. The results are presented as a baseline along with findings on how using an ensemble of chapters can significantly improve results in deep learning methods. The motivation behind the ensemble of chapters method is discussed as the compositionality of subtexts which make up a larger work and contribute to the overall genre.
Joseph Worsham, Jugal K. Kalita
COLING2
2018 Parallel Attention Mechanisms in Neural Machine Translation
abstract
Recent papers in neural machine translation have proposed the strict use of attention mechanisms over previous standards such as recurrent and convolutional neural networks (RNNs and CNNs). We propose that by running traditionally? stacked encoding branches from encoder-decoder attention-focused architectures in parallel, that even more sequential operations can be removed from the model and thereby decrease training time. In particular, we modify the recently published attention-based architecture called Transformer by Google, by replacing sequential attention modules with parallel ones, reducing the amount of training time and substantially improving BLEU scores at the same time. Experiments over the English to German and English to French translation tasks show that our model establishes a new state of the art.
Julian R. Medina, Jugal K. Kalita
ICMLA2
2018 Exploring Sentence Vector Spaces through Automatic Summarization
abstract
Given vector representations for individual words, it is necessary to compute vector representations of sentences for many applications in a compositional manner, often using artificial neural networks. Relatively little work has explored the internal structure and properties of such sentence vectors. In this paper, we explore the properties of sentence vectors in the context of automatic summarization. In particular, we show that cosine similarity between sentence vectors and document vectors is strongly correlated with sentence importance and that vector semantics can identify and correct gaps between the sentences chosen so far and the document. In addition, we identify specific dimensions which are linked to effective summaries. To our knowledge, this is the first time specific dimensions of sentence embeddings have been connected to sentence properties. We also compare the features of different methods of sentence embeddings. Many of these insights have applications in uses of sentence embeddings far beyond summarization.
Adly Templeton, Jugal K. Kalita
ICMLA2
2018 A detection framework for semantic code clones and obfuscated code
Abdullah Sheneamer, Swarup Roy, Jugal K. Kalita
Expert Syst. Appl.3
2018 A survey of detection methods for XSS attacks
Upasana Sarmah, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
J. Netw. Comput. Appl.3
2018 Index-Based Network Aligner of Protein-Protein Interaction Networks
abstract
Network Alignment over graph-structured data has received considerable attention in many recent applications. Global network alignment tries to uniquely find the best mapping for a node in one network to only one node in another network. The mapping is performed according to some matching criteria that depend on the nature of data. In molecular biology, functional orthologs, protein complexes, and evolutionary conserved pathways are some examples of information uncovered by global network alignment. Current techniques for global network alignment suffer from several drawbacks, e.g., poor performance and high memory requirements. We address these problems by proposing IBNAL, Indexes-Based Network ALigner, for better alignment quality and faster results. To accelerate the alignment step, IBNAL makes use of a novel clique-based index and is able to align large networks in seconds. IBNAL produces a higher topological quality alignment and comparable biological match in alignment relative to other state-of-the-art aligners even though topological fit is primarily used to match nodes. IBNAL's results confirm and give another evidence that homology information is more likely to be encoded in network topology than sequence information.
Ahed Elmsallati, Abdulghani Msalati, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Performing local network alignment by ensembling global aligners
abstract
Interactions among proteins are important mechanisms in living cells. The whole set of interactions is often referred to as a protein-protein interaction network (PIN). Comparison among such networks may discover conserved (or disrupted) patterns of interactions among species. Such comparison is performed using network alignment algorithms. They help analyse PPI networks for a better understanding of biological processes such as finding conserved regions between species, giving us insight into their evolution. However, there is no best aligner or standard evaluation measure to assess the quality of alignments. In this work, we use several aligners to produce an ensembled result, which can further improve individual aligners' alignment quality. Two basic ensemble approaches are used: One by finding majority node mappings from aligners and another by combining their results into one final alignment. These alignments are then evaluated based on three scoring schemes: Gene Ontology Consistency (GOC), Node Coverage (NCV) and Generalised Symmetric Substructure Score (GS3) using IsoBase PPI networks. Results show that the majority voting based ensemble scheme performs well in GS3while the ensemble by the union of the decision by different aligners produces satisfactory outcomes in comparison in GOC and NCV scores.
Hazel N. Manners, Ahed Elmsallati, Pietro H. Guzzi, Swarup Roy, Jugal K. Kalita
BIBM5
2017 Schemes for Labeling Semantic Code Clones using Machine Learning
abstract
Machine learning approaches built to identify code clones fail to perform well due to insufficient training samples and have been restricted only up to Type-III clones. A majority of the publicly available code clone corpora are incomplete in nature and lack labeled samples for semantic or Type-IV clones. We present here two schemes for labeling all types of clones including Type-IV clones. We restrict our study to Java code only. First, we use an unsupervised approach to label Type-IV clones and validate them using expert Java programmers. Next, we present a supervised scheme for labeling (or classifying) unknown samples based on labeled samples derived from our first scheme. We evaluate the performance of our schemes using six well-known Java code clone corpora and report on the quality of produced clones in terms of kappa agreement, mean error and accuracy scores. Results show that both schemes produce high quality code clones facilitating future use of machine learning in detecting clones of Type-IV.
Abdullah Sheneamer, Hanan Hazazi, Swarup Roy, Jugal K. Kalita
ICMLA4
2016 Sentiment Analysis of Restaurant Reviews on Yelp with Incremental Learning
abstract
Sentiment analysis of customer reviews has a crucial impact on a business's development strategy. Despite the fact that a repository of reviews evolves over time, sentiment analysis often relies on offline solutions where training data is collected before the model is built. If we want to avoid retraining the entire model from time to time, incremental learning becomes the best alternative solution for this task. In this work, we present a variant of online random forests to perform sentiment analysis on customers' reviews. Our model is able to achieve accuracy similar to offline methods and comparable to other online models.
Tri Doan, Jugal K. Kalita
ICMLA2
2016 Semantic Clone Detection Using Machine Learning
abstract
If two fragments of source code are identical to each other, they are called code clones. Code clones introduce difficulties in software maintenance and cause bug propagation. In this paper, we present a machine learning framework to automatically detect clones in software, which is able to detect Types-3 and the most complicated kind of clones, Type-4 clones. Previously used traditional features are often weak in detecting the semantic clones The novel aspects of our approach are the extraction of features from abstract syntax trees (AST) and program dependency graphs (PDG), representation of a pair of code fragments as a vector and the use of classification algorithms. The key benefit of this approach is that our approach can find both syntactic and semantic clones extremely well. Our evaluation indicates that using our new AST and PDG features is a viable methodology, since they improve detecting clones on the IJaDataset 2.0.
Abdullah Sheneamer, Jugal K. Kalita
ICMLA2
2016 Automatic Algorithm Selection in Computational Software Using Machine Learning
abstract
Computational software programs, such as Maple and Mathematica, heavily rely on superfunctions and meta-algorithms to select the optimal algorithm for a given task. These meta-algorithms may require intensive mathematical proof to formulate, incur large computational overhead, or fail to consistently select the best algorithm. Machine learning demonstrates a promising alternative for automatic algorithm selection by easing the design process and overhead while also attaining high accuracy in selection. In a case study on the resultant superfunction, a trained neural network is able to select the best algorithm out of the four available 86% of the time in Maple and 78% of the time in Mathematica. When used as a replacement for pre-existing meta-algorithms, the neural network brings about a 68% runtime improvement in Maple and 49% improvement in Mathematica. Random forests, k-nearest neighbors, and both linear and RBF kernel SVMs are also compared to the neural network model, the latter of which offers the best performance out of the tested machine learning methods.
Matthew C. Simpson, Qing Yi, Jugal K. Kalita
ICMLA3
2016 A multi-step outlier-based anomaly detection approach to network-wide traffic
Monowar Bhuyan, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Inf. Sci.3
2016 E-LDAT: a lightweight system for DDoS flooding attack detection and IP traceback using extended entropy metric
abstract
Distributed denial-of-service DDoS attacks cause havoc by exploiting threats to Internet services. In this paper, we propose E-LDAT, a lightweight extended-entropy metric-based system for both DDoS flooding attack detection and IP Internet Protocol traceback. It aims to identify DDoS attacks effectively by measuring the metric difference between legitimate traffic and attack traffic. IP traceback is performed using the metric values for an attack sample detected by the detection scheme. The method uses a generalized entropy metric with packet intensity computation on the sampled network traffic with respect to time. The E-LDAT system has been evaluated using several real-world DDoS datasets and outperforms competing methods when detecting four classes of DDoS flooding attacks, including constant rate, pulsing rate, increasing rate and subgroup attacks. The IP traceback model is also evaluated using NetFlow data in near real-time and performs well in large-scale attack networks with zombies. Copyright © 2016 John Wiley & Sons, Ltd.
Monowar Bhuyan, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Secur. Commun. Networks3
2016 FFSc: a novel measure for low-rate and high-rate DDoS attack detection using multivariate data analysis
abstract
Abstract A Distributed Denial of Service (DDoS) attack is a major security threat for networks and Internet services. Attackers can generate attack traffic similar to normal network traffic using sophisticated attacking tools. In such a situation, many intrusion detection systems fail to identify DDoS attack in real time. However, DDoS attack traffic behaves differently from legitimate network traffic in terms of traffic features. Statistical properties of various features can be analyzed to distinguish the attack traffic from legitimate traffic. In this paper, we introduce a statistical measure called Feature Feature score for multivariate data analysis to distinguish DDoS attack traffic from normal traffic. We extract three basic parameters of network traffic, namely, entropy of source IPs, variation of source IPs, and packet rate to analyze the behavior of network traffic for attack detection. The method is validated using CAIDA DDoS 2007 and MIT DARPA datasets. Copyright © 2016 John Wiley & Sons, Ltd.
Nazrul Hoque, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Secur. Commun. Networks3
2016 Global Alignment of Protein-Protein Interaction Networks: A Survey
abstract
In this paper, we survey algorithms that perform global alignment of networks or graphs. Global network alignment aligns two or more given networks to find the best mapping from nodes in one network to nodes in other networks. Since graphs are a common method of data representation, graph alignment has become important with many significant applications. Protein-protein interactions can be modeled as networks and aligning these networks of protein interactions has many applications in biological research. In this survey, we review algorithms for global pairwise alignment highlighting various proposed approaches, and classify them based on their methodology. Evaluation metrics that are used to measure the quality of the resulting alignments are also surveyed. We discuss and present a comparison between selected aligners on the same datasets and evaluate using the same evaluation metrics. Finally, a quick overview of the most popular databases of protein interaction networks is presented focusing on datasets that have been used recently.
Ahed Elmsallati, Connor Clark, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 Automatically Creating a Large Number of New Bilingual Dictionaries
abstract
This paper proposes approaches to automatically createa large number of new bilingual dictionaries for low resource languages, especially resource-poor and endangered languages, from a single input bilingual dictionary. Our algorithms produce translations of wordsin a source language to plentiful target languages using available Wordnets and a machine translator (MT). Since our approaches rely on just one input dictionary, available Wordnets and an MT, they are applicable toany bilingual dictionary as long as one of the two languagesis English or has a Wordnet linked to the Princeton Wordnet. Starting with 5 available bilingual dictionaries,we create 48 new bilingual dictionaries. Of these, 30 pairs of languages are not supported by the popular MTs: Google and Bing.
Khang Nhut Lam, Feras Al Tarouti, Jugal K. Kalita
AAAI3
2015 MODULA: A network module based local protein interaction network alignment method
abstract
Biological networks are usually used to model interactions among biological macromolecules in a cells. For instance protein-protein interaction networks (PIN) are used to model and analyse the set of interactions among proteins. The comparison of networks may result in the identification of conserved patterns of interactions corresponding to biological relevant entities such as protein complexes and pathways. Several algorithms, known as network alignment algorithms, have been proposed to unravel relations between different species at the interactome level. Algorithms may be categorized in two main classes: merge and mine and mine and merge. Algorithms belonging to the first class initially merge input network into a single integrated and then mine such networks. Conversely algorithms belonging to the second class initially analyze separately two input networks then integrate such results. In this paper we present MODULA (Network Module based PPI Aligner), a novel approach for local network alignment that belong to the second class. The algorithm at first identifies compact modules from input networks. Modules of both networks are then matched using functional knowledge. Then it uses high scoring pairs of modules as seeds to build a bigger alignment. In order to asses MODULA we compared it to the state of the art local alignment algorithms over a rather extensive and updated dataset.
Pietro H. Guzzi, Pierangelo Veltri, Swarup Roy, Jugal K. Kalita
BIBM4
2015 A multiobjective memetic algorithm for PPI network alignment
abstract
MOTIVATION: There recently has been great interest in aligning protein-protein interaction (PPI) networks to identify potentially orthologous proteins between species. It is thought that the topological information contained in these networks will yield better orthology predictions than sequence similarity alone. Recent work has found that existing aligners have difficulty making use of both topological and sequence similarity when aligning, with either one or the other being better matched. This can be at least partially attributed to the fact that existing aligners try to combine these two potentially conflicting objectives into a single objective. RESULTS: We present Optnetalign, a multiobjective memetic algorithm for the problem of PPI network alignment that uses extremely efficient swap-based local search, mutation and crossover operations to create a population of alignments. This algorithm optimizes the conflicting goals of topological and sequence similarity using the concept of Pareto dominance, exploring the tradeoff between the two objectives as it runs. This allows us to produce many high-quality candidate alignments in a single run. Our algorithm produces alignments that are much better compromises between topological and biological match quality than previous work, while better characterizing the diversity of possible good alignments between two networks. Our aligner's results have several interesting implications for future research on alignment evaluation, the design of network alignment objectives and the interpretation of alignment results. AVAILABILITY AND IMPLEMENTATION: The C++ source code to our program, along with compilation and usage instructions, is available at https://github.com/crclark/optnetaligncpp/
Connor Clark, Jugal K. Kalita
Bioinform.2
2015 Network defense: Approaches, methods and techniques
Rup Kumar Deka, Kausthav Pratim Kalita, D. K. Bhattacharya, Jugal K. Kalita
J. Netw. Comput. Appl.4
2015 An empirical evaluation of information metrics for low-rate and high-rate DDoS attack detection
Monowar Bhuyan, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Pattern Recognit. Lett.3
2014 A comparison of algorithms for the pairwise alignment of biological networks
abstract
MOTIVATION: As biological inquiry produces ever more network data, such as protein-protein interaction networks, gene regulatory networks and metabolic networks, many algorithms have been proposed for the purpose of pairwise network alignment-finding a mapping from the nodes of one network to the nodes of another in such a way that the mapped nodes can be considered to correspond with respect to both their place in the network topology and their biological attributes. This technique is helpful in identifying previously undiscovered homologies between proteins of different species and revealing functionally similar subnetworks. In the past few years, a wealth of different aligners has been published, but few of them have been compared with one another, and no comprehensive review of these algorithms has yet appeared. RESULTS: We present the problem of biological network alignment, provide a guide to existing alignment algorithms and comprehensively benchmark existing algorithms on both synthetic and real-world biological data, finding dramatic differences between existing algorithms in the quality of the alignments they produce. Additionally, we find that many of these tools are inconvenient to use in practice, and there remains a need for easy-to-use cross-platform tools for performing network alignment.
Connor Clark, Jugal K. Kalita
Bioinform.2
2014 Reconstruction of gene co-expression network from microarray data using local expression patterns
abstract
BACKGROUND: Biological networks connect genes, gene products to one another. A network of co-regulated genes may form gene clusters that can encode proteins and take part in common biological processes. A gene co-expression network describes inter-relationships among genes. Existing techniques generally depend on proximity measures based on global similarity to draw the relationship between genes. It has been observed that expression profiles are sharing local similarity rather than global similarity. We propose an expression pattern based method called GeCON to extract Gene CO-expression Network from microarray data. Pair-wise supports are computed for each pair of genes based on changing tendencies and regulation patterns of the gene expression. Gene pairs showing negative or positive co-regulation under a given number of conditions are used to construct such gene co-expression network. We construct co-expression network with signed edges to reflect up- and down-regulation between pairs of genes. Most existing techniques do not emphasize computational efficiency. We exploit a fast correlogram matrix based technique for capturing the support of each gene pair to construct the network. RESULTS: We apply GeCON to both real and synthetic gene expression data. We compare our results using the DREAM (Dialogue for Reverse Engineering Assessments and Methods) Challenge data with three well known algorithms, viz., ARACNE, CLR and MRNET. Our method outperforms other algorithms based on in silico regulatory network reconstruction. Experimental results show that GeCON can extract functionally enriched network modules from real expression data. CONCLUSIONS: In view of the results over several in-silico and real expression datasets, the proposed GeCON shows satisfactory performance in predicting co-expression network in a computationally inexpensive way. We further establish that a simple expression pattern matching is helpful in finding biologically relevant gene network. In future, we aim to introduce an enhanced GeCON to identify Protein-Protein interaction network complexes by incorporating variable density concept.
Swarup Roy, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
BMC Bioinform.3
2014 Detecting Distributed Denial of Service Attacks: Methods, Tools and Future Directions
abstract
Distributed denial of service (DDoS) attack is a coordinated attack, generally performed on a massive scale on the availability of services of a target system or network resources. Owing to the continuous evolution of new attacks and ever-increasing number of vulnerable hosts on the Internet, many DDoS attack detection or prevention mechanisms have been proposed. In this paper, we present a comprehensive survey of DDoS attacks, detection methods and tools used in wired networks. The paper also highlights open issues, research challenges and possible solutions in this area.
Monowar Bhuyan, Hirak J. Kashyap, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Comput. J.4
2014 Extracting and Displaying Temporal and Geospatial Entities from Articles on Historical Events
abstract
This paper discusses a system that extracts and displays temporal and geospatial entities in text. The first task involves identification of all events in a document followed by identification of important events using a classifier. The second task involves identifying named entities associated with the document. In particular, we extract geospatial named entities. We disambiguate the set of geospatial named entities and geocode them to determine the correct coordinates for each place name, often called grounding. We resolve ambiguity based on sentence and article context. Finally, we present a user with the key events and their associated people, places and organizations within a document in terms of a timeline and a map. For purposes of testing, we use Wikipedia articles about historical events, such as those describing wars, battles and invasions. We focus on extracting major events from the articles, although our ideas and tools can be easily used with articles from other sources such as news articles. We use several existing tools such as Evita, Google Maps, publicly available implementationsofSupportVectorMachines,HiddenMarkovModelandConditionalRandomField, and the MIT SIMILE Timeline.
Rachel Chasin, Daryl Woodward, Jeremy Witmer, Jugal K. Kalita
Comput. J.4
2014 MLH-IDS: A Multi-Level Hybrid Intrusion Detection Method
abstract
With the growth of networked computers and associated applications, intrusion detection has become essential to keeping networks secure. A number of intrusion detection methods have been developed for protecting computers and networks using conventional statistical methods as well as data mining methods. Data mining methods for misuse and anomaly-based intrusion detection, usually encompass supervised, unsupervised and outlier methods. It is necessary that the capabilities of intrusion detection methods be updated with the creation of new attacks. This paper proposes a multi-level hybrid intrusion detection method that uses a combination of supervised, unsupervised and outlier-based methods for improving the efficiency of detection of new and old attacks. The method is evaluated with a captured real-time flow and packet dataset called the Tezpur University intrusion detection system (TUIDS) dataset, a distributed denial of service dataset, and the benchmark intrusion dataset called the knowledge discovery and data mining Cup 1999 dataset and the new version of KDD (NSL-KDD) dataset. Experimental results are compared with existing multi-level intrusion detection methods and other classifiers. The performance of our method is very good.
Prasanta Gogoi, Dhruba Kumar Bhattacharyya, Bhogeswar Borah, Jugal K. Kalita
Comput. J.4
2014 Summarization of Twitter Microblogs
abstract
Owing to the sheer volume of text generated by a microblog site like Twitter, it is often difficult to fully understand what is being said about various topics. This paper presents algorithms for summarizing microblog documents. Initially, we present algorithms that produce single-document summaries but later extend them to produce summaries containing multiple documents. We evaluate the generated summaries by comparing them to both manually produced summaries and, for the multiple-post summaries, to the summarization results of some of the leading traditional summarization systems.
Beaux Sharifi, David I. Inouye, Jugal K. Kalita
Comput. J.3
2014 MIFS-ND: A mutual information-based feature selection method
Nazrul Hoque, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Expert Syst. Appl.3
2014 Network attacks: Taxonomy, tools and systems
Nazrul Hoque, Monowar Bhuyan, Ram Charan Baishya, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
J. Netw. Comput. Appl.5
2014 Stemming resource-poor Indian languages
abstract
Stemming is a basic method for morphological normalization of natural language texts. In this study, we focus on the problem of stemming several resource-poor languages from Eastern India, viz., Assamese, Bengali, Bishnupriya Manipuri and Bodo. While Assamese, Bengali and Bishnupriya Manipuri are Indo-Aryan, Bodo is a Tibeto-Burman language. We design a rule-based approach to remove suffixes from words. To reduce over-stemming and under-stemming errors, we introduce a dictionary of frequent words. We observe that, for these languages a dominant amount of suffixes are single letters creating problems during suffix stripping. As a result, we introduce an HMM-based hybrid approach to classify the mis-matched last character. For each word, the stem is extracted by calculating the most probable path in four HMM states. At each step we measure the stemming accuracy for each language. We obtain 94% accuracy for Assamese and Bengali and 87%, and 82% for Bishnupriya Manipuri and Bodo, respectively, using the hybrid approach. We compare our work with Morfessor [Creutz and Lagus 2005]. As of now, there is no reported work on stemming for Bishnupriya Manipuri and Bodo. Our results on Assamese and Bengali show significant improvement over prior published work [Sarkar and Bandyopadhyay 2008; Sharma et al. 2002, 2003].
Navanath Saharia, Utpal Sharma, Jugal K. Kalita
ACM Trans. Asian Lang. Inf. Process.3
2014 Shifting-and-Scaling Correlation Based Biclustering Algorithm
abstract
The existence of various types of correlations among the expressions of a group of biologically significant genes poses challenges in developing effective methods of gene expression data analysis. The initial focus of computational biologists was to work with only absolute and shifting correlations. However, researchers have found that the ability to handle shifting-and-scaling correlation enables them to extract more biologically relevant and interesting patterns from gene microarray data. In this paper, we introduce an effective shifting-and-scaling correlation measure named Shifting and Scaling Similarity (SSSim), which can detect highly correlated gene pairs in any gene expression data. We also introduce a technique named Intensive Correlation Search (ICS) biclustering algorithm, which uses SSSim to extract biologically significant biclusters from a gene expression data set. The technique performs satisfactorily with a number of benchmarked gene expression data sets when evaluated in terms of functional categories in Gene Ontology database.
Hasin Afzal Ahmed, Priyakshi Mahanta, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
IEEE ACM Trans. Comput. Biol. Bioinform.4
2013 An Improved Stemming Approach Using HMM for a Highly Inflectional Language
Navanath Saharia, Kishori M. Konwar, Utpal Sharma, Jugal K. Kalita
CICLing (1)4
2013 Better Twitter Summaries?
Joel Judd, Jugal K. Kalita
HLT-NAACL2
2013 Creating Reverse Bilingual Dictionaries
Khang Nhut Lam, Jugal K. Kalita
HLT-NAACL2
2013 CoBi: Pattern Based Co-Regulated Biclustering of Gene Expression Data
Swarup Roy, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Pattern Recognit. Lett.3
2013 Design and Evaluation of Soft Keyboards for Brahmic Scripts
abstract
Despite being spoken by a large percentage of the world, Indic languages in general lack user-friendly and efficient methods for text input. These languages have poor or no support for typing. Soft keyboards, because of their ease of installation and lack of reliance on specific hardware, are a promising solution as an input device for many languages. Developing an acceptable soft keyboard requires the frequency analysis of characters in order to design a layout that minimizes text-input time. This article proposes the use of various development techniques, layout variations, and evaluation methods for the creation of soft keyboards for Brahmic scripts. We propose that using optimization techniques such as genetic algorithms and multi-objective Pareto optimization to develop multi-layer keyboards will increase the speed at which text can be entered.
Lauren Hinkle, Albert Brouillette, Sujay Jayakar, Leigh Gathings, Miguel Lezcano, Jugal K. Kalita
ACM Trans. Asian Lang. Inf. Process.6
2013 Cutting Plane Training for Linear Support Vector Machines
abstract
Support Vector Machines (SVMs) have been shown to achieve high performance on classification tasks across many domains, and a great deal of work has been dedicated to developing computationally efficient training algorithms for linear SVMs. One approach [1] approximately minimizes risk through use of cutting planes, and is improved by [2], [3]. We build upon this work, presenting a modification to the algorithm developed by Franc and Sonnenburg [2]. We demonstrate empirically that our changes can reduce cutting plane training time by up to 40 percent, and discuss how changes in data sets and parameter settings affect the effectiveness of our method.
Nicholas A. Arnosti, Jugal K. Kalita
IEEE Trans. Knowl. Data Eng.2
2012 Deterministic Approach for Biclustering of Co-Regulated Genes from Gene Expression Data
abstract
This paper presents an expression pattern based biclustering technique for grouping both positively and negatively regulated genes together as co-regulated genes from microarray expression data. Most interesting variants of this problem are NP-complete requiring either large computational effort or the use of lossy heuristics to short circuit the calculation. Our approach deterministically finds all biclusters using a non-greedy approach in polynomial time. Various real datasets have been used for experiments and results are excellent.
Swarup Roy, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
KES3
2012 Summarization of Historical Articles Using Temporal Event Clustering
James Gung, Jugal K. Kalita
HLT-NAACL2
2012 An effective method for network module extraction from microarray data
abstract
BACKGROUND: The development of high-throughput Microarray technologies has provided various opportunities to systematically characterize diverse types of computational biological networks. Co-expression network have become popular in the analysis of microarray data, such as for detecting functional gene modules. RESULTS: This paper presents a method to build a co-expression network (CEN) and to detect network modules from the built network. We use an effective gene expression similarity measure called NMRS (Normalized mean residue similarity) to construct the CEN. We have tested our method on five publicly available benchmark microarray datasets. The network modules extracted by our algorithm have been biologically validated in terms of Q value and p value. CONCLUSIONS: Our results show that the technique is capable of detecting biologically significant network modules from the co-expression network. Biologist can use this technique to find groups of genes with similar functionality based on their expression information.
Priyakshi Mahanta, Hasin Afzal Ahmed, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
BMC Bioinform.4
2011 GERC: Tree Based Clustering for Gene Expression Data
abstract
Measurement of gene expression using DNA micro arrays have revolutionized biological and medical research. This paper presents a divisive clustering algorithm that produces a tree of genes called GERC tree along with the generated clusters. Unlike a dendrogram, a GERC tree is a general tree and it is an ample resource for biological information about the genes in a data set. The leaves of the tree represent the desired clusters. The clustering method was tested with several real-life data sets and the proposed method has been found satisfactory.
Hasin Afzal Ahmed, Priyakshi Mahanta, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
BIBE4
2011 Surveying Port Scans and Their Detection Methodologies
abstract
Scanning of ports on a computer occurs frequently on the Internet. An attacker performs port scans of Internet protocol addresses to find vulnerable hosts to compromise. However, it is also useful for system administrators and other network defenders to detect port scans as possible preliminaries to more serious attacks. It is a very difficult task to recognize instances of malicious port scanning. In general, a port scan may be an instance of a scan by attackers or an instance of a scan by network defenders. In this survey, we present research and development trends in this area. Our presentation includes a discussion of common port scan attacks. We provide a comparison of port scan methods based on type, mode of detection, mechanism used for detection and other characteristics. This survey also reports on the available data sets and evaluation criteria for port scan detection approaches.
Monowar Bhuyan, Dhruba Kumar Bhattacharyya, Jugal K. Kalita
Comput. J.3
2011 A Survey of Outlier Detection Methods in Network Anomaly Identification
abstract
The detection of outliers has gained considerable interest in data mining with the realization that outliers can be the key discovery to be made from very large databases. Outliers arise due to various reasons such as mechanical faults, changes in system behavior, fraudulent behavior, human error and instrument error. Indeed, for many applications the discovery of outliers leads to more interesting and useful results than the discovery of inliers. Detection of outliers can lead to identification of system faults so that administrators can take preventive measures before they escalate. It is possible that anomaly detection may enable detection of new attacks. Outlier detection is an important anomaly detection approach. In this paper, we present a comprehensive survey of well-known distance-based, density-based and other techniques for outlier detection and compare them. We provide definitions of outliers and discuss their detection based on supervised and unsupervised learning in the context of network anomaly detection.
Prasanta Gogoi, Dhruba Kumar Bhattacharyya, Bhogeswar Borah, Jugal K. Kalita
Comput. J.4
2010 Summarizing Microblogs Automatically
Beaux Sharifi, Mark-Anthony Hutton, Jugal K. Kalita
HLT-NAACL3
2008 Acquisition of Morphology of an Indic Language from Text Corpus
abstract
This article describes an approach to unsupervised learning of morphology from an unannotated corpus for a highly inflectional Indo-European language called Assamese spoken by about 30 million people. Although Assamese is one of Indias national languages, it utterly lacks computational linguistic resources. There exists no prior computational work on this language spoken widely in northeast India. The work presented is pioneering in this respect. In this article, we discuss salient issues in Assamese morphology where the presence of a large number of suffixal determiners, sandhi, samas, and the propensity to use suffix sequences make approximately 50% of the words used in written and spoken text inflected. We implement methods proposed by Gaussier and Goldsmith on acquisition of morphological knowledge, and obtain F-measure performance below 60%. This motivates us to present a method more suitable for handling suffix sequences, enabling us to increase the F-measure performance of morphology acquisition to almost 70%. We describe how we build a morphological dictionary for Assamese from the text corpus. Using the morphological knowledge acquired and the morphological dictionary, we are able to process small chunks of data at a time as well as a large corpus. We achieve approximately 85% precision and recall during the analysis of small chunks of coherent text.
Utpal Sharma, Jugal K. Kalita, Rajib K. Das
ACM Trans. Asian Lang. Inf. Process.2
2003 The Significance of Temporal-Difference Learning in Self-Play Training TD-Rummy versus EVO-rummy
Clifford Kotnik, Jugal K. Kalita
ICML2
2003 Ant algorithms for the optimal restoration of distribution feeders during cold load pickup
abstract
The ant colony algorithm is a new technique for combinatorial optimization borrowed from swarm intelligence. This paper outlines an ant colony algorithm to compute the optimal order of restoring sections in a power distribution system. Restoration of distribution feeders after long interruptions creates cold load pickup conditions due to loss of diversity among the loads. The distribution system load may have to be restored step-by-step using sectionalizing switches under such conditions to prevent overheating of substation transformer. The restoration time is dependent on the order in which sections are restored. Results obtained using this method for two test cases are presented including a comparison with the simulated annealing algorithm.
Indira Mohanty, Jugal K. Kalita, Sanjoy Das, Anil Pahwa, Erik C. Buehler
SIS2
2002 Efficient handling of high-dimensional feature spaces by randomized classifier ensembles
abstract
Handling massive datasets is a difficult problem not only due to prohibitively large numbers of entries but in some cases also due to the very high dimensionality of the data. Often, severe feature selection is performed to limit the number of attributes to a manageable size, which unfortunately can lead to a loss of useful information. Feature space reduction may well be necessary for many stand-alone classifiers, but recent advances in the area of ensemble classifier techniques indicate that overall accurate classifier aggregates can be learned even if each individual classifier operates on incomplete "feature view" training data, i.e., such where certain input attributes are excluded. In fact, by using only small random subsets of features to build individual component classifiers, surprisingly accurate and robust models can be created. In this work we demonstrate how these types of architectures effectively reduce the feature space for submodels and groups of sub-models, which lends itself to efficient sequential and/or parallel implementations. Experiments with a randomized version of Adaboost are used to support our arguments, using the text classification task as an example.
Alek Kolcz, Xiaomei Sun, Jugal K. Kalita
KDD3
2001 Summarization as Feature Selection for Text Categorization
abstract
We address the problem of evaluating the effectiveness of summarization techniques for the task of document categorization. It is argued that for a large class of automatic categorization algorithms, extraction-based document categorization can be viewed as a particular form of feature selection performed on the full text of the document and, in this context, its impact can be compared with state-of-the-art feature selection techniques especially devised to provide good categorization performance. Such a framework provides for a better assessment of the expected performance of a categorizer if the compression rate of the summarizer is known.
Alek Kolcz, Vidya Prabakarmurthi, Jugal K. Kalita
CIKM3
2000 Parsing and Interpretation in the Minimalist Paradigm
abstract
In this paper, we discuss how recent theoretical linguistic research focusing on the Minimalist Program (MP) (Cho95, Mar95, Zwa94) can be used to guide the parsing of a useful range of natural language sentences and the building of a logical representation in a principles‐based manner. We discuss the components of the MP and give an example derivation. We then propose parsing algorithms that recreate the derivation structure starting with a lexicon and the surface form of a sentence. Given the approximated derivation structure, MP principles are applied to generate a logical form, which leads to linguistically based algorithms for determining possible meanings for sentences that are ambiguous due to quantifier scope.
James S. Williams, Jugal K. Kalita
Comput. Intell.2
1997 An Informal Semantic Analysis of Motion Verbs Based on Physical Primitives
abstract
A representation scheme for verbs and prepositions specifying path and locative information is developed. The representation emphasizes the implementability of the underlying semantic primitives. The primitives pertain to mechanical characteristics such as geometric relationships among objects, force or motion characteristics implied by verbs, and their prepositional modifiers. This representation has been used to animate the performance of tasks underlying natural language imperatives.
Jugal K. Kalita, Joel C. Lee
Comput. Intell.1
1997 Situation assessment and prediction in intelligence domains
Lisa A. Jesse, Jugal K. Kalita
Knowl. Based Syst.2
1991 Interpreting Prepositions Physically
Jugal K. Kalita, Norman I. Badler
AAAI1
1989 Automatically Generating Natural Language Reports
abstract
In this paper, we describe a system which generates natural language status reports for a set of inter-related processes at various stages of progress. The system has three modules—a rule-based domain knowledge representation module, and elaborate text planning module, and a surface generation module. The knowledge representation module models a set of processes that are encountered in a typical office environment, using a body of production rules which are explicitly sequenced using an augmented Petri net mechanism. The system employs an interval-based temporal network for storing historical information. A text planning module traverses this network to search for events which need to be mentioned in a coherent report describing the current status of the system. The planner combines similar information for succinct presentation whenever applicable. It also employs discourse focus techniques and a simple notion of view transforms for the generation of good quality text. Finally, an available surface generation module which has been suitably augmented is used to produce well-structured textual reports for our chosen domain.
Jugal K. Kalita
Int. J. Man Mach. Stud.1
1986 Summarizing Natural Language Database Responses
Jugal K. Kalita, Marlene L. Jones, Gordon I. McCalla
Comput. Linguistics1
1984 A Response to the Need for Summary Responses
abstract
In this paper we argue that natural language interfaces to databases should be able to produce summary responses as well as listing actual data. We describe a system (incorporating a number of heuristics and a knowledge base built on top of the database) that has been developed to generate such summary responses. It is largely domain-independent, has been tested on many examples, and handles a wide variety of situations where summary responses would be useful.
Jugal K. Kalita, Marlene J. Colbourn, Gordon I. McCalla
COLING1