Boian S. Alexandrov

dblp:41/11017 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
21since 2021 · last 2025
0000-0001-8636-4603ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Systems, architecture and hardware · 4 · 3 since 2021Security and privacy · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 LoRID: Low-Rank Iterative Diffusion for Adversarial Purification
abstract
This work presents an information-theoretic examination of diffusion-based purification methods, the state-of-the-art adversarial defenses that utilize diffusion models to remove malicious perturbations in adversarial examples. By theoretically characterizing the inherent purification errors associated with the Markov-based diffusion purifications, we introduce LoRID, a novel Low-Rank Iterative Diffusion purification method designed to remove adversarial perturbation with low intrinsic purification errors. LoRID centers around a multi-stage purification process that leverages multiple rounds of diffusion-denoising loops at the early time-steps of the diffusion models, and the integration of Tucker decomposition, an extension of matrix factorization, to remove adversarial noise at high-noise regimes. Consequently, LoRID increases the effective diffusion time-steps and overcomes strong adversarial attacks, achieving superior robustness performance in CIFAR-10/100, CelebA-HQ, and ImageNet datasets under both white-box and grey-box settings.
Geigh Zollicoffer, Minh N. Vu, Ben Nebgen, Juan Castorena, Boian S. Alexandrov, Manish Bhattarai
AAAI5
2025 Topic Modeling and Link-Prediction for Material Property Discovery
abstract
Link prediction is a key network analysis technique that infers missing or future relations between nodes in a graph, based on observed patterns of connectivity. Scientific literature networks and knowledge graphs are typically large, sparse, and noisy, and often contain missing links, potential but unobserved connections, between concepts, entities, or methods. Here, we present an AI-driven hierarchical link prediction framework that integrates matrix factorization to infer hidden associations and steer discovery in complex material domains. Our method combines Hierarchical Nonnegative Matrix Factorization (HNMFk), Boolean matrix factorization (BNMFk) with automatic model selection. These discrete factors are then fused with Logistic matrix factorization (LMF), we use to construct a three-level topic tree from a 46,862-document corpus focused on 73 transition-metal dichalcogenides (TMDs). This class of materials has been studied in a variety of physics fields and has a multitude of current and potential applications.
Ryan Barron, Maksim Ekin Eren, Valentin G. Stanev, Cynthia Matuszek, Boian S. Alexandrov
DocEng5
2025 Bridging Legal Knowledge and AI: Retrieval-Augmented Generation with Vector Stores, Knowledge Graphs, and Hierarchical Non-negative Matrix Factorization
abstract
Agentic Generative AI, powered by Large Language Models (LLMs) and enhanced with Retrieval-Augmented Generation (RAG), Knowledge Graphs (KGs), and Vector Stores (VSs), represents a transformative technology applicable across specialized domains such as legal systems, research, recommender systems, cybersecurity, and global security, including proliferation research. This technology excels at inferring relationships within vast unstructured or semi-structured datasets. The legal domain we focus on here comprises inherently complex data characterized by extensive, interrelated, and semi-structured knowledge systems with complex relations. It comprises constitutions, statutes, regulations, and case law. Extracting insights and navigating the intricate networks of legal documents and their relations is crucial for effective legal research and decision-making. Here, we introduce a generative AI system, a jurisdiction-specific legal information retrieval that integrates RAG, VS, and KG, constructed via Hierarchical Non-Negative Matrix Factorization (HNMFk), to enhance information retrieval and AI reasoning and minimize hallucinations. In the legal system, these technologies empower AI agents to identify and analyze complex connections among cases, statutes, and legal precedents, uncovering hidden relationships and predicting legal trends—challenging tasks essential for ensuring justice and improving operational efficiency. Our system employs web scraping techniques to systematically collect legal texts, such as statutes, constitutional provisions, and case law, from publicly accessible platforms like Justia. It bridges the gap between traditional keyword-based searches and contextual understanding by leveraging advanced semantic representations, hierarchical relationships, and latent topic discovery. This approach is demonstrated in legal document clustering, summarization, and cross-referencing tasks. The framework marks a significant step toward augmenting legal research with scalable, interpretable, and accurate retrieval methods for semi-structured data, advancing the intersection of computational law and artificial intelligence.
Ryan Barron, Maksim Ekin Eren, Olga M. Serafimova, Cynthia Matuszek, Boian S. Alexandrov
ICAIL5
2025 Topological Signatures of Adversaries in Multimodal Alignments
abstract
Multimodal Machine Learning systems, particularly those aligning text and image data like CLIP/BLIP models, have become increasingly prevalent, yet remain susceptible to adversarial attacks. While substantial research has addressed adversarial robustness in unimodal contexts, defense strategies for multimodal systems are underexplored. This work investigates the topological signatures that arise between image and text embeddings and shows how adversarial attacks disrupt their alignment, introducing distinctive signatures. We specifically leverage persistent homology and introduce two novel Topological-Contrastive losses based on Total Persistence and Multi-scale kernel methods to analyze the topological signatures introduced by adversarial perturbations. We observe a pattern of monotonic changes in the proposed topological losses emerging in a wide range of attacks on image-text alignments, as more adversarial samples are introduced in the data. By designing an algorithm to back-propagate these signatures to input samples, we are able to integrate these signatures into Maximum Mean Discrepancy tests, creating a novel class of tests that leverage topological signatures for better adversarial detection.
Minh Nhat Vu, Geigh Zollicoffer, Huy Quang Mai, Ben Nebgen, Boian S. Alexandrov, Manish Bhattarai
ICML5
2025 Specificity and tunability of efflux pumps: A new role for the proton gradient?
abstract
Efflux pumps that transport antibacterial drugs out of bacterial cells have broad specificity, commonly leading to broad spectrum resistance and limiting treatment strategies for infections. It remains unclear how efflux pumps can maintain this broad spectrum specificity to diverse drug molecules while limiting the efflux of other cytoplasmic content. We have investigated the origins of this broad specificity using theoretical models informed by the experimentally determined structural and kinetic properties of efflux pumps. We developed a set of mathematical models describing operation of efflux pumps as a discrete cyclic stochastic process across a network of states characterizing pump conformations and the presence/absence of bound ligands and protons. These include a minimal three-state model that lends itself to clear analytic calculations as well as a five-state model that relaxes some of the simpler model's most strict assumptions. We found that the pump specificity is determined not solely by the drug affinity to the pump-as is commonly assumed-but it is also directly affected by the periplasmic pH and the transmembrane potential. Therefore, changes to the proton concentration gradient and voltage drop across the membrane can influence how effective the pump is at extruding a particular drug molecule. Furthermore, we found that while both the proton concentration gradient across the membrane and the transmembrane potential contribute to the thermodynamic force driving the pump, their effects on the efflux enter not strictly in a combined proton motive force. Rather, they have two distinguishable effects on the overall throughput. These results highlight the unexpected effects of thermodynamic driving forces out of equilibrium and illustrate how efflux pump structure and function are conducive to the emergence of multidrug resistance.
Matthew Gerry, Duncan Kirby, Boian S. Alexandrov, Dvira Segal, Anton Zilman
PLoS Comput. Biol.3
2024 TopicTag: Automatic Annotation of NMF Topic Models Using Chain of Thought and Prompt Tuning with LLMs
abstract
Topic modeling is a technique for organizing and extracting themes from large collections of unstructured text. Non-negative matrix factorization (NMF) is a common unsupervised approach that decomposes a term frequency-inverse document frequency (TF-IDF) matrix to uncover latent topics and segment the dataset accordingly. While useful for highlighting patterns and clustering documents, NMF does not provide explicit topic labels, necessitating subject matter experts (SMEs) to assign labels manually. We present a methodology for automating topic labeling in documents clustered via NMF with automatic model determination (NMFk). By leveraging the output of NMFk and employing prompt engineering, we utilize large language models (LLMs) to generate accurate topic labels. Our case study on over 34,000 scientific abstracts on Knowledge Graphs demonstrates the effectiveness of our method in enhancing knowledge management and document organization.
Selma Wanna, Nick Solovyev 0001, Ryan Barron, Maksim Ekin Eren, Manish Bhattarai, Kim Ø. Rasmussen, Boian S. Alexandrov
DocEng7
2024 Tensor Train Low-rank Approximation (TT-LoRA): Democratizing AI with Accelerated LLMs
abstract
In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing (NLP) tasks, such as question-answering, sentiment analysis, text summarization, and machine translation. However, the ever-growing complexity of LLMs demands immense computational resources, hindering the broader research and application of these models. To address this, various parameter-efficient fine-tuning strategies, such as Low-Rank Approximation (LoRA) and Adapters, have been developed. Despite their potential, these methods often face limitations in compressibility. Specifically, LoRA struggles to scale effectively with the increasing number of trainable parameters in modern large scale LLMs. Additionally, Low-Rank Economic Tensor-Train Adaptation (LoRETTA), which utilizes tensor train decomposition, has not yet achieved the level of compression necessary for fine-tuning very large scale models with limited resources. This paper introduces Tensor Train Low-Rank Approximation (TT-LoRA), a novel parameter-efficient fine-tuning (PEFT) approach that extends LoRETTA with optimized tensor train (TT) decomposition integration. By eliminating Adapters and traditional LoRA-based structures, TT-LoRA achieves greater model compression without compromising downstream task performance, along with reduced inference latency and computational overhead. We conduct an exhaustive parameter search to establish benchmarks that highlight the tradeoff between model compression and performance. Our results demonstrate significant compression of LLMs while maintaining comparable performance to larger models, facilitating their deployment on resource-constraint platforms.
Afia Anjum, Maksim Ekin Eren, Ismael Boureima, Boian S. Alexandrov, Manish Bhattarai
ICMLA4
2024 Domain-Specific Retrieval-Augmented Generation Using Vector Stores, Knowledge Graphs, and Tensor Factorization
abstract
Large Language Models (LLMs) are pre-trained on large-scale corpora and excel in numerous general natural language processing (NLP) tasks, such as question answering (QA). Despite their advanced language capabilities, when it comes to domain-specific and knowledge-intensive tasks, LLMs suffer from hallucinations, knowledge cut-offs, and lack of knowledge attributions. Additionally, fine tuning LLMs' intrinsic knowledge to highly specific domains is an expensive and time consuming process. The retrieval-augmented generation (RAG) process has recently emerged as a method capable of optimization of LLM responses, by referencing them to a predetermined ontology. It was shown that using a Knowledge Graph (KG) ontology for RAG improves the QA accuracy, by taking into account relevant sub-graphs that preserve the information in a structured manner. In this paper, we introduce SMART-SLIC, a highly domain-specific LLM framework, that integrates RAG with KG and a vector store (VS) that store factual domain specific information. Importantly, to avoid hallucinations in the KG, we build these highly domain-specific KGs and VSs without the use of LLMs, but via NLP, data mining, and nonnegative tensor factorization with automatic model selection. Pairing our RAG with a domain-specific: (i) KG (containing structured information), and (ii) VS (containing unstructured information) enables the development of domain-specific chat-bots that attribute the source of information, mitigate hallucinations, lessen the need for fine-tuning, and excel in highly domain-specific question answering tasks. We pair SMART-SLIC with chain-of-thought prompting agents. The framework is designed to be generalizable to adapt to any specific or specialized domain. In this paper, we demonstrate the question answering capabilities of our framework on a corpus of scientific publications on malware analysis and anomaly detection.
Ryan Barron, Ves Grantcharov, Selma Wanna, Maksim Ekin Eren, Manish Bhattarai, Nick Solovyev 0001, George Tompkins, Charles K. Nicholas, Kim Ø. Rasmussen, Cynthia Matuszek, Boian S. Alexandrov
ICMLA11
2024 Distributed out-of-memory NMF on CPU/GPU architectures
abstract
Abstract We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density $$10^{-6}$$ 10 - 6 .
Ismael Boureima, Manish Bhattarai, Maksim Ekin Eren, Erik Skau, Philip Romero, Stephan J. Eidenbenz, Boian S. Alexandrov
J. Supercomput.7
2024 Correction to: Distributed out-of-memory NMF on CPU/GPU architectures
Ismael Boureima, Manish Bhattarai, Maksim Ekin Eren, Erik Skau, Philip Romero, Stephan J. Eidenbenz, Boian S. Alexandrov
J. Supercomput.7
2023 Robust Adversarial Defense by Tensor Factorization
abstract
As machine learning techniques become increasingly prevalent in data analysis, the threat of adversarial attacks has surged, necessitating robust defense mechanisms. Among these defenses, methods exploiting low-rank approximations for input data preprocessing and neural network (NN) parameter factorization have shown potential. Our work advances this field further by integrating the tensorization of input data with low-rank decomposition and tensorization of NN parameters to enhance adversarial defense. The proposed approach demonstrates significant defense capabilities, maintaining robust accuracy even when subjected to the strongest known auto-attacks. Evaluations against leading-edge robust performance benchmarks reveal that our results not only hold their ground against the best defensive methods available but also exceed all current defense strategies that rely on tensor factorizations. This study underscores the potential of integrating tensorization and low-rank decomposition as a robust defense against adversarial attacks in machine learning.
Manish Bhattarai, Mehmet Cagri Kaymak, Ryan Barron, Ben Nebgen, Kim Ø. Rasmussen, Boian S. Alexandrov
ICMLA6
2023 Interactive Distillation of Large Single-Topic Corpora of Scientific Papers
abstract
Highly specific datasets of scientific literature are important for both research and education. However, it is difficult to build such datasets at scale. A common approach is to build these datasets reductively by applying topic modeling on an established corpus and selecting specific topics. A more robust but time-consuming approach is to build the dataset constructively in which a subject matter expert (SME) handpicks documents. This method does not scale and is prone to error as the dataset grows. Here we showcase a new tool, based on machine learning, for constructively generating targeted datasets of scientific literature. Given a small initial “core” corpus of papers, we build a citation network of documents. At each step of the citation network, we generate text embeddings and visualize the embeddings through dimensionality reduction. Papers are kept in the dataset if they are “similar” to the core or are otherwise pruned through human-in-the-loop selection. Additional insight into the papers is gained through SUb-topic modeling using SeNMFk. We demonstrate our new tool for literature review by applying it to two different fields in machine learning.
Nick Solovyev 0001, Ryan Barron, Manish Bhattarai, Maksim Ekin Eren, Kim Ø. Rasmussen, Boian S. Alexandrov
ICMLA6
2023 MalwareDNA: Simultaneous Classification of Malware, Malware Families, and Novel Malware
abstract
Malware is one of the most dangerous and costly cyber threats to national security and a crucial factor in modern cyber-space. However, the adoption of machine learning (ML) based solutions against malware threats has been relatively slow. Shortcomings in the existing ML approaches are likely contributing to this problem. The majority of current ML approaches ignore real-world challenges such as the detection of novel malware. In addition, proposed ML approaches are often designed either for malware/benign-ware classification or malware family classification. Here we introduce and showcase preliminary capabilities of a new method that can perform precise identification of novel malware families, while also unifying the capability for malware/benign-ware classification and malware family classification into a single framework.
Maksim Ekin Eren, Manish Bhattarai, Kim Ø. Rasmussen, Boian S. Alexandrov, Charles K. Nicholas
ISI4
2023 Examining DNA breathing with pyDNA-EPBD
abstract
MOTIVATION: The two strands of the DNA double helix locally and spontaneously separate and recombine in living cells due to the inherent thermal DNA motion. This dynamics results in transient openings in the double helix and is referred to as "DNA breathing" or "DNA bubbles." The propensity to form local transient openings is important in a wide range of biological processes, such as transcription, replication, and transcription factors binding. However, the modeling and computer simulation of these phenomena, have remained a challenge due to the complex interplay of numerous factors, such as, temperature, salt content, DNA sequence, hydrogen bonding, base stacking, and others. RESULTS: We present pyDNA-EPBD, a parallel software implementation of the Extended Peyrard-Bishop-Dauxois (EPBD) nonlinear DNA model that allows us to describe some features of DNA dynamics in detail. The pyDNA-EPBD generates genomic scale profiles of average base-pair openings, base flipping probability, DNA bubble probability, and calculations of the characteristically dynamic length indicating the number of base pairs statistically significantly affected by a single point mutation using the Markov Chain Monte Carlo algorithm. AVAILABILITY AND IMPLEMENTATION: pyDNA-EPBD is supported across most operating systems and is freely available at https://github.com/lanl/pyDNA_EPBD. Extensive documentation can be found at https://lanl.github.io/pyDNA_EPBD/.
Anowarul Kabir, Manish Bhattarai, Kim Ø. Rasmussen, Amarda Shehu, Anny Usheva, Alan R. Bishop, Boian S. Alexandrov
Bioinform.7
2023 Distributed non-negative RESCAL with automatic model selection for exascale data
abstract
With the boom in the development of computer hardware and software, social media, IoT platforms, and communications, there has been exponential growth in the volume of data produced worldwide. Among these data, relational datasets are growing in popularity as they provide unique insights regarding the evolution of communities and their interactions. Relational datasets are naturally non-negative, sparse, and extra-large. Relational data usually contain triples (subject, relation, object) and are represented as graphs/multigraphs, called knowledge graphs, which need to be embedded into a low-dimensional dense vector space. Among various embedding models, RESCAL allows the learning of relational data to extract the posterior distributions over the latent variables and to make predictions of missing relations. However, RESCAL is computationally demanding and requires a fast and distributed implementation to analyze extra-large real-world datasets. Here we introduce a distributed non-negative RESCAL algorithm for heterogeneous CPU/GPU architectures with automatic selection of the number of latent communities (model selection), called pyDRESCALk. We demonstrate the correctness of pyDRESCALk with real-world and large synthetic tensors and the efficacy showing near-linear scaling that concurs with the theoretical complexities. Finally, pyDRESCALk determines the number of latent communities in an 11-terabyte dense and 9-exabyte sparse synthetic tensor.
Manish Bhattarai, Namita Kharat, Ismael Boureima, Erik Skau, Ben Nebgen, Hristo N. Djidjev, Sanjay V. Rajopadhye, James P. Smith, Boian S. Alexandrov
J. Parallel Distributed Comput.9
2023 Semi-Supervised Classification of Malware Families Under Extreme Class Imbalance via Hierarchical Non-Negative Matrix Factorization with Automatic Model Selection
abstract
Identification of the family to which a malware specimen belongs is essential in understanding the behavior of the malware and developing mitigation strategies. Solutions proposed by prior work, however, are often not practicable due to the lack of realistic evaluation factors. These factors include learning under class imbalance, the ability to identify new malware, and the cost of production-quality labeled data. In practice, deployed models face prominent, rare, and new malware families. At the same time, obtaining a large quantity of up-to-date labeled malware for training a model can be expensive. In this article, we address these problems and propose a novel hierarchical semi-supervised algorithm, which we call the HNMFk Classifier , that can be used in the early stages of the malware family labeling process. Our method is based on non-negative matrix factorization with automatic model selection, that is, with an estimation of the number of clusters. With HNMFk Classifier , we exploit the hierarchical structure of the malware data together with a semi-supervised setup, which enables us to classify malware families under conditions of extreme class imbalance. Our solution can perform abstaining predictions, or rejection option, which yields promising results in the identification of novel malware families and helps with maintaining the performance of the model when a low quantity of labeled data is used. We perform bulk classification of nearly 2,900 both rare and prominent malware families, through static analysis, using nearly 388,000 samples from the EMBER-2018 corpus. In our experiments, we surpass both supervised and semi-supervised baseline models with an F1 score of 0.80.
Maksim Ekin Eren, Manish Bhattarai, Robert J. Joyce, Edward Raff, Charles K. Nicholas, Boian S. Alexandrov
ACM Trans. Priv. Secur.6
2022 SeNMFk-SPLIT: large corpora topic modeling by semantic non-negative matrix factorization with automatic model selection
abstract
As the amount of text data continues to grow, topic modeling is serving an important role in understanding the content hidden by the overwhelming quantity of documents. One popular topic modeling approach is non-negative matrix factorization (NMF), an unsupervised machine learning (ML) method. Recently, Semantic NMF with automatic model selection (SeNMFk) has been proposed as a modification to NMF. In addition to heuristically estimating the number of topics, SeNMFk also incorporates the semantic structure of the text. This is performed by jointly factorizing the term frequency-inverse document frequency (TF-IDF) matrix with the co-occurrence/word-context matrix, the values of which represent the number of times two words co-occur in a predetermined window of the text. In this paper, we introduce a novel distributed method, SeNMFk-SPLIT, for semantic topic extraction suitable for large corpora. Contrary to SeNMFk, our method enables the joint factorization of large documents by decomposing the word-context and term-document matrices separately. We demonstrate the capability of SeNMFk-SPLIT by applying it to the entire artificial intelligence (AI) and ML scientific literature uploaded on arXiv.
Maksim Ekin Eren, Nick Solovyev 0001, Manish Bhattarai, Kim Ø. Rasmussen, Charles K. Nicholas, Boian S. Alexandrov
DocEng6
2022 One-Shot Federated Group Collaborative Filtering
abstract
Non-negative matrix factorization (NMF) with missing-value completion is a well-known effective Collaborative Filtering (CF) method used to provide personalized user recommendations. However, traditional CF relies on a privacy-invasive collection of user data to build a central recommender model. One-shot federated learning has recently emerged as a method to mitigate the privacy problem while addressing the traditional communication bottleneck of federated learning. In this paper, we present the first one-shot federated CF implementation, named One-FedCF, for groups of users or collaborating organizations. In our solution, the clients first apply local CF in-parallel to build distinct, client-specific recommenders. Then, the privacy-preserving local item patterns and biases from each client are shared with the processor to perform joint factorization in order to extract the global item patterns. Extracted patterns are then aggregated to each client to build the local models via information retrieval transfer. In our experiments, we demonstrate our approach with two MovieLens datasets and show results competitive with the state-of-the-art federated recommender systems at a substantial decrease in the number of communications.
Maksim Ekin Eren, Manish Bhattarai, Nick Solovyev 0001, Luke E. Richards, Roberto Yus, Charles K. Nicholas, Boian S. Alexandrov
ICMLA7
2022 Improved Protein Decoy Selection via Non-Negative Matrix Factorization
abstract
A central challenge in protein modeling research and protein structure prediction in particular is known as decoy selection. The problem refers to selecting biologically-active/native tertiary structures among a multitude of physically-realistic structures generated by template-free protein structure prediction methods. Research on decoy selection is active. Clustering-based methods are popular, but they fail to identify good/near-native decoys on datasets where near-native decoys are severely under-sampled by a protein structure prediction method. Reasonable progress is reported by methods that additionally take into account the internal energy of a structure and employ it to identify basins in the energy landscape organizing the multitude of decoys. These methods, however, incur significant time costs for extracting basins from the landscape. In this paper, we propose a novel decoy selection method based on non-negative matrix factorization. We demonstrate that our method outperforms energy landscape-based methods. In particular, the proposed method addresses both the time cost issue and the challenge of identifying good decoys in a sparse dataset, successfully recognizing near-native decoys for both easy and hard protein targets.
Nasrin Akhter 0001, Kazi Lutful Kabir, Gopinath Chennupati, Raviteja Vangara, Boian S. Alexandrov, Hristo N. Djidjev, Amarda Shehu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 Factorization of Binary Matrices: Rank Relations, Uniqueness and Model Selection of Boolean Decomposition
abstract
The application of binary matrices are numerous. Representing a matrix as a mixture of a small collection of latent vectors via low-rank decomposition is often seen as an advantageous method to interpret and analyze data. In this work, we examine the factorizations of binary matrices using standard arithmetic (real and nonnegative) and logical operations (Boolean and ℤ 2 ). We examine the relationships between the different ranks, and discuss when factorization is unique. In particular, we characterize when a Boolean factorization X = W ∧ H has a unique W , a unique H (for a fixed W ), and when both W and H are unique, given a rank constraint. We introduce a method for robust Boolean model selection, called BMF k , and show on numerical examples that BMF k not only accurately determines the correct number of Boolean latent features but reconstruct the pre-determined factors accurately.
Derek DeSantis, Erik Skau, Duc Phan Minh Truong, Boian S. Alexandrov
ACM Trans. Knowl. Discov. Data4
2021 COVID-19 multidimensional kaggle literature organization
abstract
The unprecedented outbreak of Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2), or COVID-19, continues to be a significant worldwide problem. As a result, a surge of new COVID-19 related research has followed suit. The growing number of publications requires document organization methods to identify relevant information. In this paper, we expand upon our previous work with clustering the CORD-19 dataset by applying multi-dimensional analysis methods. Tensor factorization is a powerful unsupervised learning method capable of discovering hidden patterns in a document corpus. We show that a higher-order representation of the corpus allows for the simultaneous grouping of similar articles, relevant journals, authors with similar research interests, and topic keywords. These groupings are identified within and among the latent components extracted via tensor decomposition. We further demonstrate the application of this method with a publicly available interactive visualization of the dataset.
Maksim Ekin Eren, Nick Solovyev 0001, Chris Hamer, Renee McDonald, Boian S. Alexandrov, Charles K. Nicholas
DocEng5
2020 Decoy Selection in Protein Structure Determination via Symmetric Non-negative Matrix Factorization
abstract
The so-called dark proteome, referring to regions of the protein universe that remain inaccessible by either wet-or dry-laboratory methods, continues to spur computational research in protein structure determination. An outstanding challenge relates to the ability to discriminate relevant tertiary structure(s) among many structures, also referred to as decoys, that are computed for a protein of interest. The problem is known as decoy selection. While prime for investigation as an inference problem, the decoy datasets generated in silico are sparse and highly imbalanced towards the negative class (irrelevant structures). These characteristics continue to challenge both supervised and unsupervised learning approaches to this problem. In this paper, we propose a novel decoy selection method based on symmetric non-negative matrix factorization in a graph clustering setting. The method is evaluated on two datasets, a benchmark dataset of ensembles of decoys for a varied list of protein molecules, and a dataset of decoy ensembles for targets drawn from the recent CASP competitions. The evaluation demonstrates that the proposed method outperforms several state-of-the-art decoy selection methods. This performance, as well as the method's computational expediency, suggest that the proposed method advances the state of the art in decoy selection and, in particular, our the ability to tackle inherent challenges related to imbalanced datasets.
Kazi Lutful Kabir, Gopinath Chennupati, Raviteja Vangara, Hristo N. Djidjev, Boian S. Alexandrov, Amarda Shehu
BIBM5
2020 Semantic Nonnegative Matrix Factorization with Automatic Model Determination for Topic Modeling
abstract
Non-negative Matrix Factorization (NMF) models the topics of a text corpus by decomposing the matrix of term frequency-inverse document frequency (TF-IDF) representation, X, into two low-rank non-negative matrices: W , representing the topics and H, mapping the documents onto space of topics. One challenge, common to all topic models, is the determination of the number of latent topics (aka model determination). Determining the correct number of topics is important: underestimating the number of topics results in a poor topic separation, under-fitting, while overestimating leads to noisy topics, over-fitting. Here, we introduce SeNMFk, a semantic-assisted NMF-based topic modeling method, which incorporates semantic correlations in NMF by using a word-context matrix, and employs a method for determination of the number of latent topics. SeNMFk first creates a random ensemble of matrices based on the initial TF-IDF matrix and a word-context matrix, and then applies a coupled factorization to acquire sets of stable coherent topics that are robust to noise. The latent dimension is determined based on the stability of these topics. We show that SeNMFk accurately determines the number of high-quality topics in benchmark text corpora, which leads to an accurate document clustering.
Raviteja Vangara, Erik Skau, Gopinath Chennupati, Hristo N. Djidjev, Thomas Tierney, James P. Smith, Manish Bhattarai, Valentin G. Stanev, Boian S. Alexandrov
ICMLA9
2020 Multi-Dimensional Anomalous Entity Detection via Poisson Tensor Factorization
abstract
As the attack surfaces of large enterprise networks grow, anomaly detection systems based on statistical user behavior analysis play a crucial role in identifying malicious activities. Previous work has shown that link prediction algorithms based on non-negative matrix factorization learn highly accurate predictive models of user actions. However, most statistical link prediction models have been constructed on bipartite graphs, and fail to capture the nuanced, multi-faceted details of a user's activity profile. This paper establishes a new benchmark for red team event detection on the Los Alamos National Laboratory Unified Host and Network Dataset by applying a tensor factorization model that exploits the multi-dimensional and sparse structure of user authentication logs. We show that learning patterns of normal activity across multiple dimensions in one unified statistical framework yields improved detection of penetration testing events. We further show operational value by developing fusion methods that can identify anomalous users, source devices, and destination devices in the network.
Maksim Ekin Eren, Juston Moore, Boian S. Alexandrov
ISI3
2020 Distributed non-negative matrix factorization with determination of the number of latent features
Gopinath Chennupati, Raviteja Vangara, Erik Skau, Hristo N. Djidjev, Boian S. Alexandrov
J. Supercomput.5
2019 Non-Negative Matrix Factorization for Selection of Near-Native Protein Tertiary Structures
abstract
Identifying biologically-active protein structure(s) from an ensemble of computed three-dimensional structures is a major challenge. Clustering-based methods are time-consuming and often under perform on structure datasets that are highly imbalanced. Energy landscape-based methods improve performance over imbalanced datasets but incur significant time costs. In this paper we propose a novel method based on non-negative matrix factorization. The method outperforms energy landscape-based clustering methods, addressing both time costs and challenges with imbalanced structure datasets.
Nasrin Akhter 0001, Raviteja Vangara, Gopinath Chennupati, Boian S. Alexandrov, Hristo N. Djidjev, Amarda Shehu
BIBM4
2016 The role of structural parameters in DNA cyclization
abstract
BACKGROUND: The intrinsic bendability of DNA plays an important role with relevance for myriad of essential cellular mechanisms. The flexibility of a DNA fragment can be experimentally and computationally examined by its propensity for cyclization, quantified by the Jacobson-Stockmayer J factor. In this study, we use a well-established coarse-grained three-dimensional model of DNA and seven distinct sets of experimentally and computationally derived conformational parameters of the double helix to evaluate the role of structural parameters in calculating DNA cyclization. RESULTS: We calculate the cyclization rates of 86 DNA sequences with previously measured J factors and lengths between 57 and 325 bp as well as of 20,000 randomly generated DNA sequences with lengths between 350 and 4000 bp. Our comparison with experimental data is complemented with analysis of simulated data. CONCLUSIONS: Our data demonstrate that all sets of parameters yield very similar results for longer DNA fragments, regardless of the nucleotide sequence, which are in agreement with experimental measurements. However, for DNA fragments shorter than 100 bp, all sets of parameters performed poorly yielding results with several orders of magnitude difference from the experimental measurements. Our data show that DNA cyclization rates calculated using conformational parameters based on nucleosome packaging data are most similar to the experimental measurements. Overall, our study provides a comprehensive large-scale assessment of the role of structural parameters in calculating DNA cyclization rates.
Ludmil B. Alexandrov, Alan R. Bishop, Kim Ø. Rasmussen, Boian S. Alexandrov
BMC Bioinform.4
2013 Binding of Nucleoid-Associated Protein Fis to DNA Is Regulated by DNA Breathing Dynamics
abstract
Physicochemical properties of DNA, such as shape, affect protein-DNA recognition. However, the properties of DNA that are most relevant for predicting the binding sites of particular transcription factors (TFs) or classes of TFs have yet to be fully understood. Here, using a model that accurately captures the melting behavior and breathing dynamics (spontaneous local openings of the double helix) of double-stranded DNA, we simulated the dynamics of known binding sites of the TF and nucleoid-associated protein Fis in Escherichia coli. Our study involves simulations of breathing dynamics, analysis of large published in vitro and genomic datasets, and targeted experimental tests of our predictions. Our simulation results and available in vitro binding data indicate a strong correlation between DNA breathing dynamics and Fis binding. Indeed, we can define an average DNA breathing profile that is characteristic of Fis binding sites. This profile is significantly enriched among the identified in vivo E. coli Fis binding sites. To test our understanding of how Fis binding is influenced by DNA breathing dynamics, we designed base-pair substitutions, mismatch, and methylation modifications of DNA regions that are known to interact (or not interact) with Fis. The goal in each case was to make the local DNA breathing dynamics either closer to or farther from the breathing profile characteristic of a strong Fis binding site. For the modified DNA segments, we found that Fis-DNA binding, as assessed by gel-shift assay, changed in accordance with our expectations. We conclude that Fis binding is associated with DNA breathing dynamics, which in turn may be regulated by various nucleotide modifications.
Kristy Nowak-Lovato, Ludmil B. Alexandrov, Afsheen Banisadr, Amy L. Bauer, Alan R. Bishop, Anny Usheva, Fangping Mu, Elizabeth Hong-Geller, Kim Ø. Rasmussen, William S. Hlavacek, Boian S. Alexandrov
PLoS Comput. Biol.11
2009 Toward a Detailed Description of the Thermally Induced Dynamics of the Core Promoter
abstract
Establishing the general and promoter-specific mechanistic features of gene transcription initiation requires improved understanding of the sequence-dependent structural/dynamic features of promoter DNA. Experimental data suggest that a spontaneous dsDNA strand separation at the transcriptional start site is likely to be a requirement for transcription initiation in several promoters. Here, we use Langevin molecular dynamic simulations based on the Peyrard-Bishop-Dauxois nonlinear model of DNA (PBD LMD) to analyze the strand separation (bubble) dynamics of 80-bp-long promoter DNA sequences. We derive three dynamic criteria, bubble probability, bubble lifetime, and average strand separation, to characterize bubble formation at the transcriptional start sites of eight mammalian gene promoters. We observe that the most stable dsDNA openings do not necessarily coincide with the most probable openings and the highest average strand displacement, underscoring the advantages of proper molecular dynamic simulations. The dynamic profiles of the tested mammalian promoters differ significantly in overall profile and bubble probability, but the transcriptional start site is often distinguished by large (longer than 10 bp) and long-lived transient openings in the double helix. In support of these results are our experimental transcription data demonstrating that an artificial bubble-containing DNA template is transcribed bidirectionally by human RNA polymerase alone in the absence of any other transcription factors.
Boian S. Alexandrov, Vladimir Gelev, Sang Wook Yoo, Alan R. Bishop, Kim Ø. Rasmussen, Anny Usheva
PLoS Comput. Biol.1