Bahrad A. Sokhansanj

dblp:23/3382 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0002-5050-5926ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100% Computational science and engineering · 0%
Artificial intelligence
1 paper
Representation and self-supervised learning · 50% Generative modeling · 50%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis
DNA sequence analysis
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › sequence analysis › sequence modeling
genomic sequence modeling
0.912025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology › gene regulation › regulatory element discovery
regulatory element prediction
0.312025
Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis · NeurIPS 2025
Bioinformatics and computational biology
genome engineering
0.012000
Genomic engineering: moving beyond DNA sequence to function · Proc. IEEE 2000
Bioinformatics and computational biology
genomics
0.012000
Genomic engineering: moving beyond DNA sequence to function · Proc. IEEE 2000
Bioinformatics and computational biology › genomics
DNA sequencing
0.012000
Genomic engineering: moving beyond DNA sequence to function · Proc. IEEE 2000

Methods — techniques the papers use, named apart from their topics

transition-matrix loss · 1.7transformer · 1.7n-gram statistics · 1.7
YearPublicationVenuePosition
2025 Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis
abstract
Transformers have revolutionized nucleotide sequence analysis, yet capturing long‑range dependencies remains challenging. Recent studies show that autoregressive transformers often exhibit Markovian behavior by relying on fixed-length context windows for next-token prediction. However, standard self-attention mechanisms are computationally inefficient for long sequences due to their quadratic complexity and do not explicitly enforce global transition consistency. We introduce CARMANIA (Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis), a self-supervised pretraining framework that augments next-token (NT) prediction with a transition-matrix (TM) loss. The TM loss aligns predicted token transitions with empirically derived n-gram statistics from each input sequence, encouraging the model to capture higher-order dependencies beyond local context. This integration enables CARMANIA to learn organism-specific sequence structures that reflect both evolutionary constraints and functional organization. We evaluate CARMANIA across diverse genomic tasks, including regulatory element prediction, functional gene classification, taxonomic inference, antimicrobial resistance detection, and biosynthetic gene cluster classification. CARMANIA outperforms the previous best long-context model by at least 7\%, matches state-of-the-art on shorter sequences (exceeding prior results on 20/40 tasks while running $\sim$2.5$\times$ faster), and shows particularly strong improvements on enhancer and housekeeping gene classification tasks—including up to a 34\% absolute gain in Matthews correlation coefficient (MCC) for enhancer prediction. The TM loss boosts accuracy in 33 of 40 tasks, especially where local motifs or regulatory patterns drive prediction. This enables more effective modeling of sequence-dependent biological features while maintaining robustness across non-coding and low-signal regions. Code available at https://github.com/EESI/carmania.
Mohammadsaleh Refahi, Mahdi Abavisani, Bahrad A. Sokhansanj, James R. Brown, Gail L. Rosen
NeurIPS3
2021 Learning, visualizing and exploring 16S rRNA structure using an attention-based deep neural network
abstract
Recurrent neural networks with memory and attention mechanisms are widely used in natural language processing because they can capture short and long term sequential information for diverse tasks. We propose an integrated deep learning model for microbial DNA sequence data, which exploits convolutional neural networks, recurrent neural networks, and attention mechanisms to predict taxonomic classifications and sample-associated attributes, such as the relationship between the microbiome and host phenotype, on the read/sequence level. In this paper, we develop this novel deep learning approach and evaluate its application to amplicon sequences. We apply our approach to short DNA reads and full sequences of 16S ribosomal RNA (rRNA) marker genes, which identify the heterogeneity of a microbial community sample. We demonstrate that our implementation of a novel attention-based deep network architecture, Read2Pheno, achieves read-level phenotypic prediction. Training Read2Pheno models will encode sequences (reads) into dense, meaningful representations: learned embedded vectors output from the intermediate layer of the network model, which can provide biological insight when visualized. The attention layer of Read2Pheno models can also automatically identify nucleotide regions in reads/sequences which are particularly informative for classification. As such, this novel approach can avoid pre/post-processing and manual interpretation required with conventional approaches to microbiome sequence classification. We further show, as proof-of-concept, that aggregating read-level information can robustly predict microbial community properties, host phenotype, and taxonomic classification, with performance at least comparable to conventional approaches. An implementation of the attention-based deep learning network is available at https://github.com/EESI/sequence_attention (a python package) and https://github.com/EESI/seq2att (a command line tool).
Zhengqiao Zhao, Stephen Woloszynek, Felix Agbavor, Joshua Chang Mell, Bahrad A. Sokhansanj, Gail L. Rosen
PLoS Comput. Biol.5
2020 Genetic grouping of SARS-CoV-2 coronavirus sequences using informative subtype markers for pandemic spread visualization
abstract
We propose an efficient framework for genetic subtyping of SARS-CoV-2, the novel coronavirus that causes the COVID-19 pandemic. Efficient viral subtyping enables visualization and modeling of the geographic distribution and temporal dynamics of disease spread. Subtyping thereby advances the development of effective containment strategies and, potentially, therapeutic and vaccine strategies. However, identifying viral subtypes in real-time is challenging: SARS-CoV-2 is a novel virus, and the pandemic is rapidly expanding. Viral subtypes may be difficult to detect due to rapid evolution; founder effects are more significant than selection pressure; and the clustering threshold for subtyping is not standardized. We propose to identify mutational signatures of available SARS-CoV-2 sequences using a population-based approach: an entropy measure followed by frequency analysis. These signatures, Informative Subtype Markers (ISMs), define a compact set of nucleotide sites that characterize the most variable (and thus most informative) positions in the viral genomes sequenced from different individuals. Through ISM compression, we find that certain distant nucleotide variants covary, including non-coding and ORF1ab sites covarying with the D614G spike protein mutation which has become increasingly prevalent as the pandemic has spread. ISMs are also useful for downstream analyses, such as spatiotemporal visualization of viral dynamics. By analyzing sequence data available in the GISAID database, we validate the utility of ISM-based subtyping by comparing spatiotemporal analyses using ISMs to epidemiological studies of viral transmission in Asia, Europe, and the United States. In addition, we show the relationship of ISMs to phylogenetic reconstructions of SARS-CoV-2 evolution, and therefore, ISMs can play an important complementary role to phylogenetic tree-based analysis, such as is done in the Nextstrain project. The developed pipeline dynamically generates ISMs for newly added SARS-CoV-2 sequences and updates the visualization of pandemic spatiotemporal dynamics, and is available on Github at https://github.com/EESI/ISM (Jupyter notebook), https://github.com/EESI/ncov_ism (command line tool) and via an interactive website at https://covid19-ism.coe.drexel.edu/.
Zhengqiao Zhao, Bahrad A. Sokhansanj, Charvi Malhotra, Kitty Zheng, Gail L. Rosen
PLoS Comput. Biol.2
2009 Mining, Modeling, and Evaluation of Subnetworks From Large Biomolecular Networks and Its Comparison Study
abstract
In this paper, we present a novel method to mine, model, and evaluate a regulatory system executing cellular functions that can be represented as a biomolecular network. Our method consists of two steps. First, a novel scale-free network clustering approach is applied to such a biomolecular network to obtain various subnetworks. Second, computational models are generated for the subnetworks and simulated to predict their behavior in the cellular context. We discuss and evaluate some of the advanced computational modeling approaches, in particular, state-space modeling, probabilistic Boolean network modeling, and fuzzy logic modeling. The modeling and simulation results represent hypotheses that are tested against high-throughput biological datasets (microarrays and/or genetic screens) under normal and perturbation conditions. Experimental results on time-series gene expression data for the human cell cycle indicate that our approach is promising for subnetwork mining and simulation from large biomolecular networks.
Xiaohua Hu 0001, Michael Kwok-Po Ng, Fang-Xiang Wu, Bahrad A. Sokhansanj
IEEE Trans. Inf. Technol. Biomed.4
2007 Accelerated search for biomolecular network models to interpret high-throughput experimental data
abstract
BACKGROUND: The functions of human cells are carried out by biomolecular networks, which include proteins, genes, and regulatory sites within DNA that encode and control protein expression. Models of biomolecular network structure and dynamics can be inferred from high-throughput measurements of gene and protein expression. We build on our previously developed fuzzy logic method for bridging quantitative and qualitative biological data to address the challenges of noisy, low resolution high-throughput measurements, i.e., from gene expression microarrays. We employ an evolutionary search algorithm to accelerate the search for hypothetical fuzzy biomolecular network models consistent with a biological data set. We also develop a method to estimate the probability of a potential network model fitting a set of data by chance. The resulting metric provides an estimate of both model quality and dataset quality, identifying data that are too noisy to identify meaningful correlations between the measured variables. RESULTS: Optimal parameters for the evolutionary search were identified based on artificial data, and the algorithm showed scalable and consistent performance for as many as 150 variables. The method was tested on previously published human cell cycle gene expression microarray data sets. The evolutionary search method was found to converge to the results of exhaustive search. The randomized evolutionary search was able to converge on a set of similar best-fitting network models on different training data sets after 30 generations running 30 models per generation. Consistent results were found regardless of which of the published data sets were used to train or verify the quantitative predictions of the best-fitting models for cell cycle gene dynamics. CONCLUSION: Our results demonstrate the capability of scalable evolutionary search for fuzzy network models to address the problem of inferring models based on complex, noisy biomolecular data sets. This approach yields multiple alternative models that are consistent with the data, yielding a constrained set of hypotheses that can be used to optimally design subsequent experiments.
Suman Datta, Bahrad A. Sokhansanj
BMC Bioinform.2
2007 A Novel Approach for Mining and Fuzzy Simulation of Subnetworks From Large Biomolecular Networks
abstract
Understanding the biomolecular network implementing cellular function goes beyond the old dogma of "one gene: one function"; only through comprehensive system understanding can we predict the impact of genetic variation in the population, design effective disease therapeutics, and evaluate the potential side-effects of therapies. In this paper, we present a novel method to model the regulatory system that executes a cellular function, which can be represented as a biomolecular network. Our method consists of three steps. First, the biomolecular network is derived using data-mining approaches to extend the initial conceptual biomolecular network from the literature search, etc. Secondly, once the whole biomolecular network structure is complete, a novel scale-free network clustering approach is applied to obtain various subnetworks. Lastly, fuzzy rule based models are generated for the subnetworks and simulations are run to predict their behavior in the cellular context. The modeling results represent hypotheses that are tested against high-throughput data sets (microarrays and/or genetic screens) for both the natural system and perturbations. If computational results do not match experimental or previously published results, then new hypotheses are formed and they feed back into the data-mining and analyzing step to refine the biomolecular network for the next iteration. This is repeated until a good match between modeling and data is obtained. Notably, the dynamic modeling component of this method depends on the automated network structure generation of the first component and the subnetwork clustering, which are both essential to make the solution tractable. Experimental results on human gene interaction networks and gene expression time series data for the human cell cycle indicate that our approach is promising for subnetwork mining and simulation from large biomolecular networks, as it produces a better convergence between continuous modeling and experiments.
Xiaohua Hu 0001, Bahrad A. Sokhansanj, Daniel Duanqing Wu, Yuchun Tang
IEEE Trans. Fuzzy Syst.2
2004 Linear fuzzy gene network models obtained from microarray data by exhaustive search
abstract
BACKGROUND: Recent technological advances in high-throughput data collection allow for experimental study of increasingly complex systems on the scale of the whole cellular genome and proteome. Gene network models are needed to interpret the resulting large and complex data sets. Rationally designed perturbations (e.g., gene knock-outs) can be used to iteratively refine hypothetical models, suggesting an approach for high-throughput biological system analysis. We introduce an approach to gene network modeling based on a scalable linear variant of fuzzy logic: a framework with greater resolution than Boolean logic models, but which, while still semi-quantitative, does not require the precise parameter measurement needed for chemical kinetics-based modeling. RESULTS: We demonstrated our approach with exhaustive search for fuzzy gene interaction models that best fit transcription measurements by microarray of twelve selected genes regulating the yeast cell cycle. Applying an efficient, universally applicable data normalization and fuzzification scheme, the search converged to a small number of models that individually predict experimental data within an error tolerance. Because only gene transcription levels are used to develop the models, they include both direct and indirect regulation of genes. CONCLUSION: Biological relationships in the best-fitting fuzzy gene network models successfully recover direct and indirect interactions predicted from previous knowledge to result in transcriptional correlation. Fuzzy models fit on one yeast cell cycle data set robustly predict another experimental data set for the same system. Linear fuzzy gene networks and exhaustive rule search are the first steps towards a framework for an integrated modeling and experiment approach to high-throughput "reverse engineering" of complex biological systems.
Bahrad A. Sokhansanj, J. Patrick Fitch, Judy N. Quong, Andrew A. Quong
BMC Bioinform.1
2001 Matrix formulation of a universal microbial transcript profiling system
abstract
DNA chips and microarrays are used to profile gene transcription. Unfortunately, the initial fabrication cost for a chip and the reagent costs to amplify thousands of open reading frames for a microarray are over $100K for a typical 4 Mbase bacterial genome. To avoid these expensive steps, a matrix formulation of a universal hybrid chip-microarray approach to transcript profiling is demonstrated for synthetic data. Initial considerations for application to the 4.3 Mbase bacterium Yersinia pestis are also presented. This approach can be applied to arbitrary bacteria by recalculating a matrix and pseudoinverse. This approach avoids the large upfront expenses associated with DNA chips and microarrays.
J. Patrick Fitch, Jefferson Ng, Bahrad A. Sokhansanj
ICASSP3
2000 Genomic engineering: moving beyond DNA sequence to function
abstract
The DNA sequence of the human genome has been determined. This significant scientific milestone has been a multidisciplinary effort sponsored by both public and private investments. Engineering and computer science contributions have been essential to success, especially with the accelerated schedules of the Human Genome Project. A tutorial summary of the biological and health implications of sequencing the human genome is presented together with examples of how genome data from human and other organisms are already used. Engineering contributions to sequencing are identified as well as predictions of how engineering methods may contribute to the "postsequencing" era of biology.
J. Patrick Fitch, Bahrad A. Sokhansanj
Proc. IEEE2