Aniruddha Datta

dblp:07/2455 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Test-Fleet Scheduling in Complex Validation and Production Environments
abstract
We present a solution to the complex design-automation problem of scheduling test operations in a validation laboratory or production facility. Our goal is to maximize the utilization of a fleet of test stations and minimize the overall test time for a set of products. We consider the realistic scenario where tests can have dependency graphs, implying that some tests must be completed and passed before others can proceed. We also consider a mix of product types that require different kinds of tests and a mix of testers, which implies that each product can only be tested only on a specific set of testers. To ensure scalability and flexibility, we have formulated this scheduling problem as a “partially observable stochastic game”, a multi-agent extension of a partially observable Markov decision process. We have implemented multi-agent reinforcement learning agents to maximize parallelization in a manner that speeds up both training and inferencing. We present scheduling results for synthetic test cases as well as real-life data from a production facility.
Aniruddha Datta, Bhanu Vikas Yaganti, Mate Palocska, Andrew Dove, Arik Peltz, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.1
2024 Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge
abstract
Recent advances in Hierarchical Multi-label Classification (HMC), particularly neurosymbolic-based approaches, have demonstrated improved consistency and accuracy by enforcing constraints on a neural model during training. However, such work assumes the existence of such constraints a-priori. In this paper, we relax this strong assumption and present an approach based on Error Detection Rules (EDR) that allow for learning explainable rules about the failure modes of machine learning models. We show that these rules are not only effective in detecting when a machine learning classifier has made an error but also can be leveraged as constraints for HMC, thereby allowing the recovery of explainable constraints even if they are not provided. We show that our approach is effective in detecting machine learning errors and recovering constraints, is noise tolerant, and can function as a source of knowledge for neurosymbolic models on multiple datasets, including a newly introduced military vehicle recognition dataset.
Joshua Shay Kricheli, Khoa Vo 0002, Aniruddha Datta, Spencer Ozgur, Paulo Shakarian
CIKM3
2024 Test-Fleet Optimization Using Machine Learning
abstract
We present a solution to the complex problem of scheduling test operations in a validation lab or production facility. Our goal is to maximize the utilization of a fleet of test stations and minimize the overall test time for a set of products. We consider the realistic scenario where tests can have dependency graphs, implying that some tests must be completed and passed before others can proceed. We also consider a mix of product types that require different kinds of tests and a mix of testers, which implies that each product can only be tested only on a specific set of testers. To ensure scalability and flexibility, we have formulated this scheduling problem as a “partially observable stochastic game”, a multi-agent extension of a partially observable Markov decision process. We have implemented multi-agent reinforcement learning agents to maximize parallelization in a manner that speeds up both training and inferencing. We present scheduling results for synthetic test cases as well as real-life data from a production facility.
Aniruddha Datta, Bhanu Vikas Yaganti, Andrew Dove, Arik Peltz, Krishnendu Chakrabarty
ETS1
2023 Precision Targeting of Non-Small Cell Lung Cancer: Identifying Optimal Drug Targets and FDA-Approved Combinations for Enhanced Therapeutic Efficacy
abstract
This study proposes a Boolean network model to identify optimal drug targets and select the most effective FDA-approved drug combinations for Non-Small Cell Lung Cancer (NSCLC). The Boolean network models the signaling pathways in NSCLC to capture the intricate molecular interactions driving tumor progression. We evaluate the model by employing the size difference (SD) score, which reflects the degree of cell dysregulation due to gene mutations and allows us to identify optimal drug targets in NSCLC cells to address this dysregulation. Specifically, leveraging the FDA-approved drug database, we identified the robust drug or drug combination for 1, 2, and 3 mutations that maximize tumor cell death and minimize cell proliferation for NSCLC-associated gene mutations. Our findings provide a strong foundation for personalized therapeutic strategies and hold promise for advancing precision oncology to effectively combat NSCLC.
Pranabesh Bhattacharjee, Aditya Lahiri, Norman Peter Reeves, Aniruddha Datta
BIBE4
2023 Combination Supplements for Endometrial Cancer
abstract
This paper suggests an optimal supplement combination (EGCG, Curcumin, Melatonin, Aspirin, and Baicalein) for treating endometrial cancer using a Boolean model. Endometrial cancer affects the uterine lining, with a high incidence in 2023. Surgery, chemotherapy, or radiation may not be options for some patients, leading to increased interest in dietary supplements. However, using ineffective supplements can cause health problems. Our study focuses on safer, less toxic supplements, offering the most efficient combination.
Madhurima Mondal, Aditya Lahiri, Haswanth Vundavilli, Giuseppe Del Priore, Norman Peter Reeves, Aniruddha Datta
BIBE6
2022 In Silico Modeling of the Induction of Apoptosis by Cryptotanshinone in Osteosarcoma Cell Lines
abstract
Osteosarcoma (OS) is the most common primary malignant bone tumor of both children and pet canines. Its characteristic genomic instability and complexity coupled with the dearth of knowledge about its etiology has made improvement in the current treatment difficult. We use the existing literature about the biological pathways active in OS and combine it with the current research involving natural compounds to identify new targets and design more effective drug therapies. The key components of these pathways are modeled as a Boolean network with multiple inputs and multiple outputs. The combinatorial circuit is employed to theoretically predict the efficacies of various drugs in combination with Cryptotanshinone. We show that the action of the herbal drug, Cryptotanshinone on OS cell lines induces apoptosis by increasing sensitivity to TNF-related apoptosis-inducing ligand (TRAIL) through its multi-pronged action on STAT3, DRP1 and DR5. The Boolean framework is used to detect additional drug intervention points in the pathway that could amplify the action of Cryptotanshinone.
Radhika Saraf, Aniruddha Datta, Chao Sima, Jianping Hua, Rosana Lopes, Michael L. Bittner, Tasha Miller, Heather M. Wilson-Robles
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Integrative Network Modeling Highlights the Crucial Roles of Rho-GDI Signaling Pathway in the Progression of non-Small Cell Lung Cancer
abstract
Non-small cell lung cancer (NSCLC) is the most prevalent form of lung cancer and a leading cause of cancer-related deaths worldwide. Using an integrative approach, we analyzed a publicly available merged NSCLC transcriptome dataset using machine learning, protein-protein interaction (PPI) networks and bayesian modeling to pinpoint key cellular factors and pathways likely to be involved with the onset and progression of NSCLC. First, we generated multiple prediction models using various machine learning classifiers to classify NSCLC and healthy cohorts. Our models achieved prediction accuracies ranging from 0.83 to 1.0, with XGBoost emerging as the best performer. Next, using functional enrichment analysis (and gene co-expression network analysis with WGCNA) of the machine learning feature-selected genes, we determined that genes involved in Rho GTPase signaling that modulate actin stability and cytoskeleton were likely to be crucial in NSCLC. We further assembled a PPI network for the feature-selected genes that was partitioned using Markov clustering to detect protein complexes functionally relevant to NSCLC. Finally, we modeled the perturbations in RhoGDI signaling using a bayesian network; our simulations suggest that aberrations in ARHGEF19 and/or RAC2 gene activities contributed to impaired MAPK signaling and disrupted actin and cytoskeleton organization and were arguably key contributors to the onset of tumorigenesis in NSCLC. We hypothesize that targeted measures to restore aberrant ARHGEF19 and/or RAC2 functions could conceivably rescue the cancerous phenotype in NSCLC. Our findings offer promising avenues for early predictive biomarker discovery, targeted therapeutic intervention and improved clinical outcomes in NSCLC.
Saransh Gupta, Haswanth Vundavilli, Rodolfo S. Allendes Osorio, Mari N. Itoh, Attayeb Mohsen, Aniruddha Datta, Kenji Mizuguchi, Lokesh P. Tripathi
IEEE J. Biomed. Health Informatics6
2020 Robust Recurrent CNV Detection in the Presence of Inter-Subject Variability
abstract
The study of recurrent copy number variations (CNVs) plays an important role in understanding the onset and evolution of complex diseases such as cancer. Array-based comparative genomic hybridization (aCGH) is a widely used microarray based technology for identifying CNVs. However, due to high noise levels and inter-sample variability, detecting recurrent CNVs from aCGH data remains a challenging topic. This paper proposes a novel method for identification of the recurrent CNVs. In the proposed method, the noisy aCGH data is modeled as the superposition of three matrices: a full-rank matrix of weighted piece-wise generating signals accounting for the clean aCGH data, a Gaussian noise matrix to model the inherent experimentation errors and other sources of error, and a sparse matrix to capture the sparse inter-sample (sample-specific) variations. We demonstrated the ability of our method to separate accurately recurrent CNVs from sample-specific variations and noise in both simulated (artificial) data and real data. The proposed method produced more accurate results than current state-of-the-art methods used in recurrent CNV detection and exhibited robustness to noise and sample-specific variations.
Mustafa Alshawaqfeh 0001, Ahmad Al Kawam, Erchin Serpedin, Aniruddha Datta
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 A Gaussian Mixture-Model Exploiting Pathway Knowledge for Dissecting Cancer Heterogeneity
abstract
In this work, we develop a systematic approach for applying pathway knowledge to a multivariate Gaussian mixture model for dissecting a heterogeneous cancer tissue. The downstream transcription factors are selected as observables from available partial pathway knowledge in such a way that the subpopulations produce some differential behavior in response to the drugs selected in the upstream. For each subpopulation, each unique (drug, observable) pair is considered as a unique dimension of a multivariate Gaussian distribution. Expectation-maximization (EM) algorithm with hill-climbing is then used to rank the most probable estimates of the mixture composition based on the log-likelihood value. A major contribution of this work is to examine the efficacy of the EM based approach in estimating the composition of experimental mixture sets from cell-by-cell measurements collected on a dynamic cell imaging platform. Towards this end, we apply the algorithm on hourly data collected for two different mixture compositions of A2058, HCT116, and SW480 cell lines for three scenarios: untreated, Lapatinib-treated, and Temsirolimus-treated. Additionally, we show how this methodology can provide a basis for comparing the killing rate of different drugs for a heterogeneous cancer tissue. This obviously has important implications for designing efficient drugs for treating heterogeneous malignant tumors.
Rajan Kapoor, Aniruddha Datta, Chao Sima, Jianping Hua, Rosana Lopes, Michael L. Bittner
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 In Silico Design and Experimental Validation of Combination Therapy for Pancreatic Cancer
abstract
The number of deaths associated with Pancreatic Cancer has been on the rise in the United States making it an especially dreaded disease. The overall prognosis for pancreatic cancer patients continues to be grim because of the complexity of the disease at the molecular level involving the potential activation/inactivation of several diverse signaling pathways. In this paper, we first model the aberrant signaling in pancreatic cancer using a multi-fault Boolean Network. Thereafter, we theoretically evaluate the efficacy of different drug combinations by simulating this boolean network with drugs at the relevant intervention points and arrive at the most effective drug(s) to achieve cell death. The simulation results indicate that drug combinations containing Cryptotanshinone, a traditional Chinese herb derivative, result in considerably enhanced cell death. These in silico results are validated using wet lab experiments we carried out on Human Pancreatic Cancer (HPAC) cell lines.
Haswanth Vundavilli, Aniruddha Datta, Chao Sima, Jianping Hua, Rosana Lopes, Michael L. Bittner
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 Cryptotanshinone Induces Cell Death in Lung Cancer by Targeting Aberrant Feedback Loops
abstract
Signaling pathways oversee highly efficient cellular mechanisms such as growth, division, and death. These processes are controlled by robust negative feedback loops that inhibit receptor-mediated growth factor pathways. Specifically, the ERK, the AKT, and the S6K feedback loops attenuate signaling via growth factor receptors and other kinase receptors to regulate cell growth. Irregularity in any of these supervised processes can lead to uncontrolled cell proliferation and possibly Cancer. These irregularities primarily occur as mutated genes, and an exhaustive search of the perfect drug combination by performing experiments can be both costly and complex. Hence, in this paper, we model the Lung Cancer pathway as a Modified Boolean Network that incorporates feedback. By simulating this network, we theoretically predict the drug combinations that achieve the desired goal for the majority of mutations. Our theoretical analysis identifies Cryptotanshinone, a traditional Chinese herb derivative, as a potent drug component in the fight against cancer. We validated these theoretical results using multiple wet lab experiments carried out on H2073 and SW900 lung cancer cell lines.
Haswanth Vundavilli, Aniruddha Datta, Chao Sima, Jainping Hua, Rosana Lopes, Michael L. Bittner
IEEE J. Biomed. Health Informatics2
2019 Emergence of DSS efforts in genomics: Past contributions and challenges
Arun Sen, Ahmad Al Kawam, Aniruddha Datta
Decis. Support Syst.3
2018 A Bayesian approach to determine the composition of heterogeneous cancer tissue
abstract
BACKGROUND: Cancer Tissue Heterogeneity is an important consideration in cancer research as it can give insights into the causes and progression of cancer. It is known to play a significant role in cancer cell survival, growth and metastasis. Determining the compositional breakup of a heterogeneous cancer tissue can also help address the therapeutic challenges posed by heterogeneity. This necessitates a low cost, scalable algorithm to address the challenge of accurate estimation of the composition of a heterogeneous cancer tissue. METHODS: In this paper, we propose an algorithm to tackle this problem by utilizing the data of accurate, but high cost, single cell line cell-by-cell observation methods in low cost aggregate observation method for heterogeneous cancer cell mixtures to obtain their composition in a Bayesian framework. RESULTS: The algorithm is analyzed and validated using synthetic data and experimental data. The experimental data is obtained from mixtures of three separate human cancer cell lines, HCT116 (Colorectal carcinoma), A2058 (Melanoma) and SW480 (Colorectal carcinoma). CONCLUSION: The algorithm provides a low cost framework to determine the composition of heterogeneous cancer tissue which is a crucial aspect in cancer research.
Ashish Katiyar, Anwoy Kumar Mohanty, Jianping Hua, Chao Sima, Rosana Lopes, Aniruddha Datta, Michael L. Bittner
BMC Bioinform.6
2018 Simulating variance heterogeneity in quantitative genome wide association studies
abstract
BACKGROUND: Analyzing Variance heterogeneity in genome wide association studies (vGWAS) is an emerging approach for detecting genetic loci involved in gene-gene and gene-environment interactions. vGWAS analysis detects variability in phenotype values across genotypes, as opposed to typical GWAS analysis, which detects variations in the mean phenotype value. RESULTS: A handful of vGWAS analysis methods have been recently introduced in the literature. However, very little work has been done for evaluating these methods. To enable the development of better vGWAS analysis methods, this work presents the first quantitative vGWAS simulation procedure. To that end, we describe the mathematical framework and algorithm for generating quantitative vGWAS phenotype data from genotype profiles. Our simulation model accounts for both haploid and diploid genotypes under different modes of dominance. Our model is also able to simulate any number of genetic loci causing mean and variance heterogeneity. CONCLUSIONS: We demonstrate the utility of our simulation procedure through generating a variety of genetic loci types to evaluate common GWAS and vGWAS analysis methods. The results of this evaluation highlight the challenges current tools face in detecting GWAS and vGWAS loci.
Ahmad Al Kawam, Mustafa Alshawaqfeh 0001, James J. Cai, Erchin Serpedin, Aniruddha Datta
BMC Bioinform.5
2018 Examining De Novo Transcriptome Assemblies via a Quality Assessment Pipeline
abstract
New de novo transcriptome assembly and annotation methods provide an incredible opportunity to study the transcriptome of organisms that lack an assembled and annotated genome. There are currently a number of de novo transcriptome assembly methods, but it has been difficult to evaluate the quality of these assemblies. In order to assess the quality of the transcriptome assemblies, we composed a workflow of multiple quality check measurements that in combination provide a clear evaluation of the assembly performance. We presented novel transcriptome assemblies and functional annotations for Pacific Whiteleg Shrimp (Litopenaeus vannamei ), a mariculture species with great national and international interest, and no solid transcriptome/genome reference. We examined Pacific Whiteleg transcriptome assemblies via multiple metrics, and provide an improved gene annotation. Our investigations show that assessing the quality of an assembly purely based on the assembler's statistical measurements can be misleading; we propose a hybrid approach that consists of statistical quality checks and further biological-based evaluations.
Noushin Ghaffari, Osama A. Arshad, Hyundoo Jeong, John Thiltges, Michael F. Criscitiello, Byung-Jun Yoon, Aniruddha Datta, Charles D. Johnson
IEEE ACM Trans. Comput. Biol. Bioinform.7
2018 Deep Sequencing Data Analysis
abstract
This paper discussed the recent advances in Deep Sequencing Data Analysis for systems biology research. Deep sequencing technologies have been primarily applied to genomic sequencing but have recently been applied for transcriptomic profiling or mapping histone modifications. Deep Sequencing technology shows clear advantages over existing profiling technologies in terms of amount of sequence coverage, revealing new transcriptomic insights, measurement of expression of different transcript isoforms and accuracy of defining transcription level. However, being a relatively newer method for transcriptomic profiling, standardized approaches for analysis of deep sequencing expression data are still being developed. The analysis and application of deep sequencing data presents enormous challenges in the areas of machine learning, signal processing, systems theory and statistics. The emphasis of the special issue is on the latest computational challenges and finding rigorous and novel engineering approaches to tackle structural and functional systems biology problems using deep sequencing technologies
Bijoy K. Ghosh, Aniruddha Datta, Ranadip Pal
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Understanding the Bioinformatics Challenges of Integrating Genomics Into Healthcare
abstract
Genomic data is paving the way towards personalized healthcare. By unveiling genetic disease-contributing factors, genomic data can aid in the detection, diagnosis, and treatment of a wide range of complex diseases. Integrating genomic data into healthcare is riddled with a wide range of challenges spanning social, ethical, legal, educational, economic, and technical aspects. Bioinformatics is a core integration aspect presenting an overwhelming number of unaddressed challenges. In this paper we tackle the fundamental bioinformatics integration concerns including: genomic data generation, storage, representation, and utilization in conjunction with clinical data. We divide the bioinformatics challenges into a series of seven intertwined integration aspects spanning the areas of informatics, knowledge management, and communication. For each aspect, we provide a detailed discussion of the current research directions, outstanding challenges, and possible resolutions. This paper seeks to help narrow the gap between the genomic applications, which are being predominantly utilized in research settings, and the clinical adoption of these applications.
Ahmad Al Kawam, Arun Sen, Aniruddha Datta, Nancy Dickey
IEEE J. Biomed. Health Informatics3
2017 Application of big data analytics in process safety and risk management
abstract
In recent years, there has been an increasing interest in the field of big data analytics. It has been established that there exist large amounts of data in the energy industry1. However, there is a need to develop methods combining domain knowledge to transform this data into meaningful information to return business intelligence. The existing literature on big data analytics focuses on applications in various fields such as healthcare, aviation industry, finance, energy industry, and supply chain. However, within the energy industry, the application of big data analytics in process safety and risk management is in the nascent stages. The objective of this study is to discuss the potential of big data analytics in the area of process safety and risk management in the energy industry. The paper outlines the systemic framework with different stakeholders, data sources, challenges, and discusses the benefits of big data analytics in process safety. Four case studies with different applications ranging from incident database analysis, predictive modeling for pump failures, dynamic risk mapping of operating plant, and image analysis to gain insights are demonstrated. It is concluded that the application of big data analytics would provide valuable insights for more informed policy, strategic, and operational risk decision-making leading to a safer and more reliable industry.
Pankaj Goel, Aniruddha Datta, M. Sam Mannan
IEEE BigData2
2017 Towards targeted combinatorial therapy design for the treatment of castration-resistant prostate cancer
abstract
BACKGROUND: Prostate cancer is one of the most prevalent cancers in males in the United States and amongst the leading causes of cancer related deaths. A particularly virulent form of this disease is castration-resistant prostate cancer (CRPC), where patients no longer respond to medical or surgical castration. CRPC is a complex, multifaceted and heterogeneous malady with limited standard treatment options. RESULTS: The growth and progression of prostate cancer is a complicated process that involves multiple pathways. The signaling network comprising the integral constituents of the signature pathways involved in the development and progression of prostate cancer is modeled as a combinatorial circuit. The failures in the gene regulatory network that lead to cancer are abstracted as faults in the equivalent circuit and the Boolean circuit model is then used to design therapies tailored to counteract the effect of each molecular abnormality and to propose potentially efficacious combinatorial therapy regimens. Furthermore, stochastic computational modeling is utilized to identify potentially vulnerable components in the network that may serve as viable candidates for drug development. CONCLUSION: The results presented herein can aid in the design of scientifically well-grounded targeted therapies that can be employed for the treatment of prostate cancer patients.
Osama A. Arshad, Aniruddha Datta
BMC Bioinform.2
2017 A Survey of Software and Hardware Approaches to Performing Read Alignment in Next Generation Sequencing
abstract
Computational genomics is an emerging field that is enabling us to reveal the origins of life and the genetic basis of diseases such as cancer. Next Generation Sequencing (NGS) technologies have unleashed a wealth of genomic information by producing immense amounts of raw data. Before any functional analysis can be applied to this data, read alignment is applied to find the genomic coordinates of the produced sequences. Alignment algorithms have evolved rapidly with the advancement in sequencing technology, striving to achieve biological accuracy at the expense of increasing space and time complexities. Hardware approaches have been proposed to accelerate the computational bottlenecks created by the alignment process. Although several hardware approaches have achieved remarkable speedups, most have overlooked important biological features, which have hampered their widespread adoption by the genomics community. In this paper, we provide a brief biological introduction to genomics and NGS. We discuss the most popular next generation read alignment tools and algorithms. Furthermore, we provide a comprehensive survey of the hardware implementations used to accelerate these algorithms.
Ahmad Al Kawam, Sunil P. Khatri, Aniruddha Datta
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Hypoxia Stress Response Pathways: Modeling and Targeted Therapy
abstract
Hypoxia is a consequence of the decrease in the oxygen reaching the tissues of the body. It is a prominent feature of most solid tumors and is known to promote malignant progression, metastatic capacity, resistance to chemotherapy, and leads to poor patient prognosis. When a cell is under hypoxic stress, a cascade of cell signals is initiated through a family of transcription factors named as hypoxia inducible factors (HIFs). During hypoxia, HIF stabilizes and enters the nucleus and binds to the DNA via the hypoxia response element (HRE) and leads to the translation of downstream genes. The decision of adaptation or cell death depends on the extent of hypoxic stress faced by the cells. Proper understanding of hypoxic stress response is critical for understanding the mechanism of tumor cell adaptation to hypoxia and to develop efficient therapeutic interventions. In this paper, we develop a Boolean network model with targeted drug intervention in a cell that mimics persistent hypoxia. This hypoxic pathway is combined with pathways that help the cell adapt to the situation or undergo cell death. It is linked to apoptosis, cell survival, and energy production via the p53/Mdm2, PI3k/Akt/mTOR, and Glycolysis/TCA cycle pathways, respectively. In this model, we have incorporated eight known anticancer drugs that target these pathways. Through simulations, we have identified drug combinations that provided overall benefits to the cell in comparison to the no intervention case. Where applicable, the behavior predicted by this model is in agreement with experimental observations from the published literature.
Sriram Sridharan, Rajani Varghese, Vijayanagaram Venkatraj, Aniruddha Datta
IEEE J. Biomed. Health Informatics4
2016 Detecting Cell Growth and Drug Response in Heterogeneous Populations: A Dynamic Imaging Approach
abstract
Tumor heterogeneity has been increasingly recognized as one of the potentially contributing factors in explaining drug resistance. In order to gain better understanding of heterogeneous cancer cell populations and different cells' responses to various drugs, we use fluorescent proteins to mark the cells and a live-cell dynamic imaging platform to collect cell-by-cell measurements. After addressing the issue of fluorescent reporter variance in a Bayesian inference framework, we decompose the different cell types in the mixture and calculate their proportions and counts over time responding to different drug treatments. Additionally, the drug response as characterized by the cell death rate was also computed for these cells, and their different sensitivity was demonstrated. Overall, this work represents an important advancement toward evaluating cancer heterogeneity and drug responses in heterogeneous cancer cell populations.
Chao Sima, Jianping Hua, Rosana Lopes, Aniruddha Datta, Michael L. Bittner
BIBE4
2016 A measurement-based control design approach for efficient cancer chemotherapy
Sofiane Khadraoui, Fouzi Harrou, Hazem N. Nounou, Mohamed N. Nounou, Aniruddha Datta, Shankar P. Bhattacharyya
Inf. Sci.5
2016 Using Boolean Logic Modeling of Gene Regulatory Networks to Exploit the Links Between Cancer and Metabolism for Therapeutic Purposes
abstract
The uncontrolled cell proliferation that is characteristically associated with cancer is usually accompanied by alterations in the genome and cell metabolism. Indeed, the phenomenon of cancer cells metabolizing glucose using a less efficient anaerobic process even in the presence of normal oxygen levels, termed the Warburg effect, is currently considered to be one of the hallmarks of cancer. Diabetes, much like cancer, is defined by significant metabolic changes. Recent epidemiological studies have shown that diabetes patients treated with the antidiabetic drug Metformin have significantly lowered risk of cancer as compared to patients treated with other antidiabetic drugs. We utilize a Boolean logic model of the pathways commonly mutated in cancer to not only investigate the efficacy of Metformin for cancer therapeutic purposes but also demonstrate how Metformin in concert with other cancer drugs could provide better and less toxic clinical outcomes as compared to using cancer drugs alone.
Osama A. Arshad, Priyadharshini S. Venkatasubramani, Aniruddha Datta, Jijayanagaram Venkatraj
IEEE J. Biomed. Health Informatics3
2016 A Conjugate Exponential Model for Cancer Tissue Heterogeneity
abstract
The diagnosis and treatment of cancer is made difficult by the heterogeneous nature of the cell population. Determining its compositional breakup from measurements of various measurable traits (such as gene expression measurements) is an important problem in the field of cancer diagnosis and treatment. In addition, the computational aspect of the problem also needs attention. The processing of the collected data must be done as efficiently as possible in terms of computational speed and memory requirements. The use of Markov chain Monte Carlo methods is time consuming, and hence, other methods need to be used for the analysis. In this paper, we develop a model for heterogeneous cancer tissue, which uses quantitative polymerase chain reaction gene expression data to determine the compositional breakup of cell populations in the heterogeneous tissue. We develop a fast algorithm for the model using variational methods and demonstrate its use on synthetic and real-world gene expression data collected from fibroblasts and compare the performance of the algorithm with other methods such as Markov chain Monte Carlo and expectation maximization.
Anwoy Kumar Mohanty, Aniruddha Datta, Vijayanagaram Venkatraj
IEEE J. Biomed. Health Informatics2
2013 Parameter Estimation of Biological Phenomena: An Unscented Kalman Filter Approach
abstract
Recent advances in high-throughput technologies for biological data acquisition have spurred a broad interest in the construction of mathematical models for biological phenomena. The development of such mathematical models relies on the estimation of unknown parameters of the system using the time-course profiles of different metabolites in the system. One of the main challenges in the parameter estimation of biological phenomena is the fact that the number of unknown parameters is much more than the number of metabolites in the system. Moreover, the available metabolite measurements are corrupted by noise. In this paper, a new parameter estimation algorithm is developed based on the stochastic estimation framework for nonlinear systems, namely the unscented Kalman filter (UKF). A new iterative UKF algorithm with covariance resetting is developed in which the UKF algorithm is applied iteratively to the available noisy time profiles of the metabolites. The proposed estimation algorithm is applied to noisy time-course data synthetically produced from a generic branched pathway as well as real time-course profile for the Cad system of E. coli. The simulation results demonstrate the effectiveness of the proposed scheme.
Nader Meskin, Hazem N. Nounou, Mohamed N. Nounou, Aniruddha Datta
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 Wavelet-based Multiscale Filtering of Genomic Data
abstract
Measured biological data are a rich source of information about the biological phenomena they represent. For example, time-series genomic or metabolic micro array data can be used to construct dynamic genetic regulatory network models, which can be used to better understand the biological system and to design intervention strategies to cure or manage major diseases. Unfortunately, biological measurements are usually highly contaminated with errors that mask the important features in the data. Therefore, these noisy measurements need to be filtered to enhance their usefulness in practice. Wavelet-based multiscale filtering has been shown to be a powerful data analysis and denoising tool. In this work, different batch as well as online multiscale filtering techniques are used to filter biological data contaminated with white noise. The performances of these multiscale filtering techniques are demonstrated and compared to those of some conventional low pass filters using simulated time series metabolic data. The results of this comparative study show that significant improvement can be achieved using multiscale filtering over conventional filtering methods.
Mohamed N. Nounou, Hazem N. Nounou, Nader Meskin, Aniruddha Datta
ASONAM4
2012 Multiscale Denoising of Biological Data: A Comparative Analysis
abstract
Measured microarray genomic and metabolic data are a rich source of information about the biological systems they represent. For example, time-series biological data can be used to construct dynamic genetic regulatory network models, which can be used to design intervention strategies to cure or manage major diseases. Also, copy number data can be used to determine the locations and extent of aberrations in chromosome sequences. Unfortunately, measured biological data are usually contaminated with errors that mask the important features in the data. Therefore, these noisy measurements need to be filtered to enhance their usefulness in practice. Wavelet-based multiscale filtering has been shown to be a powerful denoising tool. In this work, different batch as well as online multiscale filtering techniques are used to denoise biological data contaminated with white or colored noise. The performances of these techniques are demonstrated and compared to those of some conventional low-pass filters using two case studies. The first case study uses simulated dynamic metabolic data, while the second case study uses real copy number data. Simulation results show that significant improvement can be achieved using multiscale filtering over conventional filtering techniques.
Mohamed N. Nounou, Hazem N. Nounou, Nader Meskin, Aniruddha Datta, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 Fuzzy Intervention in Biological Phenomena
abstract
An important objective of modeling biological phenomena is to develop therapeutic intervention strategies to move an undesirable state of a diseased network toward a more desirable one. Such transitions can be achieved by the use of drugs to act on some genes/metabolites that affect the undesirable behavior. Due to the fact that biological phenomena are complex processes with nonlinear dynamics that are impossible to perfectly represent with a mathematical model, the need for model-free nonlinear intervention strategies that are capable of guiding the target variables to their desired values often arises. In many applications, fuzzy systems have been found to be very useful for parameter estimation, model development and control design of nonlinear processes. In this paper, a model-free fuzzy intervention strategy (that does not require a mathematical model of the biological phenomenon) is proposed to guide the target variables of biological systems to their desired values. The proposed fuzzy intervention strategy is applied to three different biological models: a glycolytic-glycogenolytic pathway model, a purine metabolism pathway model, and a generic pathway model. The simulation results for all models demonstrate the effectiveness of the proposed scheme.
Hazem N. Nounou, Mohamed N. Nounou, Nader Meskin, Aniruddha Datta, Edward R. Dougherty
IEEE ACM Trans. Comput. Biol. Bioinform.4
2011 Cancer therapy design based on pathway logic
abstract
MOTIVATION: Cancer encompasses various diseases associated with loss of cell cycle control, leading to uncontrolled cell proliferation and/or reduced apoptosis. Cancer is usually caused by malfunction(s) in the cellular signaling pathways. Malfunctions occur in different ways and at different locations in a pathway. Consequently, therapy design should first identify the location and type of malfunction to arrive at a suitable drug combination. RESULTS: We consider the growth factor (GF) signaling pathways, widely studied in the context of cancer. Interactions between different pathway components are modeled using Boolean logic gates. All possible single malfunctions in the resulting circuit are enumerated and responses of the different malfunctioning circuits to a 'test' input are used to group the malfunctions into classes. Effects of different drugs, targeting different parts of the Boolean circuit, are taken into account in deciding drug efficacy, thereby mapping each malfunction to an appropriate set of drugs.
Ritwik Layek, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.2
2009 Adaptive intervention in probabilistic boolean networks
abstract
MOTIVATION: A basic problem of translational systems biology is to utilize gene regulatory networks as a vehicle to design therapeutic intervention strategies to beneficially alter network and, therefore, cellular dynamics. One strain of research has this problem from the perspective of control theory via the design of optimal Markov chain decision processes, mainly in the framework of probabilistic Boolean networks (PBNs). Full optimization assumes that the network is accurately modeled and, to the extent that model inference is inaccurate, which can be expected for gene regulatory networks owing to the combination of model complexity and a paucity of time-course data, the designed intervention strategy may perform poorly. We desire intervention strategies that do not assume accurate full-model inference. RESULTS: This article demonstrates the feasibility of applying on-line adaptive control to improve intervention performance in genetic regulatory networks modeled by PBNs. It shows via simulations that when the network is modeled by a member of a known family of PBNs, an adaptive design can yield improved performance in terms of the average cost. Two algorithms are presented, one better suited for instantaneously random PBNs and the other better suited for context-sensitive PBNs with low switching probability between the constituent BNs.
Ritwik Layek, Aniruddha Datta, Ranadip Pal, Edward R. Dougherty
Bioinform.2
2006 Intervention in a family of Boolean networks
abstract
MOTIVATION: Intervention in a gene regulatory network is used to avoid undesirable states, such as those associated with a disease. Several types of intervention have been studied in the framework of a probabilistic Boolean network (PBN), which is a collection of Boolean networks in which the gene state vector transitions according to the rules of one of the constituent networks and where network choice is governed by a selection distribution. The theory of automatic control has been applied to find optimal strategies for manipulating external control variables that affect the transition probabilities to desirably affect dynamic evolution over a finite time horizon. In this paper we treat a case in which we lack the governing probability structure for Boolean network selection, so we simply have a family of Boolean networks, but where these networks possess a common attractor structure. This corresponds to the situation in which network construction is treated as an ill-posed inverse problem in which there are many Boolean networks created from the data under the constraint that they all possess attractor structures matching the data states, which are assumed to arise from sampling the steady state of the real biological network. RESULTS: Given a family of Boolean networks possessing a common attractor structure composed of singleton attractors, a control algorithm is derived by minimizing a composite finite-horizon cost function that is a weighted average over all the individual networks, the idea being that we desire a control policy that on average suits the networks because these are viewed as equivalent relative to the data. The weighting for each network at any time point is taken to be proportional to the instantaneous estimated probability of that network being the underlying network governing the state transition. The results are applied to a family of Boolean networks derived from gene-expression data collected in a study of metastatic melanoma, the intent being to devise a control strategy that reduces the WNT5A gene's action in affecting biological regulation. AVAILABILITY: The software is available on request. SUPPLEMENTARY INFORMATION: The supplementary Information is available at http://ee.tamu.edu/~edward/tree
Ashish Choudhury, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.2
2005 Intervention in context-sensitive probabilistic Boolean networks
abstract
MOTIVATION: Intervention in a gene regulatory network is used to help it avoid undesirable states, such as those associated with a disease. Several types of intervention have been studied in the framework of a probabilistic Boolean network (PBN), which is essentially a finite collection of Boolean networks in which at any discrete time point the gene state vector transitions according to the rules of one of the constituent networks. For an instantaneously random PBN, the governing Boolean network is randomly chosen at each time point. For a context-sensitive PBN, the governing Boolean network remains fixed for an interval of time until a binary random variable determines a switch. The theory of automatic control has been previously applied to find optimal strategies for manipulating external (control) variables that affect the transition probabilities of an instantaneously random PBN to desirably affect its dynamic evolution over a finite time horizon. This paper extends the methods of external control to context-sensitive PBNs. RESULTS: This paper treats intervention via external control variables in context-sensitive PBNs by extending the results for instantaneously random PBNs in several directions. First, and most importantly, whereas an instantaneously random PBN yields a Markov chain whose state space is composed of gene vectors, each state of the Markov chain corresponding to a context-sensitive PBN is composed of a pair, the current gene vector occupied by the network and the current constituent Boolean network. Second, the analysis is applied to PBNs with perturbation, meaning that random gene perturbation is permitted at each instant with some probability. Third, the (mathematical) influence of genes within the network is used to choose the particular gene with which to intervene. Lastly, PBNs are designed from data using a recently proposed inference procedure that takes steady-state considerations into account. The results are applied to a context-sensitive PBN derived from gene-expression data collected in a study of metastatic melanoma, the intent being to devise a control strategy that reduces the WNT5A gene's action in affecting biological regulation, since the available data suggest that disruption of this influence could reduce the chance of a melanoma metastasizing.
Ranadip Pal, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.2
2005 Boolean relationships among genes responsive to ionizing radiation in the NCI 60 ACDS
abstract
MOTIVATION: An early use of gene-expression data coming from microarrays was to discover non-linear multivariate intergene relationships. Pursuing this direction, the motivation for this paper is 2-fold: (1) to discover and elucidate multivariate logical predictive relations among gene expressions in a dataset arising from radiation studies using the NCI 60 Anti-Cancer Drug Screen (ACDS) cell lines; and (2) to demonstrate how these logical relations based on coarse quantization reflect corresponding relations in the continuous data. RESULTS: Using the coefficient of determination, a large number of logical relationships have been discovered among genes in the NCI 60 ACDS cell lines. Moreover, these relationships can be seen directly in the original continuous data, and many are robust relative to the thresholds used to obtain the logical data from the continuous data. A key observation is that a number of intergene relationships appear to be considerably stronger when p53 is functional as compared to when it is not, which is consistent with earlier findings in the literature. AVAILABILITY: The appendix is available at http://gsp.tamu.edu/Publications/supplement.htm CONTACT: [email protected].
Ranadip Pal, Aniruddha Datta, Albert J. Fornace Jr., Michael L. Bittner, Edward R. Dougherty
Bioinform.2
2005 Generating Boolean networks with a prescribed attractor structure
abstract
MOTIVATION: Dynamical modeling of gene regulation via network models constitutes a key problem for genomics. The long-run characteristics of a dynamical system are critical and their determination is a primary aspect of system analysis. In the other direction, system synthesis involves constructing a network possessing a given set of properties. This constitutes the inverse problem. Generally, the inverse problem is ill-posed, meaning there will be many networks, or perhaps none, possessing the desired properties. Relative to long-run behavior, we may wish to construct networks possessing a desirable steady-state distribution. This paper addresses the long-run inverse problem pertaining to Boolean networks (BNs). RESULTS: The long-run behavior of a BN is characterized by its attractors. The rest of the state transition diagram is partitioned into level sets, the j-th level set being composed of all states that transition to one of the attractor states in exactly j transitions. We present two algorithms for the attractor inverse problem. The attractors are specified, and the sizes of the predictor sets and the number of levels are constrained. Algorithm complexity and performance are analyzed. The algorithmic solutions have immediate application. Under the assumption that sampling is from the steady state, a basic criterion for checking the validity of a designed network is that there should be concordance between the attractor states of the model and the data states. This criterion can be used to test a design algorithm: randomly select a set of states to be used as data states; generate a BN possessing the selected states as attractors, perhaps with some added requirements such as constraints on the number of predictors and the level structure; apply the design algorithm; and check the concordance between the attractor states of the designed network and the data states. AVAILABILITY: The software and supplementary material is available at http://gsp.tamu.edu/Publications/BNs/bn.htm
Ranadip Pal, Ivan Ivanov 0001, Aniruddha Datta, Michael L. Bittner, Edward R. Dougherty
Bioinform.3
2004 External control in Markovian genetic regulatory networks: the imperfect information case
abstract
Probabilistic Boolean Networks, which form a subclass of Markovian Genetic Regulatory Networks, have been recently introduced as a rule-based paradigm for modeling gene regulatory networks. In an earlier paper, we introduced external control into Markovian Genetic Regulatory networks. More precisely, given a Markovian genetic regulatory network whose state transition probabilities depend on an external (control) variable, a Dynamic Programming-based procedure was developed by which one could choose the sequence of control actions that minimized a given performance index over a finite number of steps. The control algorithm of that paper, however, could be implemented only when one had perfect knowledge of the states of the Markov Chain. This paper presents a control strategy that can be implemented in the imperfect information case, and makes use of the available measurements which are assumed to be probabilistically related to the states of the underlying Markov Chain.
Aniruddha Datta, Ashish Choudhury, Michael L. Bittner, Edward R. Dougherty
Bioinform.1
2003 External Control in Markovian Genetic Regulatory Networks
Aniruddha Datta, Ashish Choudhury, Michael L. Bittner, Edward R. Dougherty
Mach. Learn.1
1996 A modified model reference adaptive control scheme for rigid robots
abstract
This paper seeks to improve the zero-state performance of a standard model reference adaptive control scheme for rigid robots by augmenting the usual model reference control law with an auxiliary component synthesized from the tracking error. A condition involving a certain design parameter is given which can be used to decide a priori whether the modified scheme can potentially outperform its unmodified counterpart. If the joint accelerations are available for measurement, then it is shown that this condition can always be satisfied and the zero-state performance of the modified scheme can in fact be arbitrarily improved.
Aniruddha Datta, Ming-Tzu Ho
IEEE Trans. Robotics Autom.1
1991 Robust adaptive control: a unified approach
abstract
A complete tutorial review of the entire field is presented, beginning with simple instability examples to identify the causes of nonrobust behavior in adaptive control. Some of the mathematical groundwork is presented, and the theory for the design and analysis of adaptive laws is developed. Commonly used adaptive controller structures are discussed, highlighting their particular robustness properties. Particular attention is paid to model reference, pole placement, and linear quadratic controller structures. Designs and analyses of model reference, pole placement, and linear quadratic controllers, based on combining the corresponding controller structures with the various robust adaptive laws, are presented. Suggestions for future research are given.>
Petros A. Ioannou, Aniruddha Datta
Proc. IEEE2