Suzanne Sindi

dblp:57/6502 · also Suzanne S. Sindi · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0003-2742-4332ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2024 A new look at TFPI inhibition of factor X activation
abstract
Blood coagulation is a vital physiological process involving a complex network of biochemical reactions, which converge to form a blood clot that repairs vascular injury. This process unfolds in three phases: initiation, amplification, and propagation, ultimately leading to thrombin formation. Coagulation begins when tissue factor (TF) is exposed on an injured vessel's wall. The first step is when activated factor VII (VIIa) in the plasma binds to TF, forming complex TF:VIIa, which activates factor X. Activated factor X (Xa) is necessary for coagulation, so the regulation of its activation is crucial. Tissue Factor Pathway Inhibitor (TFPI) is a critical regulator of the initiation phase as it inhibits the activation of factor X. While previous studies have proposed two pathways-direct and indirect binding-for TFPI's inhibitory role, the specific biochemical reactions and their rates remain ambiguous. Many existing mathematical models only assume an indirect pathway, which may be less effective under physiological flow conditions. In this study, we revisit datasets from two experiments focused on activated factor X formation in the presence of TFPI. We employ an adaptive Metropolis method for parameter estimation to reinvestigate a previously proposed biochemical scheme and corresponding rates for both inhibition pathways. Our findings show that both pathways are essential to replicate the static experimental results. Previous studies have suggested that flow itself makes a significant contribution to the inhibition of factor X activation. We added flow to this model with our estimated parameters to determine the contribution of the two inhibition pathways under these conditions. We found that direct binding of TFPI is necessary for inhibition under flow. The indirect pathway has a weaker inhibitory effect due to removal of solution phase inhibitory complexes by flow.
Fabian Santiago, Shannon Bride, Dougald Monroe, Karin Leiderman, Suzanne Sindi
PLoS Comput. Biol.6
2023 Inferring gene regulatory networks using transcriptional profiles as dynamical attractors
abstract
Genetic regulatory networks (GRNs) regulate the flow of genetic information from the genome to expressed messenger RNAs (mRNAs) and thus are critical to controlling the phenotypic characteristics of cells. Numerous methods exist for profiling mRNA transcript levels and identifying protein-DNA binding interactions at the genome-wide scale. These enable researchers to determine the structure and output of transcriptional regulatory networks, but uncovering the complete structure and regulatory logic of GRNs remains a challenge. The field of GRN inference aims to meet this challenge using computational modeling to derive the structure and logic of GRNs from experimental data and to encode this knowledge in Boolean networks, Bayesian networks, ordinary differential equation (ODE) models, or other modeling frameworks. However, most existing models do not incorporate dynamic transcriptional data since it has historically been less widely available in comparison to "static" transcriptional data. We report the development of an evolutionary algorithm-based ODE modeling approach (named EA) that integrates kinetic transcription data and the theory of attractor matching to infer GRN architecture and regulatory logic. Our method outperformed six leading GRN inference methods, none of which incorporate kinetic transcriptional data, in predicting regulatory connections among TFs when applied to a small-scale engineered synthetic GRN in Saccharomyces cerevisiae. Moreover, we demonstrate the potential of our method to predict unknown transcriptional profiles that would be produced upon genetic perturbation of the GRN governing a two-state cellular phenotypic switch in Candida albicans. We established an iterative refinement strategy to facilitate candidate selection for experimentation; the experimental results in turn provide validation or improvement for the model. In this way, our GRN inference approach can expedite the development of a sophisticated mathematical model that can accurately describe the structure and dynamics of the in vivo GRN.
Ruihao Li 0005, Jordan C. Rozum, Morgan M. Quail, Mohammad N. Qasim, Suzanne Sindi, Clarissa J. Nobile, Réka Albert, Aaron D. Hernday
PLoS Comput. Biol.5
2022 ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders
abstract
MOTIVATION: Single-cell RNA sequencing (scRNAseq) technologies allow for measurements of gene expression at a single-cell resolution. This provides researchers with a tremendous advantage for detecting heterogeneity, delineating cellular maps or identifying rare subpopulations. However, a critical complication remains: the low number of single-cell observations due to limitations by rarity of subpopulation, tissue degradation or cost. This absence of sufficient data may cause inaccuracy or irreproducibility of downstream analysis. In this work, we present Automated Cell-Type-informed Introspective Variational Autoencoder (ACTIVA): a novel framework for generating realistic synthetic data using a single-stream adversarial variational autoencoder conditioned with cell-type information. Within a single framework, ACTIVA can enlarge existing datasets and generate specific subpopulations on demand, as opposed to two separate models [such as single-cell GAN (scGAN) and conditional scGAN (cscGAN)]. Data generation and augmentation with ACTIVA can enhance scRNAseq pipelines and analysis, such as benchmarking new algorithms, studying the accuracy of classifiers and detecting marker genes. ACTIVA will facilitate analysis of smaller datasets, potentially reducing the number of patients and animals necessary in initial studies. RESULTS: We train and evaluate models on multiple public scRNAseq datasets. In comparison to GAN-based models (scGAN and cscGAN), we demonstrate that ACTIVA generates cells that are more realistic and harder for classifiers to identify as synthetic which also have better pair-wise correlation between genes. Data augmentation with ACTIVA significantly improves classification of rare subtypes (more than 45% improvement compared with not augmenting and 4% better than cscGAN) all while reducing run-time by an order of magnitude in comparison to both models. AVAILABILITY AND IMPLEMENTATION: The codes and datasets are hosted on Zenodo (https://doi.org/10.5281/zenodo.5879639). Tutorials are available at https://github.com/SindiLab/ACTIVA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Abbas-Ali Heydari, Oscar A. Davalos, Lihong Zhao, Katrina K. Hoyer, Suzanne Sindi
Bioinform.5
2022 A structured model and likelihood approach to estimate yeast prion propagon replication rates and their asymmetric transmission
abstract
Prion proteins cause a variety of fatal neurodegenerative diseases in mammals but are generally harmless to Baker's yeast (Saccharomyces cerevisiae). This makes yeast an ideal model organism for investigating the protein dynamics associated with these diseases. The rate of disease onset is related to both the replication and transmission kinetics of propagons, the transmissible agents of prion diseases. Determining the kinetic parameters of propagon replication in yeast is complicated because the number of propagons in an individual cell depends on the intracellular replication dynamics and the asymmetric division of yeast cells within a growing yeast cell colony. We present a structured population model describing the distribution and replication of prion propagons in an actively dividing population of yeast cells. We then develop a likelihood approach for estimating the propagon replication rate and their transmission bias during cell division. We first demonstrate our ability to correctly recover known kinetic parameters from simulated data, then we apply our likelihood approach to estimate the kinetic parameters for six yeast prion variants using propagon recovery data. We find that, under our modeling framework, all variants are best described by a model with an asymmetric transmission bias. This demonstrates the strength of our framework over previous formulations assuming equal partitioning of intracellular constituents during cell division.
Fabian Santiago, Suzanne Sindi
PLoS Comput. Biol.2
2020 A unifying model for the propagation of prion proteins in yeast brings insight into the [PSI+] prion
abstract
The use of yeast systems to study the propagation of prions and amyloids has emerged as a crucial aspect of the global endeavor to understand those mechanisms. Yeast prion systems are intrinsically multi-scale: the molecular chemical processes are indeed coupled to the cellular processes of cell growth and division to influence phenotypical traits, observable at the scale of colonies. We introduce a novel modeling framework to tackle this difficulty using impulsive differential equations. We apply this approach to the [PSI+] yeast prion, which is associated with the misconformation and aggregation of Sup35. We build a model that reproduces and unifies previously conflicting experimental observations on [PSI+] and thus sheds light onto characteristics of the intracellular molecular processes driving aggregate replication. In particular our model uncovers a kinetic barrier for aggregate replication at low densities, meaning the change between prion or prion-free phenotype is a bi-stable transition. This result is based on the study of prion curing experiments, as well as the phenomenon of colony sectoring, a phenotype which is often ignored in experimental assays and has never been modeled. Furthermore, our results provide further insight into the effect of guanidine hydrochloride (GdnHCl) on Sup35 aggregates. To qualitatively reproduce the GdnHCl curing experiment, aggregate replication must not be completely inhibited, which suggests the existence of a mechanism different than Hsp104-mediated fragmentation. Those results are promising for further development of the [PSI+] model, but also for extending the use of this novel framework to other yeast prion or amyloid systems.
Paul Lemarre, Laurent Pujo-Menjouet, Suzanne Sindi
PLoS Comput. Biol.3
2019 Fine Tuning Sparsity Penalties to Improve Structural Variant Detection
abstract
Genomic variation shared by members of the same species that are longer than a single nucleotide are commonly called structural variants (SVs). Though relatively rare, they represent an increasingly important class of variation as SVs have been associated with diseases and susceptibility to some types of cancer. Computational approaches for detecting SVs often involve parameters that describe certain relevant biological phenomena. In our work, such parameters relate the incidence of inherited and novel SVs to probabilistic models of observing these SVs. In the work presented here, we investigate the sensitivity of our computational framework to these parameters. In particular, we demonstrate the robustness of our method by identifying a wide range of parameter values that lead to high-accuracy SV predictions in simulated data.
Hansell Perez, Melissa Spence, Roummel F. Marcia, Suzanne Sindi
BIBM5
2018 Detecting Novel Structural Variants In Genomes By Leveraging Parent-Child Relatedness
Melissa Spence, Mario Banuelos, Roummel F. Marcia, Suzanne Sindi
BIBM4
2018 Negative Binomial Optimization for Biomedical Structural Variant Signal Reconstruction
abstract
Structural variants (SVs) - novel adjacencies in an individual's genome - lead to genomic diversity across all organisms. When DNA fragments of an unknown genome are compared to a reference genome, errors in sequencing and mapping obscure true genomic rearrangements. When the sequencing coverage is low, this may lead to high false positive rates in predicted SVs. In this paper, we propose a novel maximum likelihood approach to SV prediction incorporating low-coverage sequencing data and coverage distribution. Specifically, we address mean and variance assumptions proposed by Poisson models and develop a Negative Binomial framework which reflects a more accurate representation of DNA fragments in an individual's genome. We incorporate both sparsity and inheritance in our model with an ℓ1penalty and linear constraints, respectively. We validate our model on both simulated and real genomic data of related individuals. Moreover, our results indicate an improvement on thresholding observations of candidate variants.
Mario Banuelos, Suzanne Sindi, Roummel F. Marcia
ICASSP2
2016 Sparse signal recovery methods for variant detection in next-generation sequencing data
abstract
Recent advances in high-throughput sequencing technologies, have led to the collection of vast quantities of genomic data., Structural variants (SVs) - rearrangements of the genome, larger than one letter such as inversions, insertions, deletions, and duplications - are an important source of genetic, variation and have been implicated in some genetic diseases., However, inferring SVs from sequencing data has proven to, be challenging because true SVs are rare and are prone to, low-coverage noise. In this paper, we attempt to mitigate the, deleterious effects of low-coverage sequences by following a, maximum likelihood approach to SV prediction. Specifically, we model the noise using Poisson statistics and constrain, the solution with a sparsity-promoting ℓ1penalty since SV, instances should be rare. In addition, because offspring SVs, inherit SVs from their parents, we incorporate familial relationships, in the optimization problem formulation to increase, the likelihood of detecting true SV occurrences. Numerical, results are presented to validate our proposed approach.
Mario Banuelos, Rubi Almanza, Lasith Adhikari, Suzanne Sindi, Roummel F. Marcia
ICASSP4
2014 Characterization of structural variants with single molecule and hybrid sequencing approaches
abstract
Abstract Motivation : Structural variation is common in human and cancer genomes. High-throughput DNA sequencing has enabled genome-scale surveys of structural variation. However, the short reads produced by these technologies limit the study of complex variants, particularly those involving repetitive regions. Recent ‘third-generation’ sequencing technologies provide single-molecule templates and longer sequencing reads, but at the cost of higher per-nucleotide error rates. Results : We present MultiBreak-SV, an algorithm to detect structural variants (SVs) from single molecule sequencing data, paired read sequencing data, or a combination of sequencing data from different platforms. We demonstrate that combining low-coverage third-generation data from Pacific Biosciences (PacBio) with high-coverage paired read data is advantageous on simulated chromosomes. We apply MultiBreak-SV to PacBio data from four human fosmids and show that it detects known SVs with high sensitivity and specificity. Finally, we perform a whole-genome analysis on PacBio data from a complete hydatidiform mole cell line and predict 1002 high-probability SVs, over half of which are confirmed by an Illumina-based assembly. Availability and implementation : MultiBreak-SV is available at http://compbio.cs.brown.edu/software/ . Contact : [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Anna M. Ritz, Ali Bashir, Suzanne Sindi, David Hsu, Iman Hajirasouliha, Benjamin J. Raphael
Bioinform.3
2012 Identification of polymorphic inversions from genotypes
abstract
BACKGROUND: Polymorphic inversions are a source of genetic variability with a direct impact on recombination frequencies. Given the difficulty of their experimental study, computational methods have been developed to infer their existence in a large number of individuals using genome-wide data of nucleotide variation. Methods based on haplotype tagging of known inversions attempt to classify individuals as having a normal or inverted allele. Other methods that measure differences between linkage disequilibrium attempt to identify regions with inversions but unable to classify subjects accurately, an essential requirement for association studies. RESULTS: We present a novel method to both identify polymorphic inversions from genome-wide genotype data and classify individuals as containing a normal or inverted allele. Our method, a generalization of a published method for haplotype data 1, utilizes linkage between groups of SNPs to partition a set of individuals into normal and inverted subpopulations. We employ a sliding window scan to identify regions likely to have an inversion, and accumulation of evidence from neighboring SNPs is used to accurately determine the inversion status of each subject. Further, our approach detects inversions directly from genotype data, thus increasing its usability to current genome-wide association studies (GWAS). CONCLUSIONS: We demonstrate the accuracy of our method to detect inversions and classify individuals on principled-simulated genotypes, produced by the evolution of an inversion event within a coalescent model 2. We applied our method to real genotype data from HapMap Phase III to characterize the inversion status of two known inversions within the regions 17q21 and 8p23 across 1184 individuals. Finally, we scan the full genomes of the European Origin (CEU) and Yoruba (YRI) HapMap samples. We find population-based evidence for 9 out of 15 well-established autosomic inversions, and for 52 regions previously predicted by independent experimental methods in ten (9+1) individuals 34. We provide efficient implementations of both genotype and haplotype methods as a unified R package inveRsion.
Alejandro Cáceres, Suzanne Sindi, Benjamin J. Raphael, Mario Cáceres, Juan R. González
BMC Bioinform.2
2009 Identification and Frequency Estimation of Inversion Polymorphisms from Haplotype Data
Suzanne Sindi, Benjamin J. Raphael
RECOMB1
2009 A geometric approach for classification and comparison of structural variants
abstract
MOTIVATION: Structural variants, including duplications, insertions, deletions and inversions of large blocks of DNA sequence, are an important contributor to human genome variation. Measuring structural variants in a genome sequence is typically more challenging than measuring single nucleotide changes. Current approaches for structural variant identification, including paired-end DNA sequencing/mapping and array comparative genomic hybridization (aCGH), do not identify the boundaries of variants precisely. Consequently, most reported human structural variants are poorly defined and not readily compared across different studies and measurement techniques. RESULTS: We introduce Geometric Analysis of Structural Variants (GASV), a geometric approach for identification, classification and comparison of structural variants. This approach represents the uncertainty in measurement of a structural variant as a polygon in the plane, and identifies measurements supporting the same variant by computing intersections of polygons. We derive a computational geometry algorithm to efficiently identify all such intersections. We apply GASV to sequencing data from nine individual human genomes and several cancer genomes. We obtain better localization of the boundaries of structural variants, distinguish genetic from putative somatic structural variants in cancer genomes, and integrate aCGH and paired-end sequencing measurements of structural variants. This work presents the first general framework for comparing structural variants across multiple samples and measurement techniques, and will be useful for studies of both genetic structural variants and somatic rearrangements in cancer. AVAILABILITY: http://cs.brown.edu/people/braphael/software.html .
Suzanne Sindi, Elena Helman, Ali Bashir, Benjamin J. Raphael
Bioinform.1
2007 Handbook of Computational Molecular Biology: Edited by Srinivas Aluru
abstract
The ‘Handbook of Computational Molecular Biology’, edited by Srinivas Aluru, is a sizable 1104 pages and represents the work of over 60 contributors. In addition to providing an excellent survey and overview of many topics, the text does not shy away from delving into the details. This makes the handbook valuable to those new to the field as well as more experienced researchers. Of course the task of covering many different areas while maintaining usability and readability is not easy. Inevitably, no matter how large the book, there will have to be topics left out. And certainly with the progression of the field, topics will fall in and out of fashion and new topics will emerge. None of these seemingly unavoidable facts diminish the need for a comprehensive resource. The handbook has done a good job of developing major areas in the field while maintaining a well-organized structure. The handbook comprises 38 chapters organized into eight disjoint sections/parts. Each section covers a distinct relatively broad area in the field and can be read independently. The chapters in each section proceed from more introductory material in earlier chapters to more advanced and recent topics in the later chapters. This structure ensures that even a reader with relatively little background will be able to access material from each of the broad areas, while the more advanced topics ensure a more experienced researcher will find the handbook useful. References are provided at the end of each chapter. Indeed, there appears to have been great care taken by the authors to provide an extensive list of references. There is also a global index making it possible to locate concepts that are common in several sections. Frequently the authors include pseudo-code versions of the algorithms they discuss. This feature greatly assists a reader interested in implementing their own version of the algorithms. The authors also take great care (as will be noted later in this review) to include references to available software utilizing the discussed algorithms for those readers who may have data they want to analyze. For those interested in using the handbook as a textbook, Srinivas Aluru gives suggestions on adapting the book for that purpose in the preface. The diversity and depth of topics covered mean that the text can be utilized for either an introductory or advanced course. While exercises are not formally included, the pseudo-code of the algorithms provides a good starting point for a computer science course while the survey of existing software will be useful for a course with less emphasis on programming. Although the text is designed for readers without experience in computational biology, those with a background in computer science will be at an advantage, especially in parts II and VII of the handbook (which cover string data structures, and bioinformatics databases and data mining, respectively). (For example, several chapters in these sections assume some familiarity with dynamic programming.) The authors of the introductory sections in these parts take care to provide illustrative figures and examples to assist the less experienced reader. Part I of the text is on Sequence Alignment; the first chapter includes a brief introduction to molecular biology. Although biological notions are developed throughout the book as needed, a reader new to computational biology would benefit by referring to a more complete introduction. Earlier chapters in this first section give a detailed study of global and local alignment for nucleotide and protein sequences. Later chapters explore more recent work in similarity-based gene recognition and multiple and parametric sequence alignment. Throughout the later chapters of Part I, the theory of Hidden Markov Models (HMMs) is developed. For readers not experienced with HMMs (or those interested in more detail) references to more complete theory are cited. The second part covers string data structures and their applications to computational biology. While not all readers will be familiar with data structures like suffix trees and suffix arrays, the authors are careful to provide proofs of the necessary properties. Applications using these data structures, such as repeat detection and identification of promoters and regulatory sequences are well-developed. In this section, as well as throughout the text, pseudo-code is provided for the algorithms discussed. These will greatly assist the interested researcher or student in writing their own software. For those readers not planning to develop their own code, references to available software packages, such as MUMmer and REP-uter that implement these structures are provided. The third part discusses genome assembly and clustering of expressed sequence tags (ESTs). This section begins by surveying current algorithms in genome assembly and proceeds to more recent topics such as comparative assembly and genome reconstruction. The chapter on genome assembly compares the methods of eight assembly programs used today as well as develops the theory behind shotgun sequence assembly. The second chapter in this section details the assembly of the human genome. By discussing the use of bacterial artificial chromosomes (BACs) and physical maps of chromosomes, this chapter gives a ‘real world’ look at genome assembly on a large scale. The final chapters in this section discuss clustering algorithms for ESTs and sequence assembly. Part IV covers a variety of relatively recent topics in genome-wide analysis. Global alignment and multiple global alignment of long sequences are discussed and (as with genome assembly packages) current software tools are compared. An interested reader will not only gain an understanding of some of the major computational issues in the area, but will be able to make a more informed choice about what software to utilize. Later chapters in this part discuss computational analysis of alternative splicing, human genetic linkage analysis and haplotype inference. The fifth section explores a variety of topics in another relatively long-standing area of computational biology, phylogenetics. The first chapter in this section gives an extensive general overview of the problem of phylogenetic reconstruction. Maximum likelihood and maximum parsimony methods are developed and compared. More generalized notions of phylogenetic trees are developed such as consensus trees and supertrees. Later chapters discuss the problem of large-scale phylogenetic reconstruction. The next two parts of the handbook discuss recent topics in computational biology in the area of systems biology. Part VI develops theory related to microarrays such as microarray design, data analysis, data storage and retrieval. There are two chapters dedicated to clustering algorithms, the second of which is a survey of biclustering algorithms. The final two chapters discuss the rapidly expanding area of identification of gene regulatory networks from microarray data and the modeling of such networks. Both deterministic and stochastic models are developed. The models discussed are systems of ordinary differential equations (although an example of a partial differential equation model is given). In addition, there is a survey of modeling packages (such as CellML and E-Cell) that can be used by researchers from knowledge of the biological system without formally developing a system of differential equations. The chapters in section seven cover a variety of topics in computational structural biology. The chapters in this section are less unified than in previous sections. The first three discuss protein structure prediction. The other chapters discuss processing reconstructed 3D maps of molecular complexes, detection of distant homologs and the use of parallel supercomputers in biomoleular modeling. The more recent problem of RNA structure prediction is not discussed in this section. The final section again will perhaps be of most interest to the computer scientist. These last four chapters in the text cover classic string searching problems such as finding approximate matches of a query string in a database, searching for motifs in a sequence and data mining. In spite of the large number of topics covered, there will be those whose area is left out, such as RNA secondary structure and modeling genome evolution. Although the handbook may not cover a particular topic a researcher is interested in, the impressive depth and breath make this text useful for nearly all interested in the field. Computational molecular biology is a rapidly growing field drawing in students and researchers from the biological sciences, mathematics, physics and computer science. Both the interdisciplinary nature of the field combined with its rapid growth pose difficultly in finding textbooks that are accessible to a broad audience and advanced enough to be useful to researchers. The ‘Handbook for Computational Molecular Biology’ will be a great resource to those interested in this exciting field.
Suzanne Sindi
Briefings Bioinform.1