Panayiotis V. Benos

dblp:78/2659 · also Panayiotis Takis Benos · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0003-3172-3132ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 Learning Identifiable Factorized Causal Representations of Cellular Responses
abstract
The study of cells and their responses to genetic or chemical perturbations promises to accelerate the discovery of therapeutics targets. However, designing adequate and insightful models for such data is difficult because the response of a cell to perturbations essentially depends on contextual covariates (e.g., genetic background or type of the cell). There is therefore a need for models that can identify interactions between drugs and contextual covariates. This is crucial for discovering therapeutics targets, as such interactions may reveal drugs that affect certain cell types but not others. We tackle this problem with a novel Factorized Causal Representation (FCR) learning method, an identifiable deep generative model that reveals causal structure in single-cell perturbation data from several cell lines. FCR learns multiple cellular representations that are disentangled, comprised of covariate-specific (Z_x), treatment-specific (Z_t) and interaction-specific (Z_tx) representations. Based on recent advances of non-linear ICA theory, we prove the component-wise identifiability of Z_tx and block-wise identifiability of Z_t and Z_x. Then, we present our implementation of FCR, and empirically demonstrate that FCR outperforms state-of-the-art baselines in various tasks across four single-cell datasets.
Haiyi Mao, Romain Lopez, Jan-Christian Hütter, David Richmond, Panayiotis V. Benos
NeurIPS6
2024 CellularPotts.jl: simulating multiscale cellular models in Julia
abstract
SUMMARY: CellularPotts.jl is a software package written in Julia to simulate biological cellular processes such as division, adhesion, and signaling. Accurately modeling and predicting these simple processes is crucial because they facilitate more complex biological phenomena related to important disease states like tumor growth, wound healing, and infection. Here we take advantage of Cellular Potts Modeling to simulate cellular interactions and combine them with differential equations to model dynamic cell signaling patterns. These models are advantageous over other approaches because they retain spatial information about each cell while remaining computationally efficient at larger scales. Users of this package define three key inputs to create valid model definitions: a 2- or 3-dimensional space, a table describing the cells to be positioned in that space, and a list of model penalties that dictate cell behaviors. Models can then be evolved over time to collect statistics, simulated repeatedly to investigate how changing a specific property impacts cellular behavior, and visualized using any of the available plotting libraries in Julia. AVAILABILITY AND IMPLEMENTATION: The CellularPotts.jl package is released under the MIT license and is available at https://github.com/RobertGregg/CellularPotts.jl. An archived version of the code (v0.3.2) at time of submission can also be found at https://doi.org/10.5281/zenodo.10407783.
Robert W. Gregg, Panayiotis V. Benos
Bioinform.2
2024 A hybrid constrained continuous optimization approach for optimal causal discovery from biological data
abstract
MOTIVATION: Understanding causal effects is a fundamental goal of science and underpins our ability to make accurate predictions in unseen settings and conditions. While direct experimentation is the gold standard for measuring and validating causal effects, the field of causal graph theory offers a tantalizing alternative: extracting causal insights from observational data. Theoretical analysis has shown that this is indeed possible, given a large dataset and if certain conditions are met. However, biological datasets, frequently, do not meet such requirements but evaluation of causal discovery algorithms is typically performed on synthetic datasets, which they meet all requirements. Thus, real-life datasets are needed, in which the causal truth is reasonably known. In this work we first construct such a large-scale real-life dataset and then we perform on it a comprehensive benchmarking of various causal discovery methods. RESULTS: We find that the PC algorithm is particularly accurate at estimating causal structure, including the causal direction which is critical for biological applicability. However, PC does only produces cause-effect directionality, but not estimates of causal effects. We propose PC-NOTEARS (PCnt), a hybrid solution, which includes the PC output as an additional constraint inside the NOTEARS optimization. This approach combines PC algorithm's strengths in graph structure prediction with the NOTEARS continuous optimization to estimate causal effects accurately. PCnt achieved best aggregate performance across all structural and effect size metrics. AVAILABILITY AND IMPLEMENTATION: https://github.com/zhu-yh1/PC-NOTEARS.
Yuehua Zhu, Panayiotis V. Benos, Maria Chikina
Bioinform.2
2024 Enabling the clinical application of artificial intelligence in genomics: a perspective of the AMIA Genomics and Translational Bioinformatics Workgroup
abstract
OBJECTIVE: Given the importance AI in genomics and its potential impact on human health, the American Medical Informatics Association-Genomics and Translational Biomedical Informatics (GenTBI) Workgroup developed this assessment of factors that can further enable the clinical application of AI in this space. PROCESS: A list of relevant factors was developed through GenTBI workgroup discussions in multiple in-person and online meetings, along with review of pertinent publications. This list was then summarized and reviewed to achieve consensus among the group members. CONCLUSIONS: Substantial informatics research and development are needed to fully realize the clinical potential of such technologies. The development of larger datasets is crucial to emulating the success AI is achieving in other domains. It is important that AI methods do not exacerbate existing socio-economic, racial, and ethnic disparities. Genomic data standards are critical to effectively scale such technologies across institutions. With so much uncertainty, complexity and novelty in genomics and medicine, and with an evolving regulatory environment, the current focus should be on using these technologies in an interface with clinicians that emphasizes the value each brings to clinical decision-making.
Nephi Walton, Radhakrishnan Nagarajan, Chen Wang 0001, Murat Sincan, Robert R. Freimuth, David B. Everman, Derek C. Walton, Scott McGrath, Dominick J. Lemas, Panayiotis V. Benos, Alexander V. Alekseyenko, Qianqian Song 0002, Ece D. Gamsiz Uzun, Casey Overby Taylor, Alper Uzun, Thomas N. Person, Nadav Rappoport, Zhongming Zhao, Marc S. Williams
J. Am. Medical Informatics Assoc.10
2023 An intrinsically interpretable neural network architecture for sequence-to-function learning
abstract
MOTIVATION: Sequence-based deep learning approaches have been shown to predict a multitude of functional genomic readouts, including regions of open chromatin and RNA expression of genes. However, a major limitation of current methods is that model interpretation relies on computationally demanding post hoc analyses, and even then, one can often not explain the internal mechanics of highly parameterized models. Here, we introduce a deep learning architecture called totally interpretable sequence-to-function model (tiSFM). tiSFM improves upon the performance of standard multilayer convolutional models while using fewer parameters. Additionally, while tiSFM is itself technically a multilayer neural network, internal model parameters are intrinsically interpretable in terms of relevant sequence motifs. RESULTS: We analyze published open chromatin measurements across hematopoietic lineage cell-types and demonstrate that tiSFM outperforms a state-of-the-art convolutional neural network model custom-tailored to this dataset. We also show that it correctly identifies context-specific activities of transcription factors with known roles in hematopoietic differentiation, including Pax5 and Ebf1 for B-cells, and Rorc for innate lymphoid cells. tiSFM's model parameters have biologically meaningful interpretations, and we show the utility of our approach on a complex task of predicting the change in epigenetic state as a function of developmental transition. AVAILABILITY AND IMPLEMENTATION: The source code, including scripts for the analysis of key findings, can be found at https://github.com/boooooogey/ATAConv, implemented in Python.
Ali Tugrul Balci, Mark Maher Ebeid, Panayiotis V. Benos, Dennis Kostka, Maria Chikina
Bioinform.3
2021 A Pipeline for Integrated Theory and Data-Driven Modeling of Biomedical Data
abstract
Genome sequencing technologies have the potential to transform clinical decision making and biomedical research by enabling high-throughput measurements of the genome at a granular level. However, to truly understand mechanisms of disease and predict the effects of medical interventions, high-throughput data must be integrated with demographic, phenotypic, environmental, and behavioral data from individuals. Further, effective knowledge discovery methods must infer relationships between these data types. We recently proposed a pipeline (CausalMGM) to achieve this. CausalMGM uses probabilistic graphical models to infer the relationships between variables in the data; however, CausalMGM's graphical structure learning algorithm can only handle small datasets efficiently. We propose a new methodology (piPref-Div) that selects the most informative variables for CausalMGM, enabling it to scale. We validate the efficacy of piPref-Div against other feature selection methods and demonstrate how the use of the full pipeline improves breast cancer outcome prediction and provides biologically interpretable views of gene expression data.
Vineet K. Raghu, Xiaoyu Ge, Arun Balajiee Lekshmi Narayanan, Daniel J. Shirer, Isha Das, Panayiotis V. Benos, Panos K. Chrysanthis
IEEE ACM Trans. Comput. Biol. Bioinform.6
2020 Interpretable Factors in scRNA-seq Data with Disentangled Generative Models
abstract
Single-cell RNA sequencing (scRNA-seq) experiments measure transcriptional profiles that encode diverse sources of variation, both biological and technical, and complex coordinations among them. PCA is a popular method for interpretable dimension reduction. However, it assumes a linear mapping between data and latent components and this may not be warranted in complex data. Single cell variational inference (scVI) offers a nonlinear method for latent space mapping, but the latent factors are not in general interpretable. In light of disentangled representations that learn independent data generative factors of single cell data in an unsupervised way, we propose factor variational inference (factorVI). FactorVI learns the disentangled factors among biologically relevant latent variables directly by penalizing correlations between them. We evaluate the factorVI through clustering in publicly available datasets and we observe high accuracy. We also propose biological interpretation of the latent factors.
Haiyi Mao, Matthew J. Broerman, Panayiotis V. Benos
BIBE3
2020 Causal network perturbations for instance-specific analysis of single cell and disease samples
abstract
MOTIVATION: Complex diseases involve perturbation in multiple pathways and a major challenge in clinical genomics is characterizing pathway perturbations in individual samples. This can lead to patient-specific identification of the underlying mechanism of disease thereby improving diagnosis and personalizing treatment. Existing methods rely on external databases to quantify pathway activity scores. This ignores the data dependencies and that pathways are incomplete or condition-specific. RESULTS: ssNPA is a new approach for subtyping samples based on deregulation of their gene networks. ssNPA learns a causal graph directly from control data. Sample-specific network neighborhood deregulation is quantified via the error incurred in predicting the expression of each gene from its Markov blanket. We evaluate the performance of ssNPA on liver development single-cell RNA-seq data, where the correct cell timing is recovered; and two TCGA datasets, where ssNPA patient clusters have significant survival differences. In all analyses ssNPA consistently outperforms alternative methods, highlighting the advantage of network-based approaches. AVAILABILITY AND IMPLEMENTATION: http://www.benoslab.pitt.edu/Software/ssnpa/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kristina Buschur, Maria Chikina, Panayiotis V. Benos
Bioinform.3
2020 An improvement of ComiR algorithm for microRNA target prediction by exploiting coding region sequences of mRNAs
abstract
MicroRNA are small non-coding RNAs that post-transcriptionally regulate the expression levels of messenger RNAs. MicroRNA regulation activity depends on the recognition of binding sites located on mRNA molecules. ComiR is a web tool realized to predict the targets of a set of microRNAs, starting from their expression profile. ComiR was trained with the information regarding binding sites in the 3'utr region, by using a reliable dataset containing the targets of endogenously expressed microRNA in D. melanogaster S2 cells. This dataset was obtained by comparing the results from two different experimental approaches, i.e., inhibition, and immunoprecipitation of the AGO1 protein--a component of the microRNA induced silencing complex.In this work, we tested whether including coding region binding sites in ComiR algorithm improves the performance of the tool in predicting microRNA targets. We focused the analysis on the D. melanogaster species and updated the ComiR underlying database with the currently available releases of mRNA and microRNA sequences. As a result, we find that ComiR algorithm trained with the information related to the coding regions is more efficient in predicting the microRNA targets, with respect to the algorithm trained with 3'utr information. On the other hand, we show that 3'utr based predictions can be seen as complementary to the coding region based predictions, which suggests that both predictions, from 3'utr and coding regions, should be considered in comprehensive analysis.Furthermore, we observed that the lists of targets obtained by analyzing data from one experimental approach only, that is, inhibition or immunoprecipitation of AGO1, are not reliable enough to test the performance of our microRNA target prediction algorithm. Further analysis will be conducted to investigate the effectiveness of the tool with data from other species, provided that validated datasets, as obtained from the comparison of RISC proteins inhibition and immunoprecipitation experiments, will be available for the same samples. Finally, we propose to upgrade the existing ComiR web-tool by including the coding region based trained model, available together with the 3'utr based one.
Giorgio Bertolazzi, Panayiotis V. Benos, Michele Tumminello, Claudia Coronnello
BMC Bioinform.2
2019 Mixed graphical models for integrative causal analysis with application to chronic lung disease diagnosis and prognosis
abstract
MOTIVATION: Integration of data from different modalities is a necessary step for multi-scale data analysis in many fields, including biomedical research and systems biology. Directed graphical models offer an attractive tool for this problem because they can represent both the complex, multivariate probability distributions and the causal pathways influencing the system. Graphical models learned from biomedical data can be used for classification, biomarker selection and functional analysis, while revealing the underlying network structure and thus allowing for arbitrary likelihood queries over the data. RESULTS: In this paper, we present and test new methods for finding directed graphs over mixed data types (continuous and discrete variables). We used this new algorithm, CausalMGM, to identify variables directly linked to disease diagnosis and progression in various multi-modal datasets, including clinical datasets from chronic obstructive pulmonary disease (COPD). COPD is the third leading cause of death and a major cause of disability and thus determining the factors that cause longitudinal lung function decline is very important. Applied on a COPD dataset, mixed graphical models were able to confirm and extend previously described causal effects and provide new insights on the factors that potentially affect the longitudinal lung function decline of COPD patients. AVAILABILITY AND IMPLEMENTATION: The CausalMGM package is available on http://www.causalmgm.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrew J. Sedgewick, Kristina Buschur, Ivy Shi, Joseph D. Ramsey, Vineet K. Raghu, Dimitris V. Manatakis, Yingze Zhang, Jessica Bon, Divay Chandra, Chad Karoleski, Frank C. Sciurba, Peter Spirtes, Clark Glymour, Panayiotis V. Benos
Bioinform.14
2018 ECCB 2018: The 17th European Conference on Computational Biology
abstract
This volume of Bioinformatics includes the proceedings papers of the 17th European Conference in Computational Biology (ECCB), an annual international Conference for research in computational biology and bioinformatics. The Conference is being held jointly with ISMB in the odd-numbered years and independently in the even-numbered years. This year, the 17th ECCB (ECCB 2018) will take place in Athens, at the Stavros Niarchos Foundation Cultural Center (SNFCC), from September 8 to 12, 2018. SNFCC, which was opened in June 2017, is a multifunctional arts, education and entertainment complex located on the edge of Faliro Bay, 4.5 km south of the center of Athens. Information on the ECCB 2018 can be found at eccb18.org and will be later archived at www.ebi.ac.uk/eccb/2018/. With more than 1,000 participants from academia and industry, ECCB is the leading European Conference on computational biology and bioinformatics and the second largest internationally, next to ISMB (Intelligent Systems in Molecular Biology). The work presented in ECCB is related to all domains of the field of computational biology and bioinformatics, ranging from molecular level to systems biology level. The proceedings papers present new computational methodologies and tools for addressing challenging problems in the field, including research in single cell high-throughput data, microbiome data, 3 D organization of the genome, and others. The Conference also includes Highlights presentations, which showcase important papers published over the last year. Since 2016, it also features an Applications track that presents computational biology work applied in industry, clinics, governmental organizations, and other fields beyond academia. ECCB meetings are held every year in a different European country. In odd-numbered years, ECCB is co-organized with the ISMB Conference, which is held in Europe. Previous meetings were organized in Prague, Czech Republic (Beerenwinkel and Bromberg, 2017, ISMB/ECCB 2017); The Hague, Netherlands (Heringa and Reinders, ECCB 2016); Dublin, Ireland (Moreau and Beerenwinkel, 2015, ISMB/ECCB 2015); Strasbourg, France (Devignes and Moreau, 2014, ECCB 2014); Berlin, Germany (Ben-Tal, 2013, ISMB/ECCB 2013); Basel, Switzerland (Schwede and Iber, 2012, ECCB 2012); Vienna, Austria (Gaasterland and Vingron, 2011, ISMB/ECCB 2011); Ghent, Belgium (Moreau and Heringa, 2010, ECCB 2010); Stockholm, Sweden (Gusfield and Tramontano, 2009, ISMB/ECCB 2009); Cagliari, Italy (Tramontano, 2008, ECCB 2008); Vienna, Austria (Lengauer et al., 2009, ISMB/ECCB 2007); Eilat, Israel (Wolfson and Safer, 2006, ECCB 2006); Madrid, Spain (Guigo et al., 2005, ECCB 2005); Glasgow, United Kingdom (Thornton et al., 2004, ISMB/ECCB 2004); Paris, France (Lenhof and Sagot, 2003, ECCB 2003); and Saarbrücken, Germany (Lengauer, 2002). ECCB 2018 continues the 16-year tradition and is held under the auspices of the Hellenic Society for Computational Biology and Bioinformatics (HSCBB; www.hscbb.gr), which promotes bioinformatics research and training in Greece since 2009. HSCBB organizes an annual conference in a different Greek city each time (e.g., Athens, Thessaloniki, Alexandroupolis, Patras, Lamia, Heraklion), aiming to provide exposure to new field developments to graduate students and researchers. The broad participation in the annual HSCBB conferences (more than 130 participants each year) show that the bioinformatics community in Greek universities and research centers has now matured significantly. HSCBB has also a strong international presence in the field of computational biology. It is an observer in the Greek ELIXIR consortium and is affiliated with the International Society for Computational Biology (ISCB) and Global Organization for Bioinformatics Learning, Education & Training (GOBLET). The organization of ECCB 2018, along with the European Student Council Symposium 2018 (see below), inspired the formation this year of the Greek chapter of the ISCB Student Council. ECCB 2018 received a record high number of 280 applications for proceedings talks from institutions from 48 countries. The submissions were organized in five themes, according to their topic: (1) Data (organization, integration, knowledge discovery, multi-scale modeling), (2) Genes (expression, function, editing, geno/phenotype), (3) Genome (sequence analysis, evolution, phylogeny, microbiome), (4) Proteins and Structural Biology (structure, function, alterations, drug design), and (5) Systems (molecular pathways, signaling, metabolomics). All proceedings submissions were subjected to a rigorous peer-review process (2-4 reviews per paper), organized by the members of the Programme Committees of the corresponding theme. The Programme Committees consisted of a Theme Chair (Michael Krauthammer, Roderic Guigo, Martin Vingron, Ivet Bahar, Alfonso Valencia, respectively) and two to four co-chairs (depending on the number of submissions assigned to this Theme). For the review, the EasyChair system (www.easychair.org) was used by 398 reviewers, recruited by the Programme Committees. Reviewers evaluated the impact and reproducibility of the presented research, as well as its suitability for the ECCB audience. Once the review was completed, the committee chairs and co-chairs selected 48 papers to be included in the ECCB 2018 proceedings (acceptance rate: 17%), with the number of papers accepted in each being proportional to the number of submissions initially assigned to this Theme. These papers required minor revisions and the authors had two weeks to modify them accordingly. All Proceedings Track papers and any supplementary files accompanying them are freely available in the electronic form of Oxford University Press journal Bioinformatics, as a special issue of September 2018. Together with the Proceedings track ECCB 2018 features 24 Highlights talks about scientific work already published in high impact science journals. They are coordinately presented and managed across the five themes with the proceeding presentations. We continued with newly established Application and ELIXIR tracks with 15 Applications talks, and 12 ELIXIR talks, all of which were selected after review. The aim of the Application track is to give a voice to those who apply computational biology in industry, clinics, governmental organizations, and other fields beyond academia. The aim of the ELIXIR Track is to showcase to the community the latest outputs and services from across the initiative. The presentations focus on developments relating to services and infrastructure within ELIXIR. ECCB 2018 also features a Poster Track, with 617 accepted posters, in which researchers present their latest findings in the corresponding thematic areas (Data, Genes, Genome, Proteins and Structural Biology, Systems). Seven distinguished keynote speakers will present their work: Prof. Bonnie Berger from the Massachusetts Institute of Technology (MIT), Prof. Christos Davatzikos from University of Pennsylvania, Prof. John Ioannidis from Stanford University, Prof. Manolis Kellis from MIT, Prof. Jose Onuchic from Rice University, Dr. Janet Thornton, Director Emeritus of the European Bioinformatics Institute (EMBL-EBI), and Dr. Eleftheria Zeggini from the German Research Center for Environmental Health (GmbH). Furthermore, ECCB 2018 is hosting for the first time three special invited talks before the lunch breaks and a joint keynote talk. The invited speakers will present the funding opportunities of the European Research Council (ERC) funding mechanisms (Dr. Maria Siomos) and a talk on the ISCB Student Council Internship Program (Farzana Rahman). The joint keynote will be about ELIXIR: A Common European Infrastructure for Bioinformatics Research presented by the ELIXIR director Dr. Niklas Blomberg. For the first time in ECCB 2018, ELIXIR organizes eight short workshops, namely ELIXIR TeSS Usability Study, Beacons, BioSchemas, Galaxy, OpenEbench Fnd Bio.tools, Research under the GDPR, Implementation of DMPs and Data Stewardship in practice, Secure access to your services using ELIXIR AAI. The 5th European Student Council Symposium (ESCS) is taking place before the main conference organized by the Student Council of the ISCB (chair Daniele Parisi from KU Leuven and co-chair Yvonne Saara Gladbach from Rostock University). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. For the first time the ISCB Student Council will present also an invited talk as stated above. We are really thankful to the ISCB Student Council for their enthusiasm, hard work and genuine scientific interest, which has made the ESCS an inseparable part of ECCB. During the ECCB 2018, ten exhibitor booths of all sizes will be set up on the Exhibit Hall floor, next to the central Conference venue and close to the Poster session and lunch area, namely: ELIXIR, BioExcel, EMBL-EBI, ERC, ISCB and ISCB Student Council, Goblet, Oxford University Press, Cambridge University Press, Peer Community In. The booths will be grouped together in a way that provides an enhanced networking environment for our exhibitors and delegates. Exhibitors will showcase the latest trends in computational biology and bioinformatics technology, in scientific literature, as well as in modelling and simulation. We invite participants to visit the exhibition area and support these community-minded organizations, who deliver a strong message in supporting the Computational Biology and Bioinformatics scientific field. The weekend before the conference 8 workshops (including two from Special Interest Groups – SIGs) and 13 Tutorials will take place. The aim of workshops is to provide participants the opportunity to discuss different perspectives on the cutting edge of a selected research field, through presenting technical issues, exchanging research ideas and sharing practical experiences. The two SIGs and six workshops preceding the ECCB 2018 main conference were selected out of a total of 12 applications, running all but one for a single day: (SIG-1) BioExcel 2nd SIG Meeting: Advanced Simulations for Biomolecular Research organized by Rossen Apostolov (KTH Royal Institute of Technology in Stockholm, Sweden), Zoe Cournia (Biomedical Research Foundation of the Academy of Athens, Greece), Vera Matser, (EMBL-EBI, UK), Anastas Mishev (UKIM Saints Cyril and Methodius University of Skopje, Republic of Macedonia), Hrachya Astsatryan (Institute for Informatics and Automation Problems, Armenia), Adam Carter (EPCC, UK) (SIG-2): Intrinsically disordered proteins: advances and state-of-the-art of the field organized by Silvio Tosatto (University of Padua, Italy), Zsuzsanna Dosztanyi (Eötvös Loránd University, Hungary), Norman Davey (University College Dublin, Ireland), Damiano Piovesan (University of Padua, Italy) (W1) Computational Epigenomics a two-day workshop organized by Yassen Assenov (German Cancer Research Center, Germany), Guido Sanguinetti (University of Edinburgh, UK), Jordana Bell (King’s College London, UK), Jörn Walter (Saarland University, Germany), Christoph Bock (CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, Austria) and Verena Wolf (Saarland University, Germany) (W2) BioNetVisA 2018 workshop: From biological network reconstruction to data visualization and analysis in molecular biology and medicine organized by Inna Kuperstein (Institut Curie, France), Emmanuel Barillot (Institut Curie, France), Andrei Zinovyev (Institut Curie, France), Luis Cristobal Monraz Gomez (Institut Curie, France), Hioraki Kitano (RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Institute for Chemical Research, Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain), Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, Netherlands), Patrick Kemmeren (Princess Maxima Center for Pediatric Oncology, Utrecht, Netherlands) (W3) CPW2018: Computational Pathology Workshop—second edition organized by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Netherlands), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Memorial Sloan Kettering Cancer Center, USA), Raphaël Marée (Université de Liège, Belgium), David Ameisen (IRIF, CNRS and Université Paris Diderot, France), Paul Van Diest (UMC Utrecht, Netherlands), Jeffrey Fine (Magee-Womens Hospital of UPMC) (W4) Interactive visualizations to guide health diagnostics and personalized medicine organized by Dimitrios Tzovaras, Konstantinos Votis and Kostas Stamatopoulos (all at Information Technologies Institute of the Center for Research and Technology Hellas, Greece) (W5) Logical modelling of cellular networks organized by Anna Niarakis (Univ Evry, Université Paris-Saclay, France) and Denis Thieffry (Ecole Normale Supérieure, Paris, France). (W6) Recent Computational Advances in Metagenomics organized by Valentin Loux, Mahendra Mariadassou, Pierre Peterlongo and Sophie Schbath (all at INRA, France) The purpose of the ECCB tutorial program is to provide participants with lectures and hands-on training to the most important and emerging topics in bioinformatics and computational biology research. The tutorials offer skills ranging from early and basic steps of computational analysis in recently introduced topics to advanced computational skills in important established topics. A total of thirteen ECCB 2018 tutorials running for half or a full day will be held at ECCB 2018. These were selected out of sixteen applications: (T1) Automated Machine Learning for Bioinformatics and Computational Biology organized by Ioannis Tsamardinos (University of Crete, Greece; Gnosis Data Analysis, Greece), Kleio Maria Verrou (University of Crete, Greece) and Vincenzo Lagani (Ilia State University, Georgia; Gnosis Data Analysis, Greece) (T2) Computational Mass Spectrometry with OpenMS—From Algorithms to Integrated Workflows organized by Julianus Pfeuffer (Freie Universität Berlin, Germany) and Timo Sachsenberg (Universität Tübingen, Germany) (T3)—Write your own R script to jointly analyze RNA-, ATAC-seq, DNA methylation, SNPs and find drug targets in signal transduction networks using TRANSFAC® organized by Philip Stegmaier (geneXplain GmbH, Germany), Olga Kel-Margoulis (geneXplain GmbH, Germany), Alexander Kel (geneXplain GmbH, Germany; Institute of Systems Biology, Russia) (T4) Deep learning for predicting protein–DNA and –RNA binding organized by Yaron Orenstein (Ben-Gurion University, Israel) (T5) Delving into non-coding RNA with RNAcentral and Rfam. Organized by Anton Petrov and Ioanna Kalvari (both at EMBL-EBI, UK) (T6) DIANA Tools and Databases: In silico investigation of miRNA functions organized by Artemis Hatzigeorgiou, Spyros Tastsoglou, Dimitra Karagkouni, Nikos Perdikopanis, Giorgos Skoufos, Ioannis Kavakiotis (all at University of Thessaly, Greece) (T7) Exploring programmatic access to Protein sequence, function and structure with UniProt and PDBe organized by Andrew Nightingale and Mihaly Varadi (both at EMBL-EBI, UK) (T8) Fifth International Hands-on Tutorial on Logical Modelling: Exploring the dynamics of biological systems organized by Tomas Helikar (University of Nebraska, USA) and Juilee Thakar (University of Rochester Medical Center, USA) (T9) From an application to a fully integrated workflow—Comprehensive software engineering with SeqAn organized by René Rahn and Hannes Hauswedell (both at Freie Universität Berlin, Germany) (T10) Hands-on on Protein Function Prediction with Machine Learning and Interactive Analytics organized by Rabie Saidi and Tunca Dogan (EMBL-EBI, UK) (T11) Single-cell RNA-Seq Data Analysis organized by Panagiotis Papasaikas (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) and Atul Sethi (Friedrich Miescher Institute for Biomedical Research (FMI), Basel, Switzerland; University Hospital Basel, University of Basel, Switzerland; Swiss Institute of Bioinformatics, Switzerland) (T12) Modern and scalable tools for efficient analysis of very large metagenomic datasets organized by Alexander Sczyrba (Bielefeld University, Germany), Christian Henke (Bielefeld University, Germany), Clovis Galiez (MPI for Biophysical Chemistry, Germany), Milot Mirdita (PI for Biophysical Chemistry, Germany) and Johannes Soeding (MPI for Biophysical Chemistry, Germany) (T13) User Experience Design for Computational Biologists organized by Nikiforos Karamanis and Xavier Watkins (both at EMBL-EBI, UK) Finally, following the tradition of ECCB, a number of travel fellowships, sponsored by ISCB, were given to young participants. We received 103 applications from scientists from 31 countries. After careful review and consideration, we awarded 14 travel fellowships, mainly to PhD students from nine countries that had an accepted proceedings paper. Volunteering is another important way for young scientists to visit and support the conference, as well as connect with other young scientists. The skills, talent and dedication of the volunteers are expected to largely contribute to the overall quality of the Conference. We made our best to accept all 91 applicants from 16 countries who were eligible to attend the Conference at a very low registration fee. Next to the volunteers, we would like to thank all the people who helped through their work to make this Conference a success. The Theme Chairs and Co-chairs as well all the reviewers who have been the heart of the Conference and helped selecting excellent papers, making thus this Conference a success. We are truly indebted to them. We are also grateful to the ECCB Steering Committee for their advice and continuous the of the organization of ECCB 2018. This conference would be at all the support from Anna Steering Committee and Chair ECCB our application and advice about to make a of was It was our to and inspired by for science with to in a will be in our and in our During the one year of the conference the support of Steering Committee Chair ECCB ECCB and Yves ECCB 2010, ECCB was Yves and for your continuous and We are also grateful to the ISCB Society for Computational for their and in ECCB 2018 at the international level and for This Conference would be the support of our and ELIXIR level and BioExcel was our level Cambridge University Press, Oxford University Press, Community and were ISCB, the Student Council of the ISCB, and who as exhibitors at ECCB 2018 also A thank to all for their and for to make ECCB 2018 an and Conference. A number of people to the organization of ECCB 2018. They beyond their of for making this conference really Maria a as was for the the the venue the and was for and Panagiotis was for the technical support of the registration Dr. organized the members of the DIANA also this Spyros Tastsoglou, Dimitra Karagkouni, Perdikopanis, Dimitrios Finally, the of a conference is its presentations, the conference for its participants. We thank all of for to ECCB 2018. We will the and in the city of Athens, the city and were of
Artemis G. Hatzigeorgiou, Pantelis G. Bagos, Panayiotis V. Benos, Christoforos Nikolaou, Yves Moreau, Ioannis Kavakiotis
Bioinform.3
2018 piMGM: incorporating multi-source priors in mixed graphical models for learning disease networks
abstract
Motivation: Learning probabilistic graphs over mixed data is an important way to combine gene expression and clinical disease data. Leveraging the existing, yet imperfect, information in pathway databases for mixed graphical model (MGM) learning is an understudied problem with tremendous potential applications in systems medicine, the problems of which often involve high-dimensional data. Results: We present a new method, piMGM, which can learn with accuracy the structure of probabilistic graphs over mixed data by appropriately incorporating priors from multiple experts with different degrees of reliability. We show that piMGM accurately scores the reliability of prior information from a given expert even at low sample sizes. The reliability scores can be used to determine active pathways in healthy and disease samples. We tested piMGM on both simulated and real data from TCGA, and we found that its performance is not affected by unreliable priors. We demonstrate the applicability of piMGM by successfully using prior information to identify pathway components that are important in breast cancer and improve cancer subtype classification. Availability and implementation: http://www.benoslab.pitt.edu/manatakisECCB2018.html. Supplementary information: Supplementary data are available at Bioinformatics online.
Dimitris V. Manatakis, Vineet K. Raghu, Panayiotis V. Benos
Bioinform.3
2017 Integrated Theory-and Data-Driven Feature Selection in Gene Expression Data Analysis
abstract
The exponential growth of high dimensional biological data has led to a rapid increase in demand for automated approaches for knowledge production. Existing methods rely on two general approaches to address this challenge: 1) the Theory-driven approach, which utilizes prior accumulated knowledge, and 2) the Data-driven approach, which solely utilizes the data to deduce scientific knowledge. Both of these approaches alone suffer from bias toward past/present knowledge, as they fail to incorporate all of the current knowledge that is available to make new discoveries. In this paper, we show how an integrated method can effectively address the high dimensionality of big biological data, which is a major problem for pure data-driven analysis approaches. We realize our approach in a novel two-step analytical workflow that incorporates a new feature selection paradigm as the first step to handling high-throughput gene expression data analysis and that utilizes graphical causal modeling as the second step to handle the automatic extraction of causal relationships. Our results, on real-world clinical datasets from The Cancer Genome Atlas (TCGA), demonstrate that our method is capable of intelligently selecting genes for learning effective causal networks.
Vineet K. Raghu, Xiaoyu Ge, Panos K. Chrysanthis, Panayiotis V. Benos
ICDE4
2016 Learning mixed graphical models with separate sparsity parameters and stability-based model selection
abstract
BACKGROUND: Mixed graphical models (MGMs) are graphical models learned over a combination of continuous and discrete variables. Mixed variable types are common in biomedical datasets. MGMs consist of a parameterized joint probability density, which implies a network structure over these heterogeneous variables. The network structure reveals direct associations between the variables and the joint probability density allows one to ask arbitrary probabilistic questions on the data. This information can be used for feature selection, classification and other important tasks. RESULTS: We studied the properties of MGM learning and applications of MGMs to high-dimensional data (biological and simulated). Our results show that MGMs reliably uncover the underlying graph structure, and when used for classification, their performance is comparable to popular discriminative methods (lasso regression and support vector machines). We also show that imposing separate sparsity penalties for edges connecting different types of variables significantly improves edge recovery performance. To choose these sparsity parameters, we propose a new efficient model selection method, named Stable Edge-specific Penalty Selection (StEPS). StEPS is an expansion of an earlier method, StARS, to mixed variable types. In terms of edge recovery, StEPS selected MGMs outperform those models selected using standard techniques, including AIC, BIC and cross-validation. In addition, we use a heuristic search that is linear in size of the sparsity value search space as opposed to the cubic grid search required by other model selection methods. We applied our method to clinical and mRNA expression data from the Lung Genomics Research Consortium (LGRC) and the learned MGM correctly recovered connections between the diagnosis of obstructive or interstitial lung disease, two diagnostic breathing tests, and cigarette smoking history. Our model also suggested biologically relevant mRNA markers that are linked to these three clinical variables. CONCLUSIONS: MGMs are able to accurately recover dependencies between sets of continuous and discrete variables in both simulated and biomedical datasets. Separation of sparsity penalties by edge type is essential for accurate network edge recovery. Furthermore, our stability based method for model selection determines sparsity parameters faster and more accurately (in terms of edge recovery) than other model selection methods. With the ongoing availability of comprehensive clinical and biomedical datasets, MGMs are expected to become a valuable tool for investigating disease mechanisms and answering an array of critical healthcare questions.
Andrew J. Sedgewick, Ivy Shi, Rory M. Donovan, Panayiotis V. Benos
BMC Bioinform.4
2015 The center for causal discovery of biomedical knowledge from big data
abstract
The Big Data to Knowledge (BD2K) Center for Causal Discovery is developing and disseminating an integrated set of open source tools that support causal modeling and discovery of biomedical knowledge from large and complex biomedical datasets. The Center integrates teams of biomedical and data scientists focused on the refinement of existing and the development of new constraint-based and Bayesian algorithms based on causal Bayesian networks, the optimization of software for efficient operation in a supercomputing environment, and the testing of algorithms and software developed using real data from 3 representative driving biomedical projects: cancer driver mutations, lung disease, and the functional connectome of the human brain. Associated training activities provide both biomedical and data scientists with the knowledge and skills needed to apply and extend these tools. Collaborative activities with the BD2K Consortium further advance causal discovery tools and integrate tools and resources developed by other centers.
Gregory F. Cooper, Ivet Bahar, Michael J. Becich, Panayiotis V. Benos, Jeremy M. Berg, Jeremy U. Espino, Clark Glymour, Rebecca S. Jacobson, Michelle Kienholz, Adrian V. Lee, Xinghua Lu 0001, Richard Scheines
J. Am. Medical Informatics Assoc.4
2012 Novel Modeling of Combinatorial miRNA Targeting Identifies SNP with Potential Role in Bone Density
abstract
MicroRNAs (miRNAs) are post-transcriptional regulators that bind to their target mRNAs through base complementarity. Predicting miRNA targets is a challenging task and various studies showed that existing algorithms suffer from high number of false predictions and low to moderate overlap in their predictions. Until recently, very few algorithms considered the dynamic nature of the interactions, including the effect of less specific interactions, the miRNA expression level, and the effect of combinatorial miRNA binding. Addressing these issues can result in a more accurate miRNA:mRNA modeling with many applications, including efficient miRNA-related SNP evaluation. We present a novel thermodynamic model based on the Fermi-Dirac equation that incorporates miRNA expression in the prediction of target occupancy and we show that it improves the performance of two popular single miRNA target finders. Modeling combinatorial miRNA targeting is a natural extension of this model. Two other algorithms show improved prediction efficiency when combinatorial binding models were considered. ComiR (Combinatorial miRNA targeting), a novel algorithm we developed, incorporates the improved predictions of the four target finders into a single probabilistic score using ensemble learning. Combining target scores of multiple miRNAs using ComiR improves predictions over the naïve method for target combination. ComiR scoring scheme can be used for identification of SNPs affecting miRNA binding. As proof of principle, ComiR identified rs17737058 as disruptive to the miR-488-5p:NCOA1 interaction, which we confirmed in vitro. We also found rs17737058 to be significantly associated with decreased bone mineral density (BMD) in two independent cohorts indicating that the miR-488-5p/NCOA1 regulatory axis is likely critical in maintaining BMD in women. With increasing availability of comprehensive high-throughput datasets from patients ComiR is expected to become an essential tool for miRNA-related studies.
Claudia Coronnello, Ryan J. Hartmaier, Arshi Arora, Luai Huleihel, Kusum V. Pandit, Abha S. Bais, Michael Butterworth, Naftali Kaminski, Gary D. Stormo, Steffi Oesterreich, Panayiotis V. Benos
PLoS Comput. Biol.11
2009 HHMMiR: efficient de novo prediction of microRNAs using hierarchical hidden Markov models
abstract
BACKGROUND: MicroRNAs (miRNAs) are small non-coding single-stranded RNAs (20-23 nts) that are known to act as post-transcriptional and translational regulators of gene expression. Although, they were initially overlooked, their role in many important biological processes, such as development, cell differentiation, and cancer has been established in recent times. In spite of their biological significance, the identification of miRNA genes in newly sequenced organisms is still based, to a large degree, on extensive use of evolutionary conservation, which is not always available. RESULTS: We have developed HHMMiR, a novel approach for de novo miRNA hairpin prediction in the absence of evolutionary conservation. Our method implements a Hierarchical Hidden Markov Model (HHMM) that utilizes region-based structural as well as sequence information of miRNA precursors. We first established a template for the structure of a typical miRNA hairpin by summarizing data from publicly available databases. We then used this template to develop the HHMM topology. CONCLUSION: Our algorithm achieved average sensitivity of 84% and specificity of 88%, on 10-fold cross-validation of human miRNA precursor data. We also show that this model, trained on human sequences, works well on hairpins from other vertebrate as well as invertebrate species. Furthermore, the human trained model was able to correctly classify ~97% of plant miRNA precursors. The success of this approach in such a diverse set of species indicates that sequence conservation is not necessary for miRNA prediction. This may lead to efficient prediction of miRNA genes in virtually any organism.
Sabah Kadri, Veronica F. Hinman, Panayiotis V. Benos
BMC Bioinform.3
2009 Extracting biologically significant patterns from short time series gene expression data
abstract
BACKGROUND: Time series gene expression data analysis is used widely to study the dynamics of various cell processes. Most of the time series data available today consist of few time points only, thus making the application of standard clustering techniques difficult. RESULTS: We developed two new algorithms that are capable of extracting biological patterns from short time point series gene expression data. The two algorithms, ASTRO and MiMeSR, are inspired by the rank order preserving framework and the minimum mean squared residue approach, respectively. However, ASTRO and MiMeSR differ from previous approaches in that they take advantage of the relatively few number of time points in order to reduce the problem from NP-hard to linear. Tested on well-defined short time expression data, we found that our approaches are robust to noise, as well as to random patterns, and that they can correctly detect the temporal expression profile of relevant functional categories. Evaluation of our methods was performed using Gene Ontology (GO) annotations and chromatin immunoprecipitation (ChIP-chip) data. CONCLUSION: Our approaches generally outperform both standard clustering algorithms and algorithms designed specifically for clustering of short time series gene expression data. Both algorithms are available at http://www.benoslab.pitt.edu/astro/.
Alain B. Tchagang, Kevin V. Bui, Thomas McGinnis, Panayiotis V. Benos
BMC Bioinform.4
2008 Biological evaluation of biclustering algorithms using Gene Ontology and chIP-chip data
abstract
In this paper, we propose a new framework for assessing the biological significance of the outputs of any biclustering algorithm. The framework relies on the p-value computed by a Fisher's exact test on a 2x2 contingency table derived from gene ontology (GO) enrichment level and chromatin immunoprecipitation (ChIP) data enrichment level. We illustrate the framework using our published robust biclustering algorithm (RoBA), the Cheng and Church (CC) algorithm, and a well-defined set of yeast cell cycle gene expression data and chip-chip data. Our evaluation also shows that the biclusters identified by RoBA are biologically more homogeneous than the ones identified by the Cheng and Church (CC) algorithm.
Alain B. Tchagang, Ahmed H. Tewfik, Panayiotis V. Benos
ICASSP3
2007 DNA Familial Binding Profiles Made Easy: Comparison of Various Motif Alignment and Clustering Strategies
abstract
Transcription factor (TF) proteins recognize a small number of DNA sequences with high specificity and control the expression of neighbouring genes. The evolution of TF binding preference has been the subject of a number of recent studies, in which generalized binding profiles have been introduced and used to improve the prediction of new target sites. Generalized profiles are generated by aligning and merging the individual profiles of related TFs. However, the distance metrics and alignment algorithms used to compare the binding profiles have not yet been fully explored or optimized. As a result, binding profiles depend on TF structural information and sometimes may ignore important distinctions between subfamilies. Prediction of the identity or the structural class of a protein that binds to a given DNA pattern will enhance the analysis of microarray and ChIP-chip data where frequently multiple putative targets of usually unknown TFs are predicted. Various comparison metrics and alignment algorithms are evaluated (a total of 105 combinations). We find that local alignments are generally better than global alignments at detecting eukaryotic DNA motif similarities, especially when combined with the sum of squared distances or Pearson's correlation coefficient comparison metrics. In addition, multiple-alignment strategies for binding profiles and tree-building methods are tested for their efficiency in constructing generalized binding models. A new method for automatic determination of the optimal number of clusters is developed and applied in the construction of a new set of familial binding profiles which improves upon TF classification accuracy. A software tool, STAMP, is developed to host all tested methods and make them publicly available. This work provides a high quality reference set of familial binding profiles and the first comprehensive platform for analysis of DNA profiles. Detecting similarities between DNA motifs is a key step in the comparative study of transcriptional regulation, and the work presented here will form the basis for tool and method development for future transcriptional modeling studies.
Shaun Mahony, Philip E. Auron, Panayiotis V. Benos
PLoS Comput. Biol.3
2006 Self-organizing neural networks to support the discovery of DNA-binding motifs
Shaun Mahony, Panayiotis V. Benos, Terry J. Smith, Aaron Golden
Neural Networks2