EDBT 2026 Demo / reviewers in the wild / expert
Pietro Pinoli
dblp:118/4967
· DBLP profile ↗
29ranked-venue papers
7as first author
9since 2021 · last 2022
0000-0001-9786-2851ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 4Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | MCTK: a Multi-modal Conversational Troubleshooting Kit for supporting users in web applicationsabstractConversational Interfaces for user assistance are becoming persuasive. Today, though, most chatbots are not integrated into the application in which they are placed, but only superimposed, with no communication between the conversational and the graphical interface. We propose Multi-modal Conversational Troubleshooting Kit (MCTK), a Python package to easily integrate a conversational agent for troubleshooting in web applications. MCTK is multi-modal: once the system recognizes the problem the user is encountering, the textual solution in the chat is coupled with visual hints in the GUI. On top of that, MCTK is easy to configure and offers separation of concerns: dialogue designers can work on the conversation without the necessity of modifying the code, and vice versa. Giulio Antonio Abbo, Pietro Crovari, Sara Pidò, Pietro Pinoli, Franca Garzotto |
AVI | 4 |
| 2022 | Conceptual models and databases for searching the genome
Anna Bernasconi 0002, Pietro Pinoli |
EDBT | 2 |
| 2022 | ViruClust: direct comparison of SARS-CoV-2 genomes and genetic variants in space and timeabstractMOTIVATION: The ongoing evolution of SARS-CoV-2 and the rapid emergence of variants of concern at distinct geographic locations have relevant implications for the implementation of strategies for controlling the COVID-19 pandemic. Combining the growing body of data and the evidence on potential functional implications of SARS-CoV-2 mutations can suggest highly effective methods for the prioritization of novel variants of potential concern, e.g. increasing in frequency locally and/or globally. However, these analyses may be complex, requiring the integration of different data and resources. We claim the need for a streamlined access to up-to-date and high-quality genome sequencing data from different geographic regions/countries, and the current lack of a robust and consistent framework for the evaluation/comparison of the results. RESULTS: To overcome these limitations, we developed ViruClust, a novel tool for the comparison of SARS-CoV-2 genomic sequences and lineages in space and time. ViruClust is made available through a powerful and intuitive web-based user interface. Sophisticated large-scale analyses can be executed with a few clicks, even by users without any computational background. To demonstrate potential applications of our method, we applied ViruClust to conduct a thorough study of the evolution of the most prevalent lineage of the Delta SARS-CoV-2 variant, and derived relevant observations. By allowing the seamless integration of different types of functional annotations and the direct comparison of viral genomes and genetic variants in space and time, ViruClust represents a highly valuable resource for monitoring the evolution of SARS-CoV-2, facilitating the identification of variants and/or mutations of potential concern. AVAILABILITY AND IMPLEMENTATION: ViruClust is openly available at http://gmql.eu/viruclust/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Luca Cilibrasi, Pietro Pinoli, Anna Bernasconi 0002, Arif Canakoglu, Matteo Chiara, Stefano Ceri |
Bioinform. | 2 |
| 2022 | GeCoAgent: A Conversational Agent for Empowering Genomic Data Extraction and AnalysisabstractWith the availability of reliable and low-cost DNA sequencing, human genomics is relevant to a growing number of end-users, including biologists and clinicians. Typical interactions require applying comparative data analysis to huge repositories of genomic information for building new knowledge, taking advantage of the latest findings in applied genomics for healthcare. Powerful technology for data extraction and analysis is available, but broad use of the technology is hampered by the complexity of accessing such methods and tools. This work presents GeCoAgent, a big-data service for clinicians and biologists. GeCoAgent uses a dialogic interface, animated by a chatbot, for supporting the end-users’ interaction with computational tools accompanied by multi-modal support. While the dialogue progresses, the user is accompanied in extracting the relevant data from repositories and then performing data analysis, which often requires the use of statistical methods or machine learning. Results are returned using simple representations (spreadsheets and graphics), while at the end of a session the dialogue is summarized in textual format. The innovation presented in this article is concerned with not only the delivery of a new tool but also our novel approach to conversational technologies, potentially extensible to other healthcare domains or to general data science. Pietro Crovari, Sara Pidò, Pietro Pinoli, Anna Bernasconi 0002, Arif Canakoglu, Franca Garzotto, Stefano Ceri |
ACM Trans. Comput. Heal. | 3 |
| 2022 | Investigating Deep Learning Based Breast Cancer Subtyping Using Pan-Cancer and Multi-Omic DataabstractBreast Cancer comprises multiple subtypes implicated in prognosis. Existing stratification methods rely on the expression quantification of small gene sets. Next Generation Sequencing promises large amounts of omic data in the next years. In this scenario, we explore the potential of machine learning and, particularly, deep learning for breast cancer subtyping. Due to the paucity of publicly available data, we leverage on pan-cancer and non-cancer data to design semi-supervised settings. We make use of multi-omic data, including microRNA expressions and copy number alterations, and we provide an in-depth investigation of several supervised and semi-supervised architectures. Obtained accuracy results show simpler models to perform at least as well as the deep semi-supervised approaches on our task over gene expression data. When multi-omic data types are combined together, performance of deep models shows little (if any) improvement in accuracy, indicating the need for further analysis on larger datasets of multi-omic data as and when they become available. From a biological perspective, our linear model mostly confirms known gene-subtype annotations. Conversely, deep approaches model non-linear relationships, which is reflected in a more varied and still unexplored set of representative omic features that may prove useful for breast cancer subtyping. Francisco Cristovao, Silvia Cascianelli, Arif Canakoglu, Mark J. Carman, Luca Nanni, Pietro Pinoli, Marco Masseroli |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | Predicting Drug Synergism by Means of Non-Negative Matrix Tri-FactorizationabstractTraditional drug experiments to find synergistic drug pairs are time-consuming and expensive due to the numerous possible combinations of drugs that have to be examined. Thus, computational methods that can give suggestions for synergistic drug investigations are of great interest. Here, we propose a Non-negative Matrix Tri-Factorization (NMTF) based approach that leverages the integration of different data types for predicting synergistic drug pairs in multiple specific cell lines. Our computational framework relies on a network-based representation of available data about drug synergism, which also allows integrating genomic information about cell lines. We computationally evaluate the performances of our method in finding missing relationships between synergistic drug pairs and cell lines, and in computing synergy scores between drug pairs in a specific cell line, as well as we estimate the benefit of adding cell line genomic data to the network. Our approach obtains very good performance (Average Precision Score equal to 0.937, Pearson's correlation coefficient equal to 0.760) when cell line genomic data and rich data about synergistic drugs in a cell line are considered. Finally, we systematically searched our top-scored predictions in the available literature and in the NCI ALMANAC, a well-known database of drug combination experiments, proving the goodness of our findings. Pietro Pinoli, Gaia Ceddia, Stefano Ceri, Marco Masseroli |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | A review on viral data sources and search systems for perspective mitigation of COVID-19abstractWith the outbreak of the COVID-19 disease, the research community is producing unprecedented efforts dedicated to better understand and mitigate the effects of the pandemic. In this context, we review the data integration efforts required for accessing and searching genome sequences and metadata of SARS-CoV2, the virus responsible for the COVID-19 disease, which have been deposited into the most important repositories of viral sequences. Organizations that were already present in the virus domain are now dedicating special interest to the emergence of COVID-19 pandemics, by emphasizing specific SARS-CoV2 data and services. At the same time, novel organizations and resources were born in this critical period to serve specifically the purposes of COVID-19 mitigation while setting the research ground for contrasting possible future pandemics. Accessibility and integration of viral sequence data, possibly in conjunction with the human host genotype and clinical data, are paramount to better understand the COVID-19 disease and mitigate its effects. Few examples of host-pathogen integrated datasets exist so far, but we expect them to grow together with the knowledge of COVID-19 disease; once such datasets will be available, useful integrative surveillance mechanisms can be put in place by observing how common variants distribute in time and space, relating them to the phenotypic impact evidenced in the literature. Anna Bernasconi 0002, Arif Canakoglu, Marco Masseroli, Pietro Pinoli, Stefano Ceri |
Briefings Bioinform. | 4 |
| 2021 | Federated sharing and processing of genomic datasets for tertiary data analysisabstractMOTIVATION: With the spreading of biological and clinical uses of next-generation sequencing (NGS) data, many laboratories and health organizations are facing the need of sharing NGS data resources and easily accessing and processing comprehensively shared genomic data; in most cases, primary and secondary data management of NGS data is done at sequencing stations, and sharing applies to processed data. Based on the previous single-instance GMQL system architecture, here we review the model, language and architectural extensions that make the GMQL centralized system innovatively open to federated computing. RESULTS: A well-designed extension of a centralized system architecture to support federated data sharing and query processing. Data is federated thanks to simple data sharing instructions. Queries are assigned to execution nodes; they are translated into an intermediate representation, whose computation drives data and processing distributions. The approach allows writing federated applications according to classical styles: centralized, distributed or externalized. AVAILABILITY: The federated genomic data management system is freely available for non-commercial use as an open source project at http://www.bioinformatics.deib.polimi.it/FederatedGMQLsystem/. CONTACT: {arif.canakoglu, pietro.pinoli}@polimi.it. Arif Canakoglu, Pietro Pinoli, Andrea Gulino, Luca Nanni, Marco Masseroli, Stefano Ceri |
Briefings Bioinform. | 2 |
| 2021 | Identifying collateral and synthetic lethal vulnerabilities within the DNA-damage responseabstractBACKGROUND: A pair of genes is defined as synthetically lethal if defects on both cause the death of the cell but a defect in only one of the two is compatible with cell viability. Ideally, if A and B are two synthetic lethal genes, inhibiting B should kill cancer cells with a defect on A, and should have no effects on normal cells. Thus, synthetic lethality can be exploited for highly selective cancer therapies, which need to exploit differences between normal and cancer cells. RESULTS: In this paper, we present a new method for predicting synthetic lethal (SL) gene pairs. As neighbouring genes in the genome have highly correlated profiles of copy number variations (CNAs), our method clusters proximal genes with a similar CNA profile, then predicts mutually exclusive group pairs, and finally identifies the SL gene pairs within each group pairs. For mutual-exclusion testing we use a graph-based method which takes into account the mutation frequencies of different subjects and genes. We use two different methods for selecting the pair of SL genes; the first is based on the gene essentiality measured in various conditions by means of the "Gene Activity Ranking Profile" GARP score; the second leverages the annotations of gene to biological pathways. CONCLUSIONS: This method is unique among current SL prediction approaches, it reduces false-positive SL predictions compared to previous methods, and it allows establishing explicit collateral lethality relationship of gene pairs within mutually exclusive group pairs. Pietro Pinoli, Sriganesh Srihari, Limsoon Wong, Stefano Ceri |
BMC Bioinform. | 1 |
| 2020 | Empowering Virus Sequence Research Through Conceptual Modeling
Anna Bernasconi 0002, Arif Canakoglu, Pietro Pinoli, Stefano Ceri |
ER | 3 |
| 2020 | Matrix Factorization-based Technique for Drug Repurposing PredictionsabstractClassical drug design methodologies are hugely costly and time-consuming, with approximately 85% of the new proposed molecules failing in the first three phases of the FDA drug approval process. Thus, strategies to find alternative indications for already approved drugs that leverage computational methods are of crucial relevance. We previously demonstrated the efficacy of the Non-negative Matrix Tri-Factorization, a method that allows exploiting both data integration and machine learning, to infer novel indications for approved drugs. In this work, we present an innovative enhancement of the NMTF method that consists of a shortest-path evaluation of drug-protein pairs using the protein-to-protein interaction network. This approach allows inferring novel protein targets that were never considered as drug targets before, increasing the information fed to the NMTF method. Indeed, this novel advance enables the investigation of drug-centric predictions, simultaneously identifying therapeutic classes, protein targets and diseases associated with a particular drug. To test our methodology, we applied the NMTF and shortest-path enhancement methods to an outdated collection of data and compared the predictions against the most updated version, obtaining very good performance, with an Average Precision Score of 0.82. The data enhancement strategy allowed increasing the number of putative protein targets from 3,691 to 15,295, while the predictive performance of the method is slightly increased. Finally, we also validated our top-scored predictions according to the literature, finding relevant confirmation of predicted interactions between drugs and protein targets, as well as of predicted annotations between drugs and both therapeutic classes and diseases. Gaia Ceddia, Pietro Pinoli, Stefano Ceri, Marco Masseroli |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Deleterious Impact of Mutational Processes on Transcription Factor Binding Sites in Human CancerabstractSomatic mutations occurring in many cancer types are associated with well-understood processes, such as exposure to tobacco smoking or to ultraviolet (UV) light, but also with mutational processes of so far unknown etiology. Mutational processes can be described in terms of so-called mutational signatures, most often represented as vectors of mutation probabilities which indicate what mutation types are preferentially induced by the mutational processes. In this paper we propose a framework to identify which mutational processes are more likely to harm binding sites of a given transcription factor. Our method starts from the binding site motif and assigns to each mutational signature both a hit score, i.e., the likelihood that the mutational process mutates a binding sequence in at least one nucleotide, and a measure of deleteriousness, i.e., the likelihood that a binding site can be disrupted by mutations belonging to the signature. In a final step, the determined scores can be adjusted according to the strengths with which individual mutational signatures have contributed to the observed mutational load of a tumor. We apply the method to CTCF, a transcription factor that is a core architectural protein dictating the dimensional structure of the genome. Our analysis concentrates on melanoma (skin cancer), for which we show that our framework predicts the disruption of CTCF binding sites by specific UV-light associated mutational signatures, confirming our biological expectations. Pietro Pinoli, Eirini Stamoulakatou, Stefano Ceri, Rosario M. Piro |
BIBE | 1 |
| 2019 | Analysis and Visualization of Mutation Enrichments for Selected Genomic Regions and Cancer TypesabstractSeveral studies highlight the relevance of somatic mutations in non-coding regions of the genome which exhibit common interesting behaviors. MutViz is a tool for the identification of mutation enrichments on arbitrary sets of user-defined regions; for a variety of cancer types, it contains preloaded mutations from public datasets, well organized within an effective database organization. MutViz provides a user-friendly interface helping the user in providing sets of regions as input and in obtaining their fast exploration as output, together with simple statistical testing of novel hypotheses. Andrea Gulino, Eirini Stamoulakatou, Arif Canakoglu, Pietro Pinoli |
BIBM | 4 |
| 2019 | Non-negative Matrix Tri-Factorization for Data Integration and Network-based Drug RepositioningabstractDrug discovery is a high cost and high risk process, thus finding new uses for approved drugs, i.e. drug repositioning, via computational methods has become increasingly interesting. In this study, we present a new network-based approach for predicting potential new indications for existing drugs through their connections with other biological entities. For this aim, we first built a large network integrating drugs, proteins, biological pathways and drugs' categories as nodes of the network, and connections between such nodes as links of the network. Our method leverages the Non-Negative Matrix Tri-Factorization reconstruction of adjacency matrices in order to predict novel category-drug links, i.e. a new category (or use)associated with a drug, taking the entire network information into account. We tested our method on a set of 1,120 drugs labeled with ten categories; when we hide to the method the 10% of the drug-category associations, it was able to infer those missing values with a recall of 60% and a precision of 70%. Precision and recall remain higher than a Random Classifier in case of larger percentage of hidden links, demonstrating the robustness of the method. Also, we were able to predict novel drug-label associations not yet reported in the repository. Finally, we favorably compared our method with a state of the art method for drug repositioning; the NMTF method achieved an average precision score of 0.68 vs. the 0.55 score of the state of the art method. Gaia Ceddia, Pietro Pinoli, Stefano Ceri, Marco Masseroli |
CIBCB | 2 |
| 2019 | Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing dataabstractMOTIVATION: We previously proposed a paradigm shift in genomic data management, based on the Genomic Data Model (GDM) for mediating existing data formats and on the GenoMetric Query Language (GMQL) for supporting, at a high level of abstraction, data extraction and the most common data-driven computations required by tertiary data analysis of Next Generation Sequencing datasets. Here, we present a new GMQL-based system with enhanced accessibility, portability, scalability and performance. RESULTS: The new system has a well-designed modular architecture featuring: (i) an intermediate representation supporting many different implementations (including Spark, Flink and SciDB); (ii) a high-level technology-independent repository abstraction, supporting different repository technologies (e.g., local file system, Hadoop File System, database or others); (iii) several system interfaces, including a user-friendly Web-based interface, a Web Service interface, and a programmatic interface for Python language. Biological use case examples, using public ENCODE, Roadmap Epigenomics and TCGA datasets, demonstrate the relevance of our work. AVAILABILITY AND IMPLEMENTATION: The GMQL system is freely available for non-commercial use as open source project at: http://www.bioinformatics.deib.polimi.it/GMQLsystem/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marco Masseroli, Arif Canakoglu, Pietro Pinoli, Abdulrahman Kaitoua, Andrea Gulino, Olha Horlova, Luca Nanni, Anna Bernasconi 0002, Stefano Perna, Eirini Stamoulakatou, Stefano Ceri |
Bioinform. | 3 |
| 2019 | PyGMQL: scalable data extraction and analysis for heterogeneous genomic datasetsabstractBACKGROUND: With the growth of available sequenced datasets, analysis of heterogeneous processed data can answer increasingly relevant biological and clinical questions. Scientists are challenged in performing efficient and reproducible data extraction and analysis pipelines over heterogeneously processed datasets. Available software packages are suitable for analyzing experimental files from such datasets one by one, but do not scale to thousands of experiments. Moreover, they lack proper support for metadata manipulation. RESULTS: We present PyGMQL, a novel software for the manipulation of region-based genomic files and their relative metadata, built on top of the GMQL genomic big data management system. PyGMQL provides a set of expressive functions for the manipulation of region data and their metadata that can scale to arbitrary clusters and implicitly apply to thousands of files, producing millions of regions. PyGMQL provides data interoperability, distribution transparency and query outsourcing. The PyGMQL package integrates scalable data extraction over the Apache Spark engine underlying the GMQL implementation with native Python support for interactive data analysis and visualization. It supports data interoperability, solving the impedance mismatch between executing set-oriented queries and programming in Python. PyGMQL provides distribution transparency (the ability to address a remote dataset) and query outsourcing (the ability to assign processing to a remote service) in an orthogonal way. Outsourced processing can address cloud-based installations of the GMQL engine. CONCLUSIONS: PyGMQL is an effective and innovative tool for supporting tertiary data extraction and analysis pipelines. We demonstrate the expressiveness and performance of PyGMQL through a sequence of biological data analysis scenarios of increasing complexity, which highlight reproducibility, expressive power and scalability. Luca Nanni, Pietro Pinoli, Arif Canakoglu, Stefano Ceri |
BMC Bioinform. | 2 |
| 2019 | Metadata management for scientific databasesabstractMost scientific databases consist of datasets (or sources) which in turn include samples (or files) with an identical structure (or schema). In many cases, samples are associated with rich metadata, describing the process that leads to building them (e.g.: the experimental conditions used during sample generation). Metadata are typically used in scientific computations just for the initial data selection; at most, metadata about query results is recovered after executing the query, and associated with its results by post-processing. In this way, a large body of information that could be relevant for interpreting query results goes unused during query processing. In this paper, we present ScQL, a new algebraic relational language, whose operations apply to objects consisting of data–metadatapairs, by preserving such one-to-one correspondence throughout the computation. We formally define each operation and we describe an optimization, called meta-first , that may significantly reduce the query processing overhead by anticipating the use of metadata for selectively loading into the execution environment only those input samples that contribute to the result samples. In ScQL, metadata have the same relevance as data, and contribute to building query results; in this way, the resulting samples are systematically associated with metadata about either the specific input samples involved or about query processing, thereby yielding a new form of metadata provenance . We present many examples of use of ScQL, relative to several application domains, and we demonstrate the effectiveness of the meta-first optimization. • ScQL: a metadata aware query language for databases of scientific samples. • ScQL is compiled to a lower level representation suitable for optimization. • The novel meta-first optimization exploits metadata so as to speed up query evaluation. Pietro Pinoli, Stefano Ceri, Davide Martinenghi, Luca Nanni |
Inf. Syst. | 1 |
| 2018 | DLA: a Distributed, Location-based and Apriori-based Algorithm for Biological Sequence Pattern MiningabstractWith the rapid growth of genomic data, the need for scalable data mining algorithms has increased. Frequent contiguous sequence mining is a technique that can help biologists to better understand the function and structure of our DNA, by capturing the common characteristics among related sequences. Many sequence mining algorithms have been developed over time. However, most of them suffer from scaling issues when dealing with big data or give no warranty for the completeness of their result. In this paper, we propose a distributed sequential pattern mining algorithm implemented on Apache Spark. Specifically, the algorithm exploits the Apriori Property and information about each patterns location within the original sequence, to drastically reduce the number of candidates at each iteration. Experimental results on real-world datasets confirm our performance expectations, showing a better scalability when compared to other distributed solutions. Eirini Stamoulakatou, Andrea Gulino, Pietro Pinoli |
IEEE BigData | 3 |
| 2018 | Demonstration of GenoMetric Query LanguageabstractIn the last ten years, genomic computing has made gigantic steps due to Next Generation Sequencing (NGS), a high-throughput, massively parallel technology; the cost of producing a complete human sequence dropped to 1000 US$ in 2015 and is expected to drop below 100 US$ by 2020. Several new methods have recently become available for extracting heterogeneous datasets from the genome, revealing data signals such as variations from a reference sequence, levels of expression of coding regions, or protein binding enrichments ('peaks') with their statistical or geometric properties. Huge collections of such datasets are made available by large international consortia. Stefano Ceri, Arif Canakoglu, Andrea Gulino, Abdulrahman Kaitoua, Marco Masseroli, Luca Nanni, Pietro Pinoli |
CIKM | 7 |
| 2017 | Evaluating Genomic Big Data Operations on SciDB and Spark
Simone Cattani, Stefano Ceri, Abdulrahman Kaitoua, Pietro Pinoli |
ICWE | 4 |
| 2017 | Framework for Supporting Genomic OperationsabstractNext Generation Sequencing (NGS) is a family of technologies for reading the DNA or RNA, capable of producing whole genome sequences at an impressive speed, and causing a revolution of both biological research and medical practice. In this exciting scenario, while a huge number of specialized bio-informatics programs extract information from sequences, there is an increasing need for a new generation of systems and frameworks capable of integrating such information, providing holistic answers to the needs of biologists and clinicians. To respond to this need, we developed GMQL, a new query language for genomic data management that operates on heterogeneous genomic datasets. In this paper, we focus on three domain-specific operations of GMQL used for the efficient processing of operations on genomic regions, and we describe their efficient implementation; the paper develops a theory of binning strategies as a generic approach to parallel execution of genomic operations, and then describes how binning is embedded into two efficient implementations of the operations using Flink and Spark, two emerging frameworks for data management on the cloud. Abdulrahman Kaitoua, Pietro Pinoli, Michele Bertoni, Stefano Ceri |
IEEE Trans. Computers | 2 |
| 2017 | Data Management for Heterogeneous Genomic DatasetsabstractNext Generation Sequencing (NGS), a family of technologies for reading DNA and RNA, is changing biological research, and will soon change medical practice, by quickly providing sequencing data and high-level features of numerous individual genomes in different biological and clinical conditions. The availability of millions of whole genome sequences may soon become the biggest and most important "big data" problem of mankind. In this exciting framework, we recently proposed a new paradigm to raise the level of abstraction in NGS data management, by introducing a GenoMetric Query Language (GMQL) and demonstrating its usefulness through several biological query examples. Leveraging on that effort, here we motivate and formalize GMQL operations, especially focusing on the most characteristic and domain-specific ones. Furthermore, we address their efficient implementation and illustrate the architecture of the new software system that we have developed for their execution on big genomic data in a cloud computing environment, providing the evaluation of its performance. The new system implementation is available for download at the GMQL website (http://www.bioinformatics.deib.polimi.it/GMQL/); GMQL can also be tested through a set of predefined queries on ENCODE and Roadmap Epigenomics data at http://www.bioinformatics.deib.polimi.it/GMQL/queries/. Stefano Ceri, Abdulrahman Kaitoua, Marco Masseroli, Pietro Pinoli, Francesco Venco |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | Data Management for Next Generation Genomic ComputingabstractNext-generation sequencing (NGS) has dramatically reduced the cost and time of reading the DNA. Huge investments are targeted to sequencing the DNA of large populations, and repositories of well-curated sequence data are being collected. Answers to fundamental biomedical problems are hidden in these data, e.g. how cancer arises, how driving mutations occur, how much cancer is dependent on environment. So far, the bio-informatics research community has been mostly challenged by primary analysis (production of sequences in the form of short DNA segments, or ''reads'') and secondary analysis (alignment of reads to a reference genome and search for specific features on the reads); yet, the most important emerging problem is the so-called tertiary analysis, concerned with multi-sample processing of heterogeneous information. Tertiary analysis is responsible of sense making, e.g., discovering how heterogeneous regions interact with each other. \nThis new scenario creates an opportunity for rethinking genomic computing through the lens of fundamental data management. We propose an essential data model, using few general abstractions that guarantee interoperability between existing data formats, and a new-generation query language inspired by classic relational algebra and extended with orthogonal, domain-specific abstractions for genomics. They open doors to the seamless integration of descriptive statistics and high-level data analysis (e.g., DNA region clustering and extraction of regulatory networks). In this vision, computational efficiency is achieved by using parallel computing on both clusters and public clouds; the technology is applicable to federated repositories, and can be exploited for providing integrated access to curated data, made available by large consortia, through user-friendly search services. Our most far-fetching vision is to move towards an Internet of Genomes exploiting data indexing and crawling. Stefano Ceri, Abdulrahman Kaitoua, Marco Masseroli, Pietro Pinoli, Francesco Venco |
EDBT | 4 |
| 2015 | Evaluating cloud frameworks on genomic applicationsabstractWe are developing a new, holistic data management system for genomics, which uses cloud-based computing for querying thousands of heterogeneous genomic datasets. In our project, it is essential to leverage upon a modern cloud computing framework, so as to encode our query expressions into high-level operations provided by the framework. After releasing our first implementation using Pig and Hadoop 1, we are currently targeting Spark and Flink, two emerging frameworks for general-purpose big data analytics. While Spark appears to have a stronger critical mass, Flink supports high-level optimization for data management operations; both systems appear suited to support our domain-specific data management operations. In this paper, we focus on a comparison of the two frameworks at work based upon three typical genomic applications, stemming from our data management requirements and needs; we describe the coding of the genomic applications using Flink and Spark, discuss their common aspects and differences, and comparatively evaluate the performance and scalability of the implementations over datasets consisting of billions of genomic regions. Michele Bertoni, Stefano Ceri, Abdulrahman Kaitoua, Pietro Pinoli |
IEEE BigData | 4 |
| 2015 | GenoMetric Query Language: a novel approach to large-scale genomic data managementabstractMOTIVATION: Improvement of sequencing technologies and data processing pipelines is rapidly providing sequencing data, with associated high-level features, of many individual genomes in multiple biological and clinical conditions. They allow for data-driven genomic, transcriptomic and epigenomic characterizations, but require state-of-the-art 'big data' computing strategies, with abstraction levels beyond available tool capabilities. RESULTS: We propose a high-level, declarative GenoMetric Query Language (GMQL) and a toolkit for its use. GMQL operates downstream of raw data preprocessing pipelines and supports queries over thousands of heterogeneous datasets and samples; as such it is key to genomic 'big data' analysis. GMQL leverages a simple data model that provides both abstractions of genomic region data and associated experimental, biological and clinical metadata and interoperability between many data formats. Based on Hadoop framework and Apache Pig platform, GMQL ensures high scalability, expressivity, flexibility and simplicity of use, as demonstrated by several biological query examples on ENCODE and TCGA datasets. AVAILABILITY AND IMPLEMENTATION: The GMQL toolkit is freely available for non-commercial use at http://www.bioinformatics.deib.polimi.it/GMQL/. Marco Masseroli, Pietro Pinoli, Francesco Venco, Abdulrahman Kaitoua, Vahid Jalili, Fernando Palluzzi, Heiko Müller 0003, Stefano Ceri |
Bioinform. | 2 |
| 2015 | Computational algorithms to predict Gene Ontology annotationsabstractBACKGROUND: Gene function annotations, which are associations between a gene and a term of a controlled vocabulary describing gene functional features, are of paramount importance in modern biology. Datasets of these annotations, such as the ones provided by the Gene Ontology Consortium, are used to design novel biological experiments and interpret their results. Despite their importance, these sources of information have some known issues. They are incomplete, since biological knowledge is far from being definitive and it rapidly evolves, and some erroneous annotations may be present. Since the curation process of novel annotations is a costly procedure, both in economical and time terms, computational tools that can reliably predict likely annotations, and thus quicken the discovery of new gene annotations, are very useful. METHODS: We used a set of computational algorithms and weighting schemes to infer novel gene annotations from a set of known ones. We used the latent semantic analysis approach, implementing two popular algorithms (Latent Semantic Indexing and Probabilistic Latent Semantic Analysis) and propose a novel method, the Semantic IMproved Latent Semantic Analysis, which adds a clustering step on the set of considered genes. Furthermore, we propose the improvement of these algorithms by weighting the annotations in the input set. RESULTS: We tested our methods and their weighted variants on the Gene Ontology annotation sets of three model organism genes (Bos taurus, Danio rerio and Drosophila melanogaster ). The methods showed their ability in predicting novel gene annotations and the weighting procedures demonstrated to lead to a valuable improvement, although the obtained results vary according to the dimension of the input annotation set and the considered algorithm. CONCLUSIONS: Out of the three considered methods, the Semantic IMproved Latent Semantic Analysis is the one that provides better results. In particular, when coupled with a proper weighting policy, it is able to predict a significant number of novel annotations, demonstrating to actually be a helpful tool in supporting scientists in the curation process of gene functional annotations. Pietro Pinoli, Davide Chicco, Marco Masseroli |
BMC Bioinform. | 1 |
| 2014 | Latent Dirichlet Allocation based on Gibbs Sampling for gene function predictionabstractGene function annotations are key elements in biology and bioinformatics. A typical annotation is the association between a gene and a feature term that describes a functional feature of the gene by using a controlled vocabulary term (e.g. a Gene Ontology (GO) feature term). Unfortunately, available annotations contain errors and biologically validated ones are incomplete by definition, since new knowledge is continuously discovered. Thus, computational algorithms which are able to provide ranked lists of predicted new gene annotations are an excellent contribution to the bioinformatics research. Here, we propose two variants of the known Latent Dirichlet Allocation (LDA) algorithm applied to the prediction of gene annotations. LDA is a very efficient machine learning method built on a set of multinomial probability distributions over a set of topics, given a document (a gene, in our case), and on a set of multinomial probability distributions over a set of words (feature terms, in our case), given a topic. In topic modeling, a topic can be considered as a latent meta-category of words, and a document as a mixture of topics. Our two LDA variants use the collapsed Gibbs Sampling method during the training phase, with two distinct initialization approaches to adapt the LDA mathematical model to the biomolecular annotation scenario. Using six outdated datasets of GO annotations of human and brown rat genes, we compared the annotations predicted by our methods to the ones given by the truncated Singular Value Decomposition (tSVD) method previously developed; then, we validated them by using the annotations available in an updated version of the same datasets. Obtained results show the efficiency of our new proposed algorithms. Pietro Pinoli, Davide Chicco, Marco Masseroli |
CIBCB | 1 |
| 2013 | Enhanced probabilistic latent semantic analysis with weighting schemes to predict genomic annotationsabstractGenomic annotations with functional controlled terms, such as the Gene Ontology (GO) ones, are paramount in modern biology. Yet, they are known to be incomplete, since the current biological knowledge is far to be definitive. In this scenario, computational methods that are able to support and quicken the curation of these annotations can be very useful. In a previous work, we discussed the benefits of using the Probabilistic Latent Semantic Analysis algorithm in order to predict novel GO annotations, compared to some Singular Value Decomposition (SVD) based approaches. In this paper, we propose a further enhancement of that method, which aims at weighting the available associations between genes and functional terms before using them as input to the predictive system. The tests that we performed on the annotations of human genes to GO functional terms showed the efficacy of our approach. Pietro Pinoli, Davide Chicco, Marco Masseroli |
BIBE | 1 |
| 2012 | Probabilistic Latent Semantic Analysis for prediction of Gene Ontology annotationsabstractConsistency and completeness of biomolecular annotations is a keypoint of correct interpretation of biological experiments. Yet, the associations between genes (or proteins) and features correctly annotated are just some of all the existing ones. As time goes by, they increase in number and become more useful, but they remain incomplete and some of them incorrect. To support and quicken their time-consuming curation procedure and to improve consistence of available annotations, computational methods that are able to supply a ranked list of predicted annotations are hence extremely useful. Starting from a previous work on the automatic prediction of Gene Ontology (GO) annotations based on the Singular Value Decomposition of the annotation matrix, where every matrix element corresponds to the association of a gene with a feature, we propose the use of a modified Probabilistic Latent Semantic Analysis (pLSA) algorithm, named pLSAnorm, to better perform such prediction. pLSA is a statistical technique from the natural language processing field, which has not been used in bioinformatics annotation prediction yet; it takes advantage of the latent information contained in the analyzed data co-occurrences. We proved the effectiveness of the pLSAnorm prediction method by performing k-fold cross-validation of the GO annotations of two organisms, Gallus gallus and Bos taurus. Obtained results demonstrate the efficacy of our approach. Marco Masseroli, Davide Chicco, Pietro Pinoli |
IJCNN | 3 |