Andrea Gulino

dblp:205/9955 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0003-0201-9461ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Distributed and cloud data management · 50% Graph data management · 25% Query processing and optimization · 25%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 92% Computational finance and economics · 8%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics
genomic data management
0.822019
Optimal Binning for Genomics · IEEE Trans. Computers 2019
Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data · Bioinform. 2019
Distributed and cloud data management
distributed query processing
0.512021
Distributed Company Control in Company Shareholding Graphs · ICDE 2021
Graph data management
graph query processing
0.512021
Distributed Company Control in Company Shareholding Graphs · ICDE 2021
Query processing and optimization
parallel query processing
0.512021
Distributed Company Control in Company Shareholding Graphs · ICDE 2021
Bioinformatics and computational biology
genomics
0.412019
Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data · Bioinform. 2019
Bioinformatics and computational biology › metagenomics
metagenomic binning
0.412019
Optimal Binning for Genomics · IEEE Trans. Computers 2019
Distributed and cloud data management
data partitioning
0.412019
Optimal Binning for Genomics · IEEE Trans. Computers 2019
Bioinformatics and computational biology › genomics
next-generation sequencing data analysis
0.112019
Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data · Bioinform. 2019
Cloud and datacenter computing › cloud data management
cloud data processing
0.112019
Optimal Binning for Genomics · IEEE Trans. Computers 2019

Methods — techniques the papers use, named apart from their topics

query optimization · 1.1mathematical modeling · 1.1query partitioning · 1.0parallel execution · 1.0spark · 0.8genometric query language · 0.8flink · 0.8SciDB · 0.8
YearPublicationVenuePosition
2021 Distributed Company Control in Company Shareholding Graphs
abstract
The Company Control Problem is of central importance to banks, financial intermediaries, financial intelligence units, regulatory and supervisory authorities such as the Central Banks. It consists in understanding who takes decisions in a large company network, that is, who controls the majority of votes for each single company. This has an impact on a large number of business areas, with examples including evaluation of creditworthiness, economic analysis of the control dispersion, anti-money laundering, prevention of potentially hostile takeovers, evaluation of risks, and shock propagation.This paper is based on our experience with the Central Bank of Italy and presents an approach to the solution of the company control problem in distributed settings, especially relevant, as large and distributed ownership graphs reflect European-size applications where scalability is paramount.In particular, we formalize the problem as query answering on a large distributed database. We study how independent subqueries can be executed in each partition and the partial results assembled at a master site to produce the answer. We study the formal properties of the problem, that is not easily parallelizable, and then present a method that supports parallelism at best.We present a thorough experimental evaluation of our approach with the Italian company graph of the Bank of Italy and the European Register of Financial Intermediaries and Affiliates as well as many artificial graphs to fully assess scalability.
Andrea Gulino, Stefano Ceri, Georg Gottlob, Emanuel Sallinger, Luigi Bellomarini
ICDE1
2021 Federated sharing and processing of genomic datasets for tertiary data analysis
abstract
MOTIVATION: With the spreading of biological and clinical uses of next-generation sequencing (NGS) data, many laboratories and health organizations are facing the need of sharing NGS data resources and easily accessing and processing comprehensively shared genomic data; in most cases, primary and secondary data management of NGS data is done at sequencing stations, and sharing applies to processed data. Based on the previous single-instance GMQL system architecture, here we review the model, language and architectural extensions that make the GMQL centralized system innovatively open to federated computing. RESULTS: A well-designed extension of a centralized system architecture to support federated data sharing and query processing. Data is federated thanks to simple data sharing instructions. Queries are assigned to execution nodes; they are translated into an intermediate representation, whose computation drives data and processing distributions. The approach allows writing federated applications according to classical styles: centralized, distributed or externalized. AVAILABILITY: The federated genomic data management system is freely available for non-commercial use as an open source project at http://www.bioinformatics.deib.polimi.it/FederatedGMQLsystem/. CONTACT: {arif.canakoglu, pietro.pinoli}@polimi.it.
Arif Canakoglu, Pietro Pinoli, Andrea Gulino, Luca Nanni, Marco Masseroli, Stefano Ceri
Briefings Bioinform.3
2020 Performance Prediction for Data-driven Workflows on Apache Spark
abstract
Spark is an in-memory framework for implementing distributed applications of various types. Predicting the execution time of Spark applications is an important but challenging problem that has been tackled in the past few years by several studies; most of them achieving good prediction accuracy on simple applications (e.g. known ML algorithms or SQL-based applications). In this work, we consider complex data-driven workflow applications, in which the execution and data flow can be modeled by Directly Acyclic Graphs (DAGs). Workflows can be made of an arbitrary combination of known tasks, each applying a set of Spark operations to their input data. By adopting a hybrid approach, combining analytical and machine learning (ML) models, trained on small DAGs, we can predict, with good accuracy, the execution time of unseen workflows of higher complexity and size. We validate our approach through an extensive experimentation on real-world complex applications, comparing different ML models and choices of feature sets.
Andrea Gulino, Arif Canakoglu, Stefano Ceri, Danilo Ardagna
MASCOTS1
2019 Analysis and Visualization of Mutation Enrichments for Selected Genomic Regions and Cancer Types
abstract
Several studies highlight the relevance of somatic mutations in non-coding regions of the genome which exhibit common interesting behaviors. MutViz is a tool for the identification of mutation enrichments on arbitrary sets of user-defined regions; for a variety of cancer types, it contains preloaded mutations from public datasets, well organized within an effective database organization. MutViz provides a user-friendly interface helping the user in providing sets of regions as input and in obtaining their fast exploration as output, together with simple statistical testing of novel hypotheses.
Andrea Gulino, Eirini Stamoulakatou, Arif Canakoglu, Pietro Pinoli
BIBM1
2019 Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data
abstract
MOTIVATION: We previously proposed a paradigm shift in genomic data management, based on the Genomic Data Model (GDM) for mediating existing data formats and on the GenoMetric Query Language (GMQL) for supporting, at a high level of abstraction, data extraction and the most common data-driven computations required by tertiary data analysis of Next Generation Sequencing datasets. Here, we present a new GMQL-based system with enhanced accessibility, portability, scalability and performance. RESULTS: The new system has a well-designed modular architecture featuring: (i) an intermediate representation supporting many different implementations (including Spark, Flink and SciDB); (ii) a high-level technology-independent repository abstraction, supporting different repository technologies (e.g., local file system, Hadoop File System, database or others); (iii) several system interfaces, including a user-friendly Web-based interface, a Web Service interface, and a programmatic interface for Python language. Biological use case examples, using public ENCODE, Roadmap Epigenomics and TCGA datasets, demonstrate the relevance of our work. AVAILABILITY AND IMPLEMENTATION: The GMQL system is freely available for non-commercial use as open source project at: http://www.bioinformatics.deib.polimi.it/GMQLsystem/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marco Masseroli, Arif Canakoglu, Pietro Pinoli, Abdulrahman Kaitoua, Andrea Gulino, Olha Horlova, Luca Nanni, Anna Bernasconi 0002, Stefano Perna, Eirini Stamoulakatou, Stefano Ceri
Bioinform.5
2019 Optimal Binning for Genomics
abstract
Genome sequencing is expected to be the most prolific source of big data in the next decade; millions of whole genome datasets will open new opportunities for biological research and personalized medicine. Genome sequences are abstracted in the form of interesting regions, describing abnormalities of the genome. The parallel execution on the cloud of complex operations for joining and mapping billions of genomic regions is increasingly important. Genome binning, i.e., partitioning of the genome into small-size segments, adapts classic data partitioning methods to genomics; region distributions to bins must reflect operation-specific correctness rules. As a consequence, determining the optimal bin size for such operations is a complex mathematical problem, whose solution requires careful modeling. The main result of this paper is the mathematical formulation and solution of the optimal binning problem for join and map operations in the context of GMQL, a query language over genomic regions; the model is validated by experiments showing its accuracy and sensitivity to the variation of operations' parameters. We also optimize sequences of operations by inheriting the binning between two consecutive operations and we show the deployment of GMQL and the tuning of the proposed model on different cloud computing systems.
Andrea Gulino, Abdulrahman Kaitoua, Stefano Ceri
IEEE Trans. Computers1
2018 DLA: a Distributed, Location-based and Apriori-based Algorithm for Biological Sequence Pattern Mining
abstract
With the rapid growth of genomic data, the need for scalable data mining algorithms has increased. Frequent contiguous sequence mining is a technique that can help biologists to better understand the function and structure of our DNA, by capturing the common characteristics among related sequences. Many sequence mining algorithms have been developed over time. However, most of them suffer from scaling issues when dealing with big data or give no warranty for the completeness of their result. In this paper, we propose a distributed sequential pattern mining algorithm implemented on Apache Spark. Specifically, the algorithm exploits the Apriori Property and information about each patterns location within the original sequence, to drastically reduce the number of candidates at each iteration. Experimental results on real-world datasets confirm our performance expectations, showing a better scalability when compared to other distributed solutions.
Eirini Stamoulakatou, Andrea Gulino, Pietro Pinoli
IEEE BigData2
2018 Demonstration of GenoMetric Query Language
abstract
In the last ten years, genomic computing has made gigantic steps due to Next Generation Sequencing (NGS), a high-throughput, massively parallel technology; the cost of producing a complete human sequence dropped to 1000 US$ in 2015 and is expected to drop below 100 US$ by 2020. Several new methods have recently become available for extracting heterogeneous datasets from the genome, revealing data signals such as variations from a reference sequence, levels of expression of coding regions, or protein binding enrichments ('peaks') with their statistical or geometric properties. Huge collections of such datasets are made available by large international consortia.
Stefano Ceri, Arif Canakoglu, Andrea Gulino, Abdulrahman Kaitoua, Marco Masseroli, Luca Nanni, Pietro Pinoli
CIKM3