EDBT 2026 Demo / reviewers in the wild / expert
Andrea Gulino
dblp:205/9955
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0003-0201-9461ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2Systems, architecture and hardware · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Distributed and cloud data management · 50% Graph data management · 25% Query processing and optimization · 25% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 92% Computational finance and economics · 8% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics
genomic data management |
0.8 | 2 | 2019 | Optimal Binning for Genomics · IEEE Trans. Computers 2019 Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data · Bioinform. 2019 |
Distributed and cloud data management
distributed query processing |
0.5 | 1 | 2021 | Distributed Company Control in Company Shareholding Graphs · ICDE 2021 |
Graph data management
graph query processing |
0.5 | 1 | 2021 | Distributed Company Control in Company Shareholding Graphs · ICDE 2021 |
Query processing and optimization
parallel query processing |
0.5 | 1 | 2021 | Distributed Company Control in Company Shareholding Graphs · ICDE 2021 |
Bioinformatics and computational biology
genomics |
0.4 | 1 | 2019 | Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data · Bioinform. 2019 |
Bioinformatics and computational biology › metagenomics
metagenomic binning |
0.4 | 1 | 2019 | Optimal Binning for Genomics · IEEE Trans. Computers 2019 |
Distributed and cloud data management
data partitioning |
0.4 | 1 | 2019 | Optimal Binning for Genomics · IEEE Trans. Computers 2019 |
Bioinformatics and computational biology › genomics
next-generation sequencing data analysis |
0.1 | 1 | 2019 | Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data · Bioinform. 2019 |
Cloud and datacenter computing › cloud data management
cloud data processing |
0.1 | 1 | 2019 | Optimal Binning for Genomics · IEEE Trans. Computers 2019 |
Methods — techniques the papers use, named apart from their topics
query optimization · 1.1mathematical modeling · 1.1query partitioning · 1.0parallel execution · 1.0spark · 0.8genometric query language · 0.8flink · 0.8SciDB · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Distributed Company Control in Company Shareholding GraphsabstractThe Company Control Problem is of central importance to banks, financial intermediaries, financial intelligence units, regulatory and supervisory authorities such as the Central Banks. It consists in understanding who takes decisions in a large company network, that is, who controls the majority of votes for each single company. This has an impact on a large number of business areas, with examples including evaluation of creditworthiness, economic analysis of the control dispersion, anti-money laundering, prevention of potentially hostile takeovers, evaluation of risks, and shock propagation.This paper is based on our experience with the Central Bank of Italy and presents an approach to the solution of the company control problem in distributed settings, especially relevant, as large and distributed ownership graphs reflect European-size applications where scalability is paramount.In particular, we formalize the problem as query answering on a large distributed database. We study how independent subqueries can be executed in each partition and the partial results assembled at a master site to produce the answer. We study the formal properties of the problem, that is not easily parallelizable, and then present a method that supports parallelism at best.We present a thorough experimental evaluation of our approach with the Italian company graph of the Bank of Italy and the European Register of Financial Intermediaries and Affiliates as well as many artificial graphs to fully assess scalability. Andrea Gulino, Stefano Ceri, Georg Gottlob, Emanuel Sallinger, Luigi Bellomarini |
ICDE | 1 |
| 2021 | Federated sharing and processing of genomic datasets for tertiary data analysisabstractMOTIVATION: With the spreading of biological and clinical uses of next-generation sequencing (NGS) data, many laboratories and health organizations are facing the need of sharing NGS data resources and easily accessing and processing comprehensively shared genomic data; in most cases, primary and secondary data management of NGS data is done at sequencing stations, and sharing applies to processed data. Based on the previous single-instance GMQL system architecture, here we review the model, language and architectural extensions that make the GMQL centralized system innovatively open to federated computing. RESULTS: A well-designed extension of a centralized system architecture to support federated data sharing and query processing. Data is federated thanks to simple data sharing instructions. Queries are assigned to execution nodes; they are translated into an intermediate representation, whose computation drives data and processing distributions. The approach allows writing federated applications according to classical styles: centralized, distributed or externalized. AVAILABILITY: The federated genomic data management system is freely available for non-commercial use as an open source project at http://www.bioinformatics.deib.polimi.it/FederatedGMQLsystem/. CONTACT: {arif.canakoglu, pietro.pinoli}@polimi.it. Arif Canakoglu, Pietro Pinoli, Andrea Gulino, Luca Nanni, Marco Masseroli, Stefano Ceri |
Briefings Bioinform. | 3 |
| 2020 | Performance Prediction for Data-driven Workflows on Apache SparkabstractSpark is an in-memory framework for implementing distributed applications of various types. Predicting the execution time of Spark applications is an important but challenging problem that has been tackled in the past few years by several studies; most of them achieving good prediction accuracy on simple applications (e.g. known ML algorithms or SQL-based applications). In this work, we consider complex data-driven workflow applications, in which the execution and data flow can be modeled by Directly Acyclic Graphs (DAGs). Workflows can be made of an arbitrary combination of known tasks, each applying a set of Spark operations to their input data. By adopting a hybrid approach, combining analytical and machine learning (ML) models, trained on small DAGs, we can predict, with good accuracy, the execution time of unseen workflows of higher complexity and size. We validate our approach through an extensive experimentation on real-world complex applications, comparing different ML models and choices of feature sets. Andrea Gulino, Arif Canakoglu, Stefano Ceri, Danilo Ardagna |
MASCOTS | 1 |
| 2019 | Analysis and Visualization of Mutation Enrichments for Selected Genomic Regions and Cancer TypesabstractSeveral studies highlight the relevance of somatic mutations in non-coding regions of the genome which exhibit common interesting behaviors. MutViz is a tool for the identification of mutation enrichments on arbitrary sets of user-defined regions; for a variety of cancer types, it contains preloaded mutations from public datasets, well organized within an effective database organization. MutViz provides a user-friendly interface helping the user in providing sets of regions as input and in obtaining their fast exploration as output, together with simple statistical testing of novel hypotheses. Andrea Gulino, Eirini Stamoulakatou, Arif Canakoglu, Pietro Pinoli |
BIBM | 1 |
| 2019 | Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing dataabstractMOTIVATION: We previously proposed a paradigm shift in genomic data management, based on the Genomic Data Model (GDM) for mediating existing data formats and on the GenoMetric Query Language (GMQL) for supporting, at a high level of abstraction, data extraction and the most common data-driven computations required by tertiary data analysis of Next Generation Sequencing datasets. Here, we present a new GMQL-based system with enhanced accessibility, portability, scalability and performance. RESULTS: The new system has a well-designed modular architecture featuring: (i) an intermediate representation supporting many different implementations (including Spark, Flink and SciDB); (ii) a high-level technology-independent repository abstraction, supporting different repository technologies (e.g., local file system, Hadoop File System, database or others); (iii) several system interfaces, including a user-friendly Web-based interface, a Web Service interface, and a programmatic interface for Python language. Biological use case examples, using public ENCODE, Roadmap Epigenomics and TCGA datasets, demonstrate the relevance of our work. AVAILABILITY AND IMPLEMENTATION: The GMQL system is freely available for non-commercial use as open source project at: http://www.bioinformatics.deib.polimi.it/GMQLsystem/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marco Masseroli, Arif Canakoglu, Pietro Pinoli, Abdulrahman Kaitoua, Andrea Gulino, Olha Horlova, Luca Nanni, Anna Bernasconi 0002, Stefano Perna, Eirini Stamoulakatou, Stefano Ceri |
Bioinform. | 5 |
| 2019 | Optimal Binning for GenomicsabstractGenome sequencing is expected to be the most prolific source of big data in the next decade; millions of whole genome datasets will open new opportunities for biological research and personalized medicine. Genome sequences are abstracted in the form of interesting regions, describing abnormalities of the genome. The parallel execution on the cloud of complex operations for joining and mapping billions of genomic regions is increasingly important. Genome binning, i.e., partitioning of the genome into small-size segments, adapts classic data partitioning methods to genomics; region distributions to bins must reflect operation-specific correctness rules. As a consequence, determining the optimal bin size for such operations is a complex mathematical problem, whose solution requires careful modeling. The main result of this paper is the mathematical formulation and solution of the optimal binning problem for join and map operations in the context of GMQL, a query language over genomic regions; the model is validated by experiments showing its accuracy and sensitivity to the variation of operations' parameters. We also optimize sequences of operations by inheriting the binning between two consecutive operations and we show the deployment of GMQL and the tuning of the proposed model on different cloud computing systems. Andrea Gulino, Abdulrahman Kaitoua, Stefano Ceri |
IEEE Trans. Computers | 1 |
| 2018 | DLA: a Distributed, Location-based and Apriori-based Algorithm for Biological Sequence Pattern MiningabstractWith the rapid growth of genomic data, the need for scalable data mining algorithms has increased. Frequent contiguous sequence mining is a technique that can help biologists to better understand the function and structure of our DNA, by capturing the common characteristics among related sequences. Many sequence mining algorithms have been developed over time. However, most of them suffer from scaling issues when dealing with big data or give no warranty for the completeness of their result. In this paper, we propose a distributed sequential pattern mining algorithm implemented on Apache Spark. Specifically, the algorithm exploits the Apriori Property and information about each patterns location within the original sequence, to drastically reduce the number of candidates at each iteration. Experimental results on real-world datasets confirm our performance expectations, showing a better scalability when compared to other distributed solutions. Eirini Stamoulakatou, Andrea Gulino, Pietro Pinoli |
IEEE BigData | 2 |
| 2018 | Demonstration of GenoMetric Query LanguageabstractIn the last ten years, genomic computing has made gigantic steps due to Next Generation Sequencing (NGS), a high-throughput, massively parallel technology; the cost of producing a complete human sequence dropped to 1000 US$ in 2015 and is expected to drop below 100 US$ by 2020. Several new methods have recently become available for extracting heterogeneous datasets from the genome, revealing data signals such as variations from a reference sequence, levels of expression of coding regions, or protein binding enrichments ('peaks') with their statistical or geometric properties. Huge collections of such datasets are made available by large international consortia. Stefano Ceri, Arif Canakoglu, Andrea Gulino, Abdulrahman Kaitoua, Marco Masseroli, Luca Nanni, Pietro Pinoli |
CIKM | 3 |