EDBT 2026 Demo / reviewers in the wild / expert
Simona E. Rombo
dblp:r/SimonaERombo · also Simona Ester Rombo
· DBLP profile ↗
38ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-3833-835XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-authorTheory of computation · 6 · 3 first-authorArtificial intelligence and machine learning · 5Human-computer interaction and ubiquitous computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BioSet2Vec: extraction of k-mer dictionaries from multiple sets of biological sequences via big data technologiesabstractBACKGROUND: In several contexts involving large collections of sets of biological sequences, a relevant problem is that of selecting significant groups of k-mers that characterize one set with regards to the others in the same collection. RESULTS: Here a software framework is proposed implementing a novel methodology for the extraction of k-mer dictionaries, from multiple sets of biological sequences. It has been implemented according to the most recent technologies for Big Data analytics, with the perspective of allowing its usage with a variety of input datasets of any size. In particular, two different packages are provided. The first is BioFt, enabling the extraction of recurrent patterns based on k-mers frequency and the computation of other metrics from information retrieval, here specialized for biological sequences. The second package BioSet2Vec, instead, extends the functionality of BioFt by allowing the creation of dictionaries according to different criteria. CONCLUSIONS: The framework has been validated on three different case studies: (1) the characterization of different chromatin states; (2) the study of association between different diseases and related genes; (3) the analysis of genomes of different organisms. All tests performed on the considered datasets have shown the potentialities of the proposed approach. Ylenia Galluzzo, Raffaele Giancarlo, Simona E. Rombo, Filippo Utro |
BMC Bioinform. | 3 |
| 2024 | Neighborhood based computational approaches for the prediction of lncRNA-disease associationsabstractMOTIVATION: Long non-coding RNAs (lncRNAs) are a class of molecules involved in important biological processes. Extensive efforts have been provided to get deeper understanding of disease mechanisms at the lncRNA level, guiding towards the detection of biomarkers for disease diagnosis, treatment, prognosis and prevention. Unfortunately, due to costs and time complexity, the number of possible disease-related lncRNAs verified by traditional biological experiments is very limited. Computational approaches for the prediction of disease-lncRNA associations allow to identify the most promising candidates to be verified in laboratory, reducing costs and time consuming. RESULTS: We propose novel approaches for the prediction of lncRNA-disease associations, all sharing the idea of exploring associations among lncRNAs, other intermediate molecules (e.g., miRNAs) and diseases, suitably represented by tripartite graphs. Indeed, while only a few lncRNA-disease associations are still known, plenty of interactions between lncRNAs and other molecules, as well as associations of the latters with diseases, are available. A first approach presented here, NGH, relies on neighborhood analysis performed on a tripartite graph, built upon lncRNAs, miRNAs and diseases. A second approach (CF) relies on collaborative filtering; a third approach (NGH-CF) is obtained boosting NGH by collaborative filtering. The proposed approaches have been validated on both synthetic and real data, and compared against other methods from the literature. It results that neighborhood analysis allows to outperform competitors, and when it is combined with collaborative filtering the prediction accuracy further improves, scoring a value of AUC equal to 0966. AVAILABILITY: Source code and sample datasets are available at: https://github.com/marybonomo/LDAsPredictionApproaches.git. Mariella Bonomo, Simona E. Rombo |
BMC Bioinform. | 2 |
| 2023 | Discriminative pattern discovery for the characterization of different network populationsabstractMOTIVATION: An interesting problem is to study how gene co-expression varies in two different populations, associated with healthy and unhealthy individuals, respectively. To this aim, two important aspects should be taken into account: (i) in some cases, pairs/groups of genes show collaborative attitudes, emerging in the study of disorders and diseases; (ii) information coming from each single individual may be crucial to capture specific details, at the basis of complex cellular mechanisms; therefore, it is important avoiding to miss potentially powerful information, associated with the single samples. RESULTS: Here, a novel approach is proposed, such that two different input populations are considered, and represented by two datasets of edge-labeled graphs. Each graph is associated to an individual, and the edge label is the co-expression value between the two genes associated to the nodes. Discriminative patterns among graphs belonging to different sample sets are searched for, based on a statistical notion of 'relevance' able to take into account important local similarities, and also collaborative effects, involving the co-expression among multiple genes. Four different gene expression datasets have been analyzed by the proposed approach, each associated to a different disease. An extensive set of experiments show that the extracted patterns significantly characterize important differences between healthy and unhealthy samples, both in the cooperation and in the biological functionality of the involved genes/proteins. Moreover, the provided analysis confirms some results already presented in the literature on genes with a central role for the considered diseases, still allowing to identify novel and useful insights on this aspect. AVAILABILITY AND IMPLEMENTATION: The algorithm has been implemented using the Java programming language. The data underlying this article and the code are available at https://github.com/CriSe92/DiscriminativeSubgraphDiscovery. Fabio Fassetti, Simona E. Rombo, Cristina Serrao |
Bioinform. | 2 |
| 2022 | Topological ranks reveal functional knowledge encoded in biological networks: a comparative analysisabstractMOTIVATION: Biological networks topology yields important insights into biological function, occurrence of diseases and drug design. In the last few years, different types of topological measures have been introduced and applied to infer the biological relevance of network components/interactions, according to their position within the network structure. Although comparisons of such measures have been previously proposed, to what extent the topology per se may lead to the extraction of novel biological knowledge has never been critically examined nor formalized in the literature. RESULTS: We present a comparative analysis of nine outstanding topological measures, based on compact views obtained from the rank they induce on a given input biological network. The goal is to understand their ability in correctly positioning nodes/edges in the rank, according to the functional knowledge implicitly encoded in biological networks. To this aim, both internal and external (gold standard) validation criteria are taken into account, and six networks involving three different organisms (yeast, worm and human) are included in the comparison. The results show that a distinct handful of best-performing measures can be identified for each of the considered organisms, independently from the reference gold standard. AVAILABILITY: Input files and code for the computation of the considered topological measures and K-haus distance are available at https://gitlab.com/MaryBonomo/ranking. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online. Mariella Bonomo, Raffaele Giancarlo, Daniele Greco, Simona E. Rombo |
Briefings Bioinform. | 4 |
| 2022 | DIAMIN: a software library for the distributed analysis of large-scale molecular interaction networksabstractBACKGROUND: Huge amounts of molecular interaction data are continuously produced and stored in public databases. Although many bioinformatics tools have been proposed in the literature for their analysis, based on their modeling through different types of biological networks, several problems still remain unsolved when the problem turns on a large scale. RESULTS: We propose DIAMIN, that is, a high-level software library to facilitate the development of applications for the efficient analysis of large-scale molecular interaction networks. DIAMIN relies on distributed computing, and it is implemented in Java upon the framework Apache Spark. It delivers a set of functionalities implementing different tasks on an abstract representation of very large graphs, providing a built-in support for methods and algorithms commonly used to analyze these networks. DIAMIN has been tested on data retrieved from two of the most used molecular interactions databases, resulting to be highly efficient and scalable. As shown by different provided examples, DIAMIN can be exploited by users without any distributed programming experience, in order to perform various types of data analysis, and to implement new algorithms based on its primitives. CONCLUSIONS: The proposed DIAMIN has been proved to be successful in allowing users to solve specific biological problems that can be modeled relying on biological networks, by using its functionalities. The software is freely available and this will hopefully allow its rapid diffusion through the scientific community, to solve both specific data analysis and more complex tasks. Lorenzo Di Rocco, Umberto Ferraro Petrillo, Simona E. Rombo |
BMC Bioinform. | 3 |
| 2021 | Integrative bioinformatics and omics data source interoperability in the next-generation sequencing era - EditorialabstractWith the advent of high-throughput and next-generation sequencing (NGS) technologies [1], huge amounts of ‘omics’ data (i.e. data from genomics, proteomics, pharmacogenomics, metagenomics, etc.) are continuously produced. Combining and integrating diverse omics data types is important in order to investigate the molecular machinery of complex diseases, with the hope for better disease prevention and treatment [2]. Experimental data repositories of omics data are publicly available, with the main aim of fostering the cooperation among research groups and laboratories all over the world. However, despite their openness, the effective integrated use of available public sources is hampered by the heterogeneity, complexity and large size of data stored therein. The main issues to be addressed when approaching omics data integration are related to the difficulty in managing and analyzing these data. Indeed, specific and multidisciplinary competences are required, and combining data of different types is not a simple task. Both the extensional (i.e. the real data) and intensional (i.e. the corresponding metadata) levels may be involved in this integration process, according to the specific problem under consideration. In the last few years, information systems researchers have made significant efforts in the proposal of effective methodologies for the integration of structured and semi-structured data formats [3]. However, omics data are often unstructured. This pushes toward the study of how data source integration can be successfully performed when structured, semi-structured and unstructured data sources coexist. This themed issue provides an extensive overview of the main challenges related to omics data integration and the methods that have been recently proposed in order to address them. It comprises nine manuscripts, each dealing with one of four central key issues, as detailed below. Understanding how the direct or indirect relationships among cellular components may impact the occurrence and progress of disorders and diseases is an important issue, which requires omics data integration to be addressed. Manuscripts of this group start from the assumption that, confirmed by several studies in the literature, the occurrence and progress of many diseases have genetic causes. For example, genetic variations have direct effects on individual phenotypes, possibly causing the production of partially or totally dysfunctional proteins. With this regards, Galano-Frutos, García-Cebollada and Sancho in Molecular Dynamics Simulations for Genetic Interpretation in Protein Coding Regions: Where we Are, Where to Go and When observe that predicting whether the replacement of one amino acid residue with another will be tolerated or cause disease is a key factor. In particular, first they review existing prediction tools based on evolutionary information and simple physical–chemical properties. Then, they describe more recent and accurate methods, such as full-atom molecular dynamics simulation in explicit solvent and discuss how these methods can be used in order to interpret human genetic variations at a large scale. Another important aspect is related to data coming from ‘single cell analysis’, which may be used in order to understand differences in healthy/unhealthy populations. In Computational methods for the integrative analysis of single cell data, Forcato, Romano and Bicciato describe the computational methods for the integrative analysis of single-cell genomic data. They mainly focus on the integration of single-cell RNA sequencing datasets and on the joint analysis of multimodal signals from individual cells. Omics data are represented in a wide variety of notations and formats, often with different levels of quality. This intrinsic heterogeneity in both data and repositories makes difficult their effective combination for producing new knowledge and may hamper their correct use and exploitation. ‘Genomic data integration’ is the topic of The road towards data integration in human genomics: players, steps and interactions by Bernasconi, Canakoglu, Masseroli and Ceri. In this manuscript, the authors first describe a technological pipeline from data production to data integration. Then, they propose a taxonomy of genomic data players and apply it to about 30 important players. They specifically focus on integrator players and evaluate the computational environment for data integration purposes provided by them. The role of ‘conceptual models’ to support the efficient management of genomic data is discussed in Using Conceptual Modeling to Improve Genome Data Management by Pastor, León Palacio, Reyes Román, García S. and Casamayor. The authors describe a solution that helps researchers to organize, store and process information and, at the same time, focuses only on relevant data minimizing the information overload in clinical research context. An overview of available ‘patient-level datasets’ containing both genotypic and phenotypic data is presented in GenoPheno: cataloging large-scale phenotypic and next generation sequencing data within human datasets by Gutiérrez-Sacristán, De Niz, Kothari,Won Kong, Mandl and Avillach. In this manuscript, the authors describe a dynamic, online catalog for consultation, contribution and revision by the research community. It consists of 30 datasets and was created by them with the purpose of making it publicly available. A survey on ‘machine learning’ methods operating on the cloud for gene regulation studies is presented in Machine learning-based analysis of multi-omics data on the cloud for investigating gene regulations by Oh, Park, Kim and Chae. The authors describe these methods, categorize them according to five different goals and summarize them in terms of multiomics input types. They explain the positive role that the cloud can play for the analysis of multiomics data. They also discuss some important issues to address when machine learning-based approaches operating on the cloud are adopted for the analysis of gene regulations. Structured sparsity regularization for analyzing high-dimensional omics data by Vinga focuses on ‘structured regularizers’ and ‘penalty functions’, when applied to omics data. The author analyzes their potential in identifying disease’s molecular signature, in order to create high-performance clinical decision support systems and, ultimately, favor personalized healthcare. Microbial communities and viral populations have a crucial role in the environment and in human health. In Comparison of Microbiome Samples: Methods and Computational Challenges, Comin, Di Camillo, Pizzi and Vandin provide a study on ‘metagenomic’ NGS datasets. These authors compare datasets from three different viewpoints, namely: (i) species identification and quantification; (ii) efficient computation of distances between metagenomic sample datasets; (iii) identification of metagenomics features associated with a phenotype. In Epidemiological Data Analysis of Viral Quasispecies in the Next-Generation Sequencing Era, Knyazev, Hughes, Skums and Zelikovsky deal with the analysis of intrahost RNA viral populations. In particular, they examine bioinformatics tools that: (i) characterize the complexity of intrahost viral population; (ii) support epidemiological analysis in inferring drug-resistant mutations, infection age and patient linkage; (iii) support surveillance systems for fast response and outbreak control. Hopefully, this themed issue will represent a springboard for fruitful collaborations among researchers from multidisciplinary areas, which could give a significant boost to the advancement of knowledge in different fields through a more effective analysis of omics data. The Editors are grateful to both the Editor-in-Chief and the Publisher for having trusted this project and for having supported them in all their needs. Many thank also to all the authors and reviewers, whose expertise and effort allowed the realization of this themed issue. Simona E. Rombo is Associate Professor in Computer Science at the Department of Mathematics and Computer Science of University of Palermo. Her main research interests include Bioinformatics, algorithms and methodologies for network analysis, Big Data analytics. She is the Principal Investigator of several national and international research projects in these fields, and she is cofounder of a spin-off working on decision support for Precision Medicine. SER has been visiting scientist at different research institutes, among which the Department of Computer Science at Purdue University and the College of Computing at Georgia Institute of Technology. Domenico Ursino received the MSc Degree in Computer Engineering from the University of Calabria in July 1995. He received the PhD in System Engineering and Computer Science from the University of Calabria in January 2000. From January 2005 to December 2017 he was an Associate Professor at the University Mediterranea of Reggio Calabria. From January 2018 he is a Full Professor at the Polytechnic University of Marche. His research interests include Social Network Analysis, Social Internetworking, Source and Data Integration, Innovation Management, Multiple Internet of Things scenarios, Knowledge Extraction and Representation, Biomedical Applications, Recommender Systems, Data Lakes. In these research fields, he published more than 200 papers. Pora Kim is an assistant professor in the School of Biomedical Informatics, The University of Texas Health Science Center at Houston. Her research interest includes bioinformatics and cancer genomics. Simona E. Rombo, Domenico Ursino |
Briefings Bioinform. | 1 |
| 2019 | Customer recommendation based on profile matching and customized campaigns in on-line social networksabstractWe propose a general framework for the recommendation of possible customers (users) to advertisers (e.g., brands) based on the comparison between On-Line Social Network profiles. In particular, we associate suitable categories and subcategories to both user and brand profiles in the considered On-line Social Network. When categories involve posts and comments, the comparison is based on word embedding, and this allows to take into account the similarity between the topics of particular interest for a brand and the user preferences. Furthermore, user personal information, such as age, job or genre, are used for targeting specific advertising campaigns. Results on real Facebook dataset show that the proposed approach is successful in identifying the most suitable set of users to be used as target for a given advertisement campaign. Mariella Bonomo, Gaspare Ciaccio, Andrea De Salve, Simona E. Rombo |
ASONAM | 4 |
| 2019 | FEDRO: a software tool for the automatic discovery of candidate ORFs in plants with c →u RNA editingabstractBACKGROUND: RNA editing is an important mechanism for gene expression in plants organelles. It alters the direct transfer of genetic information from DNA to proteins, due to the introduction of differences between RNAs and the corresponding coding DNA sequences. Software tools successful for the search of genes in other organisms not always are able to correctly perform this task in plants organellar genomes. Moreover, the available software tools predicting RNA editing events utilise algorithms that do not account for events which may generate a novel start codon. RESULTS: We present FEDRO, a Java software tool implementing a novel strategy to generate candidate Open Reading Frames (ORFs) resulting from Cytidine to Uridine (c→u) editing substitutions which occur in the mitochondrial genome (mtDNA) of a given input plant. The goal is to predict putative proteins of plants mitochondria that have not been yet annotated. In order to validate the generated ORFs, a screening is performed by checking for sequence similarity or presence in active transcripts of the same or similar organisms. We illustrate the functionalities of our framework on a model organism. CONCLUSIONS: The proposed tool may be used also on other organisms and genomes. FEDRO is publicly available at http://math.unipa.it/rombo/FEDRO . Fabio Fassetti, Claudia Giallombardo, Ofelia Leone, Luigi Palopoli 0001, Simona E. Rombo, Adolfo Saiardi |
BMC Bioinform. | 5 |
| 2019 | Analyzing big datasets of genomic sequences: fast and scalable collection of k-mer statisticsabstractBACKGROUND: Distributed approaches based on the MapReduce programming paradigm have started to be proposed in the Bioinformatics domain, due to the large amount of data produced by the next-generation sequencing techniques. However, the use of MapReduce and related Big Data technologies and frameworks (e.g., Apache Hadoop and Spark) does not necessarily produce satisfactory results, in terms of both efficiency and effectiveness. We discuss how the development of distributed and Big Data management technologies has affected the analysis of large datasets of biological sequences. Moreover, we show how the choice of different parameter configurations and the careful engineering of the software with respect to the specific framework under consideration may be crucial in order to achieve good performance, especially on very large amounts of data. We choose k-mers counting as a case study for our analysis, and Spark as the framework to implement FastKmer, a novel approach for the extraction of k-mer statistics from large collection of biological sequences, with arbitrary values of k. RESULTS: One of the most relevant contributions of FastKmer is the introduction of a module for balancing the statistics aggregation workload over the nodes of a computing cluster, in order to overcome data skew while allowing for a full exploitation of the underlying distributed architecture. We also present the results of a comparative experimental analysis showing that our approach is currently the fastest among the ones based on Big Data technologies, while exhibiting a very good scalability. CONCLUSIONS: We provide evidence that the usage of technologies such as Hadoop or Spark for the analysis of big datasets of biological sequences is productive only if the architectural details and the peculiar aspects of the considered framework are carefully taken into account for the algorithm design and implementation. Umberto Ferraro Petrillo, Mara Sorella, Giuseppe Cattaneo, Raffaele Giancarlo, Simona E. Rombo |
BMC Bioinform. | 5 |
| 2019 | DNA combinatorial messages and Epigenomics: The case of chromatin organization and nucleosome occupancy in eukaryotic genomes
Raffaele Giancarlo, Simona E. Rombo, Filippo Utro |
Theor. Comput. Sci. | 2 |
| 2018 | An Integrative Framework for the Construction of Big Functional Networks
Claudia Giallombardo, Salvatore Morfea, Simona E. Rombo |
BIBM | 3 |
| 2018 | In vitro versus in vivo compositional landscapes of histone sequence preferences in eucaryotic genomesabstractMotivation: Although the nucleosome occupancy along a genome can be in part predicted by in vitro experiments, it has been recently observed that the chromatin organization presents important differences in vitro with respect to in vivo. Such differences mainly regard the hierarchical and regular structures of the nucleosome fiber, whose existence has long been assumed, and in part also observed in vitro, but that does not apparently occur in vivo. It is also well known that the DNA sequence has a role in determining the nucleosome occupancy. Therefore, an important issue is to understand if, and to what extent, the structural differences in the chromatin organization between in vitro and in vivo have a counterpart in terms of the underlying genomic sequences. Results: We present the first quantitative comparison between the in vitro and in vivo nucleosome maps of two model organisms (S. cerevisiae and C. elegans). The comparison is based on the construction of weighted k-mer dictionaries. Our findings show that there is a good level of sequence conservation between in vitro and in vivo in both the two organisms, in contrast to the abovementioned important differences in chromatin structural organization. Moreover, our results provide evidence that the two organisms predispose themselves differently, in terms of sequence composition and both in vitro and in vivo, for the nucleosome occupancy. This leads to the conclusion that, although the notion of a genome encoding for its own nucleosome occupancy is general, the intrinsic histone k-mer sequence preferences tend to be species-specific. Availability and implementation: The files containing the dictionaries and the main results of the analysis are available at http://math.unipa.it/rombo/material. Supplementary information: Supplementary data are available at Bioinformatics online. Raffaele Giancarlo, Simona E. Rombo, Filippo Utro |
Bioinform. | 2 |
| 2018 | Efficient Algorithms for Sequence Analysis with Entropic ProfilesabstractEntropy, being closely related to repetitiveness and compressibility, is a widely used information-related measure to assess the degree of predictability of a sequence. Entropic profiles are based on information theory principles, and can be used to study the under-/over-representation of subwords, by also providing information about the scale of conserved DNA regions. Here, we focus on the algorithmic aspects related to entropic profiles. In particular, we propose linear time algorithms for their computation that rely on suffix-based data structures, more specifically on the truncated suffix tree (TST) and on the enhanced suffix array (ESA). We performed an extensive experimental campaign showing that our algorithms, beside being faster, make it possible the analysis of longer sequences, even for high degrees of resolution, than state of the art algorithms. Cinzia Pizzi, Mattia Ornamenti, Simone Spangaro, Simona E. Rombo, Laxmi Parida |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2017 | 2D Motif Basis Applied to the Classification of Digital ImagesabstractThe classification of raw data often involves the problem of selecting the appropriate set of features to represent the input data. Different types of features can be extracted from the input dataset, but only some of them are actually relevant for the classification process. Since relevant features are often unknown in real-world problems, many candidate features are usually introduced. This degrades both the speed and the predictive accuracy of the classifier due to the presence of redundancy in the set of candidate features. Recently, a special class of bidimensional motifs, i.e. 2D motif basis has been introduced in the literature. 2D motif basis showed to be powerful in capturing the relevant information of digital images, also achieving good performances for image compression. Here, we investigate the effectiveness of 2D motif basis, when they are used as features for image classification. We embed such features in a bag-of-words model, and then we apply K-Nearest Neighbour for the classification step. Results obtained on both benchmark image datasets and video frames datasets show that, despite the pixel-level nature of the considered features, the achieved accuracy is high and comparable with that of other techniques proposed in the literature. Angelo Furfaro, Maria Carmela Groccia, Simona E. Rombo |
Comput. J. | 3 |
| 2017 | Foreword: Algorithms, Strings and Theoretical Approaches in the Big Data Era - Special Issue in Honor of the 60th Birthday of Professor Raffaele Giancarlo
Simona E. Rombo, Filippo Utro |
Theor. Comput. Sci. | 1 |
| 2015 | Searching for repetitions in biological networks: methods, resources and toolsabstractWe present here a compact overview of the data, models and methods proposed for the analysis of biological networks based on the search for significant repetitions. In particular, we concentrate on three problems widely studied in the literature: 'network alignment', 'network querying' and 'network motif extraction'. We provide (i) details of the experimental techniques used to obtain the main types of interaction data, (ii) descriptions of the models and approaches introduced to solve such problems and (iii) pointers to both the available databases and software tools. The intent is to lay out a useful roadmap for identifying suitable strategies to analyse cellular data, possibly based on the joint use of different interaction data types or analysis techniques. Simona Panni, Simona E. Rombo |
Briefings Bioinform. | 2 |
| 2015 | Epigenomic k-mer dictionaries: shedding light on how sequence composition influences in vivo nucleosome positioningabstractMOTIVATION: Information-theoretic and compositional analysis of biological sequences, in terms of k-mer dictionaries, has a well established role in genomic and proteomic studies. Much less so in epigenomics, although the role of k-mers in chromatin organization and nucleosome positioning is particularly relevant. Fundamental questions concerning the informational content and compositional structure of nucleosome favouring and disfavoring sequences with respect to their basic building blocks still remain open. RESULTS: We present the first analysis on the role of k-mers in the composition of nucleosome enriched and depleted genomic regions (NER and NDR for short) that is: (i) exhaustive and within the bounds dictated by the information-theoretic content of the sample sets we use and (ii) informative for comparative epigenomics. We analize four different organisms and we propose a paradigmatic formalization of k-mer dictionaries, providing two different and complementary views of the k-mers involved in NER and NDR. The first extends well known studies in this area, its comparative nature being its major merit. The second, very novel, brings to light the rich variety of k-mers involved in influencing nucleosome positioning, for which an initial classification in terms of clusters is also provided. Although such a classification offers many insights, the following deserves to be singled-out: short poly(dA:dT) tracts are reported in the literature as fundamental for nucleosome depletion, however a global quantitative look reveals that their role is much less prominent than one would expect based on previous studies. AVAILABILITY AND IMPLEMENTATION: Dictionaries, clusters and Supplementary Material are available online at http://math.unipa.it/rombo/epigenomics/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Raffaele Giancarlo, Simona E. Rombo, Filippo Utro |
Bioinform. | 2 |
| 2014 | Entropic Profiles, Maximal Motifs and the Discovery of Significant Repetitions in Genomic Sequences
Laxmi Parida, Cinzia Pizzi, Simona E. Rombo |
WABI | 3 |
| 2014 | Compressive biological sequence analysis and archival in the era of high-throughput sequencing technologiesabstractHigh-throughput sequencing technologies produce large collections of data, mainly DNA sequences with additional information, requiring the design of efficient and effective methodologies for both their compression and storage. In this context, we first provide a classification of the main techniques that have been proposed, according to three specific research directions that have emerged from the literature and, for each, we provide an overview of the current techniques. Finally, to make this review useful to researchers and technicians applying the existing software and tools, we include a synopsis of the main characteristics of the described approaches, including details on their implementation and availability. Performance of the various methods is also highlighted, although the state of the art does not lend itself to a consistent and coherent comparison among all the methods presented here. Raffaele Giancarlo, Simona E. Rombo, Filippo Utro |
Briefings Bioinform. | 2 |
| 2014 | Algorithms and tools for protein-protein interaction networks clustering, with a special focus on population-based stochastic methodsabstractMOTIVATION: Protein-protein interaction (PPI) networks are powerful models to represent the pairwise protein interactions of the organisms. Clustering PPI networks can be useful for isolating groups of interacting proteins that participate in the same biological processes or that perform together specific biological functions. Evolutionary orthologies can be inferred this way, as well as functions and properties of yet uncharacterized proteins. RESULTS: We present an overview of the main state-of-the-art clustering methods that have been applied to PPI networks over the past decade. We distinguish five specific categories of approaches, describe and compare their main features and then focus on one of them, i.e. population-based stochastic search. We provide an experimental evaluation, based on some validation measures widely used in the literature, of techniques in this class, that are as yet less explored than the others. In particular, we study how the capability of Genetic Algorithms (GAs) to extract clusters in PPI networks varies when different topology-based fitness functions are used, and we compare GAs with the main techniques in the other categories. The experimental campaign shows that predictions returned by GAs are often more accurate than those produced by the contestant methods. Interesting issues still remain open about possible generalizations of GAs allowing for cluster overlapping. AVAILABILITY AND IMPLEMENTATION: We point out which methods and tools described here are publicly available. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Clara Pizzuti, Simona E. Rombo |
Bioinform. | 2 |
| 2014 | An evolutionary restricted neighborhood search clustering approach for PPI networks
Clara Pizzuti, Simona E. Rombo |
Neurocomputing | 2 |
| 2014 | Irredundant tandem motifs
Laxmi Parida, Cinzia Pizzi, Simona E. Rombo |
Theor. Comput. Sci. | 3 |
| 2013 | Image Classification Based on 2D Feature Motifs
Angelo Furfaro, Maria Carmela Groccia, Simona E. Rombo |
FQAS | 3 |
| 2012 | Experimental evaluation of topological-based fitness functions to detect complexes in PPI networksabstractThe detection of groups of proteins sharing common biological features is an important research issue, intensively investigated in the last few years, because of the insights it can give in understanding cell behavior. In this paper we present an extensive experimental evaluation campaign aiming at exploring the capability of Genetic Algorithms (GAs) to find clusters in protein-protein interaction networks, when different topological-based fitness functions are employed. A complete experimentation on the yeast protein-protein interaction network, along with a comparative evaluation of the effectiveness in detecting true complexes on the yeast and human networks, reveals GAs as a feasible and competitive computational technique to cope with this problem. Clara Pizzuti, Simona E. Rombo |
GECCO | 2 |
| 2012 | Characterization and Extraction of Irredundant Tandem Motifs
Laxmi Parida, Cinzia Pizzi, Simona E. Rombo |
SPIRE | 3 |
| 2012 | A Coclustering Approach for Mining Large Protein-Protein Interaction NetworksabstractSeveral approaches have been presented in the literature to cluster Protein-Protein Interaction (PPI) networks. They can be grouped in two main categories: those allowing a protein to participate in different clusters and those generating only nonoverlapping clusters. In both cases, a challenging task is to find a suitable compromise between the biological relevance of the results and a comprehensive coverage of the analyzed networks. Indeed, methods returning high accurate results are often able to cover only small parts of the input PPI network, especially when low-characterized networks are considered. We present a coclustering-based technique able to generate both overlapping and nonoverlapping clusters. The density of the clusters to search for can also be set by the user. We tested our method on the two networks of yeast and human, and compared it to other five well-known techniques on the same interaction data sets. The results showed that, for all the examples considered, our approach always reaches a good compromise between accuracy and network coverage. Furthermore, the behavior of our algorithm is not influenced by the structure of the input network, different from all the techniques considered in the comparison, which returned very good results on the yeast network, while on the human network their outcomes are rather poor. Clara Pizzuti, Simona E. Rombo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Extracting string motif bases for quorum higher than two
Simona E. Rombo |
Theor. Comput. Sci. | 1 |
| 2011 | Image Compression by 2D Motif BasisabstractApproaches to image compression and indexing based on extensions to 2D of some of the Lempel-Ziv incremental parsing techniques have been proposed in the recent past. In these approaches, an image is decomposed into a number of patches, consisting each of a square or rectangular solid block. This paper proposes image compression techniques based on patches that are not necessarily solid blocks, but are affected instead by a controlled number of undetermined or don't care pixels. Such patches are chosen from a set of candidate motifs that are extracted in turn from the image 2D motif basis, the latter consisting of a compact set of patterns that result from the autocorrelation of the image with itself. As is expected, it is found that limited indeterminacy can be traded for higher compression at the expense of negligible loss. Preliminary experiments show that this technique yields higher compression than other popular techniques such as GZIP, BZIP and JPEG. Alessia Amelio, Alberto Apostolico, Simona E. Rombo |
DCC | 3 |
| 2011 | Asymmetric Comparison and Querying of Biological NetworksabstractComparing and querying the protein-protein interaction (PPI) networks of different organisms is important to infer knowledge about conservation across species. Known methods that perform these tasks operate symmetrically, i.e., they do not assign a distinct role to the input PPI networks. However, in most cases, the input networks are indeed distinguishable on the basis of how the corresponding organism is biologically well characterized. In this paper a new idea is developed, that is, to exploit differences in the characterization of organisms at hand in order to devise methods for comparing their PPI networks. We use the PPI network (called Master) of the best characterized organism as a fingerprint to guide the alignment process to the second input network (called Slave), so that generated results preferably retain the structural characteristics of the Master network. Technically, this is obtained by generating from the Master a finite automaton, called alignment model, which is then fed with (a linearization of) the Slave for the purpose of extracting, via the Viterbi algorithm, matching subgraphs. We propose an approach able to perform global alignment and network querying, and we apply it on PPI networks. We tested our method showing that the results it returns are biologically relevant. Nicola Ferraro, Luigi Palopoli 0001, Simona Panni, Simona E. Rombo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2010 | "Master-Slave" Biological Network Alignment
Nicola Ferraro, Luigi Palopoli 0001, Simona Panni, Simona E. Rombo |
ISBRA | 4 |
| 2009 | Optimal extraction of motif patterns in 2D
Simona E. Rombo |
Inf. Process. Lett. | 1 |
| 2008 | Motif patterns in 2D
Alberto Apostolico, Laxmi Parida, Simona E. Rombo |
Theor. Comput. Sci. | 3 |
| 2007 | GRAPPIN: Bipartite GRAph Based Protein-Protein Interaction Network Similarity SearchabstractWe propose an algorithm, called BI-GRAPPIN, to search for similarities across PPI networks. The technique core consists in computing a maximum weight matching of bipartite graphs to compare the neighborhoods of pairs of proteins in different PPI networks. The idea is that proteins belonging to different networks should be matched look- ing not only at their own sequence similarity, but also at the similarity of proteins they "strongly" interact with, ei- ther directly or indirectly. We implemented the method and tested it on both real and synthetic data, showing its effec- tiveness in solving ambiguous situations and in individuat- ing functionally related proteins. Differently from previous work, the presented algorithm allows to take into account both quantitative and reliability information possibly avail- able about interactions. Valeria Fionda, Luigi Palopoli 0001, Simona Panni, Simona E. Rombo |
BIBM | 4 |
| 2007 | Protein Data Condensation for Effective Quaternary Structure Classification
Fabrizio Angiulli, Valeria Fionda, Simona E. Rombo |
IDEAL | 3 |
| 2007 | PINCoC : A Co-clustering Based Approach to Analyze Protein-Protein Interaction Networks
Clara Pizzuti, Simona E. Rombo |
IDEAL | 2 |
| 2006 | JSSPrediction: a Framework to Predict Protein Secondary Structures Using IntegrationabstractIdentifying protein secondary structures is a difficult task. Recently, a lot of software tools for protein secondary structure prediction have been produced and made available on-line, mostly with good performances. However, prediction tools work correctly for families of proteins, such that users have to know which predictor to use for a given unknown protein. We propose a framework to improve secondary structure prediction by integrating results obtained from a set of available predictors. Our contribution consists in the definition of a two phase approach: (i) select a set of predictors which have good performances with the unknown protein family, and (U) integrate the prediction results of the selected prediction tools. Experimental results are also reported Luigi Palopoli 0001, Simona E. Rombo, Giorgio Terracina, Giuseppe Tradigo, Pierangelo Veltri |
CBMS | 2 |
| 2005 | Flexible Pattern Discovery with (Extended) Disjunctive Logic Programming
Luigi Palopoli 0001, Simona E. Rombo, Giorgio Terracina |
ISMIS | 2 |
| 2004 | Discovering Representative Models in Large Time Series Databases
Simona E. Rombo, Giorgio Terracina |
FQAS | 1 |