VLDB 2026 Research / reviewers in the wild / expert
Michele Ceccarelli
dblp:34/5571
· DBLP profile ↗
45ranked-venue papers
20as first author
5since 2021 · last 2023
0000-0002-4702-6617ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 15 · 12 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorSystems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Survey of Steganography Tools at Layers 2-4 and HTTPabstractSteganography has evolved into various forms and remains an effective way to hide sensitive information. Network Steganography, also known as "Covert Channels," is popular in fields such as terrorism and security. As a result, the scientific community created a specific taxonomy to categorize it, and developed several techniques and tools to conceal communication between the parties. In this paper, we have curated a list of available software tools that can be used to create a covert channel at layers 2 to 4 of the ISO/OSI model and related to the HTTP protocol. Stefano Bistarelli, Michele Ceccarelli, Chiara Luchini, Ivan Mercanti, Francesco Santini 0001 |
ARES | 2 |
| 2023 | MOViDA: multiomics visible drug activity prediction with a biologically informed neural network modelabstractMOTIVATION: The process of drug development is inherently complex, marked by extended intervals from the inception of a pharmaceutical agent to its eventual launch in the market. Additionally, each phase in this process is associated with a significant failure rate, amplifying the inherent challenges of this task. Computational virtual screening powered by machine learning algorithms has emerged as a promising approach for predicting therapeutic efficacy. However, the complex relationships between the features learned by these algorithms can be challenging to decipher. RESULTS: We have engineered an artificial neural network model designed specifically for predicting drug sensitivity. This model utilizes a biologically informed visible neural network, thereby enhancing its interpretability. The trained model allows for an in-depth exploration of the biological pathways integral to prediction and the chemical attributes of drugs that impact sensitivity. Our model harnesses multiomics data derived from a different tumor tissue sources, as well as molecular descriptors that encapsulate the properties of drugs. We extended the model to predict drug synergy, resulting in favorable outcomes while retaining interpretability. Given the imbalanced nature of publicly available drug screening datasets, our model demonstrated superior performance to state-of-the-art visible machine learning algorithms. AVAILABILITY AND IMPLEMENTATION: MOViDA is implemented in Python using PyTorch library and freely available for download at https://github.com/Luigi-Ferraro/MOViDA. Training data, RIS score and drug features are archived on Zenodo https://doi.org/10.5281/zenodo.8180380. Luigi Ferraro, Giovanni Scala, Luigi Cerulo, Emanuele Carosati, Michele Ceccarelli |
Bioinform. | 5 |
| 2021 | A review of COVID-19 biomarkers and drug targets: resources and toolsabstractThe stratification of patients at risk of progression of COVID-19 and their molecular characterization is of extreme importance to optimize treatment and to identify therapeutic options. The bioinformatics community has responded to the outbreak emergency with a set of tools and resource to identify biomarkers and drug targets that we review here. Starting from a consolidated corpus of 27 570 papers, we adopt latent Dirichlet analysis to extract relevant topics and select those associated with computational methods for biomarker identification and drug repurposing. The selected topics span from machine learning and artificial intelligence for disease characterization to vaccine development and to therapeutic target identification. Although the way to go for the ultimate defeat of the pandemic is still long, the amount of knowledge, data and tools generated so far constitutes an unprecedented example of global cooperation to this threat. Francesca P. Caruso, Giovanni Scala, Luigi Cerulo, Michele Ceccarelli |
Briefings Bioinform. | 4 |
| 2021 | Network-based identification of key master regulators associated with an immune-silent cancer phenotypeabstractA cancer immune phenotype characterized by an active T-helper 1 (Th1)/cytotoxic response is associated with responsiveness to immunotherapy and favorable prognosis across different tumors. However, in some cancers, such an intratumoral immune activation does not confer protection from progression or relapse. Defining mechanisms associated with immune evasion is imperative to refine stratification algorithms, to guide treatment decisions and to identify candidates for immune-targeted therapy. Molecular alterations governing mechanisms for immune exclusion are still largely unknown. The availability of large genomic datasets offers an opportunity to ascertain key determinants of differential intratumoral immune response. We follow a network-based protocol to identify transcription regulators (TRs) associated with poor immunologic antitumor activity. We use a consensus of four different pipelines consisting of two state-of-the-art gene regulatory network inference techniques, regularized gradient boosting machines and ARACNE to determine TR regulons, and three separate enrichment techniques, including fast gene set enrichment analysis, gene set variation analysis and virtual inference of protein activity by enriched regulon analysis to identify the most important TRs affecting immunologic antitumor activity. These TRs, referred to as master regulators (MRs), are unique to immune-silent and immune-active tumors, respectively. We validated the MRs coherently associated with the immune-silent phenotype across cancers in The Cancer Genome Atlas and a series of additional datasets in the Prediction of Clinical Outcomes from Genomic Profiles repository. A downstream analysis of MRs specific to the immune-silent phenotype resulted in the identification of several enriched candidate pathways, including NOTCH1, TGF-$\beta $, Interleukin-1 and TNF-$\alpha $ signaling pathways. TGFB1I1 emerged as one of the main negative immune modulators preventing the favorable effects of a Th1/cytotoxic response. Raghvendra Mall, Mohamad Saad 0001, Jessica Roelands, Darawan Rinchai, Khalid Kunji, Hossam Almeer, Wouter Hendrickx, Francesco M. Marincola, Michele Ceccarelli, Davide Bedognetti |
Briefings Bioinform. | 9 |
| 2021 | Adaptive one-class Gaussian processes allow accurate prioritization of oncology drug targetsabstractMOTIVATION: The cost of drug development has dramatically increased in the last decades, with the number new drugs approved per billion US dollars spent on R&D halving every year or less. The selection and prioritization of targets is one the most influential decisions in drug discovery. Here we present a Gaussian Process model for the prioritization of drug targets cast as a problem of learning with only positive and unlabeled examples. RESULTS: Since the absence of negative samples does not allow standard methods for automatic selection of hyperparameters, we propose a novel approach for hyperparameter selection of the kernel in One Class Gaussian Processes. We compare our methods with state-of-the-art approaches on benchmark datasets and then show its application to druggability prediction of oncology drugs. Our score reaches an AUC 0.90 on a set of clinical trial targets starting from a small training set of 102 validated oncology targets. Our score recovers the majority of known drug targets and can be used to identify novel set of proteins as drug target candidates. AVAILABILITY AND IMPLEMENTATION: The matrix of features for each protein is available at: https://bit.ly/3iLgZTa. Source code implemented in Python is freely available for download at https://github.com/AntonioDeFalco/Adaptive-OCGP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antonio de Falco, Zoltán Dezsö, Francesco Ceccarelli, Luigi Cerulo, Angelo Ciaramella, Michele Ceccarelli |
Bioinform. | 6 |
| 2020 | Machine learning prediction of oncology drug targets based on protein and network propertiesabstractBACKGROUND: The selection and prioritization of drug targets is a central problem in drug discovery. Computational approaches can leverage the growing number of large-scale human genomics and proteomics data to make in-silico target identification, reducing the cost and the time needed. RESULTS: We developed a machine learning approach to score proteins to generate a druggability score of novel targets. In our model we incorporated 70 protein features which included properties derived from the sequence, features characterizing protein functions as well as network properties derived from the protein-protein interaction network. The advantage of this approach is that it is unbiased and even less studied proteins with limited information about their function can score well as most of the features are independent of the accumulated literature. We build models on a training set which consist of targets with approved drugs and a negative set of non-drug targets. The machine learning techniques help to identify the most important combination of features differentiating validated targets from non-targets. We validated our predictions on an independent set of clinical trial drug targets, achieving a high accuracy characterized by an Area Under the Curve (AUC) of 0.89. Our most predictive features included biological function of proteins, network centrality measures, protein essentiality, tissue specificity, localization and solvent accessibility. Our predictions, based on a small set of 102 validated oncology targets, recovered the majority of known drug targets and identifies a novel set of proteins as drug target candidates. CONCLUSIONS: We developed a machine learning approach to prioritize proteins according to their similarity to approved drug targets. We have shown that the method proposed is highly predictive on a validation dataset consisting of 277 targets of clinical trial drug confirming that our computational approach is an efficient and cost-effective tool for drug target discovery and prioritization. Our predictions were based on oncology targets and cancer relevant biological functions, resulting in significantly higher scores for targets of oncology clinical trial drugs compared to the scores of targets of trial drugs for other indications. Our approach can be used to make indication specific drug-target prediction by combining generic druggability features with indication specific biological functions. Zoltán Dezsö, Michele Ceccarelli |
BMC Bioinform. | 2 |
| 2020 | Deep learning predicts short non-coding RNA functions from only raw sequence dataabstractSmall non-coding RNAs (ncRNAs) are short non-coding sequences involved in gene regulation in many biological processes and diseases. The lack of a complete comprehension of their biological functionality, especially in a genome-wide scenario, has demanded new computational approaches to annotate their roles. It is widely known that secondary structure is determinant to know RNA function and machine learning based approaches have been successfully proven to predict RNA function from secondary structure information. Here we show that RNA function can be predicted with good accuracy from a lightweight representation of sequence information without the necessity of computing secondary structure features which is computationally expensive. This finding appears to go against the dogma of secondary structure being a key determinant of function in RNA. Compared to recent secondary structure based methods, the proposed solution is more robust to sequence boundary noise and reduces drastically the computational cost allowing for large data volume annotations. Scripts and datasets to reproduce the results of experiments proposed in this study are available at: https://github.com/bioinformatics-sannio/ncrna-deep. Teresa Maria Rosaria Noviello, Francesco Ceccarelli, Michele Ceccarelli, Luigi Cerulo |
PLoS Comput. Biol. | 3 |
| 2020 | Preface: In memory of Alfredo Petrosino
Michele Ceccarelli |
Pattern Recognit. Lett. | 1 |
| 2018 | Detection of long non-coding RNA homology, a comparative study on alignment and alignment-free metricsabstractBACKGROUND: Long non-coding RNAs (lncRNAs) represent a novel class of non-coding RNAs having a crucial role in many biological processes. The identification of long non-coding homologs among different species is essential to investigate such roles in model organisms as homologous genes tend to retain similar molecular and biological functions. Alignment-based metrics are able to effectively capture the conservation of transcribed coding sequences and then the homology of protein coding genes. However, unlike protein coding genes the poor sequence conservation of long non-coding genes makes the identification of their homologs a challenging task. RESULTS: In this study we compare alignment-based and alignment-free string similarity metrics and look at promoter regions as a possible source of conserved information. We show that promoter regions encode relevant information for the conservation of long non-coding genes across species and that such information is better captured by alignment-free metrics. We perform a genome wide test of this hypothesis in human, mouse, and zebrafish. CONCLUSIONS: The obtained results persuaded us to postulate the new hypothesis that, unlike protein coding genes, long non-coding genes tend to preserve their regulatory machinery rather than their transcribed sequence. All datasets, scripts, and the prediction tools adopted in this study are available at https://github.com/bioinformatics-sannio/lncrna-homologs . Teresa Maria Rosaria Noviello, Antonella Di Liddo, Giovanna M. M. Ventola, Antonietta Spagnuolo, Salvatore D'Aniello, Michele Ceccarelli, Luigi Cerulo |
BMC Bioinform. | 6 |
| 2017 | An adaptive refinement for community detection methods for disease module identification in biological networks using novel metric based on connectivity, conductance & modularityabstractDisease processes are usually driven by several genes interacting in molecular modules or pathways leading to the disease. The identification of such modules in gene or protein networks is at the core of several analysis methods in biomedical research. However, there is still a need to develop a generic framework to uncover biologically relevant modules for different types of networks. With this pretext in mind, the Disease Module Identification DREAM Challenge was initiated as an effort to systematically assess module identification methods on a panel of 6 diverse state-of-the-art genomic networks. Methods: In this paper, we propose a generic refinement method based on ideas of merging and splitting the hierarchical tree obtained from any community detection technique for constrained disease module identification in biological networks. The only constraint for a module to be considered as a candidate disease module was size of the community to be: 3 ≤ community size ≤ 100. Here, we propose a novel quality metric, called F-score, computed from several unsupervised quality metrics like modularity, conductance and connectivity to determine the quality of a graph partition at a given level of hierarchy. We also propose a quality metric, namely Inverse Confidence, which ranks and prune insignificant modules to obtain a curated list of candidate disease modules for a given biological network. The predicted modules are then evaluated on the basis of the total number of unique candidate modules that are associated with complex traits and diseases from over 200 genome-wide association study (GWAS) datasets. Results: We stood 9thout of a total of 42 teams in the competition at the offical FDR cut-off of 0.05 for identifying statistically significant disease associated modules in the 6 benchmark networks. Our proposed approach detected a total of 44 disease modules in the 6 benchmark networks in comparison to 60 for the winner of the DREAM Challenge. For several benchmark networks we were better or competitive with the winner. Raghvendra Mall, Ehsan Ullah, Khalid Kunji, Halima Bensmail, Michele Ceccarelli |
BIBM | 5 |
| 2017 | Identification of long non-coding transcripts with feature selection: a comparative studyabstractBACKGROUND: The unveiling of long non-coding RNAs as important gene regulators in many biological contexts has increased the demand for efficient and robust computational methods to identify novel long non-coding RNAs from transcripts assembled with high throughput RNA-seq data. Several classes of sequence-based features have been proposed to distinguish between coding and non-coding transcripts. Among them, open reading frame, conservation scores, nucleotide arrangements, and RNA secondary structure have been used with success in literature to recognize intergenic long non-coding RNAs, a particular subclass of non-coding RNAs. RESULTS: In this paper we perform a systematic assessment of a wide collection of features extracted from sequence data. We use most of the features proposed in the literature, and we include, as a novel set of features, the occurrence of repeats contained in transposable elements. The aim is to detect signatures (groups of features) able to distinguish long non-coding transcripts from other classes, both protein-coding and non-coding. We evaluate different feature selection algorithms, test for signature stability, and evaluate the prediction ability of a signature with a machine learning algorithm. The study reveals different signatures in human, mouse, and zebrafish, highlighting that some features are shared among species, while others tend to be species-specific. Compared to coding potential tools and similar supervised approaches, including novel signatures, such as those identified here, in a machine learning algorithm improves the prediction performance, in terms of area under precision and recall curve, by 1 to 24%, depending on the species and on the signature. CONCLUSIONS: Understanding which features are best suited for the prediction of long non-coding RNAs allows for the development of more effective automatic annotation pipelines especially relevant for poorly annotated genomes, such as zebrafish. We provide a web tool that recognizes novel long non-coding RNAs with the obtained signatures from fasta and gtf formats. The tool is available at the following url: http://www.bioinformatics-sannio.org/software/ . Giovanna M. M. Ventola, Teresa Maria Rosaria Noviello, Salvatore D'Aniello, Antonietta Spagnuolo, Michele Ceccarelli, Luigi Cerulo |
BMC Bioinform. | 5 |
| 2015 | VEGAWES: variational segmentation on whole exome sequencing for copy number detectionabstractBACKGROUND: Copy number variations are important in the detection and progression of significant tumors and diseases. Recently, Whole Exome Sequencing is gaining popularity with copy number variations detection due to low cost and better efficiency. In this work, we developed VEGAWES for accurate and robust detection of copy number variations on WES data. VEGAWES is an extension to a variational based segmentation algorithm, VEGA: Variational estimator for genomic aberrations, which has previously outperformed several algorithms on segmenting array comparative genomic hybridization data. RESULTS: We tested this algorithm on synthetic data and 100 Glioblastoma Multiforme primary tumor samples. The results on the real data were analyzed with segmentation obtained from Single-nucleotide polymorphism data as ground truth. We compared our results with two other segmentation algorithms and assessed the performance based on accuracy and time. CONCLUSIONS: In terms of both accuracy and time, VEGAWES provided better results on the synthetic data and tumor samples demonstrating its potential in robust detection of aberrant regions in the genome. Samreen Anjum, Sandro Morganella, Fulvio D'Angelo, Antonio Iavarone, Michele Ceccarelli |
BMC Bioinform. | 5 |
| 2015 | Irish: A Hidden Markov Model to detect coded information islands in free text
Luigi Cerulo, Massimiliano Di Penta, Alberto Bacchelli, Michele Ceccarelli, Gerardo Canfora |
Sci. Comput. Program. | 4 |
| 2013 | Infer gene regulatory networks from time series data with formal methodsabstractReverse engineering of regulatory relationships from genomics data is emerging as crucial to dissect the complex underlying regulatory mechanism occurring in a cell. In this paper we propose a novel reverse engineering algorithm that makes use of formal methods, usually adopted in engineering to specify and verify concurrent software and hardware systems. With a formal specification of gene regulatory hypotheses we are able to prove mathematically whether a time course experiment belongs or not to the formal specification, determining in fact whether a gene regulation exists or not. The method is capable to detect both direction and sign (inhibition/activation) of regulations whereas most of literature methods which are limited to undirected and/or unsigned relationships. The method, empirically evaluated on experimental and synthetic datasets, reaches high levels of accuracy, outperforming literature methods in terms of precision and recall, despite the computational cost increases exponentially with the size of the network. Michele Ceccarelli, Luigi Cerulo, Antonella Santone |
BIBM | 1 |
| 2013 | A Hidden Markov Model to detect coded information islands in free textabstractEmails and issue reports capture useful knowledge about development practices, bug fixing, and change activities. Extracting such a content is challenging, due to the mix-up of source code and natural language, unstructured text. Luigi Cerulo, Michele Ceccarelli, Massimiliano Di Penta, Gerardo Canfora |
SCAM | 2 |
| 2013 | A negative selection heuristic to predict new transcriptional targetsabstractBACKGROUND: Supervised machine learning approaches have been recently adopted in the inference of transcriptional targets from high throughput trascriptomic and proteomic data showing major improvements from with respect to the state of the art of reverse gene regulatory network methods. Beside traditional unsupervised techniques, a supervised classifier learns, from known examples, a function that is able to recognize new relationships for new data. In the context of gene regulatory inference a supervised classifier is coerced to learn from positive and unlabeled examples, as the counter negative examples are unavailable or hard to collect. Such a condition could limit the performance of the classifier especially when the amount of training examples is low. RESULTS: In this paper we improve the supervised identification of transcriptional targets by selecting reliable counter negative examples from the unlabeled set. We introduce an heuristic based on the known topology of transcriptional networks that in fact restores the conventional positive/negative training condition and shows a significant improvement of the classification performance. We empirically evaluate the proposed heuristic with the experimental datasets of Escherichia coli and show an example of application in the prediction of BCL6 direct core targets in normal germinal center human B cells obtaining a precision of 60%. CONCLUSIONS: The availability of only positive examples in learning transcriptional relationships negatively affects the performance of supervised classifiers. We show that the selection of reliable negative examples, a practice adopted in text mining approaches, improves the performance of such classifiers opening new perspectives in the identification of new transcriptional targets. Luigi Cerulo, Vincenzo Paduano, Pietro Zoppoli, Michele Ceccarelli |
BMC Bioinform. | 4 |
| 2012 | VegaMC: a R/bioconductor package for fast downstream analysis of large array comparative genomic hybridization datasetsabstractSUMMARY: Identification of genetic alterations of tumor cells has become a common method to detect the genes involved in development and progression of cancer. In order to detect driver genes, several samples need to be simultaneously analyzed. The Cancer Genome Atlas (TCGA) project provides access to a large amount of data for several cancer types. TGCA is an invaluable source of information, but analysis of this huge dataset possess important computational problems in terms of memory and execution times. Here, we present a R/package, called VegaMC (Vega multi-channel), that enables fast and efficient detection of significant recurrent copy number alterations in very large datasets. VegaMC is integrated with the output of the common tools that convert allele signal intensities in log R ratio and B allele frequency. It also enables the detection of loss of heterozigosity and provides in output two web pages allowing a rapid and easy navigation of the aberrant genes. Synthetic data and real datasets are used for quantitative and qualitative evaluation purposes. In particular, we demonstrate the ability of VegaMC on two large TGCA datasets: colon adenocarcinoma and glioblastoma multiforme. For both the datasets, we provide the list of aberrant genes which contain previously validated genes and can be used as basis for further investigations. AVAILABILITY: VegaMC is a R/Bioconductor Package, available at http://bioconductor.org/packages/release/bioc/html/VegaMC.html. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sandro Morganella, Michele Ceccarelli |
Bioinform. | 2 |
| 2011 | An Algorithm for Finding Gene Signatures Supervised by Survival Time Data
Stefano Maria Pagnotta, Michele Ceccarelli |
KES (1) | 2 |
| 2011 | Finding recurrent copy number alterations preserving within-sample homogeneityabstractMOTIVATION: Copy number alterations (CNAs) represent an important component of genetic variation and play a significant role in many human diseases. Development of array comparative genomic hybridization (aCGH) technology has made it possible to identify CNAs. Identification of recurrent CNAs represents the first fundamental step to provide a list of genomic regions which form the basis for further biological investigations. The main problem in recurrent CNAs discovery is related to the need to distinguish between functional changes and random events without pathological relevance. Within-sample homogeneity represents a common feature of copy number profile in cancer, so it can be used as additional source of information to increase the accuracy of the results. Although several algorithms aimed at the identification of recurrent CNAs have been proposed, no attempt of a comprehensive comparison of different approaches has yet been published. RESULTS: We propose a new approach, called Genomic Analysis of Important Alterations (GAIA), to find recurrent CNAs where a statistical hypothesis framework is extended to take into account within-sample homogeneity. Statistical significance and within-sample homogeneity are combined into an iterative procedure to extract the regions that likely are involved in functional changes. Results show that GAIA represents a valid alternative to other proposed approaches. In addition, we perform an accurate comparison by using two real aCGH datasets and a carefully planned simulation study. AVAILABILITY: GAIA has been implemented as R/Bioconductor package. It can be downloaded from the following page http://bioinformatics.biogem.it/download/gaia. CONTACT: [email protected]; [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sandro Morganella, Stefano Maria Pagnotta, Michele Ceccarelli |
Bioinform. | 3 |
| 2010 | An eclectic approach for change impact analysisabstractChange impact analysis aims at identifying software artifacts being affected by a change. In the past, this problem has been addressed by approaches relying on static, dynamic, and textual analysis. Recently, techniques based on historical analysis and association rules have been explored. This paper proposes a novel change impact analysis method based on the idea that the mutual relationships between software objects can be inferred with a statistical learning approach. We use the bivariate Granger causality test, a multivariate time series forecasting approach used to verify whether past values of a time series are useful for predicting future values of another time series. Results of a preliminary study performed on the Samba daemon show that change impact relationships inferred with the Granger causality test are complementary to those inferred with association rules. This opens the road towards the development of an eclectic impact analysis approach conceived by combining different techniques. Michele Ceccarelli, Luigi Cerulo, Gerardo Canfora, Massimiliano Di Penta |
ICSE (2) | 1 |
| 2010 | Using multivariate time series and association rules to detect logical change coupling: An empirical studyabstractIn recent years, techniques based on association rules discovery have been extensively used to determine change-coupling relations between artifacts that often changed together. Although association rules worked well in many cases, they fail to capture logical coupling relations between artifacts modified in subsequent change sets. To overcome such a limitation, we propose the use of multivariate time series analysis and forecasting, and in particular the use of Granger causality test, to determine whether a change occurred on a software artifact was consequentially related to changes occurred on some other artifacts. Results of an empirical study performed on four Java and C open source systems show that Granger causality test is able to provide a set of change couplings complementary to association rules, and a hybrid recommender built combining recommendations from association rules and Granger causality is able to achieve a higher recall than the two single techniques. Gerardo Canfora, Michele Ceccarelli, Luigi Cerulo, Massimiliano Di Penta |
ICSM | 2 |
| 2010 | VEGA: variational segmentation for copy number detectionabstractMOTIVATION: Genomic copy number (CN) information is useful to study genetic traits of many diseases. Using array comparative genomic hybridization (aCGH), researchers are able to measure the copy number of thousands of DNA loci at the same time. Therefore, a current challenge in bioinformatics is the development of efficient algorithms to detect the map of aberrant chromosomal regions. METHODS: We describe an approach for the segmentation of copy number aCGH data. Variational estimator for genomic aberrations (VEGA) adopt a variational model used in image segmentation. The optimal segmentation is modeled as the minimum of an energy functional encompassing both the quality of interpolation of the data and the complexity of the solution measured by the length of the boundaries between segmented regions. This solution is obtained by a region growing process where the stop condition is completely data driven. RESULTS: VEGA is compared with three algorithms that represent the state of the art in CN segmentation. Performance assessment is made both on synthetic and real data. Synthetic data simulate different noise conditions. Results on these data show the robustness with respect to noise of variational models and the accuracy of VEGA in terms of recall and precision. Eight mantle cell lymphoma cell lines and two samples of glioblastoma multiforme are used to evaluate the behavior of VEGA on real biological data. Comparison between results and current biological knowledge shows the ability of the proposed method in detecting known chromosomal aberrations. AVAILABILITY: VEGA has been implemented in R and is available at the address http://www.dsba.unisannio.it/Members/ceccarelli/vega in the section Download. Sandro Morganella, Luigi Cerulo, Giuseppe Viglietto, Michele Ceccarelli |
Bioinform. | 4 |
| 2010 | Learning gene regulatory networks from only positive and unlabeled dataabstractBACKGROUND: Recently, supervised learning methods have been exploited to reconstruct gene regulatory networks from gene expression data. The reconstruction of a network is modeled as a binary classification problem for each pair of genes. A statistical classifier is trained to recognize the relationships between the activation profiles of gene pairs. This approach has been proven to outperform previous unsupervised methods. However, the supervised approach raises open questions. In particular, although known regulatory connections can safely be assumed to be positive training examples, obtaining negative examples is not straightforward, because definite knowledge is typically not available that a given pair of genes do not interact. RESULTS: A recent advance in research on data mining is a method capable of learning a classifier from only positive and unlabeled examples, that does not need labeled negative examples. Applied to the reconstruction of gene regulatory networks, we show that this method significantly outperforms the current state of the art of machine learning methods. We assess the new method using both simulated and experimental data, and obtain major performance improvement. CONCLUSIONS: Compared to unsupervised methods for gene network inference, supervised methods are potentially more accurate, but for training they need a complete set of known regulatory connections. A supervised method that can be trained using only positive and unlabeled data, as presented in this paper, is especially beneficial for the task of inferring gene regulatory networks, because only an incomplete set of known regulatory connections is available in public databases such as RegulonDB, TRRD, KEGG, Transfac, and IPA. Luigi Cerulo, Charles Elkan, Michele Ceccarelli |
BMC Bioinform. | 3 |
| 2010 | TimeDelay-ARACNE: Reverse engineering of gene networks from time-course data by an information theoretic approachabstractBACKGROUND: One of main aims of Molecular Biology is the gain of knowledge about how molecular components interact each other and to understand gene function regulations. Using microarray technology, it is possible to extract measurements of thousands of genes into a single analysis step having a picture of the cell gene expression. Several methods have been developed to infer gene networks from steady-state data, much less literature is produced about time-course data, so the development of algorithms to infer gene networks from time-series measurements is a current challenge into bioinformatics research area. In order to detect dependencies between genes at different time delays, we propose an approach to infer gene regulatory networks from time-series measurements starting from a well known algorithm based on information theory. RESULTS: In this paper we show how the ARACNE (Algorithm for the Reconstruction of Accurate Cellular Networks) algorithm can be used for gene regulatory network inference in the case of time-course expression profiles. The resulting method is called TimeDelay-ARACNE. It just tries to extract dependencies between two genes at different time delays, providing a measure of these dependencies in terms of mutual information. The basic idea of the proposed algorithm is to detect time-delayed dependencies between the expression profiles by assuming as underlying probabilistic model a stationary Markov Random Field. Less informative dependencies are filtered out using an auto calculated threshold, retaining most reliable connections. TimeDelay-ARACNE can infer small local networks of time regulated gene-gene interactions detecting their versus and also discovering cyclic interactions also when only a medium-small number of measurements are available. We test the algorithm both on synthetic networks and on microarray expression profiles. Microarray measurements concern S. cerevisiae cell cycle, E. coli SOS pathways and a recently developed network for in vivo assessment of reverse engineering algorithms. Our results are compared with ARACNE itself and with the ones of two previously published algorithms: Dynamic Bayesian Networks and systems of ODEs, showing that TimeDelay-ARACNE has good accuracy, recall and F-score for the network reconstruction task. CONCLUSIONS: Here we report the adaptation of the ARACNE algorithm to infer gene regulatory networks from time-course data, so that, the resulting network is represented as a directed graph. The proposed algorithm is expected to be useful in reconstruction of small biological directed networks from time course data. Pietro Zoppoli, Sandro Morganella, Michele Ceccarelli |
BMC Bioinform. | 3 |
| 2009 | A Guideline Engine For Knowledge Management in Clinical Decision Support Systems (CDSSs)
Michele Ceccarelli, Alessandro De Stasio, Antonio Donatiello, Dante Vitale |
SEKE | 1 |
| 2009 | Virtual genetic coding and time series analysis for alternative splicing prediction in C. elegans
Michele Ceccarelli, Antonio Maratea |
Artif. Intell. Medicine | 1 |
| 2009 | A scale space approach for unsupervised feature selection in mass spectra classification for ovarian cancer detectionabstractBACKGROUND: Mass spectrometry spectra, widely used in proteomics studies as a screening tool for protein profiling and to detect discriminatory signals, are high dimensional data. A large number of local maxima (a.k.a. peaks) have to be analyzed as part of computational pipelines aimed at the realization of efficient predictive and screening protocols. With this kind of data dimensions and samples size the risk of over-fitting and selection bias is pervasive. Therefore the development of bio-informatics methods based on unsupervised feature extraction can lead to general tools which can be applied to several fields of predictive proteomics. RESULTS: We propose a method for feature selection and extraction grounded on the theory of multi-scale spaces for high resolution spectra derived from analysis of serum. Then we use support vector machines for classification. In particular we use a database containing 216 samples spectra divided in 115 cancer and 91 control samples. The overall accuracy averaged over a large cross validation study is 98.18. The area under the ROC curve of the best selected model is 0.9962. CONCLUSION: We improved previous known results on the problem on the same data, with the advantage that the proposed method has an unsupervised feature selection phase. All the developed code, as MATLAB scripts, can be downloaded from http://medeaserver.isa.cnr.it/dacierno/spectracode.htm. Michele Ceccarelli, Antonio d'Acierno, Angelo M. Facchiano |
BMC Bioinform. | 1 |
| 2009 | IRIS: a method for reverse engineering of regulatory relations in gene networksabstractBACKGROUND: The ultimate aim of systems biology is to understand and describe how molecular components interact to manifest collective behaviour that is the sum of the single parts. Building a network of molecular interactions is the basic step in modelling a complex entity such as the cell. Even if gene-gene interactions only partially describe real networks because of post-transcriptional modifications and protein regulation, using microarray technology it is possible to combine measurements for thousands of genes into a single analysis step that provides a picture of the cell's gene expression. Several databases provide information about known molecular interactions and various methods have been developed to infer gene networks from expression data. However, network topology alone is not enough to perform simulations and predictions of how a molecular system will respond to perturbations. Rules for interactions among the single parts are needed for a complete definition of the network behaviour. Another interesting question is how to integrate information carried by the network topology, which can be derived from the literature, with large-scale experimental data. RESULTS: Here we propose an algorithm, called inference of regulatory interaction schema (IRIS), that uses an iterative approach to map gene expression profile values (both steady-state and time-course) into discrete states and a simple probabilistic method to infer the regulatory functions of the network. These interaction rules are integrated into a factor graph model. We test IRIS on two synthetic networks to determine its accuracy and compare it to other methods. We also apply IRIS to gene expression microarray data for the Saccharomyces cerevisiae cell cycle and for human B-cells and compare the results to literature findings. CONCLUSIONS: IRIS is a rapid and efficient tool for the inference of regulatory relations in gene networks. A topological description of the network and a matrix of gene expression profiles are required as input to the algorithm. IRIS maps gene expression data onto discrete values and then computes regulatory functions as conditional probability tables. The suitability of the method is demonstrated for synthetic data and microarray data. The resulting network can also be embedded in a factor graph model. Sandro Morganella, Pietro Zoppoli, Michele Ceccarelli |
BMC Bioinform. | 3 |
| 2008 | KON^3: A Clinical Decision Support System, in Oncology Environment, Based on Knowledge ManagementabstractThe application of scientific methodology to clinical practice is typically realized through recommendations, policies and protocols represented as Clinical Practice Guidelines (CPG). CPG help the clinician in his choices, improving the patient care process.The representation of Guidelines and their introduction in medical information system can lead to efficient Clinical Decision Support Systems (CDSS), however this poses several interesting challenges as problems of knowledge representation, inference, workflow definition, access to unstructured data, etc. In this paper we analyze approaches and methods in computer-based CPG realization and illustrate the choices made (tools, architecture and clinical domain) to realize the KON3system to achieve a CDSS and semantic information representation, in oncology environment Michele Ceccarelli, Antonio Donatiello, Dante Vitale |
ICTAI (2) | 1 |
| 2008 | Image Registration using Non-Linear DiffusionabstractAn image registration algorithm based on mutual information maximization and non-linear diffusion is presented. It relies on a non-parametric estimation of the degree of dependency between the reference image and the template to be registered, which is intrinsically more robust against possible deformations due to imaging geometry and propagation disturbances. The approach based on non linear diffusion, nonetheless, has the advantage of producing a non-parametric discrete warping model which does not rely on a particular set of basis functions, and is therefore as much general as possible. The experimental results on simulated images have quantitatively shown the accuracy of the proposed method. Michele Ceccarelli, Maurizio di Bisceglie, Carmela Galdi, Generoso Giangregorio, Silvia Liberata Ullo |
IGARSS (5) | 1 |
| 2008 | A Fuzzy Extension of Some Classical Concordance Measures and an Efficient Algorithm for Their Computation
Michele Ceccarelli, Antonio Maratea |
KES (3) | 1 |
| 2008 | Improving fuzzy clustering of biological data by metric learning with side information
Michele Ceccarelli, Antonio Maratea |
Int. J. Approx. Reason. | 1 |
| 2007 | Microarray image gridding with stochastic search based approaches
Giuliano Antoniol, Michele Ceccarelli |
Image Vis. Comput. | 2 |
| 2007 | A Finite Markov Random Field approach to fast edge-preserving image recovery
Michele Ceccarelli |
Image Vis. Comput. | 1 |
| 2006 | An Active Contour Approach To Automatic Detection Of The Intima-Media ThicknessabstractIn this paper, we propose a snake-based approach for the automatic detection of the intima-media thickness (IMT) of the far wall of the common carotid artery. The main problems to be faced are related, from one point, to the high level of speckle noise and, from the other, to the need of an accurate segmentation of the vessel structure for diagnostic purposes. In particular, the detection of the intima layer plays a fundamental role as it represents the initial contour from where the method should start. We propose an automatic segmentation approach which makes use of a first non linear filtering based on anisotropic diffusion followed by an iterative relaxation procedure. Once the intima layer has been detected our method tries to locate an optimal initial contour to detect the wall of the artery by minimizing of a modified energy functional Michele Ceccarelli, Nicola De Luca, Sandro Morganella |
ICASSP (2) | 1 |
| 2006 | Content-based Image Retrieval by a Fuzzy Scale-space ApproachabstractImage descriptions aimed at the realization of content-based image retrieval (CBIR) should include the vagueness of both data representations and user queries. Here we show how multiscale textural gradient can be used as an efficient visual cue for image description. This feature has been already efficiently used in problems of image segmentation and texture separation. Our main idea is based on the assumption that, for image description, shape and textures should be considered together within a unified model. We report an efficient image description algorithm where the multiscale analysis is modeled by a differential morphological filter. Experiments with large image databases and comparisons with classical methods are reported. Michele Ceccarelli, Francesco Musacchia, Alfredo Petrosino |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | A Deformable Grid-Matching Approach for Microarray ImagesabstractA fundamental step of microarray image analysis is the detection of the grid structure for the accurate location of each spot, representing the state of a given gene in a particular experimental condition. This step is known as gridding and belongs to the class of deformable grid matching problems which are well known in literature. Most of the available microarray gridding approaches require human intervention; for example, to specify landmarks, some points in the spot grid, or even to precisely locate individual spots. Automating this part of the process can allow high throughput analysis. This paper focuses on the development of a fully automated procedure for the problem of automatic microarray gridding. It is grounded on the Bayesian paradigm and on image analysis techniques. The procedure has two main steps. The first step, based on the Radon transform, is aimed at generating a grid hypothesis; the second step accounts for local grid deformations. The accuracy and properties of the procedure are quantitatively assessed over a set of synthetic and real images; the results are compared with well-known methods available from the literature. Michele Ceccarelli, Giuliano Antoniol |
IEEE Trans. Image Process. | 1 |
| 2005 | Microarray image addressing based on the Radon transformabstractA fundamental step of microarray image analysis is the detection of the grid structure for the accurate localization of each spot, representing the state of a given gene in a particular experimental condition. This step is known as gridding or microarray addressing. Most of the available microarray gridding approaches require human intervention; for example, to specify landmarks, some points in the spot grid, or even to precisely locate individual spots. Automating this part of the process can allow high throughput analysis (Yang, Y, et al, 2002). This paper is aimed towards at the development fully automated procedures for the problem of automatic microarray gridding. Indeed, many of the automatic gridding approaches are based on two phases, the first aimed at the generation of an hypothesis consisting into a regular interpolating grid, whereas the second performs an adaptation of the hypothesis. Here we show that the first step can efficiently be accomplished by using the Radon transform, whereas the second step could be modeled by an iterative posterior maximization procedure (Antoniol, G and Ceccarelli, M, 2004). Giuliano Antoniol, Michele Ceccarelli, Alfredo Petrosino |
ICIP (1) | 2 |
| 2004 | Blotch removal in degraded digital video using independent component analysisabstractBlotchy noise is one of the most common and visually annoying noises noticed in digitized motion picture films. This work investigates a three-stage blotchy noise reduction scheme combining efforts of temporal median filtering, ICA unsupervised learning and deflation after motion/compensation estimation. Implementation of each module of the scheme is discussed; finally performance of the scheme is tested and compared with other methods over real short sequences and results are discussed. Michele Ceccarelli, Alfredo Petrosino |
IJCNN | 1 |
| 2002 | A parallel fuzzy scale-space approach to the unsupervised texture separation
Michele Ceccarelli, Alfredo Petrosino |
Pattern Recognit. Lett. | 1 |
| 2001 | The orientation matching approach to circular object detectionabstractThe paper reports a correlation-based method for the detection of circular objects which is capable of overcoming well-known problems arising from the use of gradient-based voting schemes. Specifically, the method is: (a) capable of detecting circular objects on the basis of both magnitude and direction of the image gradient; and (b) dealing with three-dimensional spherical objects by considering shadows depending on the direction of light. Experimental results on the accuracy of the method and comparisons with the Hough transform and the Hausdorff matching are reported. Michele Ceccarelli, Alfredo Petrosino |
ICIP (3) | 1 |
| 2000 | Unsupervised Texture Discrimination Based on Rough Fuzzy Sets and Parallel Hierarchical ClusteringabstractReports a texture separation algorithm to solve the problem of unsupervised boundary localization in textured images. The proposed algorithm is mainly characterized by the extraction of textural density gradients by a nonlinear multiple scale-space analysis of the image. Texture boundaries are extracted by segmenting the images resulting from a multiscale fuzzy gradient operation applied to detail images. The segmentation stage consists of a parallel hierarchical clustering algorithm, aimed at the minimization of a global cost functional taking into account region homogeneity and segmentation quality. Experiments and comparisons on Brodatz textures are reported. Alfredo Petrosino, Michele Ceccarelli |
ICPR | 2 |
| 1997 | Multi-feature adaptive classifiers for SAR image segmentation
Michele Ceccarelli, Alfredo Petrosino |
Neurocomputing | 1 |
| 1996 | Sequence recognition with radial basis function networks: experiments with spoken digits
Michele Ceccarelli, Joël T. Hounsou |
Neurocomputing | 1 |
| 1993 | Competitive neural networks on message-passing parallel computersabstractAbstract The paper reports two techniques for parallelizing on a MIMD multicomputer a class of learning algorithms (competitive learning) for artificial neural networks widely used in pattern recognition and understanding. The first technique presented, following the divide et impera strategy, achieves O(n/p + logP) time for n neurons and P processors interconnected as a tree. A modification of the algorithm allows the application of a systolic technique with the processors interconnected as a ring; this technique has the advantage that the communication time does not depend on the number of processors. The two techniques are also compared on the basis of predicted and measured performance on a transputer‐based MIMD machine. As the number of processors grows the advantage of the systolic approach increases. On the contrary, the divide et impera approach is more advantageous in the retrieving phase. Michele Ceccarelli, Alfredo Petrosino, Roberto Vaccaro |
Concurr. Pract. Exp. | 1 |