VLDB 2026 Research / reviewers in the wild / expert
Roberto Tagliaferri
dblp:69/6674
· DBLP profile ↗
79ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-8134-9025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 since 2021Databases, data management, data science and information retrieval · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing feature compression and reconstruction in time-series and image domains with pseudo-overlap and pseudo-grouping pooling functionsabstractDeep learning techniques are widely used for compressing and reconstructing images and time-series features, thanks to their ability to learn efficient latent representations. In this work, we introduce two families of pooling functions —pseudo-overlap and pseudo-grouping— and integrate them within an autoencoder to improve feature extraction and latent space organization. Unlike conventional pooling methods, these functions adaptively modulate feature aggregation, allowing better preservation of structural information during compression. We evaluate the proposed pooling mechanisms on both image (MNIST, FashionMNIST, CIFAR10, SVHN, CIFAR100, ImageNet) and time-series datasets (Vasicek, J-Vasicek, GMB, J-GMB, MRD, J-MRD), demonstrating enhanced reconstruction accuracy. Comparative experiments show that pseudo-overlap and pseudo-grouping outperform traditional pooling layers, especially in datasets with complex structures, highlighting the importance of carefully designing pooling operations for optimal performance. Matteo Carollo, Mikel Ferrero-Jaurrieta, Rui Paiva 0001, Francesco Bardozzo, Roberto Tagliaferri |
Knowl. Based Syst. | 5 |
| 2026 | Exploring the efficiency of GPT models in automating trial balance mapping: applications and insightsabstractAbstract This study addresses the auditors need to categorize financial statement lines from an extensive collection of charts of accounts, and to automatize their mapping from natural language to the framework specified by the European Union’s Accounting Directive. Central to the study is the assembly and analysis of an extensive dataset comprising 945 charts of accounts, each representing the financial statements of a unique company with distinct accounting practices. The dataset’s variability reflects the diverse accounting methods employed by these companies.To address document heterogeneity, the research involves the creation and deployment of a specialized algorithm for data extraction and normalization. This algorithm systematically processes the natural language text of the charts of accounts, extracting relevant financial data and normalizing it to a consistent format suitable for analysis. Following normalization, the refined dataset is used to fine-tune OpenAI’s GPT-3 model. This deep learning natural language model is adapted to map natural language statements to the specific codes defined by the European Accounting Directive, ensuring each financial statement item is accurately categorized according to the regulatory framework. The mapping process involves linking trial balance figures to their corresponding financial statement lines, ensuring the ending balances in the trial balance align with the line items in the financial statements. The specialized algorithm enhances this process by automating the extraction and normalization steps, which traditionally require significant manual effort and are prone to human error.The novel application of GPT-3 in this context and the model’s ability to handle complex, hierarchical data structures inherent in financial statements further underscores its potential in transforming traditional auditing practices. Matteo Carollo, Lucia Lauri, Francesco Bardozzo, Emanuela Mattia Cafaro, Raffaele D'Alessio, Roberto Tagliaferri |
Soft Comput. | 6 |
| 2025 | RSDiX: Lightweight and Data-Efficient VLMs for Remote Sensing through Self-DistillationabstractRemote sensing (RS) imagery plays a pivotal role in various applications, and recent deep learning models integrate nuanced linguistic information to enhance semantic understanding. This work introduces RSDiX-CLIP, a fine-tuned CLIP model that addresses intra-class similarity in RS image datasets and improves data efficiency through the OTTER self-distillation framework. Additionally, we propose RSDiX-CLIPCap, a variant of the CLIPCap framework that incorporates a pre-trained RSDiX-CLIP. Our models outperform state-of-the-art methods on zero-shot RS image classification at various model scales and attain competitive RS image captioning results, while being smaller and more data-efficient than existing methods. We also explore the impact of mixed distillation strategies and alternative contrastive learning frameworks, introducing RSDiX-CLIP-S-BERT, employing a text-only model of the Sentence-BERT family as the teacher, and RSDiX-SigLIP, built on the SigLIP contrastive learning framework. We present a novel RS captioning dataset, S2LCD, consisting of 1533 Sentinel-2 images with 7665 wide-vocabulary, diverse and detailed captions. Finally, we challenge traditional N-gram-based captioning metrics such as the BLEU score, providing statistical evidence for the higher effectiveness of semantic scores like Sentence-BERT-Similarity. These advancements aim to contribute to the data-efficiency of deep learning models for RS image-text tasks, offering promising avenues for further exploration in the field. Code & Data at https://github.com/NeuRoNeLab/RSDiX-CLIP. Andrea Terlizzi, Angelo Nazzaro, Lorenzo Bernardi, Francesco Bardozzo, Roberto Tagliaferri |
IJCNN | 5 |
| 2024 | Elegans-AI: How the connectome of a living organism could model artificial neural networksabstractThis paper introduces Elegans-AI models, a class of neural networks that leverage the connectome topology of the Caenorhabditis elegans to design deep and reservoir architectures. Utilizing deep learning models inspired by the connectome, this paper leverages the evolutionary selection process to consolidate the functional arrangement of biological neurons within their networks. The initial goal involves the conversion of natural connectomes into artificial representations. The second objective centers on embedding the complex circuitry topology of artificial connectomes into both deep learning and deep reservoir networks, highlighting their neural-dynamic short-term and long-term memory and learning capabilities. Lastly, our third objective aims to establish structural explainability by examining the heterophilic/homophilic properties within the connectome and their impact on learning capabilities. In our study, the Elegans-AI models demonstrate superior performance compared to similar models that utilize either randomly rewired artificial connectomes or simulated bio-plausible ones. Notably, these Elegans-AI models achieve a top-1 accuracy of 99.99% on both Cifar10 and Cifar100, and 99.84% on MNIST Unsup. They do this with significantly fewer learning parameters, particularly when reservoir configurations of the connectome are used. Our findings indicate a clear connection between bio-plausible network patterns, the small-world characteristic, and learning outcomes, emphasizing the significant role of evolutionary optimization in shaping the topology of artificial neural networks for improved learning performance. Francesco Bardozzo, Andrea Terlizzi, Claudio Simoncini, Pietro Liò, Roberto Tagliaferri |
Neurocomputing | 5 |
| 2022 | Cross X-AI: Explainable Semantic Segmentation of Laparoscopic Images in Relation to Depth EstimationabstractIn this work, two deep learning models, trained to segment the liver and perform depth reconstruction, are compared and analysed with their post-hoc explanation interplay. The first model (a U-Net) is designed to perform liver semantic segmentation over different subjects and scenarios. Particularly, the image pixels representing the liver are classified and separated by the surrounding pixels. Meanwhile, with the second model, a depth estimation is performed to regress the z-position of each pixel (relative depths). In general, these two models apply a sort of classification task which can be explained for each model individually and that can be combined to show additional relations and insights between the most relevant learned features. In detail, this work shows how post-hoc explainable AI systems (X-AI) based on Grad CAM and Grad CAM++ can be compared by introducing Cross X-AI (CX-AI). Typically the post-hoc explanation maps provide different visual explanations of their decisions based on the two proposed approaches. Our results show that the Grad Cam++ segmentation explanation maps present cross-learning strategies similar to disparity explanations (and vice versa). Francesco Bardozzo, Mattia delli Priscoli, Toby Collins, Antonello Forgione, Alexandre Hostettler, Roberto Tagliaferri |
IJCNN | 6 |
| 2022 | StaSiS-Net: A stacked and siamese disparity estimation network for depth reconstruction in modern 3D laparoscopy
Francesco Bardozzo, Toby Collins, Antonello Forgione, Alexandre Hostettler, Roberto Tagliaferri |
Medical Image Anal. | 5 |
| 2022 | A comparison of deep learning models for end-to-end face-based video retrieval in unconstrained videosabstractAbstract Face-based video retrieval (FBVR) is the task of retrieving videos that containing the same face shown in the query image. In this article, we present the first end-to-end FBVR pipeline that is able to operate on large datasets of unconstrained, multi-shot, multi-person videos. We adapt an existing audiovisual recognition dataset to the task of FBVR and use it to evaluate our proposed pipeline. We compare a number of deep learning models for shot detection, face detection, and face feature extraction as part of our pipeline on a validation dataset made of more than 4000 videos. We obtain 97.25% mean average precision on an independent test set, composed of more than 1000 videos. The pipeline is able to extract features from videos at $$\sim $$ ∼ 7 times the real-time speed, and it is able to perform a query on thousands of videos in less than 0.5 s. Gioele Ciaparrone, Leonardo Chiariglione, Roberto Tagliaferri |
Neural Comput. Appl. | 3 |
| 2022 | Deep learning for volatility forecasting in asset managementabstractAbstract Predicting volatility is a critical activity for taking risk- adjusted decisions in asset trading and allocation. In order to provide effective decision-making support, in this paper we investigate the profitability of a deep Long Short-Term Memory (LSTM) Neural Network for forecasting daily stock market volatility using a panel of 28 assets representative of the Dow Jones Industrial Average index combined with the market factor proxied by the SPY and, separately, a panel of 92 assets belonging to the NASDAQ 100 index. The Dow Jones plus SPY data are from January 2002 to August 2008, while the NASDAQ 100 is from December 2012 to November 2017. If, on the one hand, we expect that this evolutionary behavior can be effectively captured adaptively through the use of Artificial Intelligence (AI) flexible methods, on the other, in this setting, standard parametric approaches could fail to provide optimal predictions. We compared the volatility forecasts generated by the LSTM approach to those obtained through use of widely recognized benchmarks models in this field, in particular, univariate parametric models such as the Realized Generalized Autoregressive Conditionally Heteroskedastic (R-GARCH) and the Glosten–Jagannathan–Runkle Multiplicative Error Models (GJR-MEM). The results demonstrate the superiority of the LSTM over the widely popular R-GARCH and GJR-MEM univariate parametric methods, when forecasting in condition of high volatility, while still producing comparable predictions for more tranquil periods. Alessio Petrozziello, Luigi Troiano, Angela Serra, Ivan Jordanov, Giuseppe Storti, Roberto Tagliaferri, Michele La Rocca 0001 |
Soft Comput. | 6 |
| 2021 | A review on drug repurposing applicable to COVID-19abstractDrug repurposing involves the identification of new applications for existing drugs at a lower cost and in a shorter time. There are different computational drug-repurposing strategies and some of these approaches have been applied to the coronavirus disease 2019 (COVID-19) pandemic. Computational drug-repositioning approaches applied to COVID-19 can be broadly categorized into (i) network-based models, (ii) structure-based approaches and (iii) artificial intelligence (AI) approaches. Network-based approaches are divided into two categories: network-based clustering approaches and network-based propagation approaches. Both of them allowed to annotate some important patterns, to identify proteins that are functionally associated with COVID-19 and to discover novel drug-disease or drug-target relationships useful for new therapies. Structure-based approaches allowed to identify small chemical compounds able to bind macromolecular targets to evaluate how a chemical compound can interact with the biological counterpart, trying to find new applications for existing drugs. AI-based networks appear, at the moment, less relevant since they need more data for their application. Serena Dotolo, Anna Marabotti, Angelo M. Facchiano, Roberto Tagliaferri |
Briefings Bioinform. | 4 |
| 2021 | A multiple network-based bioinformatics pipeline for the study of molecular mechanisms in oncological diseases for personalized medicineabstractMOTIVATION: Assessment of genetic mutations is an essential element in the modern era of personalized cancer treatment. Our strategy is focused on 'multiple network analysis' in which we try to improve cancer diagnostics by using biological networks. Genetic alterations in some important hubs or in driver genes such as BRAF and TP53 play a critical role in regulating many important molecular processes. Most of the studies are focused on the analysis of the effects of single mutations, while tumors often carry mutations of multiple driver genes. The aim of this work is to define an innovative bioinformatics pipeline focused on the design and analysis of networks (such as biomedical and molecular networks), in order to: (1) improve the disease diagnosis; (2) identify the patients that could better respond to a given drug treatment; and (3) predict what are the primary and secondary effects of gene mutations involved in human diseases. RESULTS: By using our pipeline based on a multiple network approach, it has been possible to demonstrate and validate what are the joint effects and changes of the molecular profile that occur in patients with metastatic colorectal carcinoma (mCRC) carrying mutations in multiple genes. In this way, we can identify the most suitable drugs for the therapy for the individual patient. This information is useful to improve precision medicine in cancer patients. As an application of our pipeline, the clinically significant case studies of a cohort of mCRC patients with the BRAF V600E-TP53 I195N missense combined mutation were considered. AVAILABILITY: The procedures used in this paper are part of the Cytoscape Core, available at (www.cytoscape.org). Data used here on mCRC patients have been published in [55]. SUPPLEMENTARY INFORMATION: A supplementary file containing a more detailed discussion of this case study and other cases is available at the journal site as Supplementary Data. Serena Dotolo, Anna Marabotti, Anna Maria Rachiglio, Riziero Esposito Abate, Marco Benedetto, Fortunato Ciardiello, Antonella De Luca, Nicola Normanno, Angelo M. Facchiano, Roberto Tagliaferri |
Briefings Bioinform. | 10 |
| 2021 | Signal metrics analysis of oscillatory patterns in bacterial multi-omic networksabstractMOTIVATION: One of the branches of Systems Biology is focused on a deep understanding of underlying regulatory networks through the analysis of the biomolecules oscillations and their interplay. Synthetic Biology exploits gene or/and protein regulatory networks towards the design of oscillatory networks for producing useful compounds. Therefore, at different levels of application and for different purposes, the study of biomolecular oscillations can lead to different clues about the mechanisms underlying living cells. It is known that network-level interactions involve more than one type of biomolecule as well as biological processes operating at multiple omic levels. Combining network/pathway-level information with genetic information it is possible to describe well-understood or unknown bacterial mechanisms and organism-specific dynamics. RESULTS: Following the methodologies used in signal processing and communication engineering, a methodology is introduced to identify and quantify the extent of multi-omic oscillations. These are due to the process of multi-omic integration and depend on the gene positions on the chromosome. Ad hoc signal metrics are designed to allow further biotechnological explanations and provide important clues about the oscillatory nature of the pathways and their regulatory circuits. Our algorithms designed for the analysis of multi-omic signals are tested and validated on 11 different bacteria for thousands of multi-omic signals perturbed at the network level by different experimental conditions. Information on the order of genes, codon usage, gene expression and protein molecular weight is integrated at three different functional levels. Oscillations show interesting evidence that network-level multi-omic signals present a synchronized response to perturbations and evolutionary relations along taxa. AVAILABILITY AND IMPLEMENTATION: The algorithms, the code (in language R), the tool, the pipeline and the whole dataset of multi-omic signal metrics are available at: https://github.com/lodeguns/Multi-omicSignals. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Francesco Bardozzo, Pietro Liò, Roberto Tagliaferri |
Bioinform. | 3 |
| 2020 | Motor strength classification with machine learning approaches applied to anatomical neuroimagesabstractPattern recognition methods for classification are leveraged in the field of computational anatomy and neuroimaging showing high reliability and applicability. Body-brain human functions related to the motor-strength features can be discovered by data integration and analysis of 3D brain images, phenotype and behavioural information. This work is focused on the study of feature-based interplay of 3D brain structures with motor-strength information. In particular, this research introduces an ensemble of supervised machine learning approaches for a binary motor-strength classification (strong vs weak) based on 3D brain anatomical features. The proposed approach has been evaluated on 1113 case studies by obtaining well-defined features and reaching the average accuracy of 72% on the test set. Francesco Bardozzo, Sebastian Cano Uribe, Andrea G. Russo, Mateo Jiménez Castaño, Mattia delli Priscoli, Fabrizio Esposito, Roberto Tagliaferri |
IJCNN | 7 |
| 2020 | A comparative analysis of multi-backbone Mask R-CNN for surgical tools detectionabstractReal-time surgical tool segmentation and tracking based on convolutional neural networks (CNN) has gained increasing interest in the field of mini-invasive surgery. In fact, the application of this novel artificial vision technologies allows both to reduce surgical risks and to increase patient safety. Moreover, these types of models can be used both to track the tools and detect markers or external artefacts in a real-time video stream. Multiple object detection and instance segmentation can be addressed efficiently by leveraging region-based CNN models. Thus, this work provides a comparison among state-of-the-art multi-backbone Mask R-CNNs to solve these tasks. Moreover, we show that such models can serve as a basis for tracking algorithms. The models were trained and tested with a data-set of 4955 manually annotated images, validated by 3 experts in the field. We tested 12 different combinations of CNN backbones and training hyperparameters. The results show that it is possible to employ a modern CNN to tackle the surgical tool detection problem, with the best-performing Mask R-CNN configuration achieving 87% Average Precision (AP) at Intersection over Union (IOU) 0.5. Gioele Ciaparrone, Francesco Bardozzo, Mattia delli Priscoli, Juanita Londoño Kallewaard, Maycol Ruiz Zuluaga, Roberto Tagliaferri |
IJCNN | 6 |
| 2020 | Deep learning in video multi-object tracking: A survey
Gioele Ciaparrone, Francisco Luque Sánchez, Siham Tabik, Luigi Troiano, Roberto Tagliaferri, Francisco Herrera |
Neurocomputing | 5 |
| 2019 | Strong-Weak Pruning for Brain Network Identification in Connectome-Wide Neuroimaging: Application to Amyotrophic Lateral Sclerosis Disease Stage CharacterizationabstractMagnetic resonance imaging allows acquiring functional and structural connectivity data from which high-density whole-brain networks can be derived to carry out connectome-wide analyses in normal and clinical populations. Graph theory has been widely applied to investigate the modular structure of brain connections by using centrality measures to identify the "hub" of human connectomes, and community detection methods to delineate subnetworks associated with diverse cognitive and sensorimotor functions. These analyses typically rely on a preprocessing step (pruning) to reduce computational complexity and remove the weakest edges that are most likely affected by experimental noise. However, weak links may contain relevant information about brain connectivity, therefore, the identification of the optimal trade-off between retained and discarded edges is a subject of active research. We introduce a pruning algorithm to identify edges that carry the highest information content. The algorithm selects both strong edges (i.e. edges belonging to shortest paths) and weak edges that are topologically relevant in weakly connected subnetworks. The newly developed "strong-weak" pruning (SWP) algorithm was validated on simulated networks that mimic the structure of human brain networks. It was then applied for the analysis of a real dataset of subjects affected by amyotrophic lateral sclerosis (ALS), both at the early (ALS2) and late (ALS3) stage of the disease, and of healthy control subjects. SWP preprocessing allowed identifying statistically significant differences in the path length of networks between patients and healthy subjects. ALS patients showed a decrease of connectivity between frontal cortex to temporal cortex and parietal cortex and between temporal and occipital cortex. Moreover, degree of centrality measures revealed significantly different hub and centrality scores between patient subgroups. These findings suggest a widespread alteration of network topology in ALS associated with disease progression. Angela Serra, Paola Galdi, Emanuele Pesce, Michele Fratello, Francesca Trojsi, Gioacchino Tedeschi, Roberto Tagliaferri, Fabrizio Esposito |
Int. J. Neural Syst. | 7 |
| 2018 | Robust clustering of noisy high-dimensional gene expression data for patients subtypingabstractMotivation: One of the most important research areas in personalized medicine is the discovery of disease sub-types with relevance in clinical applications. This is usually accomplished by exploring gene expression data with unsupervised clustering methodologies. Then, with the advent of multiple omics technologies, data integration methodologies have been further developed to obtain better performances in patient separability. However, these methods do not guarantee the survival separability of the patients in different clusters. Results: We propose a new methodology that first computes a robust and sparse correlation matrix of the genes, then decomposes it and projects the patient data onto the first m spectral components of the correlation matrix. After that, a robust and adaptive to noise clustering algorithm is applied. The clustering is set up to optimize the separation between survival curves estimated cluster-wise. The method is able to identify clusters that have different omics signatures and also statistically significant differences in survival time. The proposed methodology is tested on five cancer datasets downloaded from The Cancer Genome Atlas repository. The proposed method is compared with the Similarity Network Fusion (SNF) approach, and model based clustering based on Student's t-distribution (TMIX). Our method obtains a better performance in terms of survival separability, even if it uses a single gene expression view compared to the multi-view approach of the SNF method. Finally, a pathway based analysis is accomplished to highlight the biological processes that differentiate the obtained patient groups. Availability and implementation: Our R source code is available online at https://github.com/angy89/RobustClusteringPatientSubtyping. Supplementary information: Supplementary data are available at Bioinformatics online. Pietro Coretto, Angela Serra, Roberto Tagliaferri |
Bioinform. | 3 |
| 2018 | Robust and sparse correlation matrix estimation for the analysis of high-dimensional genomics dataabstractMotivation: Microarray technology can be used to study the expression of thousands of genes across a number of different experimental conditions, usually hundreds. The underlying principle is that genes sharing similar expression patterns, across different samples, can be part of the same co-expression system, or they may share the same biological functions. Groups of genes are usually identified based on cluster analysis. Clustering methods rely on the similarity matrix between genes. A common choice to measure similarity is to compute the sample correlation matrix. Dimensionality reduction is another popular data analysis task which is also based on covariance/correlation matrix estimates. Unfortunately, covariance/correlation matrix estimation suffers from the intrinsic noise present in high-dimensional data. Sources of noise are: sampling variations, presents of outlying sample units, and the fact that in most cases the number of units is much larger than the number of genes. Results: In this paper, we propose a robust correlation matrix estimator that is regularized based on adaptive thresholding. The resulting method jointly tames the effects of the high-dimensionality, and data contamination. Computations are easy to implement and do not require hand tunings. Both simulated and real data are analyzed. A Monte Carlo experiment shows that the proposed method is capable of remarkable performances. Our correlation metric is more robust to outliers compared with the existing alternatives in two gene expression datasets. It is also shown how the regularization allows to automatically detect and filter spurious correlations. The same regularization is also extended to other less robust correlation measures. Finally, we apply the ARACNE algorithm on the SyNTreN gene expression data. Sensitivity and specificity of the reconstructed network is compared with the gold standard. We show that ARACNE performs better when it takes the proposed correlation matrix estimator as input. Availability and implementation: The R software is available at https://github.com/angy89/RobustSparseCorrelation. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Angela Serra, Pietro Coretto, Michele Fratello, Roberto Tagliaferri |
Bioinform. | 4 |
| 2018 | A study on multi-omic oscillations in Escherichia coli metabolic networksabstractBACKGROUND: Two important challenges in the analysis of molecular biology information are data (multi-omic information) integration and the detection of patterns across large scale molecular networks and sequences. They are are actually coupled beause the integration of omic information may provide better means to detect multi-omic patterns that could reveal multi-scale or emerging properties at the phenotype levels. RESULTS: Here we address the problem of integrating various types of molecular information (a large collection of gene expression and sequence data, codon usage and protein abundances) to analyse the E.coli metabolic response to treatments at the whole network level. Our algorithm, MORA (Multi-omic relations adjacency) is able to detect patterns which may represent metabolic network motifs at pathway and supra pathway levels which could hint at some functional role. We provide a description and insights on the algorithm by testing it on a large database of responses to antibiotics. Along with the algorithm MORA, a novel model for the analysis of oscillating multi-omics has been proposed. Interestingly, the resulting analysis suggests that some motifs reveal recurring oscillating or position variation patterns on multi-omics metabolic networks. Our framework, implemented in R, provides effective and friendly means to design intervention scenarios on real data. By analysing how multi-omics data build up multi-scale phenotypes, the software allows to compare and test metabolic models, design new pathways or redesign existing metabolic pathways and validate in silico metabolic models using nearby species. CONCLUSIONS: The integration of multi-omic data reveals that E.coli multi-omic metabolic networks contain position dependent and recurring patterns which could provide clues of long range correlations in the bacterial genome. Francesco Bardozzo, Pietro Liò, Roberto Tagliaferri |
BMC Bioinform. | 3 |
| 2018 | Consensus-based feature extraction in rs-fMRI data analysis
Paola Galdi, Michele Fratello, Francesca Trojsi, Gioacchino Tedeschi, Roberto Tagliaferri, Fabrizio Esposito |
Soft Comput. | 6 |
| 2017 | E2FM: an encrypted and compressed full-text index for collections of genomic sequencesabstractMOTIVATION: Next Generation Sequencing (NGS) platforms and, more generally, high-throughput technologies are giving rise to an exponential growth in the size of nucleotide sequence databases. Moreover, many emerging applications of nucleotide datasets-as those related to personalized medicine-require the compliance with regulations about the storage and processing of sensitive data. RESULTS: We have designed and carefully engineered E 2 FM -index, a new full-text index in minute space which was optimized for compressing and encrypting nucleotide sequence collections in FASTA format and for performing fast pattern-search queries. E 2 FM -index allows to build self-indexes which occupy till to 1/20 of the storage required by the input FASTA file, thus permitting to save about 95% of storage when indexing collections of highly similar sequences; moreover, it can exactly search the built indexes for patterns in times ranging from few milliseconds to a few hundreds milliseconds, depending on pattern length. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/montecuollo/E2FM . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ferdinando Montecuollo, Giovanni Schmid, Roberto Tagliaferri |
Bioinform. | 3 |
| 2016 | Data integration in genomics and systems biologyabstractMulti-view learning is the branch of machine learning that deals with multi modal data, i.e. with patterns represented by different sets of features. The fast spread of this learning technique is motivated by the continuing increase of real applications based on multi-view data. For example, in bioinformatics multiple experiments can be available (mRNA, miRNA and protein expression, genome wide association studies (GWAS) and others) for a set of samples. In bioinformatics multi-view approaches are useful since heterogeneous genome-wide data sources capture information on different aspects of complex biological systems. Each view provides a distinct facet of the same domain, encoding different biologically-relevant patterns. The integration of such views can provide a richer model of the underlying system than those produced by a single view alone. This paper provides a review of the literature with respect to bioinformatics, with the purpose to understand the principles and operation modes of the existing methods and their possible applications. In order to organize the proposed methods in literature and to find similarities between them, these approaches are organized according to three categories: the type of data used in the papers, the statistical problem and the stage of integration. Angela Serra, Michele Fratello, Dario Greco, Roberto Tagliaferri |
CEC | 4 |
| 2016 | CONDOP: an R package for CONdition-Dependent Operon PredictionsabstractThe use of high-throughput RNA sequencing to predict dynamic operon structures in prokaryotic genomes has recently gained popularity in bioinformatics. We provide the R implementation of a novel method that uses transcriptomic features extracted from RNA-seq transcriptome profiles to develop ensemble classifiers for condition-dependent operon predictions. The CONDOP package provides a deeper insight into RNA-seq data analysis and allows scientists to highlight the operon organization in the context of transcriptional regulation with a few lines of code. AVAILABILITY AND IMPLEMENTATION: CONDOP is implemented in R and is freely available at CRAN. CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Vittorio Fortino, Roberto Tagliaferri, Dario Greco |
Bioinform. | 2 |
| 2015 | Multi omic oscillations in bacterial pathwaysabstractWe are often able to describe what happens at almost all biological length scales, from the molecular level to the whole organism; however, putting things together in order to obtain real comprehension is difficult and less developed. The challenge is to develop novel computational intelligence frameworks that integrate the different layers of molecular information. The introduction of such methodologies would enable to discover the relation between the environmental (external) conditions and the changes in the metabolic multi omic networks (i.e. the adaptive response of the internal environment). Many molecular levels can contribute to adaptability: pathways structure, codon usage bias, transcriptomics and metabolism. Here, we develop a method that combines Probabilistic Suffix Trees and Clustering models to integrate bacterial molecular information and to map different conditions to an omic multi-dimensional objective space. This methodology allows to identify oscillations in network structure of metabolic pathways. We tested our method by considering two case studies (from Escherichia coli) with antibiotics added to the medium. These oscillations provide insights into novel emerging properties of multi omic networks and to a better understanding of antibiotic effects and their gene targeting constraints. Francesco Bardozzo, Pietro Liò, Roberto Tagliaferri |
IJCNN | 3 |
| 2015 | Impact of different metrics on multi-view clusteringabstractClustering of patients allows to find groups of subjects with similar characteristics. This categorization can facilitate diagnosis, treatment decision and prognosis prediction. Heterogeneous genome-wide data sources capture different biological aspects that can be integrated in order to better categorize the patients. Clustering methods work by comparing how patients are similar or dissimilar in a suitable similarity space. While several clustering methods have been proposed, there is no systematic comparative study concerning the impact of similarity metrics on the cluster quality. We compared seven popular similarity measures (Pearson, Spearman and Kendall Correlations; Euclidean, Canberra, Minkowski and Manhattan Distances) in conjunction with two classical single-view clustering algorithms and a late integration approach (partitioning around medoids, hierarchical clustering and matrix factorization approaches), on high dimensional multi-view cancer data coming from the TCGA repository. Performance was measured against tumour subcategories classification. Only Euclidean and Minkowski distances showed similar results in terms of clustering similarity indexes. On the other hand, an absolute best similarity measure did not emerge in terms of misclassification, but it strongly depends on the data. Angela Serra, Dario Greco, Roberto Tagliaferri |
IJCNN | 3 |
| 2015 | A multi-view genomic data simulatorabstractBACKGROUND: OMICs technologies allow to assay the state of a large number of different features (e.g., mRNA expression, miRNA expression, copy number variation, DNA methylation, etc.) from the same samples. The objective of these experiments is usually to find a reduced set of significant features, which can be used to differentiate the conditions assayed. In terms of development of novel feature selection computational methods, this task is challenging for the lack of fully annotated biological datasets to be used for benchmarking. A possible way to tackle this problem is generating appropriate synthetic datasets, whose composition and behaviour are fully controlled and known a priori. RESULTS: Here we propose a novel method centred on the generation of networks of interactions among different biological molecules, especially involved in regulating gene expression. Synthetic datasets are obtained from ordinary differential equations based models with known parameters. Our results show that the generated datasets are well mimicking the behaviour of real data, for popular data analysis methods are able to selectively identify existing interactions. CONCLUSIONS: The proposed method can be used in conjunction to real biological datasets in the assessment of data mining techniques. The main strength of this method consists in the full control on the simulated data while retaining coherence with the real biological processes. The R package MVBioDataSim is freely available to the scientific community at http://neuronelab.unisa.it/?p=1722. Michele Fratello, Angela Serra, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco |
BMC Bioinform. | 5 |
| 2015 | MVDA: a multi-view genomic data integration methodologyabstractBACKGROUND: Multiple high-throughput molecular profiling by omics technologies can be collected for the same individuals. Combining these data, rather than exploiting them separately, can significantly increase the power of clinically relevant patients subclassifications. RESULTS: We propose a multi-view approach in which the information from different data layers (views) is integrated at the levels of the results of each single view clustering iterations. It works by factorizing the membership matrices in a late integration manner. We evaluated the effectiveness and the performance of our method on six multi-view cancer datasets. In all the cases, we found patient sub-classes with statistical significance, identifying novel sub-groups previously not emphasized in literature. Our method performed better as compared to other multi-view clustering algorithms and, unlike other existing methods, it is able to quantify the contribution of single views on the final results. CONCLUSION: Our observations suggest that integration of prior information with genomic features in the subtyping analysis is an effective strategy in identifying disease subgroups. The methodology is implemented in R and the source code is available online at http://neuronelab.unisa.it/a-multi-view-genomic-data-integration-methodology/ . Angela Serra, Michele Fratello, Vittorio Fortino, Giancarlo Raiconi, Roberto Tagliaferri, Dario Greco |
BMC Bioinform. | 5 |
| 2014 | Transcriptome dynamics-based operon prediction in prokaryotesabstractBACKGROUND: Inferring operon maps is crucial to understanding the regulatory networks of prokaryotic genomes. Recently, RNA-seq based transcriptome studies revealed that in many bacterial species the operon structure vary with the change of environmental conditions. Therefore, new computational solutions that use both static and dynamic data are necessary to create condition specific operon predictions. RESULTS: In this work, we propose a novel classification method that integrates RNA-seq based transcriptome profiles with genomic sequence features to accurately identify the operons that are expressed under a measured condition. The classifiers are trained on a small set of confirmed operons and then used to classify the remaining gene pairs of the organism studied. Finally, by linking consecutive gene pairs classified as operons, our computational approach produces condition-dependent operon maps. We evaluated our approach on various RNA-seq expression profiles of the bacteria Haemophilus somni, Porphyromonas gingivalis, Escherichia coli and Salmonella enterica. Our results demonstrate that, using features depending on both transcriptome dynamics and genome sequence characteristics, we can identify operon pairs with high accuracy. Moreover, the combination of DNA sequence and expression data results in more accurate predictions than each one alone. CONCLUSION: We present a computational strategy for the comprehensive analysis of condition-dependent operon maps in prokaryotes. Our method can be used to generate condition specific operon maps of many bacterial organisms for which high-resolution transcriptome data is available. Vittorio Fortino, Olli-Pekka Smolander, Petri Auvinen, Roberto Tagliaferri, Dario Greco |
BMC Bioinform. | 4 |
| 2013 | Bioinformatic pipelines in Python with leafabstractBACKGROUND: An incremental, loosely planned development approach is often used in bioinformatic studies when dealing with custom data analysis in a rapidly changing environment. Unfortunately, the lack of a rigorous software structuring can undermine the maintainability, communicability and replicability of the process. To ameliorate this problem we propose the Leaf system, the aim of which is to seamlessly introduce the pipeline formality on top of a dynamical development process with minimum overhead for the programmer, thus providing a simple layer of software structuring. RESULTS: Leaf includes a formal language for the definition of pipelines with code that can be transparently inserted into the user's Python code. Its syntax is designed to visually highlight dependencies in the pipeline structure it defines. While encouraging the developer to think in terms of bioinformatic pipelines, Leaf supports a number of automated features including data and session persistence, consistency checks between steps of the analysis, processing optimization and publication of the analytic protocol in the form of a hypertext. CONCLUSIONS: Leaf offers a powerful balance between plan-driven and change-driven development environments in the design, management and communication of bioinformatic pipelines. Its unique features make it a valuable alternative to other related tools. Francesco Napolitano, Renato Mariani-Costantini, Roberto Tagliaferri |
BMC Bioinform. | 3 |
| 2013 | An improved combinatorial biclustering algorithm
Ekaterina Nosova, Francesco Napolitano, Roberto Amato, Sergio Cocozza, Gennaro Miele, Giancarlo Raiconi, Roberto Tagliaferri |
Neural Comput. Appl. | 7 |
| 2011 | A multi-Biclustering Combinatorial Based algorithmabstractIn the last years a large amount of information about genomes was discovered, increasing the complexity of analysis. Therefore the most advanced techniques and algorithms are required. In many cases researchers use unsupervised clustering. But the inability of clustering to solve a number of tasks requires new algorithms. So, recently, scientists turned their attention to the biclustering techniques. In this paper we propose a novel biclustering technique, that we call Combinatorial Biclustering Algorithm (BCA). This technique permits to solve the following problems: 1) classification of data with respect to rows and columns together; 2) discovering of the overlapped biclusters; 3) definition of the minimal number of rows and columns in biclusters; 4) finding all biclusters together. We apply our model to two synthetic and one real biological data sets and show the results. Ekaterina Nosova, Giancarlo Raiconi, Roberto Tagliaferri |
CIDM | 3 |
| 2011 | Beyond classical consensus clustering: The least squares approach to multiple solutions
Loredana Murino, Claudia Angelini, Italia De Feis, Giancarlo Raiconi, Roberto Tagliaferri |
Pattern Recognit. Lett. | 5 |
| 2011 | Advances in Computational Intelligence and Bioinformatics
Francesco Masulli, Roberto Tagliaferri |
Soft Comput. | 2 |
| 2010 | Gene ontology fuzzy-enrichment analysis to investigate drug mode-of-actionabstractWe investigated the possibility of gaining information on the mode of action of a set of compounds by means of Gene Ontology (GO) enrichment analysis. To this aim, we developed a new method, based on fuzzy-sets, which is able to compute sets of genes that are consistently differentially expressed when treating cells with the analyzed compounds. Then a Gene Ontology enrichment analysis is performed on these sets. The method has been tested on several different groups of drugs, whose similarity in mode of action has been predicted by a gene-expression based, unsupervised, approach and verified by searching literature. The obtained results show that GO terms that are overrepresented in these fuzzy sets provide a quick and "easy-tointerpret" view of the mode of action of the analyzed drugs. Francesco Iorio, Loredana Murino, Diego di Bernardo, Giancarlo Raiconi, Roberto Tagliaferri |
IJCNN | 5 |
| 2010 | A scalable reference-point based algorithm to efficiently search large chemical databasesabstractHight-Throughput Screening (HTS) is a powerful tool in drug discovery, but very expensive in terms of required equipment and running costs. The virtual equivalent of HTS is molecular databases with the ability to search between millions of molecules by means of a similarity measure. In this work we propose a new class of bounds, algorithms and storage strategies based on the Intersection Inequality [5] for the Tanimoto Similarity to improve state of the art performances in querying large repositories of binary fingerprints. We focus on a special case that we call the β = B algorithm. The performance of the algorithm is assessed by simulating queries over an excerpt of the ChemDB [7]. We show how the average search can be up to 37% faster than using the Bit-Bound[4] alone, depending on the amount of space dedicated to data structures needed by the algorithm. Francesco Napolitano, Roberto Tagliaferri, Pierre Baldi |
IJCNN | 2 |
| 2009 | Global optimization, Meta Clustering and consensus clustering for class predictionabstractClustering of real-world data is often ill-posed. Because of noise and intrinsic ambiguity in data, optimization models attempting to maximize a fitness function can be misled by the assumption of uniqueness of the solution. In this work we present a methodology including classic and novel techniques to approach clustering in a systematic way, with two application examples to biological data sets. The methodology is based on a process that generates multiple clustering solutions (using global optimization), performs cluster analysis on such clusterings (i.e. meta clustering) and analyzes the obtained clusterings by the appropriate application of different consensus techniques. In order to validate the method, we seek for the solutions that best match the real class labels, exploiting only a random sample of them. Finally, we guess the class labels of the remaining patterns using cluster enrichment information and verify the percentage of correct assignments for each class. The optimization of clustering objective functions together with the use of partial labeling puts the described approach in between unsupervised and semi-supervised methods. Ida Bifulco, Carmine Fedullo, Francesco Napolitano, Giancarlo Raiconi, Roberto Tagliaferri |
IJCNN | 5 |
| 2009 | Computational intelligence and machine learning in bioinformatics
Giorgio Valentini, Roberto Tagliaferri, Francesco Masulli |
Artif. Intell. Medicine | 2 |
| 2008 | High-Throughput Analysis of the Drug Mode of Action of PB28, MC18 and MC70, Three Cyclohexylpiperazine Derivative New Molecules
Vitoantonio Bevilacqua, Paolo Pannarale, Giuseppe Mastronardi, Amalia Azzariti, Stefania Tommasi, Filippo Menolascina, Francesco Iorio, Diego di Bernardo, Angelo Paradiso, Nicola A. Colabufo, Francesco Berardi, Roberto Perrone, Roberto Tagliaferri |
ICIC (2) | 13 |
| 2008 | Robust Clustering by Aggregation and Intersection Methods
Ida Bifulco, Carmine Fedullo, Francesco Napolitano, Giancarlo Raiconi, Roberto Tagliaferri |
KES (3) | 5 |
| 2008 | Using Global Optimization to Explore Multiple Solutions of Clustering Problems
Ida Bifulco, Loredana Murino, Francesco Napolitano, Giancarlo Raiconi, Roberto Tagliaferri |
KES (3) | 5 |
| 2008 | Preface
Vito Di Gesù, Roberto Tagliaferri |
Int. J. Approx. Reason. | 2 |
| 2008 | Clustering and visualization approaches for human cell cycle gene expression data analysis
Francesco Napolitano, Giancarlo Raiconi, Roberto Tagliaferri, Angelo Ciaramella, Antonino Staiano, Gennaro Miele |
Int. J. Approx. Reason. | 3 |
| 2008 | Interactive data analysis and clustering of genomic data
Angelo Ciaramella, Sergio Cocozza, Francesco Iorio, Gennaro Miele, Francesco Napolitano, Michele Pinelli, Giancarlo Raiconi, Roberto Tagliaferri |
Neural Networks | 8 |
| 2007 | Clustering, Assessment and Validation: an application to gene expression dataabstractIn this work a multi-step approach for clustering assessment, visualization and data validation is introduced. Three main approaches for data clustering are used and compared: K-means, self organizing maps and probabilistic principal surfaces. A model explorer approach with different similarity measures is used to obtain the best parameters of the methods. The approach is used to identify genes periodically expressed in tumors related to the human cell cycle. Finally, clusters are validated by using GO term information. Angelo Ciaramella, Sergio Cocozza, Francesco Iorio, Gennaro Miele, Francesco Napolitano, Michele Pinelli, Giancarlo Raiconi, Roberto Tagliaferri |
IJCNN | 8 |
| 2006 | A multi-step approach to time series analysis and gene expression clusteringabstractMOTIVATION: The huge growth in gene expression data calls for the implementation of automatic tools for data processing and interpretation. RESULTS: We present a new and comprehensive machine learning data mining framework consisting in a non-linear PCA neural network for feature extraction, and probabilistic principal surfaces combined with an agglomerative approach based on Negentropy aimed at clustering gene microarray data. The method, which provides a user-friendly visualization interface, can work on noisy data with missing points and represents an automatic procedure to get, with no a priori assumptions, the number of clusters present in the data. Cell-cycle dataset and a detailed analysis confirm the biological nature of the most significant clusters. AVAILABILITY: The software described here is a subpackage part of the ASTRONEURAL package and is available upon request from the corresponding author. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Roberto Amato, Angelo Ciaramella, Natalia Deniskina, Carmine Del Mondo, Diego di Bernardo, Ciro Donalek, Giuseppe Longo, Giuseppe Mangano, Gennaro Miele, Giancarlo Raiconi, Antonino Staiano, Roberto Tagliaferri |
Bioinform. | 12 |
| 2006 | Fuzzy relational neural network
Angelo Ciaramella, Roberto Tagliaferri, Witold Pedrycz, Antonio Di Nola |
Int. J. Approx. Reason. | 2 |
| 2006 | Improving RBF networks performance in regression tasks by means of a supervised fuzzy clustering
Antonino Staiano, Roberto Tagliaferri, Witold Pedrycz |
Neurocomputing | 2 |
| 2006 | ICA based identification of dynamical systems generating synthetic and real world time series
Angelo Ciaramella, Enza De Lauro, Salvatore De Martino, Mariarosaria Falanga, Roberto Tagliaferri |
Soft Comput. | 5 |
| 2006 | Neural Network Techniques for Proactive Password CheckingabstractThis paper deals with the access control problem. We assume that valuable resources need to be protected against unauthorized users and that, to this aim, a password-based access control scheme is employed. Such an abstract scenario captures many applicative settings. The issue we focus our attention on is the following: password-based schemes provide a certain level of security as long as users choose good passwords, i.e., passwords that are hard to guess in a reasonable amount of time. In order to force the users to make good choices, a proactive password checker can be implemented as a submodule of the access control scheme. Such a checker, any time the user chooses/changes his own password, decides on the fly whether to accept or refuse the new password, depending on its guessability. Hence, the question is: how can we get an effective and efficient proactive password checker? By means of neural networks and statistical techniques, we answer the above question, developing suitable proactive password checkers. Through a series of experiments, we show that these checkers have very good performance: error rates are comparable to those of the best existing checkers, implemented on different principles and by using other methodologies, and the memory requirements are better in several cases. It is the first time that neural network technology has been fully and successfully applied to designing proactive password checkers Angelo Ciaramella, Paolo D'Arco, Alfredo De Santis, Clemente Galdi, Roberto Tagliaferri |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2005 | BSS Toolbox for delayed and convolved mixturesabstractIn this paper a Toolbox to generate and to analyze linear, non-linear, delayed and convolved mixtures of real source signals is presented. From one hand, in fact, a simple interface based on a physical model has been implemented and stereo, dolby, delayed, and convolved mixtures can be generated. On the other hand a blind source separation analysis can be accomplished. In fact, a novel separation algorithm (i.e. APDP) is included in the Toolbox. Several experiments to separate delayed mixtures of real instruments are made. Three musical recordings of different instrumental scores are mixed and analyzed by using the Toolbox. Several comparisons with known methods are made. Angelo Ciaramella, Roberto Tagliaferri, Francesco Iorio |
IJCNN | 2 |
| 2005 | Data visualization methodologies for data mining systems in bioinformaticsabstractBioinformatics systems benefit from the use of data mining strategies to locate interesting and pertinent relationships within massive information. For example, data mining methods can ascertain and summarize the set of genes responding to a certain level of stress in an organism. Even a cursory glance through the literature in journals, reveals the persistent role of data mining in experimental biology. Integrating data mining within the context of experimental investigations is central to bioinformatics software. In this paper we describe the framework of probabilistic principal surfaces, a latent variable model which offers a large variety of appealing visualization capabilities and which can be successfully integrated in the context of microarray analysis. A preprocessing phase consisting of a nonlinear PCA neural network which seems to be very useful to deal with noisy and time dependent nature of microarray data has been added to this framework. Antonino Staiano, Angelo Ciaramella, Giancarlo Raiconi, Roberto Tagliaferri, Roberto Amato, Giuseppe Longo, Gennaro Miele, Ciro Donalek |
IJCNN | 4 |
| 2005 | The genetic development of ordinal sums
Angelo Ciaramella, Roberto Tagliaferri, Witold Pedrycz |
Fuzzy Sets Syst. | 2 |
| 2005 | A novel information geometric approach to variable selection in MLP networks
Antonio Eleuteri, Roberto Tagliaferri, Leopoldo Milano |
Neural Networks | 2 |
| 2004 | Ordinal sums by using genetic algorithmsabstractA novel approach based on the ordinal sums and genetic algorithms is introduced. The main characteristic of an ordinal sum lies in the use of different t-norms (t-conorms) defined over disjoint subintervals of the unit interval. In this approach, a genetic optimisation environment to construct the ordinal sums and to optimise subintervals, and to allocate the individual local t-norms is introduced. Different parametric and non-parametric t-norms (t-conorms) are used. Several results to demonstrate the properties of the approach are proposed. The application of the genetically designed ordinal sums in case of the Zimmermann-Zysno logic operator data is also shown. Angelo Ciaramella, Roberto Tagliaferri, Witold Pedrycz |
FUZZ-IEEE | 2 |
| 2004 | Probabilistic Principal Surfaces for Yeast Gene Microarray Data MiningabstractThe recent technological advances are producing huge data sets in almost all fields of scientific research, from astronomy to genetics. Although each research field often requires ad-hoc, fine tuned, procedures to properly exploit all the available information inherently present in the data, there is an urgent need for a new generation of general computational theories and tools capable to boost most human activities of data analysis. Here, we propose probabilistic principal surfaces (PPS) as an effective high-D data visualization and clustering tool for data mining applications, emphasizing its flexibility and generality of use in data-rich field. In order to better illustrate the potentialities of the method, we also provide a real world case-study by discussing the use of PPS for the analysis of yeast gene expression levels from microarray chips. Antonino Staiano, Lara De Vinco, Angelo Ciaramella, Giancarlo Raiconi, Roberto Tagliaferri, Roberto Amato, Giuseppe Longo, Ciro Donalek, Gennaro Miele, Diego di Bernardo |
ICDM | 5 |
| 2004 | A hierarchical Bayesian learning scheme for autoregressive neural networks: application to the CATS benchmarkabstractIn this paper, a hierarchical Bayesian learning scheme for autoregressive neural network models is shown, which overcomes the problem of identifying the separate linear and nonlinear parts modeled by the network. We show how the identification can be carried out by defining suitable priors on the parameter space, which help the learning algorithms to avoid undesired parameter configurations. An application to synthetic data is shown and we apply the method to the CATS times series prediction benchmark. Fausto Acernese, Antonio Eleuteri, Leopoldo Milano, Roberto Tagliaferri |
IJCNN | 4 |
| 2004 | ICA for modelling and generating organ pipes self-sustained tonesabstractAcoustic signals emitted by organ pipes in a variety of experimental frameworks have been recorded and analyzed by using independent component analysis. Starting from this analysis, relevant features of the signals related to single tones of the chords have been extracted. Three Landau modes are extracted with three well defined frequencies. Following the dynamical systems approach, a simple and suitable analogical model, able to reproduce the registered waveform and sound in listening, have been constructed. The conclusion is that, in first approximation, the low dimensional dynamical system representing on average the fluid-dynamical equations modelling organ pipe is constituted by three linearly coupled nonlinear oscillators in limit cycle regime. Angelo Ciaramella, Enza De Lauro, Salvatore De Martino, Mariarosaria Falanga, Roberto Tagliaferri |
IJCNN | 5 |
| 2004 | An information geometric approach to survival analysis and feature selection by neural networksabstractAn information geometric approach to survival analysis is described. It is shown how a neural network can be used to model the probability of failure of a system, and how it can be trained by minimising a suitable divergence functional in a Bayesian framework. By using the trained network, minimisation of the same divergence functional allows for fast, efficient and exact feature selection. Finally, the performance of the algorithms is illustrated on a synthetic dataset. Antonio Eleuteri, Roberto Tagliaferri, Leopoldo Milano, Michele De Laurentiis |
IJCNN | 2 |
| 2004 | Committee of spherical probabilistic principal surfacesabstractProbabilistic Principal Surfaces is a promising latent variable model which represents a powerful tool to be used in a large range of data mining applications, due to its valuable capabilities in data visualization and classification tasks. In This work we focus our attention on the latter issue, proposing two combining schemes to build an ensemble of Probabilistic Principal Surfaces which is proved to be very effective in classifying very complex artificial and real-world astronomincal data. Antonino Staiano, Roberto Tagliaferri, Guiseppe Longon, Piero Benvenuti |
IJCNN | 2 |
| 2003 | A hierarchical Bayesian learning scheme for autoregressive neural networksabstractIn this paper a hierarchical Bayesian learning scheme for autoregressive neural network models is shown, which overcomes the problem of identifying the separate linear and nonlinear parts in the network. We show how the identification can be carried out by defining suitable priors on the parameter space, which help the learning algorithms to avoid undesired parameter configurations. Some applications to synthetic data are shown to validate the proposed methodology. Fausto Acernese, Fabrizio Barone, Rosario De Rosa, Antonio Eleuteri, Leopoldo Milano, Roberto Tagliaferri |
IJCNN | 6 |
| 2003 | Amplitude and permutation indeterminacies in frequency domain convolved ICAabstractIn this paper a novel approach to solve the permutation indeterminacy in the separation of convolved mixtures in frequency domain is proposed. A fixed-point algorithm in complex domain to perform the separation of the signals for each frequency domain is used. To obtain the frequency bins a short time Fourier transform on a set of fixed frames, is considered. To solve the ambiguity of the amplitude dilation a simple method is proposed. The permutation indeterminacy is solved using an approach based on the Hungarian algorithm that solves an assignment problem and an algorithm of dynamic programming. To obtain the distances in the assignment problem, a Kullback-Leibler divergence is adopted. We shall see that this approach presents a good performance and permits to obtain a clear separation of the signals. Angelo Ciaramella, Roberto Tagliaferri |
IJCNN | 2 |
| 2003 | Survival analysis and neural networksabstractA feedforward neural network architecture for survival analysis is presented which generalizes the standard, usually linear, models described in literature. The time variable is embedded in the model and the network is able to extract its interactions with other system features. The resulting model is described in a hierarchical Bayesian framework. Experiments with synthetic and real world data show a comparison of this model with the standard ones. Antonio Eleuteri, Roberto Tagliaferri, Leopoldo Milano, Gennaro Sansone, Diego D'Agostino, Sabino De Placido, Michele De Laurentiis |
IJCNN | 2 |
| 2003 | A novel neural network-based survival analysis model
Antonio Eleuteri, Roberto Tagliaferri, Leopoldo Milano, Sabino De Placido, Michele De Laurentiis |
Neural Networks | 2 |
| 2003 | Introduction: Neural networks for analysis of complex scientific data: astronomy and geosciences
Roberto Tagliaferri, Giuseppe Longo, Bruno D'Argenio, Alberto Incoronato |
Neural Networks | 1 |
| 2003 | Neural neZtworks in astronomy
Roberto Tagliaferri, Giuseppe Longo, Leopoldo Milano, Fausto Acernese, Fabrizio Barone, Angelo Ciaramella, Rosario De Rosa, Ciro Donalek, Antonio Eleuteri, Giancarlo Raiconi, Salvatore Sessa 0002, Antonino Staiano, Alfredo Volpicelli |
Neural Networks | 1 |
| 2003 | Neural networks for blind-source separation of Stromboli explosion quakesabstractIndependent component analysis (ICA) is used to analyze the seismic signals produced by explosions of the Stromboli volcano. It has been experimentally proved that it is possible to extract the most significant components from seismometer recorders. In particular, the signal, eventually thought as generated by the source, is corresponding to the higher power spectrum, isolated by our analysis. Furthermore, the amplitude of the source signals has been found by using a simple trick and so overcoming, for this specific case, the classical problem of ICA regarding the amplitude loss of the separated signals. Fausto Acernese, Angelo Ciaramella, Salvatore De Martino, Rosario De Rosa, Mariarosaria Falanga, Roberto Tagliaferri |
IEEE Trans. Neural Networks | 6 |
| 2001 | Fuzzy Relations Neural Network: Some Preliminary ResultsabstractIn this paper, a neuro-fuzzy model is introduced. The model describes a fuzzy relational "IF-THEN" reasoning scheme using an adaptive structure based on fuzzy relations. Two training schemes for the learning of the parameters based respectively on the backpropagation algorithm and pseudo-inverse matrix technique are illustrated. The model qualities are investigated by a series of simulation examples: function approximation, classification and rule extraction. These preliminary and promising results show that the model has a good performance and that it could be used for complex systems in real world applications. Angelo Ciaramella, Roberto Tagliaferri, Witold Pedrycz |
FUZZ-IEEE | 2 |
| 2001 | Fuzzy min-max neural networks: from classification to regression
Roberto Tagliaferri, Antonio Eleuteri, M. Meneganti, Fabrizio Barone |
Soft Comput. | 1 |
| 2000 | A Hierarchical Neural Network-Based Approach to VIRGO Noise IdentificationabstractIn this paper a hierarchical neural network-based approach is presented to identify the noise in the VIRGO experiment to detect gravitational waves by means of a laser interferometer. Fausto Acernese, Fabrizio Barone, Antonio Eleuteri, F. Garufi, Leopoldo Milano, Roberto Tagliaferri |
IJCNN (6) | 6 |
| 2000 | Merging fuzzy logic, neural networks, and genetic computation in the design of a decision-support systemabstractThe main goal of evolutionary computation is to provide a near optimal technique between exploration and exploitation of a search space. This approach is based on a genetic “engine” that operates the search of the optimal solution via biological-based assumptions. Selection of the optimal maintenance interventions activity, that can be tackled with success thanks to an evolutionary approach able to correct the distresses on the road pavement, is a very complex task. This paper presents an experimental architecture that improves the evolutionary aspect with additional benefits deriving from a synergistic combination of other powerful techniques, in particular neural networks and fuzzy logic. The best rules for managing pavement maintenance activities, developed through a genetic selection, are judged by a neural network. By an appropriate introduction of simple and efficient fuzzy identifiers, the features of the distress to treat can be described in an efficient and natural way. We describe the main advantages arising from this hybrid approach discussing the applicability of the method with experimental results. © 2000 John Wiley & Sons, Inc. Vincenzo Loia, Salvatore Sessa 0002, Antonino Staiano, Roberto Tagliaferri |
Int. J. Intell. Syst. | 4 |
| 2000 | Neural Networks for Sulphur Dioxide Ground Level Concentrations Forecasting
Massimo Andretta, Antonio Eleuteri, F. Fortezza, D. Manco, L. Mingozzi, Roberto Serra, Roberto Tagliaferri |
Neural Comput. Appl. | 7 |
| 1999 | Neural nets and star/galaxy separation in wide field astronomical imagesabstractOne of the most relevant problems in the extraction of scientifically useful information from wide field astronomical images (both photographic plates and CCD frames) is the recognition of the objects against a noisy background and their classification in unresolved (starlike) and resolved (galaxies) sources. In this paper we present a neural network based method capable to perform both tasks and discuss in detail the performance of object detection in a representative celestial field. The performance of our method is compared to that of other methodologies often used within the astronomical community. Stefano Andreon, Giorgio Gargiulo, Giuseppe Longo, Roberto Tagliaferri, Nicola Capuano |
IJCNN | 4 |
| 1999 | Hybrid neural networks for frequency estimation of unevenly sampled dataabstractWe present a hybrid system composed of a neural network based estimator system and genetic algorithms. It uses an unsupervised Hebbian nonlinear neural algorithm to extract the principal components which, in turn, are used by the MUSIC frequency estimator algorithm to extract the frequencies. We generalize this method to avoid an interpolation preprocessing step and to improve the performance by using a new stop criterion to avoid over fitting. Furthermore, genetic algorithms are used to optimize the neural net weight initialization. Roberto Tagliaferri, Angelo Ciaramella, Leopoldo Milano, Fabrizio Barone |
IJCNN | 1 |
| 1999 | Automated labeling for unsupervised neural networks: a hierarchical approachabstractIn this paper a hybrid system and a hierarchical neural-net approaches are proposed to solve the automatic labeling problem for unsupervised clustering. The first method consists in the application of nonneural clustering algorithms directly to the output of a neural net; the second one is based on a multilayer organization of neural units. Both methods are a substantial improvement with respect to the most important unsupervised neural algorithms existing in literature. Experimental results are shown to illustrate clustering performance of the systems. Roberto Tagliaferri, Nicola Capuano, Giorgio Gargiulo |
IEEE Trans. Neural Networks | 1 |
| 1998 | Fuzzy neural networks for classification and detection of anomaliesabstractIn this paper, a new learning algorithm for the Simpson's fuzzy min-max neural network is presented. It overcomes some undesired properties of the Simpson's model: specifically, in it there are neither thresholds that bound the dimension of the hyperboxes nor sensitivity parameters. Our new algorithm improves the network performance: in fact, the classification result does not depend on the presentation order of the patterns in the training set, and at each step, the classification error in the training set cannot increase. The new neural model is particularly useful in classification problems as it is shown by comparison with some fuzzy neural nets cited in literature (Simpson's min-max model, fuzzy ARTMAP proposed by Carpenter, Grossberg et al. in 1992, adaptive fuzzy systems as introduced by Wang in his book) and the classical multilayer perceptron neural network with backpropagation learning algorithm. The tests were executed on three different classification problems: the first one with two-dimensional synthetic data, the second one with realistic data generated by a simulator to find anomalies in the cooling system of a blast furnace, and the third one with real data for industrial diagnosis. The experiments were made following some recent evaluation criteria known in literature and by using Microsoft Visual C++ development environment on personal computers. M. Meneganti, F. S. Saviello, Roberto Tagliaferri |
IEEE Trans. Neural Networks | 3 |
| 1997 | Response to the letter by Li and Cao
Maria Marinaro, Salvatore Rampone, Roberto Tagliaferri |
Neural Networks | 3 |
| 1996 | Outline of a linear neural network
Eduardo R. Caianiello, Maria Marinaro, Salvatore Rampone, Roberto Tagliaferri |
Neurocomputing | 4 |
| 1994 | A neural network for error correcting decoding of binary linear codes
Anna Esposito, Salvatore Rampone, Roberto Tagliaferri |
Neural Networks | 3 |
| 1992 | Neural associative memories with minimum connectivity
Eduardo R. Caianiello, Antonella De Benedictis, Alfredo Petrosino, Roberto Tagliaferri |
Neural Networks | 4 |
| 1990 | Classes of efficiently computable linear neural nets
P. Occhinegro, Roberto Tagliaferri |
Neural Networks | 2 |