Annarita D'Addabbo

dblp:30/4798 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
5since 2021 · last 2024
0000-0002-3058-5340ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 3 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Detection of Olive Trees Affected by Xylella Fastidiosa from Hyperspectral and Thermal UAV Data
abstract
We report some results of an experiment to detect early occurrence of Xylella fastidiosa (Xf) in olive trees in the Apulia Region (southern Italy), performed in the framework of a project to assess the feasibility of a service addressed to agricultural authorities. An acquisition campaign was performed in September 2022, over a Xf-affected test area, using UAVborne hyperspectral and thermal sensors. Ground data were also collected through qPCR. Results of classification through SVM provide overall accuracy values ranging from 0.76 to 0.84.
Annarita D'Addabbo, Antonella Belmonte, Fabio Bovenga, Francesco P. Lovergine, Alberto Refice, Raffaella Matarrese, Antonia Gallo, Giovanni Mita, Raied Abou Kubaa, Donato Boscia, Vincenzo Barbieri
IGARSS1
2024 Vegetation Indices Time Series to Discriminate Olive Trees Cultivars
abstract
Remotely sensed imagery allows monitoring of olive tree’s cultivars with high spatial and temporal resolutions. Sensors on aircraft or UAVs can provide very high-resolution crop classification and health maps, but they are usually adopted for limited areas and low temporal frequency. The aim of this research is to analyze the temporal trend of some vegetation indices computed from satellite images on different olive cultivars in Apulia region (southern Italy) to highlight possible differences, and identify olive trees on huge areas with high temporal repeatability.Preliminary results, obtained by using Sentinel 2 data, highlight sufficient differences in the temporal trend of some Vegetation Indices among the olive cultivars investigated in Apulia region.
Raffaella Matarrese, Andrea Guerriero, Antonella Corcella, Nicola Lucarelli, Annarita D'Addabbo, Gaetano Alessandro Vivaldi
IGARSS5
2023 Automatic Detection of Xylella Fastidiosa in Aerial Hyperspectral and Thermal Data
abstract
Xylella fastidiosa (Xf) is a plant pathogen affecting olives trees, which has been identified as the bacterium responsible of a devastating landscape transformation in Apulia Region (Italy) from 2013. Actually, it has been found to affect 679 plant species worldwide, such as almond, vine and citrus.In this paper, experimental results concerning the automatic detection of trees infected by Xf from very high resolution hyperspectral and thermal images are shown. First of all, a set of vegetation indices and plant physiological traits related to rapid changes in photosynthetic pigments and leaf processes were computed from hyperspectral data. This information together with thermal data has been used as input to a RUSBoost classifier. Trees in training and test data set were labelled by performing quantitative real time-Polymerase-Chain-Reaction (qPCR) assays.Encouraging experimental results have been obtained, with Overall Accuracies greater than 90%, also when a reduced set of features is used as input for RUSBoost.
Annarita D'Addabbo, Antonella Belmonte, Fabio Bovenga, Francesco P. Lovergine, Alberto Refice, Raffaella Matarrese, Antonia Gallo, Giovanni Mita, Raied Abou Kubaa, Donato Boscia, Claudio La Mantia, Vincenzo Barbieri
IGARSS1
2022 Multi-Frequency Sar Data for Agriculture
abstract
The study aims to consolidate and validate a suite of Earth Observation algorithms of interest for applications in agriculture. The algorithms are at different levels of maturity. Still, they share the objective of contributing to sustainable water management and food security. They deal with monitoring the soil moisture, the vegetation water content, the extent of irrigated areas and the changes in the surface roughness of agricultural fields. The paper introduces the data sets, the algorithms and discusses some examples of initial results.
Francesco Mattia, Anna Balenzano, Giuseppe Satalino, Francesco P. Lovergine, Annarita D'Addabbo, Davide Palmisano, Riccardo Grassi, Francesco Nutini, Mirco Boschetti, Georgia Verza, Michele Rinaldi, Sergio Ruggieri, Angelo Pio De Santis, Vanessa Paredes Gómez, David Alfonso Nafría García, Deodato Tapete
IGARSS5
2022 Improving Flood Monitoring Through Advanced Modeling of Sentinel-1 Multi-Temporal Stacks
abstract
Multi-temporal remotely sensed data are a precious source of information for high spatial and temporal resolution flood mapping. We present a methodology for flood mapping through processing of long time series of Sentinel-l SAR data, as well as ancillary information. A Bayesian framework is adopted to derive probabilistic maps of the presence of flood waters, through modeling of backscatter time series, based on the as-sumption that floods represent impulsive temporal anomalies. We illustrate some results on a time series of Sentinel-l data acquired from 2015 to 2021 over a test area on the Basento river watershed, Basilicata Region, in Southern Italy, recurrently subject to floods.
Alberto Refice, Annarita D'Addabbo, Francesco P. Lovergine, Fabio Bovenga, Raffaele Nutricato, Davide Oscar Nitti
IGARSS2
2019 Improving Flood Detection in Vegetated Areas through Multi-Frequency, Polarimetric and Interferometric SAR Data
abstract
The Zambezi river basin, one of the world's largest flood-plains, located in south-eastern Africa, is recurrently subject to floods [1] . It has been the subject of several studies exploiting multi-temporal SAR data to monitor its hydrological cycle and periodic inundations, e.g. [2] .
Alberto Refice, Marco Chini, Marina Zingaro, Annarita D'Addabbo
IGARSS4
2018 An Open-Source Tool for the Integration of Remotely Sensed Information and Hydro-Geomorphic Parameters for Precise Monitoring of Inundations
abstract
Multi-sensor, multi-band and multi-temporal remote sensing data can be very useful in precise flood monitoring. In this paper, we describe DAFNE, a Matlab'v-based, open source toolbox, to produce flood maps from remotely sensed and other ancillary information, through a data fusion approach. DAFNE is based on Bayesian Networks, and is composed of several independent modules, each one performing a different task. Multi-temporal and multi-sensor data can be easily handled, with the possibility of producing time series of output flood maps, and thus follow the evolution of single or recurrent flood events. Here, an application of the toolbox is illustrated to delineate a flood map, close to the peak of inundation occurred in April 2015 on the Strymonas river (Greece), from multi-band optical and SAR data.
A. Rejice, Annarita D'Addabbo, Guido Pasquariello, Francesco P. Lovergine
IGARSS2
2016 SAR/optical data fusion for flood detection
abstract
In precision flood monitoring it is important to follow the temporal evolution of an event. Often, however, sufficient temporal coverage of events spanning several days can be attained only by recurring to multi-sensor data, due to different acquisition characteristics and schedules of different types of sensors. We present an example of a successful fusion of data coming from both SAR (COSMO-SkyMed stripmap, 3-m resolution) and optical (RapidEye, multispectral, 5 m-resolution) data, covering a flood event in southern Italy. The data fusion is performed through a Bayesian network approach, a reliable means to infer probabilistic information from heterogeneous sources. Results show accordance with independent model-based flood maps reaching accuracies of up to 96%.
Annarita D'Addabbo, Alberto Refice, Guido Pasquariello, Francesco P. Lovergine
IGARSS1
2016 A Bayesian Network for Flood Detection Combining SAR Imagery and Ancillary Data
abstract
Accurate flood mapping is important for both planning activities during emergencies and as a support for the successive assessment of damaged areas. A valuable information source for such a procedure can be remote sensing synthetic aperture radar (SAR) imagery. However, flood scenarios are typical examples of complex situations in which different factors have to be considered to provide accurate and robust interpretation of the situation on the ground. For this reason, a data fusion approach of remote sensing data with ancillary information can be particularly useful. In this paper, a Bayesian network is proposed to integrate remotely sensed data, such as multitemporal SAR intensity images and interferometric-SAR coherence data, with geomorphic and other ground information. The methodology is tested on a case study regarding a flood that occurred in the Basilicata region (Italy) on December 2013, monitored using a time series of COSMO-SkyMed data. It is shown that the synergetic use of different information layers can help to detect more precisely the areas affected by the flood, reducing false alarms and missed identifications which may affect algorithms based on data from a single source. The produced flood maps are compared to data obtained independently from the analysis of optical images; the comparison indicates that the proposed methodology is able to reliably follow the temporal evolution of the phenomenon, assigning high probability to areas most likely to be flooded, in spite of their heterogeneous temporal SAR/InSAR signatures, reaching accuracies of up to 89%.
Annarita D'Addabbo, Alberto Refice, Guido Pasquariello, Francesco P. Lovergine, Domenico Capolongo, Salvatore Manfreda
IEEE Trans. Geosci. Remote. Sens.1
2015 Towards high-precision flood mapping: Multi-temporal SAR/InSAR data, Bayesian inference, and hydrologic modeling
abstract
High-resolution flood mapping is an essential step in the monitoring and prevention of inundation hazard, both to gain insight into the processes involved in the generation of flooding events, and from the practical point of view of the precise assessment of inundated areas, useful e.g. in the case of post-event recovery and insurance indemnity assessments. Synthetic Aperture Radar (SAR) data present several favourable characteristics for flood mapping, such as their relative insensitivity to the meteorological conditions during acquisitions, thanks to the use of microwaves as sensing radiation, as well as the possibility of acquiring imagery independently of solar illumination, thanks to the active nature of the radar sensors. The Italian COSMO-SkyMed (CSK) SAR constellation is particularly useful in this respect, because it allows image sequences of flooding events to be built up with short revisit times. The acquisition of several images before, during and after the event often allow a reconstruction of the flooding dynamics. Moreover, they help in interpreting the backscatter signatures of different land cover types, reducing uncertainties about the actual presence of water, which can be seriously misleading, especially over agricultural areas [1, 2]. Finally, when acquisitions are made from the same geometry, with short repeat intervals, SAR interferometry (InSAR) observables, such as the coherence or the differential InSAR phase can be exploited as additional information layers. The favorable characteristics of these next-generation sensors have been exploited by a number of researchers worldwide [3, 4, 5] to improve performances of flood mapping approaches. Recently, our group [6, 2] has used high-resolution CSK radar images for flood mapping exploiting both the intensity and the interferometric coherence, with promising results. Nevertheless, additional information can be used to improve flood detection. In case of flooding, distance from the river, terrain elevation, hydrologic information or some combination of these data can add useful information that leads to a better performance in flood detection.
Alberto Refice, Annarita D'Addabbo, Guido Pasquariello, Francesco P. Lovergine, Domenico Capolongo, Salvatore Manfreda
IGARSS2
2015 Parallel selective sampling method for imbalanced and large data classification
abstract
Several applications aim to identify rare events from very large data sets. Classification algorithms may present great limitations on large data sets and show a performance degradation due to class imbalance. Many solutions have been presented in literature to deal with the problem of huge amount of data or imbalancing separately. In this paper we assessed the performances of a novel method, Parallel Selective Sampling (PSS), able to select data from the majority class to reduce imbalance in large data sets. PSS was combined with the Support Vector Machine (SVM) classification. PSS-SVM showed excellent performances on synthetic data sets, much better than SVM. Moreover, we showed that on real data sets PSS-SVM classifiers had performances slightly better than those of SVM and RUSBoost classifiers with reduced processing times. In fact, the proposed strategy was conceived and designed for parallel and distributed computing. In conclusion, PSS-SVM is a valuable alternative to SVM and RUSBoost for the problem of classification by huge and imbalanced data, due to its accurate statistical predictions and low computational complexity.
Annarita D'Addabbo, Rosalia Maglietta
Pattern Recognit. Lett.1
2014 A Bayesian network for flood detection
abstract
We apply a Bayesian Network (BN) paradigm to the problem of monitoring flood events through synthetic aperture radar (SAR) and interferometric SAR (InSAR) data. BNs are well-founded statistical tools which help formalizing the information coming from heterogeneous sources, such as remotely sensed images, LiDAR data, and topography. The approach is tested on the fluvial floodplains of the Basilicata region (southern Italy), which have been subject to recurrent flooding events in the last years. Results show maps efficiently representing the different scattering/coherence classes with high accuracy, and also allowing separating the multitemporal dimension of the data, where available. The BN approach proves thus helpful to gain insight into the complex phenomena related to floods, possibly also with respect to comparisons with modeling data.
Annarita D'Addabbo, Alberto Refice, Guido Pasquariello, Fabio Bovenga, Maria T. Chiaradia, Davide Oscar Nitti
IGARSS1
2013 SAR and InSAR for flood monitoring: Examples with COSMO/SkyMed data
abstract
We apply high-resolution, X-band, stripmap COSMO/SkyMed data to the monitoring of a flood event in Southern Basilicata region (Italy), where a multi-temporal dataset is available, allowing interferometric processing. We show how the use of the interferometric phase information can actually help to detect precisely the areas affected by the flood, using e.g. RGB composites of various information layers derived from the data. We also present results of unsupervised clustering of the multi-temporal data, which allow to shed some light on the physical interpretation of some of the identified clusters.
Alberto Refice, Domenico Capolongo, Annarita Lepera, Guido Pasquariello, Luca Pietranera, Fabio Volpec, Annarita D'Addabbo, Fabio Bovenga
IGARSS7
2010 On the reproducibility of results of pathway analysis in genome-wide expression studies of colorectal cancers
Rosalia Maglietta, Angela Distaso, Ada Piepoli, Orazio Palumbo, Massimo Carella, Annarita D'Addabbo, Sayan Mukherjee 0001, Nicola Ancona
J. Biomed. Informatics6
2009 Association of genetic profiles to Crohn's disease by linear combinations of single nucleotide polymorphisms
Annarita D'Addabbo, Anna Latiano, Orazio Palmieri, Teresa Maria Creanza, Rosalia Maglietta, Vito Annese, Nicola Ancona
Artif. Intell. Medicine1
2009 Comparative study of gene set enrichment methods
abstract
BACKGROUND: The analysis of high-throughput gene expression data with respect to sets of genes rather than individual genes has many advantages. A variety of methods have been developed for assessing the enrichment of sets of genes with respect to differential expression. In this paper we provide a comparative study of four of these methods: Fisher's exact test, Gene Set Enrichment Analysis (GSEA), Random-Sets (RS), and Gene List Analysis with Prediction Accuracy (GLAPA). The first three methods use associative statistics, while the fourth uses predictive statistics. We first compare all four methods on simulated data sets to verify that Fisher's exact test is markedly worse than the other three approaches. We then validate the other three methods on seven real data sets with known genetic perturbations and then compare the methods on two cancer data sets where our a priori knowledge is limited. RESULTS: The simulation study highlights that none of the three method outperforms all others consistently. GSEA and RS are able to detect weak signals of deregulation and they perform differently when genes in a gene set are both differentially up and down regulated. GLAPA is more conservative and large differences between the two phenotypes are required to allow the method to detect differential deregulation in gene sets. This is due to the fact that the enrichment statistic in GLAPA is prediction error which is a stronger criteria than classical two sample statistic as used in RS and GSEA. This was reflected in the analysis on real data sets as GSEA and RS were seen to be significant for particular gene sets while GLAPA was not, suggesting a small effect size. We find that the rank of gene set enrichment induced by GLAPA is more similar to RS than GSEA. More importantly, the rankings of the three methods share significant overlap. CONCLUSION: The three methods considered in our study recover relevant gene sets known to be deregulated in the experimental conditions and pathologies analyzed. There are differences between the three methods and GSEA seems to be more consistent in finding enriched gene sets, although no method uniformly dominates over all data sets. Our analysis highlights the deep difference existing between associative and predictive methods for detecting enrichment and the use of both to better interpret results of pathway analysis. We close with suggestions for users of gene set methods.
Luca Abatangelo, Rosalia Maglietta, Angela Distaso, Annarita D'Addabbo, Teresa Maria Creanza, Sayan Mukherjee 0001, Nicola Ancona
BMC Bioinform.4
2009 Statistical assessment of discriminative features for protein-coding and non coding cross-species conserved sequence elements
abstract
Abstract Background The identification of protein coding elements in sets of mammalian conserved elements is one of the major challenges in the current molecular biology research. Many features have been proposed for automatically distinguishing coding and non coding conserved sequences, making so necessary a systematic statistical assessment of their differences. A comprehensive study should be composed of an association study, i.e. a comparison of the distributions of the features in the two classes, and a prediction study in which the prediction accuracies of classifiers trained on single and groups of features are analyzed, conditionally to the compared species and to the sequence lengths. Results In this paper we compared distributions of a set of comparative and non comparative features and evaluated the prediction accuracy of classifiers trained for discriminating sequence elements conserved among human, mouse and rat species. The association study showed that the analyzed features are statistically different in the two classes. In order to study the influence of the sequence lengths on the feature performances, a predictive study was performed on different data sets composed of coding and non coding alignments in equal number and equally long with an ascending average length. We found that the most discriminant feature was a comparative measure indicating the proportion of synonymous nucleotide substitutions per synonymous sites. Moreover, linear discriminant classifiers trained by using comparative features in general outperformed classifiers based on intrinsic ones. Finally, the prediction accuracy of classifiers trained on comparative features increased significantly by adding intrinsic features to the set of input variables, independently on sequence length (Kolmogorov-Smirnov P-value ≤ 0.05). Conclusion We observed distinct and consistent patterns for individual and combined use of comparative and intrinsic classifiers, both with respect to different lengths of sequences/alignments and with respect to error rates in the classification of coding and non-coding elements. In particular, we noted that comparative features tend to be more accurate in the classification of coding sequences – this is likely related to the fact that such features capture deviations from strictly neutral evolution expected as a consequence of the characteristics of the genetic code.
Teresa Maria Creanza, David Stephen Horner, Annarita D'Addabbo, Rosalia Maglietta, Flavio Mignone, Nicola Ancona, Graziano Pesole
BMC Bioinform.3
2008 HT-RLS: High-Throughput Web Tool for Analysis of DNA Microarray Data Using RLS classifiers
abstract
Gene expression from DNA microarray data offers biologists and pathologists the possibility to deal with the problem of disease (e. g. cancer) diagnosis and prognosis from a quantitative point of view. Microarray data provide a snapshot of the molecular status of a sample of cells in a given tissue, returning the expression levels of thousands of genes simultaneously. Several mathematical methods from learning theory, such as Regularized Least Squares (RLS) classifiers or Support Vector Machines (SVM), have been extensively adopted to classify gene expression data. These methods can be useful to answer some relevant questions such as 1) what is the right amount of data to build an accurate classifier? 2) How many and which genes are correlated with a specific pathology? The computational analysis to statistically estimate the accuracy of the chosen models is particularly time consuming, burning several days of CPU time and without high-throughput or high- performance tools becomes practically unfeasible to obtain results in a reasonable time for biomedical community. We have implemented an independent, flexible and scalable platform, for a high-throughput large-scale microarray gene expression data analysis and classification, based on R tool for statistical computing. It integrates databases and computational intensive algorithms, based on RLS classifiers and a powerful web client for data training and graphical visualization of predicted results. Our platform provides statistically significant answers to the study of the gene expression by means of microarray data and supplying useful information to relevant questions in the diagnosis and prognosis of diseases in a reasonable time. The web resource is available free of charge for academic and non-profit institutions.
Paolo D'Onorio De Meo, Danilo Carrabino, Mattia D'Antonio, Annarita D'Addabbo, Sabino Liuni, Flavio Mignone, Graziano Pesole, Nicola Ancona
CCGRID4
2008 Prediction of Crohn's Disease by Profiles of Single Nucleotide Polymorphisms
Roberto Colella, Annarita D'Addabbo, Anna Latiano, Orazio Palmieri, Vito Annese, Nicola Ancona
KES (3)2
2008 SVD Based Feature Selection and Sample Classification of Proteomic Data
Annarita D'Addabbo, Massimo Papale, Salvatore Di Paolo, Simona Magaldi, Roberto Colella, Valentina d'Onofrio, Annamaria Di Palma, Elena Ranieri, Loreto Gesualdo, Nicola Ancona
KES (3)1
2008 Statistical Assessment of MSigDB Gene Sets in Colon Cancer
Angela Distaso, Luca Abatangelo, Rosalia Maglietta, Teresa Maria Creanza, Ada Piepoli, Massimo Carella, Annarita D'Addabbo, Sayan Mukherjee 0001, Nicola Ancona
KES (2)7
2007 Selection of relevant genes in cancer diagnosis based on their prediction accuracy
Rosalia Maglietta, Annarita D'Addabbo, Ada Piepoli, Francesco Perri, Sabino Liuni, Graziano Pesole, Nicola Ancona
Artif. Intell. Medicine2
2007 A composed supervised/unsupervised approach to improve change detection from remote sensing
L. Castellana, Annarita D'Addabbo, Guido Pasquariello
Pattern Recognit. Lett.2
2006 Classification error as a measure of gene relevance in cancer diagnosis
abstract
One of the main problems in cancer diagnosis by using DNA microarray data is selecting genes relevant for the pathology by analyzing their expression profiles in tissues in two different phenotypical conditions. The question we pose is the following: how do we measure the relevance of a single gene in a given pathology? A gene is relevant for a particular disease if it is possible to correctly predict the occurrence of the pathology in new patients on the basis of expression level of this gene only. In other words, a gene is informative for the disease if its expression levels are useful for training a classifier able to generalize, that is, able to correctly predict the status of new patients. In this paper we present a selection bias free, statistically well founded method for finding relevant genes on the basis of their classification ability. We applied the method on a colon cancer data set and produced a list of relevant genes, ranked on the basis of their prediction accuracy. We found, out of more than 6500 available genes, 54 overexpressed in normal tissue and 77 overexpressed in tumor tissue having prediction accuracy greater than 7 0 % with p-value p ≤ 0.05.
Rosalia Maglietta, Annarita D'Addabbo, Ada Piepoli, Francesco Perri, Sabino Liuni, Graziano Pesole, Nicola Ancona
IJCNN2
2006 On the statistical assessment of classifiers using DNA microarray data
abstract
BACKGROUND: In this paper we present a method for the statistical assessment of cancer predictors which make use of gene expression profiles. The methodology is applied to a new data set of microarray gene expression data collected in Casa Sollievo della Sofferenza Hospital, Foggia--Italy. The data set is made up of normal (22) and tumor (25) specimens extracted from 25 patients affected by colon cancer. We propose to give answers to some questions which are relevant for the automatic diagnosis of cancer such as: Is the size of the available data set sufficient to build accurate classifiers? What is the statistical significance of the associated error rates? In what ways can accuracy be considered dependant on the adopted classification scheme? How many genes are correlated with the pathology and how many are sufficient for an accurate colon cancer classification? The method we propose answers these questions whilst avoiding the potential pitfalls hidden in the analysis and interpretation of microarray data. RESULTS: We estimate the generalization error, evaluated through the Leave-K-Out Cross Validation error, for three different classification schemes by varying the number of training examples and the number of the genes used. The statistical significance of the error rate is measured by using a permutation test. We provide a statistical analysis in terms of the frequencies of the genes involved in the classification. Using the whole set of genes, we found that the Weighted Voting Algorithm (WVA) classifier learns the distinction between normal and tumor specimens with 25 training examples, providing e = 21% (p = 0.045) as an error rate. This remains constant even when the number of examples increases. Moreover, Regularized Least Squares (RLS) and Support Vector Machines (SVM) classifiers can learn with only 15 training examples, with an error rate of e = 19% (p = 0.035) and e = 18% (p = 0.037) respectively. Moreover, the error rate decreases as the training set size increases, reaching its best performances with 35 training examples. In this case, RLS and SVM have error rates of e = 14% (p = 0.027) and e = 11% (p = 0.019). Concerning the number of genes, we found about 6000 genes (p < 0.05) correlated with the pathology, resulting from the signal-to-noise statistic. Moreover the performances of RLS and SVM classifiers do not change when 74% of genes is used. They progressively reduce up to e = 16% (p < 0.05) when only 2 genes are employed. The biological relevance of a set of genes determined by our statistical analysis and the major roles they play in colorectal tumorigenesis is discussed. CONCLUSIONS: The method proposed provides statistically significant answers to precise questions relevant for the diagnosis and prognosis of cancer. We found that, with as few as 15 examples, it is possible to train statistically significant classifiers for colon cancer diagnosis. As for the definition of the number of genes sufficient for a reliable classification of colon cancer, our results suggest that it depends on the accuracy required.
Nicola Ancona, Rosalia Maglietta, Ada Piepoli, Annarita D'Addabbo, R. Cotugno, M. Savino, Sabino Liuni, Massimo Carella, Graziano Pesole, Francesco Perri
BMC Bioinform.4
2005 Regularized Least Squares Cancer Classifiers from DNA microarray data
abstract
BACKGROUND: The advent of the technology of DNA microarrays constitutes an epochal change in the classification and discovery of different types of cancer because the information provided by DNA microarrays allows an approach to the problem of cancer analysis from a quantitative rather than qualitative point of view. Cancer classification requires well founded mathematical methods which are able to predict the status of new specimens with high significance levels starting from a limited number of data. In this paper we assess the performances of Regularized Least Squares (RLS) classifiers, originally proposed in regularization theory, by comparing them with Support Vector Machines (SVM), the state-of-the-art supervised learning technique for cancer classification by DNA microarray data. The performances of both approaches have been also investigated with respect to the number of selected genes and different gene selection strategies. RESULTS: We show that RLS classifiers have performances comparable to those of SVM classifiers as the Leave-One-Out (LOO) error evaluated on three different data sets shows. The main advantage of RLS machines is that for solving a classification problem they use a linear system of order equal to either the number of features or the number of training examples. Moreover, RLS machines allow to get an exact measure of the LOO error with just one training. CONCLUSION: RLS classifiers are a valuable alternative to SVM classifiers for the problem of cancer classification by gene expression data, due to their simplicity and low computational complexity. Moreover, RLS classifiers show generalization ability comparable to the ones of SVM classifiers also in the case the classification of new specimens involves very few gene expression levels.
Nicola Ancona, Rosalia Maglietta, Annarita D'Addabbo, Sabino Liuni, Graziano Pesole
BMC Bioinform.3
2004 Three different unsupervised methods for change detection: an application
abstract
In this work, unsupervised change detection techniques, based on three different way to compare images, are presented. Two Landsat TM registered and corrected multi-spectral images, acquired on the same geographical area on 18 May 1996 and 21 May 1997, have been used. In the first comparison technique, for each pair of corresponding pixels, the spectral change vector has been computed as the squared difference in the features vectors at the two times. In the second method, the difference image has been computed using, pixel by pixel, a chi square transformation. The third technique is based on the application of a Self-Organizing Map (SOM) neural network to clusterize the two images before comparison. The three obtained difference images has been then analyzed by using a fully automatic thresholding method exploiting the expectation-maximization (EM) algorithm. The experimental results obtained for the three difference images are comparable, showing a reliable robustness of the unsupervised approach, and only few change are detected on the analyzed scene. Moreover, the experimental results have been compared with a change detection map computed by using a supervised technique, obtaining a good agreement between unsupervised and supervised results that confirms the reliability of the considered approach. The encouraging obtained results allow to use the so-computed percentage value of changes as probability of class transitions in input to a Bayesian supervised change detection method, as presented in a companion paper by the same authors. In this framework, the unsupervised approach may be used to support supervised techniques, providing land cover transitions that can be used as guess values
Annarita D'Addabbo, Giuseppe Satalino, Guido Pasquariello, Palma Blonda
IGARSS1
2003 Extraction of urban settlements by an automatic approach on high resolution remote sensed data
abstract
Photo interpretation by human experts has been the major data source for urban planning and monitoring applications. However, images collected from a space based sensor which combines reasonably good spectral and spatial resolution, could provide an useful tool for automatic monitoring of urban area changes. With this aim, in this paper, a data fusion technique, based on a RGB - HIS transformation, has been adopted for combining high spatial resolution panchromatic satellite with multi-spectral low resolution IKONOS II images. Moreover , textural information, which characterizes urban area, has been extracted with the use of a filter for the edge extraction. An MLP classifier has been trained to produce a labelled image with great accuracy in test even if a limited training set has been used. For a photo interpreter, the results reveal a good feasibility of the classified image for monitoring the presence of changes in urban areas, useful for a cartographic updating. different years. The percentage of correctly classified objects was 89%, whereas the percentage of correct object changes was equal to 83 %, with a false alarm rate of 5%. In (3) a parallelepiped supervised classification algorithm is used to obtain a land cover map characterized by seven classes from two pan-sharpened multi-spectral images at 1m resolution. Overall accuracy values of 75-83% are obtained on test data by using a PS-MS / RGB band composition and an RGB / NIR band composition, respectively. In (4) a neuro-fuzzy classifier based on a set of IF- Then-rules is compared with a Back Propagation neural network and a Maximum Likelihood (ML) classifier to produce a land cover map from IKONOS data on the Korean peninsula, characterized by mixed composition areas. The neuro-fuzzy classifier was more accurate than other classifiers on mixed composition areas, whereas the maximum likelihood performed better on areas such as roads. As input features, the authors considered the four multi-spectral band at low spatial resolution. The usefulness of IKONOS imagery for classification of urban and suburban scenes was also investigated in (5), where a hierarchical fuzzy classification techniques is proposed to improve the classification accuracy obtained by a ML traditional classifier when using only the MS bands of IKONOS data. The ML performance of about 81% was in fact increased to 88%, by using as input to the hierarchical classifier some textural features, extracted from the PAN band, beside the ML classified image. The objective of this work was twofold. First, to validate the feasibility of the RGB-HSI data fusion approach to assimilate the information derived from IKONOS high spatial / low spectral resolution data and the spectral information from low spatial / high spectral resolution. Second, to fully exploit the contextual information of IKONOS PAN data for the automatic classification of urban settlements. Two areas in Southern Italy, characterized by a different typical landscape, were selected for the study. The first is a country area with little villages, tourist facilities and isolated holiday houses
Cristina Tarantino, Annarita D'Addabbo, L. Castellana, Guido Pasquariello, Palma Blonda, Giuseppe Satalino
IGARSS2
2002 Neural network ensemble and support vector machine classifiers for the analysis of remotely sensed data: a comparison
abstract
This paper presents a comparative evaluation between a classification strategy based on the combination of the outputs of a neural (NN) ensemble and the application of Support Vector Machine (SVM) classifiers in the analysis of remotely sensed data. Two sets of experiments have been carried out on a benchmark data set. The first set concerns the application of linear and non linear techniques to the combination of the outputs of a Multilayer Perceptron (MLP) neural network ensemble. In particular, the Bayesian and the error correlation matrix approaches are used for coefficient selection in the linear combination of the network's outputs. A MLP module is used for the non linear outputs combination. The results of linear and non linear combination schemes are compared and discussed versus the performance of SVM classifiers. The comparative analysis evidences that the nonlinear, MLP based, combination provides the best results among the different combination schemes. On the other hand, better performance can be obtained with SVM classifiers. However, the complexity of the SVM training procedure can be considered a limitation for SVMs application to real-world problems.
Guido Pasquariello, Nicola Ancona, Palma Blonda, Cristina Tarantino, Giuseppe Satalino, Annarita D'Addabbo
IGARSS6
2001 Combination of Multiple Classifiers by Fuzzy Integrals: An Application to Synthetic Aperture Radar (SAR) Data
abstract
In this work, the results obtained in the classification of a multi-source - multi-temporal remote sensed data set by means of a distributed neuro-fuzzy system are compared with the results of a traditional centralized neural classification system, based on a single multilayer perceptron (MLP) neural network module. The distributed system is composed by a set of neural classifiers, whose partial results were combined with both Sugeno and Choquet fuzzy integrals. Two classification experiments were carried out with the distributed system. In the first experiment, each neural module of the distributed system used the same learning rule but was trained with a subset of the input features, i.e., a specific spectral band. In the second experiment, the neural modules of the system were trained with the same complete set of input features available for each training pixel, but consisted of MLP networks characterized by different specific topologies or different neural algorithms. The results show that larger improvements can be obtained by combining more independent classifiers. The Choquet fuzzy integral provided better performance than Sugeno fuzzy integral. The centralized system, based on a single MLP module, provided the best classification performance.
Palma Blonda, Cristina Tarantino, Annarita D'Addabbo, Giuseppe Satalino, Guido Pasquariello
FUZZ-IEEE3