Umberto Ferraro Petrillo

dblp:41/3936 · DBLP profile ↗
← Back
41ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0002-4308-5126ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 7 first-author · 7 since 2021Theory of computation · 7 · 5 first-authorSystems, architecture and hardware · 5 · 4 since 2021Security and privacy · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1
YearPublicationVenuePosition
2026 A scalable distributed workflow for accelerating long reads self-correction
abstract
Third-Generation Sequencing (TGS) technologies have transformed genomic research by enabling the extraction of longer nucleotide sequences (referred to as long reads ) and providing deeper insights into genome structure. However, long reads are often associated with high sequencing error rates, making their correction a major challenge in many computational genomic pipelines. In this paper, we introduce HyperC , a distributed workflow designed to accelerate the execution of existing long-read self-correction tools through a hybrid parallelization strategy. By combining MPI and OpenMP, our proposal efficiently scatters and executes tasks across a distributed computing system. Optimized input data handling further reduces I/O bottlenecks and maximizes resource utilization. To assess the effectiveness of HyperC , we integrated it with CONSENT , a high-performance correction module, and conducted extensive experiments on real-world sequencing datasets. The results show significant reductions in execution time and improved scalability compared to the standalone execution of CONSENT , establishing HyperC as a robust and practical solution for high-performance genomic analysis and population-scale studies.
Riccardo Ceccaroni, Lorenzo Di Rocco, Umberto Ferraro Petrillo, Pierpaolo Brutti
Future Gener. Comput. Syst.3
2026 Distributed compressive genomics: Fundamental pattern matching primitives via spark
abstract
• We develop distributed FM-Index and CBM algorithms for scalable compressed pattern matching on large genomic collections using Apache Spark. • Weconductathoroughperformanceevaluationbasedonstandardbenchmarksandmetricsusedindistributedalgorithm analysis. • We provide a publicly available software library to easily integrate distributed compressed pattern matching into genomic data processing pipelines without requiring a deep distributed programming expertise. Compressive genomics leverages compressed data representations to enhance the efficiency of bioinformatics tasks like sequence comparison and search. Surprisingly, the fundamental operation of pattern matching on large DNA sequence collections remains unexplored in the realm of genomic analysis. However, distributed systems like Spark offer the scalability necessary to process increasingly large genomic datasets efficiently. We present the first Spark-based implementation of the FM-Index and Compressed Boyer-Moore (CBM) algorithms, evaluating their performance and providing insights into their advantages for large-scale bioinformatics applications. A comprehensive experimental study demonstrates clear performance gains over uncompressed approaches. Furthermore, we introduce SparkGeco , a distributed compressive genomics software library designed to simplify the integration of FM-Index and CBM algorithms into DNA sequence analysis pipelines within Apache Spark, thus supporting the development of efficient and scalable genomic analysis workflows. This work provides a concrete step towards high-performance, data-centric eScience solutions in computational biology.
Lorenzo Di Rocco, Umberto Ferraro Petrillo, Raffaele Giancarlo, Giuseppe Cattaneo
Future Gener. Comput. Syst.2
2025 Efficient and Scalable Alignment-Free Distributed Genotyping of SNPs and Short Indels
abstract
The growing volume of sequencing data and the ever-larger size of variants databases challenge genotyping procedures to handle massive genomics datasets efficiently. Recent alignment-free solutions leverage exclusively on the k-mers counts to speed up the analysis, but have to trade off the time gain against the memory requirements, to make the elaborations possible on a single workstation. In this paper, we present SparkGeno+, a novel alignment-free (AF) distributed pipeline for the fast and accurate genotyping of Single Nucleotide Polymorphisms (SNPs) and indels on a large scale. Starting from a previous pipeline, we identified and evaluated the performance bottlenecks that arise when performing genotyping using a standard AF approach, to develop and implement several innovations to better exploit the resources of a distributed system. The effectiveness of our proposal has been validated through an experimental analysis on widely studied datasets. The results show that the accuracy of SparkGeno+ matches the one of state-of-the-art alignment-free tools like Vargeno and MALVA. Moreover, the time performance of SparkGeno+ scales well with the number of computing units, thus allowing execution times that are in order of growth smaller than those of classical genotyping tools. This indicates SparkGeno+ to be a promising solution for large-scale genotyping applications.
Lorenzo Di Rocco, Umberto Ferraro Petrillo
IEEE Trans. Comput. Biol. Bioinform.2
2024 A distributed approach for persistent homology computation on a large scale
abstract
Abstract Persistent homology (PH) is a powerful mathematical method to automatically extract relevant insights from images, such as those obtained by high-resolution imaging devices like electron microscopes or new-generation telescopes. However, the application of this method comes at a very high computational cost that is bound to explode more because new imaging devices generate an ever-growing amount of data. In this paper, we present PixHomology , a novel algorithm for efficiently computing zero-dimensional PH on images, optimizing memory and processing time. By leveraging the Apache Spark framework, we also present a distributed version of our algorithm with several optimized variants, able to concurrently process large batches of astronomical images. Finally, we present the results of an experimental analysis showing that our algorithm and its distributed version are efficient in terms of required memory, execution time, and scalability, consistently outperforming existing state-of-the-art PH computation tools when used to process large datasets.
Riccardo Ceccaroni, Lorenzo Di Rocco, Umberto Ferraro Petrillo, Pierpaolo Brutti
J. Supercomput.3
2023 A novel image dataset for source camera identification and image based recognition systems
abstract
Abstract Multimodal emotion recognition has attracted a great deal of attention in recent years, with new interesting applications now being considered. One promising application is in the digital image forensics fields where, for example, it gives the possibility to automatically highlight subjects that are in pain, in digital images under examination, by analyzing their facial expressions. However, finding an image that represents a possible crime leaves the problem of identifying the device used to take the image open. Such a problem has been addressed by Source Camera Identification algorithms (SCI, for short). These algorithms analyze some features hidden in a target image to find traces left by the sensor that captured the image. A particularly challenging case is when the candidate source cameras for an image under investigation are of the same manufacturer and model. A fair and universal assessment of these algorithms is only possible if standard datasets are used for their benchmarking. However, our comprehensive analysis has shown that the majority of the datasets proposed so far contain a collection of images taken with different types of cameras, mostly smartphones. We fill this gap by presenting UNISA2020, a novel image dataset that contains a large collection of real-world images taken with multiple conventional digital cameras of the same type. The images in our dataset have been assembled so as to avoid artifacts that could negatively affect the identification process. To validate our dataset, we also performed a comparative experimental analysis to investigate the performance of an SCI reference algorithm when running on our dataset as well as on other SCI standard datasets.
Andrea Bruno, Paola Capasso, Giuseppe Cattaneo, Umberto Ferraro Petrillo, Riccardo Improta
Multim. Tools Appl.4
2023 Ten quick tips for bioinformatics analyses using an Apache Spark distributed computing environment
abstract
Some scientific studies involve huge amounts of bioinformatics data that cannot be analyzed on personal computers usually employed by researchers for day-to-day activities but rather necessitate effective computational infrastructures that can work in a distributed way. For this purpose, distributed computing systems have become useful tools to analyze large amounts of bioinformatics data and to generate relevant results on virtual environments, where software can be executed for hours or even days without affecting the personal computer or laptop of a researcher. Even if distributed computing resources have become pivotal in multiple bioinformatics laboratories, often researchers and students use them in the wrong ways, making mistakes that can cause the distributed computers to underperform or that can even generate wrong outcomes. In this context, we present here ten quick tips for the usage of Apache Spark distributed computing systems for bioinformatics analyses: ten simple guidelines that, if taken into account, can help users avoid common mistakes and can help them run their bioinformatics analyses smoothly. Even if we designed our recommendations for beginners and students, they should be followed by experts too. We think our quick tips can help anyone make use of Apache Spark distributed computing systems more efficiently and ultimately help generate better, more reliable scientific results.
Davide Chicco, Umberto Ferraro Petrillo, Giuseppe Cattaneo
PLoS Comput. Biol.2
2023 Using software visualization to support the teaching of distributed programming
abstract
Abstract In this paper, we introduce MARVEL, a system designed to simplify the teaching of MapReduce, a popular distributed programming paradigm, through software visualization. At its core, it allows a teacher to describe and recreate a MapReduce application by interactively requesting, through a graphical interface, the execution of a sequence of MapReduce transformations that target an input dataset. Then, the execution of each operation is illustrated on the screen by playing an appropriate graphical animation stage, highlighting aspects related to its distributed nature. The sequence of all animation stages, played back one after the other in a sequential order, results in a visualization of the whole algorithm. The content of the resulting visualization is not simulated or fictitious, but reflects the real behavior of the requested operations, thanks to the adoption of an architecture based on a real instance of a distributed system running on Apache Spark. On the teacher’s side, it is expected that by using MARVEL he/she will spend less time preparing materials and will be able to design a more interactive lesson than with electronic slides or a whiteboard. To test the effectiveness of the proposed approach on the learner side, we also conducted a small scientific experiment with a class of volunteer students who formed a control group. The results are encouraging, showing that the use of software visualization guarantees students a learning experience at least equivalent to that of conventional approaches.
Lorenzo Di Rocco, Umberto Ferraro Petrillo, Francesco Palini
J. Supercomput.2
2022 The power of word-frequency-based alignment-free functions: a comprehensive large-scale experimental analysis
abstract
MOTIVATION: Alignment-free (AF) distance/similarity functions are a key tool for sequence analysis. Experimental studies on real datasets abound and, to some extent, there are also studies regarding their control of false positive rate (Type I error). However, assessment of their power, i.e. their ability to identify true similarity, has been limited to some members of the D2 family. The corresponding experimental studies have concentrated on short sequences, a scenario no longer adequate for current applications, where sequence lengths may vary considerably. Such a State of the Art is methodologically problematic, since information regarding a key feature such as power is either missing or limited. RESULTS: By concentrating on a representative set of word-frequency-based AF functions, we perform the first coherent and uniform evaluation of the power, involving also Type I error for completeness. Two alternative models of important genomic features (CIS Regulatory Modules and Horizontal Gene Transfer), a wide range of sequence lengths from a few thousand to millions, and different values of k have been used. As a result, we provide a characterization of those AF functions that is novel and informative. Indeed, we identify weak and strong points of each function considered, which may be used as a guide to choose one for analysis tasks. Remarkably, of the 15 functions that we have considered, only four stand out, with small differences between small and short sequence length scenarios. Finally, to encourage the use of our methodology for validation of future AF functions, the Big Data platform supporting it is public. AVAILABILITY AND IMPLEMENTATION: The software is available at: https://github.com/pipp8/power_statistics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Giuseppe Cattaneo, Umberto Ferraro Petrillo, Raffaele Giancarlo, Francesco Palini, Chiara Romualdi
Bioinform.2
2022 Correction to: FASTA/Q data compressors for MapReduce-Hadoop genomics: space and time savings made easy
abstract
Following publication of the original article [1], the authors identified that the affiliations of Giuseppe Cattaneo and Raffaele Giancarlo were interchanged. The correct affiliations are given below. The correct affiliation of Giuseppe Cattaneo is: 2Dipartimento di Informatica, Università di Salerno, Fisciano, Italy. The correct affiliation of Raffaele Giancarlo is: 3Dipartimento di Matematica ed Informatica, Università di Palermo, Palermo, Italy. The original article [1] has been corrected.
Umberto Ferraro Petrillo, Francesco Palini, Giuseppe Cattaneo, Raffaele Giancarlo
BMC Bioinform.1
2022 DIAMIN: a software library for the distributed analysis of large-scale molecular interaction networks
abstract
BACKGROUND: Huge amounts of molecular interaction data are continuously produced and stored in public databases. Although many bioinformatics tools have been proposed in the literature for their analysis, based on their modeling through different types of biological networks, several problems still remain unsolved when the problem turns on a large scale. RESULTS: We propose DIAMIN, that is, a high-level software library to facilitate the development of applications for the efficient analysis of large-scale molecular interaction networks. DIAMIN relies on distributed computing, and it is implemented in Java upon the framework Apache Spark. It delivers a set of functionalities implementing different tasks on an abstract representation of very large graphs, providing a built-in support for methods and algorithms commonly used to analyze these networks. DIAMIN has been tested on data retrieved from two of the most used molecular interactions databases, resulting to be highly efficient and scalable. As shown by different provided examples, DIAMIN can be exploited by users without any distributed programming experience, in order to perform various types of data analysis, and to implement new algorithms based on its primitives. CONCLUSIONS: The proposed DIAMIN has been proved to be successful in allowing users to solve specific biological problems that can be modeled relying on biological networks, by using its functionalities. The software is freely available and this will hopefully allow its rapid diffusion through the scientific community, to solve both specific data analysis and more complex tasks.
Lorenzo Di Rocco, Umberto Ferraro Petrillo, Simona E. Rombo
BMC Bioinform.2
2021 Alignment-free Genomic Analysis via a Big Data Spark Platform
abstract
MOTIVATION: Alignment-free distance and similarity functions (AF functions, for short) are a well-established alternative to pairwise and multiple sequence alignments for many genomic, metagenomic and epigenomic tasks. Due to data-intensive applications, the computation of AF functions is a Big Data problem, with the recent literature indicating that the development of fast and scalable algorithms computing AF functions is a high-priority task. Somewhat surprisingly, despite the increasing popularity of Big Data technologies in computational biology, the development of a Big Data platform for those tasks has not been pursued, possibly due to its complexity. RESULTS: We fill this important gap by introducing FADE, the first extensible, efficient and scalable Spark platform for alignment-free genomic analysis. It supports natively eighteen of the best performing AF functions coming out of a recent hallmark benchmarking study. FADE development and potential impact comprises novel aspects of interest. Namely, (i) a considerable effort of distributed algorithms, the most tangible result being a much faster execution time of reference methods like MASH and FSWM; (ii) a software design that makes FADE user-friendly and easily extendable by Spark non-specialists; (iii) its ability to support data- and compute-intensive tasks. About this, we provide a novel and much needed analysis of how informative and robust AF functions are, in terms of the statistical significance of their output. Our findings naturally extend the ones of the highly regarded benchmarking study, since the functions that can really be used are reduced to a handful of the eighteen included in FADE. AVAILABILITYAND IMPLEMENTATION: The software and the datasets are available at https://github.com/fpalini/fade. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Umberto Ferraro Petrillo, Francesco Palini, Giuseppe Cattaneo, Raffaele Giancarlo
Bioinform.1
2021 FASTA/Q data compressors for MapReduce-Hadoop genomics: space and time savings made easy
abstract
BACKGROUND: Storage of genomic data is a major cost for the Life Sciences, effectively addressed via specialized data compression methods. For the same reasons of abundance in data production, the use of Big Data technologies is seen as the future for genomic data storage and processing, with MapReduce-Hadoop as leaders. Somewhat surprisingly, none of the specialized FASTA/Q compressors is available within Hadoop. Indeed, their deployment there is not exactly immediate. Such a State of the Art is problematic. RESULTS: We provide major advances in two different directions. Methodologically, we propose two general methods, with the corresponding software, that make very easy to deploy a specialized FASTA/Q compressor within MapReduce-Hadoop for processing files stored on the distributed Hadoop File System, with very little knowledge of Hadoop. Practically, we provide evidence that the deployment of those specialized compressors within Hadoop, not available so far, results in better space savings, and even in better execution times over compressed data, with respect to the use of generic compressors available in Hadoop, in particular for FASTQ files. Finally, we observe that these results hold also for the Apache Spark framework, when used to process FASTA/Q files stored on the Hadoop File System. CONCLUSIONS: Our Methods and the corresponding software substantially contribute to achieve space and time savings for the storage and processing of FASTA/Q files in Hadoop and Spark. Being our approach general, it is very likely that it can be applied also to FASTA/Q compression methods that will appear in the future. AVAILABILITY: The software and the datasets are available at https://github.com/fpalini/fastdoopc.
Umberto Ferraro Petrillo, Francesco Palini, Giuseppe Cattaneo, Raffaele Giancarlo
BMC Bioinform.1
2021 PNU Spoofing: a menace for biometrics authentication systems?
Andrea Bruno, Giuseppe Cattaneo, Umberto Ferraro Petrillo, Paola Capasso
Pattern Recognit. Lett.3
2019 Analyzing big datasets of genomic sequences: fast and scalable collection of k-mer statistics
abstract
BACKGROUND: Distributed approaches based on the MapReduce programming paradigm have started to be proposed in the Bioinformatics domain, due to the large amount of data produced by the next-generation sequencing techniques. However, the use of MapReduce and related Big Data technologies and frameworks (e.g., Apache Hadoop and Spark) does not necessarily produce satisfactory results, in terms of both efficiency and effectiveness. We discuss how the development of distributed and Big Data management technologies has affected the analysis of large datasets of biological sequences. Moreover, we show how the choice of different parameter configurations and the careful engineering of the software with respect to the specific framework under consideration may be crucial in order to achieve good performance, especially on very large amounts of data. We choose k-mers counting as a case study for our analysis, and Spark as the framework to implement FastKmer, a novel approach for the extraction of k-mer statistics from large collection of biological sequences, with arbitrary values of k. RESULTS: One of the most relevant contributions of FastKmer is the introduction of a module for balancing the statistics aggregation workload over the nodes of a computing cluster, in order to overcome data skew while allowing for a full exploitation of the underlying distributed architecture. We also present the results of a comparative experimental analysis showing that our approach is currently the fastest among the ones based on Big Data technologies, while exhibiting a very good scalability. CONCLUSIONS: We provide evidence that the usage of technologies such as Hadoop or Spark for the analysis of big datasets of biological sequences is productive only if the architectural details and the peculiar aspects of the considered framework are carefully taken into account for the algorithm design and implementation.
Umberto Ferraro Petrillo, Mara Sorella, Giuseppe Cattaneo, Raffaele Giancarlo, Simona E. Rombo
BMC Bioinform.1
2019 Achieving efficient source camera identification on Hadoop
Giuseppe Cattaneo, Umberto Ferraro Petrillo, Andrea F. Abate, Fabio Narducci, Silvio Barra
Multim. Tools Appl.2
2018 Partial Parallel Error Correction of Pair-end Short Reads using GPU
Umberto Ferraro Petrillo, Francesco Palini
BIBM1
2018 Using Software Visualization for Supporting the Teaching of MapReduce
Umberto Ferraro Petrillo
NSS1
2018 Informational and linguistic analysis of large genomic sequence collections via efficient Hadoop cluster algorithms
abstract
Motivation: Information theoretic and compositional/linguistic analysis of genomes have a central role in bioinformatics, even more so since the associated methodologies are becoming very valuable also for epigenomic and meta-genomic studies. The kernel of those methods is based on the collection of k-mer statistics, i.e. how many times each k-mer in {A,C,G,T}k occurs in a DNA sequence. Although this problem is computationally very simple and efficiently solvable on a conventional computer, the sheer amount of data available now in applications demands to resort to parallel and distributed computing. Indeed, those type of algorithms have been developed to collect k-mer statistics in the realm of genome assembly. However, they are so specialized to this domain that they do not extend easily to the computation of informational and linguistic indices, concurrently on sets of genomes. Results: Following the well-established approach in many disciplines, and with a growing success also in bioinformatics, to resort to MapReduce and Hadoop to deal with 'Big Data' problems, we present KCH, the first set of MapReduce algorithms able to perform concurrently informational and linguistic analysis of large collections of genomic sequences on a Hadoop cluster. The benchmarking of KCH that we provide indicates that it is quite effective and versatile. It is also competitive with respect to the parallel and distributed algorithms highly specialized to k-mer statistics collection for genome assembly problems. In conclusion, KCH is a much needed addition to the growing number of algorithms and tools that use MapReduce for bioinformatics core applications. Availability and implementation: The software, including instructions for running it over Amazon AWS, as well as the datasets are available at http://www.di-srv.unisa.it/KCH. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Umberto Ferraro Petrillo, Gianluca Roscigno, Giuseppe Cattaneo, Raffaele Giancarlo
Bioinform.1
2018 Improving the experimental analysis of tampered image detection algorithms for biometric systems
Giuseppe Cattaneo, Gianluca Roscigno, Umberto Ferraro Petrillo
Pattern Recognit. Lett.3
2017 An Efficient Implementation of the Algorithm by Lukáš et al. on Hadoop
Giuseppe Cattaneo, Umberto Ferraro Petrillo, Michele Nappi, Fabio Narducci, Gianluca Roscigno
GPC2
2017 FASTdoop: a versatile and efficient library for the input of FASTA and FASTQ files for MapReduce Hadoop bioinformatics applications
abstract
SUMMARY: MapReduce Hadoop bioinformatics applications require the availability of special-purpose routines to manage the input of sequence files. Unfortunately, the Hadoop framework does not provide any built-in support for the most popular sequence file formats like FASTA or BAM. Moreover, the development of these routines is not easy, both because of the diversity of these formats and the need for managing efficiently sequence datasets that may count up to billions of characters. We present FASTdoop, a generic Hadoop library for the management of FASTA and FASTQ files. We show that, with respect to analogous input management routines that have appeared in the Literature, it offers versatility and efficiency. That is, it can handle collections of reads, with or without quality scores, as well as long genomic sequences while the existing routines concentrate mainly on NGS sequence data. Moreover, in the domain where a comparison is possible, the routines proposed here are faster than the available ones. In conclusion, FASTdoop is a much needed addition to Hadoop-BAM. AVAILABILITY AND IMPLEMENTATION: The software and the datasets are available at http://www.di.unisa.it/FASTdoop/ . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Umberto Ferraro Petrillo, Gianluca Roscigno, Giuseppe Cattaneo, Raffaele Giancarlo
Bioinform.1
2017 Using the audio of 8-bit video games to monitor web marketing campaigns
Umberto Ferraro Petrillo
Multim. Syst.1
2017 A new distributed alignment-free approach to compare whole proteomes
Umberto Ferraro Petrillo, Concettina Guerra, Cinzia Pizzi
Theor. Comput. Sci.1
2017 An effective extension of the applicability of alignment-free biological sequence comparison algorithms with Hadoop
Giuseppe Cattaneo, Umberto Ferraro Petrillo, Raffaele Giancarlo, Gianluca Roscigno
J. Supercomput.2
2015 A PNU-Based Technique to Detect Forged Regions in Digital Images
Giuseppe Cattaneo, Umberto Ferraro Petrillo, Gianluca Roscigno, Carmine De Fusco
ACIVS2
2015 Using HTML5 to prevent detection of drive-by-download web malware
abstract
Abstract The Web is experiencing an explosive growth in the last years. New technologies are introduced at a very fast pace with the aim of narrowing the gap between web‐based applications and traditional desktop applications. The results are web applications that look and feel almost like desktop applications while retaining the advantages of being originated from the Web. However, these advancements come at a price. The same technologies used to build responsive, pleasant, and fully featured web applications can also be used to write web malware able to escape detection systems. In this article, we present new obfuscation techniques, on the basis of some of the features of the upcoming HTML5 standard, which can be used to deceive malware detection systems. The proposed techniques have been experimented on a reference set of obfuscated malware. Our results show that the malware rewritten using our obfuscation techniques goes undetected while being analyzed by a large number of detection systems. The same detection systems were able to correctly identify the same malware in its original unobfuscated form. We also provide some hints about how the existing malware detection systems can be modified in order to cope with these new techniques. Copyright © 2014 John Wiley & Sons, Ltd.
Alfredo De Santis, Giancarlo De Maio, Umberto Ferraro Petrillo
Secur. Commun. Networks3
2014 A Scalable Approach to Source Camera Identification over Hadoop
abstract
In this paper, we explore the possibility to solve a commonly-known digital image forensics problem, the Source Camera Identification (SCI) problem, using a distributed approach. The SCI problem requires to recognize the camera used to acquire a given digital image, distinguishing even among cameras of the same brand and model. The solution we present is based on the algorithm by Lukas Fridrich, as it is recognized by many as the reference solution for this problem, and is formulated according to the MapReduce paradigm, as implemented by the Hadoop framework. The first implementation we coded was straightforward to obtain as we leveraged the ability of the Hadoop framework to turn a stand-alone Java application into a distributed one with very few interventions on its original source code. However, our first experimental results with this code were not encouraging. Thus, we conducted a careful profiling activity that allowed us to pinpoint some serious performance issues arising with this vanilla porting of the algorithm. We then developed several optimizations to improve the performance of the Lukas algorithm by taking better advantage of the Hadoop framework. The out coming implementations have been subject to a thorough experimental analysis, conducted using a cluster of 33 commodity PCs and a data set of 5, 160 images. The experimental results show that the performance of our optimized implementations scale well with the number of computing nodes while exhibiting performance that are, at most, two times slower than the maximum speedup theoretically achievable.
Giuseppe Cattaneo, Gianluca Roscigno, Umberto Ferraro Petrillo
AINA3
2014 The design and implementation of a secure CAPTCHA against man-in-the-middle attacks
abstract
ABSTRACT In this paper, we propose a novel security protocol for the implementation of CAPTCHA tests that feature advance mechanisms against man‐in‐the‐middle (MITM, for short) attacks. This type of attack is fulfilled by a malicious entity, the MITM, that leverages on unaware users to mass‐solve CAPTCHA tests shielding the access to a service. The protocol that we propose uses collision‐resistant hash functions modeled as random oracles to guarantee that the solution to a CAPTCHA test solved by an end user is valid only for the server to which the user is connected to. This will prevent MITM attacks because the user is not directly connected to the server. We developed a reference implementation for our protocol that has a low impact and is easy to use, featuring a software plug‐in running in the Firefox web browser, on the client side, and a Java servlet‐based application, on the server side. Copyright © 2013 John Wiley & Sons, Ltd.
Umberto Ferraro Petrillo, Giovanni Mastroianni, Ivan Visconti
Secur. Commun. Networks1
2012 Engineering a secure mobile messaging framework
Aniello Castiglione, Giuseppe Cattaneo, Maurizio Cembalo, Alfredo De Santis, Pompeo Faruolo, Fabio Petagna, Umberto Ferraro Petrillo
Comput. Secur.7
2010 An Extensible Framework for Efficient Secure SMS
abstract
Nowadays, Short Message Service (SMS) still represents the most used mobile messaging service. SMS messages are used in many different application fields, even in cases where security features, such as authentication and confidentiality between the communicators, must be ensured. Unfortunately, the SMS technology does not provide a built-in support for any security feature. This work presents SEESMS (Secure Extensible and Efficient SMS), a software framework written in Java which allows two peers to exchange encrypted and digitally signed SMS messages. The communication between peers is secured by using public-key cryptography. The key-exchange process is implemented by using a novel and simple security protocol which minimizes the number of SMS messages to use. SEESMS supports the encryption of a communication channel through the ECIES and the RSA algorithms. The identity validation of the contacts involved in the communication is implemented through the RSA, DSA and ECDSA signature schemes. SEESMS is able to certify the phone number of the peers using the framework. Additional cryptosystems can be coded and added to SEESMS as plug-ins. Special attention has been devoted to the implementation of an efficient framework in terms of energy consumption and execution time. This efficiency is obtained in two steps. First, all the cryptosystems available in the framework are implemented using mature and fully optimized cryptographic libraries. Second, an experimental analysis was conducted to determine which combination of cryptosystems and security parameters were able to provide a better trade-off in terms of speed/security and energy consumption. This experimental analysis has also been useful to expose some serious performance issues affecting the cryptographic libraries which are commonly used to implement security features on mobile devices.
Alfredo De Santis, Aniello Castiglione, Giuseppe Cattaneo, Maurizio Cembalo, Fabio Petagna, Umberto Ferraro Petrillo
CISIS6
2010 Experimental Study of Resilient Algorithms and Data Structures
Umberto Ferraro Petrillo, Irene Finocchi, Giuseppe F. Italiano
SEA1
2010 Data Structures Resilient to Memory Faults: An Experimental Study of Dictionaries
Umberto Ferraro Petrillo, Fabrizio Grandoni 0001, Giuseppe F. Italiano
SEA1
2010 Maintaining dynamic minimum spanning trees: An experimental study
Giuseppe Cattaneo, Pompeo Faruolo, Umberto Ferraro Petrillo, Giuseppe F. Italiano
Discret. Appl. Math.3
2009 DISCERN: A collaborative visualization system for learning cryptographic protocols
abstract
In this paper we propose a novel approach to the learning of cryptographic protocols, based on a collaborative role-based visualization system, DISCERN, that helps students to understand a protocol by actively engaging them in a simulation of its execution. In DISCERN, each student shares a visual e
Giuseppe Cattaneo, Alfredo De Santis, Umberto Ferraro Petrillo
CollaborateCom3
2009 The Price of Resiliency: a Case Study on Sorting with Memory Faults
Umberto Ferraro Petrillo, Irene Finocchi, Giuseppe F. Italiano
Algorithmica1
2006 The Price of Resiliency: A Case Study on Sorting with Memory Faults
Umberto Ferraro Petrillo, Irene Finocchi, Giuseppe F. Italiano
ESA1
2005 Overcoming the obfuscation of Java programs by identifier renaming
Stelvio Cimato, Alfredo De Santis, Umberto Ferraro Petrillo
J. Syst. Softw.3
2004 Providing Privacy for Web Services by Anonymous Group Identification
abstract
In this paper we present a SOAP extension for protecting the privacy of users of a Web service. This extension allows a user to prove to a remote SOAP server to be member of a trusted group without revealing his/her identity. Our extension has been designed as to ensure interoperability among SOAP applications written in different programming languages. We developed also some implementations of our extension using different programming languages. Moreover, we conducted an extensive experimentation of our implementation to prove its feasibility in a real-world context. In sums, our work suggests that privacy can be added to Web services with very little impact on the application developer and without compromising the performance of Web services.
Giuseppe Cattaneo, Pompeo Faruolo, Umberto Ferraro Petrillo, Giuseppe Persiano
ICWS3
2004 JIVE: Java Interactive Software Visualization Environment
abstract
JIVE (Java interactive software visualization environment) is a system for the visualization of Java coded algorithms and data structures. It supports the rapid development of interactive animations through the adoption of an object oriented approach. JIVE introduces several significant innovations such as a distributed architecture able to separate transparently the visualization activity from the underlying communication needed to support it. Therefore, it becomes possible to use JIVE in a variety of scenarios ranging from debugging algorithms to software visualization in virtual classrooms environments. Moreover, JIVE uses a zoomable user interface for representing algorithms: seamless visualization of both small and large data sets is achieved by using semantic zooming. Finally, JIVE comes with a collection of already animated data types including data structures provided by the Java standard library
Giuseppe Cattaneo, Pompeo Faruolo, Umberto Ferraro Petrillo, Giuseppe F. Italiano
VL/HCC3
2002 Maintaining Dynamic Minimum Spanning Trees: An Experimental Study
Giuseppe Cattaneo, Pompeo Faruolo, Umberto Ferraro Petrillo, Giuseppe F. Italiano
ALENEX3
2001 JSEB (Java Scalable sErvices Builder): Scalable Systems for Clusters of Workstations
abstract
We present a report on JSEB (Java Scalable Service Builder) whose goal is to offer programmers a tool that can be used to efficiently add scalability and fault-tolerance to a replicated service in cluster(s) of workstations.
Maria Barra, Giuseppe Cattaneo, Umberto Ferraro Petrillo, Vittorio Scarano
ISCC3