EDBT 2026 Demo / reviewers in the wild / expert
Domenico Talia
dblp:44/5373
· DBLP profile ↗
125ranked-venue papers
20as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 72 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7Software engineering, systems software and programming languages · 5 · 4 first-authorTheory of computation · 3Computer networks · 2 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Hashtag Recommendation in Social Media With Trend Shift Detection and AdaptationabstractHashtag recommendation systems have emerged as a key tool for automatically suggesting relevant hashtags and enhancing content categorization and search. However, existing static models struggle to adapt to the highly dynamic nature of social media conversations, where new hashtags constantly emerge and existing ones undergo semantic shifts. To address these challenges, this article introduces hashtag recommendation by detecting and adapting to trend shifts (H-ADAPTS), a dynamic hashtag recommendation methodology that employs a trend-aware mechanism to detect shifts in hashtag usage—reflecting evolving trends and topics within social media conversations—and triggers efficient model adaptation based on a (small) set of recent posts. Additionally, the Apache storm framework is leveraged to support scalable and fault-tolerant analysis of high-velocity social data, enabling the timely detection of trend shifts. Experimental results from two real-world case studies, including the COVID-19 pandemic and the 2020 US presidential election, demonstrate the effectiveness of H-ADAPTS in providing timely and relevant hashtag recommendations by adapting to emerging trends, significantly outperforming existing solutions. Riccardo Cantini, Fabrizio Marozzo, Alessio Orsino, Domenico Talia, Paolo Trunfio |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Neural Topic Modeling in Social Media by Clustering Latent Hashtag RepresentationsabstractThe worldwide use of social media has generated vast volumes of user-generated content, offering valuable insights into public discourse, behavioral dynamics, and emerging trends. However, extracting meaningful topics from such data remains a significant challenge due to the informal, dynamic, and context-dependent nature of online language, where the semantics of terms and hashtags are often shaped by the specific sociocultural and temporal contexts in which they arise. To address these challenges, we propose NTM-HEC (Neural Topic Modeling via Hashtag Embedding Clustering), a novel hashtag-centric methodology for topic discovery that leverages the semantic richness encoded in hashtags, commonly used by social media users to annotate and categorize content. NTM-HEC relies on clustering low-dimensional embeddings of latent hashtag representations to uncover coherent and diverse topic structures. This enables it to fully leverage the inherently topical nature of hashtags, enhancing interpretability and improving robustness to linguistic variability and context-specificity. We evaluate the effectiveness of NTM-HEC through two case studies focused on online discourse surrounding the Russia-Ukraine conflict and the COVID-19 pandemic. In both cases, NTM-HEC outperforms competing models in topic coherence and diversity, demonstrating its ability to capture nuanced, trend-specific semantic patterns within real-world social media discussions. Riccardo Cantini, Cristian Cosentino, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
ECAI | 4 |
| 2025 | Scalable Compression of Massive Data Collections on HPC Systems
Loris Belcastro, Paolo Ferragina, Giovanni Manzini, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
Euro-Par (2) | 5 |
| 2025 | Edge-cloud solutions for big data analysis and distributed machine learning - 2abstractIn recent years, edge-cloud solutions have gained widespread adoption for efficiently collecting and analyzing IoT-generated data across various domains like urban mobility, healthcare, and smart cities. These solutions integrate resources from edge to cloud to support real-time processing and analysis tasks, reducing latency and network congestion. Big data analysis within this paradigm involves sophisticated techniques for distributed data processing, enabling applications such as predictive maintenance and smart grid management. Nevertheless, carrying out big data analysis within the edge-cloud presents several challenges, including data privacy and security, interoperability, scalability, and energy efficiency. Addressing these challenges is imperative for providing efficient and scalable solutions for data-intensive applications like federated learning, social data analysis, smart city services, and text mining. The special issue concludes with 27 scientific papers, divided into two parts for a streamlined editorial process. This editorial, as part two, presents 12 rigorously peer-reviewed papers, complementing the 15 papers covered in the previous editorial. Loris Belcastro, Jesús Carretero 0001, Domenico Talia |
Future Gener. Comput. Syst. | 3 |
| 2025 | Unmasking deception: a topic-oriented multimodal approach to uncover false information on social mediaabstractAbstract In the digital landscape, social media has emerged as a prevalent channel for global communication, connecting like-minded individuals worldwide. However, while facilitating information exchange, it is also susceptible to the dissemination of false information, posing a constant challenge to the reliability of online content. To address this issue, this paper introduces a novel methodology called TM-FID (Topic-oriented Multimodal False Information Detection), which combines false information detection and neural topic modeling within a semi-supervised multimodal approach. By jointly leveraging textual and visual information contained in online news, our approach provides insights into how false information influences specific discussion topics, thus enabling a comprehensive and fine-grained understanding of its spread and impact on social media conversation. Experimental evaluation carried out on a set of multimodal gossip-related news demonstrates the quality of the identified topics, assessed through a novel centroid-based metric, as well as the efficacy of the cross-attention mechanism used within TM-FID to accurately identify false information in multimodal news. Overall, the proposed methodology can enable effective strategies to counter the spread of false information, thereby fostering trust and confidence in the information shared on social media platforms. Riccardo Cantini, Cristian Cosentino, Irene Kilanioti, Fabrizio Marozzo, Domenico Talia |
Mach. Learn. | 5 |
| 2025 | Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judgeabstractAbstract The growing integration of Large Language Models (LLMs) into critical societal domains has raised concerns about embedded biases that can perpetuate stereotypes and undermine fairness. Such biases may stem from historical inequalities in training data, linguistic imbalances, or adversarial manipulation. Despite mitigation efforts, recent studies show that LLMs remain vulnerable to adversarial attacks that elicit biased outputs. This work proposes a scalable benchmarking framework to assess LLM robustness to adversarial bias elicitation. Our methodology involves: ( i ) systematically probing models across multiple tasks targeting diverse sociocultural biases, ( ii ) quantifying robustness through safety scores using an LLM-as-a-Judge approach, and ( iii ) employing jailbreak techniques to reveal safety vulnerabilities. To facilitate systematic benchmarking, we release a curated dataset of bias-related prompts, named CLEAR-Bias . Our analysis, identifying DeepSeek V3 as the most reliable judge LLM, reveals that bias resilience is uneven, with age, disability, and intersectional biases among the most prominent. Some small models outperform larger ones in safety, suggesting that training and architecture may matter more than scale. However, no model is fully robust to adversarial elicitation, with jailbreak attacks using low-resource languages or refusal suppression proving effective across model families. We also find that successive LLM generations exhibit slight safety gains, while models fine-tuned for the medical domain tend to be less safe than their general-purpose counterparts. Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia |
Mach. Learn. | 4 |
| 2024 | Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
Riccardo Cantini, Giada Cosenza, Alessio Orsino, Domenico Talia |
DS (1) | 4 |
| 2024 | Programming Tools for High-Performance Data AnalysisabstractIn this paper we discuss and compare some of the most popular big data programming frameworks for high performance computing (HPC) systems, such as Hadoop, Spark, Storm, Airflow, MPI, Hive, Pig and UPC++. We highlight here their support to different classes of applications, such as batch, streaming, graph-based, and query-based applications. The main features of such frameworks according to their programming model, type of parallelism, level of abstraction, verbosity, and main classes of applications are summarized. We also discuss the main factors that can influence the choice of the most appropriate framework to process and analyze big data. These include the characteristics of input data, the class of application and the infrastructure requirements. Domenico Talia, Paolo Trunfio |
HPDC | 1 |
| 2024 | TinyTTA: Efficient Test-time Adaptation via Early-exit Ensembles on Edge DevicesabstractThe increased adoption of Internet of Things (IoT) devices has led to the generation of large data streams with applications in healthcare, sustainability, and robotics. In some cases, deep neural networks have been deployed directly on these resource-constrained units to limit communication overhead, increase efficiency and privacy, and enable real-time applications. However, a common challenge in this setting is the continuous adaptation of models necessary to accommodate changing environments, i.e., data distribution shifts. Test-time adaptation (TTA) has emerged as one potential solution, but its validity has yet to be explored in resource-constrained hardware settings, such as those involving microcontroller units (MCUs). TTA on constrained devices generally suffers from i) memory overhead due to the full backpropagation of a large pre-trained network, ii) lack of support for normalization layers on MCUs, and iii) either memory exhaustion with large batch sizes required for updating or poor performance with small batch sizes. In this paper, we propose TinyTTA, to enable, for the first time, efficient TTA on constrained devices with limited memory. To address the limited memory constraints, we introduce a novel self-ensemble and batch-agnostic early-exit strategy for TTA, which enables continuous adaptation with small batch sizes for reduced memory usage, handles distribution shifts, and improves latency efficiency. Moreover, we develop the TinyTTA Engine, a first-of-its-kind MCU library that enables on-device TTA. We validate TinyTTA on a Raspberry Pi Zero 2W and an STM32H747 MCU. Experimental results demonstrate that TinyTTA improves TTA accuracy by up to 57.6\%, reduces memory usage by up to six times, and achieves faster and more energy-efficient TTA. Notably, TinyTTA is the only framework able to run TTA on MCU STM32H747 with a 512 KB memory constraint while maintaining high performance. Hong Jia, Young D. Kwon, Alessio Orsino, Ting Dang, Domenico Talia, Cecilia Mascolo |
NeurIPS | 5 |
| 2024 | A parallel machine learning-based approach for tsunami waves forecasting using regression treesabstractFollowing a seismic event, tsunami early warning systems (TEWSs) try to provide precise forecasts of the maximum height of incoming waves at designated target points along the coast. This information is crucial to trigger early warnings in areas where the impact of tsunami waves is predicted to be dangerous (or potentially cause destruction), to help the management of the potential impact of a tsunami as well as reduce environmental destruction and losses of human lives. For such a reason, it is crucial that TEWSs produce predictions with short computation time while maintaining a high prediction accuracy. This paper presents a parallel machine learning approach, based on regression trees, to discover tsunami predictive models from simulation data. In order to achieve the results in a short time, the proposed approach relies on the parallelization of the most time consuming tasks and on incremental learning executions, in order to achieve higher performances in terms of execution time, efficiency and scalability. The experimental evaluation, performed on two real tsunami cases occurred in the Western and Eastern Mediterranean basin in 2003 and 2017, shows reasonable advantages in terms of scalability and execution time, which is an important benefit in a urgent-computing scenarios. Eugenio Cesario, Salvatore Giampà, Enrico Baglione, Louise Cordrie, Jacopo Selva, Domenico Talia |
Comput. Commun. | 6 |
| 2024 | Edge-Cloud Solutions for Big Data Analysis and Distributed Machine Learning - 1
Loris Belcastro, Jesús Carretero 0001, Domenico Talia |
Future Gener. Comput. Syst. | 3 |
| 2024 | Boosting HPC data analysis performance with the ParSoDA-Py libraryabstractAbstract Developing and executing large-scale data analysis applications in parallel and distributed environments can be a complex and time-consuming task. Developers often find themselves diverted from their application logic to handle technical details about the underlying runtime and related issues. To simplify this process, ParSoDA, a Java library, has been proposed to facilitate the development of parallel data mining applications executed on HPC systems. It simplifies the process by providing built-in scalability mechanisms relying on the Hadoop and Spark frameworks. This paper presents ParSoDA-Py, the Python version of the ParSoDA library, which allows for further support of commonly used runtimes and libraries for big data analysis. After a complete library redesign, ParSoDA can be now easily integrated with other Python-based distributed runtimes for HPC systems, such as COMPSs and Apache Spark, and with the large ecosystem of Python-based data processing libraries. The paper discusses the adaptation process, which takes into consideration the new technical requirements, and evaluates both usability and scalability through some case study applications. Loris Belcastro, Salvatore Giampà, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio, Rosa M. Badia, Jorge Ejarque, Nihad Mammadli |
J. Supercomput. | 4 |
| 2023 | Unmasking COVID-19 False Information on Twitter: A Topic-Based Approach with BERT
Riccardo Cantini, Cristian Cosentino, Irene Kilanioti, Fabrizio Marozzo, Domenico Talia |
DS | 5 |
| 2023 | Using the Compute Continuum for Data Analysis: Edge-cloud Integration for Urban MobilityabstractMore and more in recent years, IT companies have adopted edge-cloud continuum solutions to efficiently perform analysis tasks on data generated by IoT devices. As an example, in the context of urban mobility, the use of edge solutions can be extremely effective in managing tasks that require real-time analysis and low response times, such as driver assistance, collision avoidance and traffic sign recognition. On the other hand, the integration with cloud systems can be convenient for tasks that require a lot of computing resources for accessing and analyzing big data collections, such as route calculations and targeted advertising. Designing and testing such hybrid edge-cloud architectures are still open issues due to their novelty, large scale, heterogeneity, and complexity. In this paper, we analyze how the compute continuum can be exploited for efficiently managing urban mobility tasks. In particular, we focus on a case study related to taxi fleets that need to find locations where they are more likely to find new passengers. Through a simulation-based approach, we demonstrate that these solutions turn out to be effective for this class of problems, especially as the number of connected vehicles increases. Loris Belcastro, Fabrizio Marozzo, Alessio Orsino, Domenico Talia, Paolo Trunfio |
PDP | 4 |
| 2022 | Convergence of HPC and Big Data in extreme-scale data analysis through the DCEx programming modelabstractHigh-level programming models can help application developers to access and use resources without the need to manage low-level architectural entities, as a parallel programming model defines a set of programming abstractions that simplify the way by which a programmer structures and expresses her/his algorithm. Early proposals of Exascale programming tools are based on the adaptation of traditional parallel programming languages and hybrid solutions. This incremental approach is too conservative, often resulting in very complex code. This paper describes the design features, the programming constructs, and the runtime mechanisms of the Data Centric programming model for Exascale systems (DCEx). DCEx is based on structuring applications into data-parallel blocks. Blocks are units of shared-and distributed-memory parallel computation, communication, and migration in the memory/storage hierarchy. Blocks and their message queues are mapped onto processes and placed in memory/storage by the DCEx runtime. Those data-parallel blocks are orchestrated by using distributed parallel patterns that simplify the development cost. DCEx aims to reach the convergence of traditional HPC programming models, mainly based on MPI, with the emerging technologies based on the data intensive paradigms. To demonstrate the potential of DCEx, we carried out an experimental evaluation developing a real-world diffusion-weighted magnetic resonance imaging data processing application in a neuroimaging research context. Francisco Javier García Blas, Javier Fernández 0001, Jesús Carretero 0001, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio, Alberto Fernández-Pena, Daniel Martín de Blas |
SBAC-PAD | 5 |
| 2022 | A fuzzy logic technique for virtual sensor networks
Luciano Caroprese, Carmela Comito, Domenico Talia, Ester Zumpano |
Future Gener. Comput. Syst. | 3 |
| 2022 | Enabling dynamic and intelligent workflows for HPC, data analytics, and AI convergence
Jorge Ejarque, Rosa M. Badia, Loïc Albertin, Giovanni Aloisio, Enrico Baglione, Yolanda Becerra 0001, Stefan Boschert, Julian R. Berlin, Alessandro D'Anca, Donatello Elia, François Exertier, Sandro Fiore, José Flich, Arnau Folch, Steven J. Gibbons, Nikolay Koldunov, Francesc Lordan, Stefano Lorito, Finn Løvholt, Jorge Macías Sánchez, Fabrizio Marozzo, Alberto Michelini, Marisol Monterrubio Velasco, Marta Pienkowska, Josep de la Puente, Anna Queralt, Enrique S. Quintana-Ortí, Juan Esteban Rodriguez, Fabrizio Romano, Jedrzej Rybicki, Miroslaw Kupczyk, Jacopo Selva, Domenico Talia, Roberto Tonini, Paolo Trunfio, Manuela Volpe |
Future Gener. Comput. Syst. | 34 |
| 2021 | Parallel extraction of Regions-of-Interest from social media dataabstractSummary Geotagged data gathered from social media can be used to discover places‐of‐interest (PoIs) that have attracted many visitors. Since a PoI is generally identified by geographical coordinates of a single point, it is hard to match it with people trajectories. Therefore, we define an area, called region‐of‐interest (RoI), represented by the boundaries of a PoI. The main goal of this study is to discover RoIs from PoIs using spatial data mining techniques. In this paper, we propose a new parallel method for extracting RoIs from social media datasets. It consists of two main steps: (i) automatic keyword extraction and data grouping and (ii) parallel RoI extraction. The first step extracts keywords identifying the PoIs; these keywords are used to group social media items according to the places they refer to. The second step uses a Parallel Clustering Approach (ParCA) of spatial dataset to identify RoIs. ParCA exploits a parallel execution of DBSCAN on subsets of data to generate subclusters on each processing node and then merge overlapping subclusters to form global clusters. ParCA was implemented using the MapReduce model. Experiments performed over a set of PoIs in the city of Rome using social media data show that our approach is highly scalable and reaches an accuracy of 79% in detecting RoIs. On a parallel computer with 50 cores, we obtained a speedup of 52 by processing large datasets divided into 32 splits, compared with the execution time registered when each dataset is not partitioned. Loris Belcastro, M. Tahar Kechadi, Fabrizio Marozzo, Luca Pastore, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Cloud Computing for Enabling Big Data Analysis Services
Domenico Talia |
CLOSER | 1 |
| 2020 | A sleep-and-wake technique for reducing energy consumption in BitTorrent networksabstractSummary File sharing is one of the leading Internet applications of P2P technology. Given the high number of computer nodes involved in peer‐to‐peer networks, reducing their aggregate energy consumption is an important challenge to be faced. In this paper, we show how the sleep‐and‐wake energy saving approach can be exploited to reduce energy consumption in BitTorrent, one of the most popular file sharing peer‐to‐peer networks. We describe BitTorrentSW, a sleep‐and‐wake approach for BitTorrent networks that allows seeders (ie, peers that hold complete files) to cyclically switch between wake and sleep modes to save energy while ensuring good file sharing performance. The decision to switch to sleep mode is taken independently by each seeder based on local information about the composition of the peer‐to‐peer network. BitTorrentSW has been evaluated through PeerSim using real BitTorrent traces. The simulation results show that, in all the configurations under analysis, the percentage of energy saved by BitTorrentSW is much higher than the percentage of increase in download time. For instance, in a network with 50% of seeders, about 20% of energy is saved using BitTorrentSW, with an increase of only 7% of the average time needed to complete a file download compared to a standard BitTorrent network in which all seeders are always powered on. Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | New Landscapes of the Data Stream Processing in the era of Fog Computing
Valeria Cardellini, Gabriele Mencagli, Domenico Talia, Massimo Torquati |
Future Gener. Comput. Syst. | 3 |
| 2019 | Spatio-temporal crime predictions in smart cities: A data-driven approach and experiments
Charles E. Catlett, Eugenio Cesario, Domenico Talia, Andrea Vinci |
Pervasive Mob. Comput. | 3 |
| 2018 | A Data-Driven Approach for Spatio-Temporal Crime Predictions in Smart CitiesabstractThe steadily increasing urbanization is causing significant economic and social transformations in urban areas and it will be posing several challenges in city management issues. In particular, given that the larger cities the higher crime rates, crime spiking is becoming one of the most important social problems in large urban areas. To handle with the increase in crimes, new technologies are enabling police departments to access growing volumes of crime-related data that can be analyzed to understand patterns and trends, finalized to an efficient deployment of police officers over the territory and more effective crime prevention. This paper presents an approach based on spatial analysis and auto-regressive models to automatically detect high-risk crime regions in urban areas and reliably forecast crime trends in each region. The final result of the algorithm is a spatio-temporal crime forecasting model, composed of a set of crime dense regions and a set of associated crime predictors, each one representing a predictive model for forecasting the number of crimes that will happen in its specific region. The experimental evaluation, performed on real-world data collected in a big area of Chicago, shows that the proposed approach achieves good accuracy in spatial and temporal crime forecasting over rolling time horizons. Charles E. Catlett, Eugenio Cesario, Domenico Talia, Andrea Vinci |
SMARTCOMP | 3 |
| 2018 | G-RoI: Automatic Region-of-Interest Detection Driven by Geotagged Social Media DataabstractGeotagged data gathered from social media can be used to discover interesting locations visited by users called Places-of-Interest (PoIs). Since a PoI is generally identified by the geographical coordinates of a single point, it is hard to match it with user trajectories. Therefore, it is useful to define an area, called Region-of-Interest ( RoI ), to represent the boundaries of the PoI’s area. RoI mining techniques are aimed at discovering ROIs from PoIs and other data. Existing RoI mining techniques are based on three main approaches: predefined shapes, density-based clustering, and grid-based aggregation. This article proposes G-RoI , a novel RoI mining technique that exploits the indications contained in geotagged social media items to discover RoIs with a high accuracy. Experiments performed over a set of PoIs in Rome and Paris using social media geotagged data, demonstrate that G-RoI in most cases achieves better results than existing techniques. In particular, the mean F 1 score is 0.34 higher than that obtained with the well-known DBSCAN algorithm in Rome RoIs and 0.23 higher in Paris RoIs. Loris Belcastro, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
ACM Trans. Knowl. Discov. Data | 3 |
| 2018 | A Workflow Management System for Scalable Data Mining on CloudsabstractThe extraction of useful information from data is often a complex process that can be conveniently modeled as a data analysis workflow. When very large data sets must be analyzed and/or complex data mining algorithms must be executed, data analysis workflows may take very long times to complete their execution. Therefore, efficient systems are required for the scalable execution of data analysis workflows, by exploiting the computing services of the Cloud platforms where data is increasingly being stored. The objective of the paper is to demonstrate how Cloud software technologies can be integrated to implement an effective environment for designing and executing scalable data analysis workflows. We describe the design and implementation of the Data Mining Cloud Framework (DMCF), a data analysis system that integrates a visual workflow language and a parallel runtime with the Software-as-a-Service (SaaS) model. DMCF was designed taking into account the needs of real data mining applications, with the goal of simplifying the development of data mining applications compared to generic workflow management systems that are not specifically designed for this domain. The result is a high-level environment that, through an integrated visual workflow language, minimizes the programming effort, making easierto domain experts the use of common patterns specifically designed forthe development and the parallel execution of data mining applications. The DMCF's visual workflow language, system architecture and runtime mechanisms are presented. We also discuss several data mining workflows developed with DMCF and the scalability obtained executing such workflows on a public Cloud. Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
IEEE Trans. Serv. Comput. | 2 |
| 2017 | A Peak Detection Method to Uncover Events from Social MediaabstractSocial networking services like Twitter and Instagram are a valuable sources of information to find out what happened or what is happening in a geographic area. This paper presents a method to catch and understand relevant events and happenings from social geo-tagged data. The proposed method consists in two main phases: (i) extraction of space-time features from social data and their modelization as time series, (ii) peak detection from time series, for identifying deviation from user normal behavior. Results of the experimental evaluation, performed over a real-word dataset of tweets, show that the proposed approach is able to accurately detect several relevant events, bounded to a geographic location and of varying importance and character, like exhibitions, festivals, competitions, and terrorist attacks such as that done at the Charlie Hebdo offices.We achieve a space accuracy up to 90%, and a time accuracy up to 95%. Carmela Comito, Deborah Falcone, Domenico Talia |
DSAA | 3 |
| 2017 | Energy-aware task allocation for small devices in wireless networksabstractSummary The continuous advances in wireless networking and mobile computing technologies have paved the way to the spreading of new classes of distributed applications running on networks of small devices such as smartphones and tablets. An issue that still prevents a wider implementation of distributed applications in wireless networks is the lack of task allocation strategies addressing both the energy constraints of small devices and the decentralized nature of wireless networks. In this paper, we focus on this twofold issue by proposing an energy‐aware scheduling strategy for allocating computational tasks over a wireless network of small devices in a decentralized but effective way. The main design principle of our scheduling strategy is finding a task allocation that prolongs the total lifetime of the network and maximizes the number of alive devices by balancing the energy load among them. A simulation analysis has been performed to assess the performance of the proposed strategy in different network and application scenarios. The results show that by using the proposed energy‐aware task allocation approach, the network lifetime is extended and the number of alive devices is significantly higher compared with alternative scheduling strategies while meeting application‐level performance constraints. Copyright © 2016 John Wiley & Sons, Ltd. Carmela Comito, Deborah Falcone, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | A data-aware scheduling strategy for workflow execution in cloudsabstractSummary As data intensive scientific computing systems become more widespread, there is a necessity of simplifying the development, deployment, and execution of complex data analysis applications for scientific discovery. The scientific workflow model is the leading approach for designing and executing data‐intensive applications in high‐performance computing infrastructures. Commonly, scientific workflows are built by a set of connected tasks arranged in a directed acyclic graph style, which communicate through storage abstractions. The Data Mining Cloud Framework (DMCF) is a system allowing users to design and execute data analysis workflows on cloud platforms, relying on cloud storage services for every I/O operation. Hercules is an in‐memory I/O solution that can be used in DMCF as an alternative to cloud storage services, providing additional performance and flexibility features. This work improves the integration between DMCF and Hercules by using a data‐aware scheduling strategy for exploiting data locality in data‐intensive workflows. This paper presents experimental results demonstrating the performance improvements achieved using the proposed data‐aware scheduling strategy in the Microsoft Azure cloud platform. In particular, with our scheduling strategy, the I/O overhead has been reduced by 55% with respect to the Azure storage, leading to a 20% reduction of the total execution time. Fabrizio Marozzo, Francisco Rodrigo Duro, Francisco Javier García Blas, Jesús Carretero 0001, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | An approach for the discovery and validation of urban mobility patterns
Eugenio Cesario, Carmela Comito, Domenico Talia |
Pervasive Mob. Comput. | 3 |
| 2017 | Energy consumption of data mining algorithms on mobile phones: Evaluation and prediction
Carmela Comito, Domenico Talia |
Pervasive Mob. Comput. | 2 |
| 2017 | Trajectory Pattern Mining for Urban Computing in the CloudabstractThe increasing pervasiveness of mobile devices along with the use of technologies like GPS, Wifi networks, RFID, and sensors, allows for the collections of large amounts of movement data. This amount of data can be analyzed to extract descriptive and predictive models that can be properly exploited to improve urban life. From a technological viewpoint, Cloud computing can play an essential role by helping city administrators to quickly acquire new capabilities and reducing initial capital costs by means of a comprehensive pay-as-you-go solution. This paper presents a workflow-based parallel approach for discovering patterns and rules from trajectory data, in a Cloud-based framework. Experimental evaluation has been carried out on both real-world and synthetic trajectory data, up to one million of trajectories. The results show that, due to the high complexity and large volumes of data involved in the application scenario, the trajectory pattern mining process takes advantage from the scalable execution environment offered by a Cloud architecture in terms of both execution time, speed-up and scale-up. Albino Altomare, Eugenio Cesario, Carmela Comito, Fabrizio Marozzo, Domenico Talia |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | HeteroPar 2014, APCIE 2014, and TASUS 2014 Special IssueabstractThese workshops were organized by members of the Nesus Cost Action IC 1305: Network for Sustainable Ultrascale Computing, which is a follow-up of COST Actions IC0804 and IC0805 1. The goal of the NESUS Action is to establish an open European research network targeting sustainable solutions for ultrascale computing aiming at cross fertilization among HPC, large-scale distributed systems, and big data management. This network aims at contributing to glue disparate researchers working across different areas and provide a meeting ground for researchers in these separate areas to exchange ideas, to identify synergies, and to pursue common activities in research topics such as sustainable software solutions (applications and system software stack), data management, energy efficiency, and resilience. The selected papers cover very important scientific issues encountered nowadays such as the following: CPU/GPU execution, system-on-chip programming, parallel algorithms taking into account various constraints (energy, communication, etc.), programming models, and so on. We really hope that the reader will enjoy this high-quality issue, and we are sure that she/he will find it highly relevant to the state-of-the-art of today's heterogeneous and parallel computing. Jesús Carretero 0001, Raimondas Ciegis, Emmanuel Jeannot, Laurent Lefèvre, Gudula Rünger, Domenico Talia, Julius Zilinskas |
Concurr. Comput. Pract. Exp. | 6 |
| 2016 | Distributed volunteer computing for solving ensemble learning problems
Eugenio Cesario, Carlo Mastroianni, Domenico Talia |
Future Gener. Comput. Syst. | 3 |
| 2016 | A distributed selectivity-driven search strategy for semi-structured data over DHT-based networks
Carmela Comito, Domenico Talia, Paolo Trunfio |
J. Parallel Distributed Comput. | 2 |
| 2016 | Mining human mobility patterns from social geo-tagged data
Carmela Comito, Deborah Falcone, Domenico Talia |
Pervasive Mob. Comput. | 3 |
| 2016 | Using Scalable Data Mining for Predicting Flight DelaysabstractFlight delays are frequent all over the world (about 20% of airline flights arrive more than 15min late) and they are estimated to have an annual cost of billions of dollars. This scenario makes the prediction of flight delays a primary issue for airlines and travelers. The main goal of this work is to implement a predictor of the arrival delay of a scheduled flight due to weather conditions. The predicted arrival delay takes into consideration both flight information (origin airport, destination airport, scheduled departure and arrival time) and weather conditions at origin airport and destination airport according to the flight timetable. Airline flight and weather observation datasets have been analyzed and mined using parallel algorithms implemented as MapReduce programs executed on a Cloud platform. The results show a high accuracy in predicting delays above a given threshold. For instance, with a delay threshold of 15min, we achieve an accuracy of 74.2% and 71.8% recall on delayed flights, while with a threshold of 60min, the accuracy is 85.8% and the delay recall is 86.9%. Furthermore, the experimental results demonstrate the predictor scalability that can be achieved performing data preparation and mining tasks as MapReduce applications on the Cloud. Loris Belcastro, Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2016 | Exploiting in-memory storage for improving workflow executions in cloud platforms
Francisco Rodrigo Duro, Fabrizio Marozzo, Francisco Javier García Blas, Domenico Talia, Paolo Trunfio |
J. Supercomput. | 4 |
| 2015 | Evaluating and predicting energy consumption of data mining algorithms on mobile devicesabstractThe pervasive availability of increasingly powerful mobile computing devices like PDAs, smartphones and wearable sensors, is widening their use in complex applications such as collaborative analysis, information sharing, and data mining in a mobile context. Energy characterization plays a critical role in determining the requirements of data-intensive applications that can be efficiently executed over mobile devices. This paper presents an experimental study of the energy consumption behaviour of representative data mining algorithms running on mobile devices. Our study reveals that, although data mining algorithms are compute- and memory-intensive, by appropriate tuning of a few parameters associated to data (e.g., data set size, number of attributes, size of produced results) those algorithms can be efficiently executed on mobile devices by saving energy and, thus, prolonging devices lifetime. Based on the outcome of this study we also proposed a machine learning approach to predict energy consumption of mobile data-intensive algorithms. Results show that a considerable accuracy is achieved when the predictor is trained with specific-algorithm features. Carmela Comito, Domenico Talia |
DSAA | 2 |
| 2015 | Energy-Aware Migration of Virtual Machines Driven by Predictive Data Mining ModelsabstractConsolidation of virtual machines (VM) is one of the key strategies used to reduce the power consumption of Cloud servers. For this reason it is extensively studied. Nevertheless, the effectiveness of a consolidation strategy strongly depends on the forecast of the VM resource needs. This paper describes the design and development of a system for energy-aware allocation of virtual machines, driven by predictive data mining models. In particular, migrations are driven by the forecast of the future computational needs (CPU, RAM) of each virtual machine, in order to efficiently allocate those on the available servers. Experimental results, performed on data of a real Cloud data centre, show encouraging benefits in terms of energy saving. Albino Altomare, Eugenio Cesario, Domenico Talia |
PDP | 3 |
| 2015 | JS4Cloud: script-based workflow programming for scalable data analysis on cloud platformsabstractSummary Workflows are an effective paradigm to model complex data analysis processes, such as knowledge discovery in databases applications, which can be efficiently executed on distributed computing systems such as a Cloud platform. Data analysis workflows can be designed through visual programming, which is a convenient design approach for high‐level users. On the other hand, script‐based workflows are a useful alternative to visual workflows, because they allow expert users to program complex applications more effectively. In order to provide Cloud users with an effective script‐based data analysis workflow formalism, we designed the JS4Cloud language. The main benefits of JS4Cloud are as follows: (i) it extends the well‐known JavaScript language while using only its basic functions (arrays, functions, and loops); (ii) it implements both a data‐driven task parallelism that automatically spawns ready‐to‐run tasks to the Cloud resources and data parallelism through an array‐based formalism; and (iii) these two types of parallelism are exploited implicitly so that workflows can be programmed in a fully sequential way, which frees users from duties like work partitioning, synchronization, and communication. We describe how JS4Cloud has been integrated within the data mining cloud framework (DMCF), a system supporting the scalable execution of data analysis workflows on Cloud platforms. In particular, we describe how data analysis workflows modeled as JS4Cloud scripts are processed by DMCF by exploiting parallelism to enable their scalable execution on Clouds. Finally, we present some data analysis workflows developed with JS4Cloud and the performance results obtained by executing such workflows on DMCF. Copyright © 2015 John Wiley & Sons, Ltd. Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Big Data Mining Services and Distributed Knowledge Discovery Applications on Clouds
Domenico Talia |
KEOD | 1 |
| 2014 | A Multi-Domain Architecture for Mining Frequent Items and Itemsets from Distributed Data Streams
Eugenio Cesario, Carlo Mastroianni, Domenico Talia |
J. Grid Comput. | 3 |
| 2014 | ServiceSs: An Interoperable Programming Framework for the Cloud
Francesc Lordan, Enric Tejedor, Jorge Ejarque, Roger Rafanell, Javier Álvarez Cid-Fuentes, Fabrizio Marozzo, Daniele Lezzi, Raül Sirvent, Domenico Talia, Rosa M. Badia |
J. Grid Comput. | 9 |
| 2013 | Using Clouds for Smart City ApplicationsabstractThe increasing pervasiveness of mobile devices along with the use of technologies like GPS, Wifi networks, RFID, etc., allows for the collections of large amounts of movement data. This amount of information can be analyzed to extract descriptive and predictive models that can be profitable exploited to improve urban life. This paper presents an integrated Cloud based framework for efficiently managing and analyzing socio-environmental data in the urban context of cities. As case study, we introduce a parallel approach for discovering patterns and rules from trajectory data. Experimental evaluation shows that the trajectory pattern mining process can take advantage from a scalable execution environment offered by a Cloud architecture. Albino Altomare, Eugenio Cesario, Carmela Comito, Fabrizio Marozzo, Domenico Talia |
CloudCom (2) | 5 |
| 2013 | Topic 9: Parallel and Distributed Programming - (Introduction)
Michael Philippsen, Domenico Talia, Ana Lucia Varbanescu |
Euro-Par | 3 |
| 2013 | Programming knowledge discovery workflows in service-oriented distributed systemsabstractSUMMARY In several scientific and business domains, very large data repositories are generated. To find interesting and useful information in those repositories, efficient data mining techniques and knowledge discovery processes must be used. The exploitation of data mining techniques in science helps scientists in hypothesis formation and gives them a support on their scientific practices, whereas in industrial processes, data mining can exploit existing data sources as a real value for companies that can take advantage from the knowledge that can be extracted from their large data sources. Data mining tasks are often composed by multiple stages that may be linked to each other to form various execution flows. Moreover, data mining tasks are often distributed because they involve data and tools located over geographically distributed environments. Therefore, it is fundamental to exploit effective paradigms, such as services and workflows, to model data mining tasks that are both multi‐staged and distributed. This paper discusses data mining services and workflows for analyzing scientific data in high‐performance distributed environments such as Grids and Clouds. We discuss how it is possible to define basic and complex services for supporting distributed data mining tasks in Grids. We also present a workflow formalism and a service‐oriented programming framework, named DIS3GNO, for designing and running distributed knowledge discovery processes in the Knowledge Grid system. DIS3GNO supports all the phases of a knowledge discovery process, including composition, execution, and results visualization. After introducing DIS3GNO, some relevant use cases implemented by it and a performance evaluation of the system are discussed.Copyright © 2012 John Wiley & Sons, Ltd. Eugenio Cesario, Marco Lackovic, Domenico Talia, Paolo Trunfio |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | A Consistency Checker for verifying the knowledge encoded into clinical DSSsabstractThe formalization and manipulation of complex and not yet assessed rules by clinicians are critical for Decision Support Systems (DSSs) performance in supporting remote monitoring of chronic patients. Sometimes, structural anomalies, such as inconsistency and redundancy, can occur. This work presents a novel system, named Consistency Checker, aimed at verifying the reliability of conditionaction clinical rules in knowledge-based DSSs. This system allows the verification of very complex rules having in their antecedent parts not only simple logical conditions, but also arithmetical expressions. Moreover, the Consistency Checker provides a new classification of the detected anomalies, aimed at fully describing the rule verification results, as well as a suitable knowledge representation formalism to encode condition-action rules in a general way. The system has been designed according to the Service Oriented Architecture and implemented as a Web Service within the CHRONIOUS Project. Eugenio Cesario, Massimo Esposito, Giuseppe De Pietro, Domenico Talia |
CBMS | 4 |
| 2012 | Fault-Tolerant Distributed Knowledge Discovery Services for GridsabstractFault tolerance is an important issue in service oriented architectures like Grid and Cloud systems, where many and heterogeneous machines are used. In this paper we present a flexible failure handling framework which extends a service-oriented architecture for Distributed Data Mining previously proposed, addressing the requirements for handling fault tolerance in service-oriented Grids. In particular, two different solutions are described and experimentally evaluated on a real Grid setting, aimed at assessing their effectiveness and performance. Eugenio Cesario, Domenico Talia |
CISIS | 2 |
| 2012 | Enabling Cloud Interoperability with COMPSs
Fabrizio Marozzo, Francesc Lordan, Roger Rafanell, Daniele Lezzi, Domenico Talia, Rosa M. Badia |
Euro-Par | 5 |
| 2012 | Topic 5: Parallel and Distributed Data Management
Domenico Talia, Alex Delis, Haimonti Dutta, Arkady B. Zaslavsky |
Euro-Par | 1 |
| 2012 | Distributed data mining patterns and services: an architecture and experimentsabstractSUMMARY Distributed data mining implements techniques for analyzing data on distributed computing systems by exploiting data distribution and parallel algorithms. The grid is a computing infrastructure for implementing distributed high‐performance applications and solving complex problems, offering effective support to the implementation and use of data mining and knowledge discovery systems. The Web Services Resource Framework has become the standard for the implementation of grid services and applications, and it can be exploited for developing high‐level services for distributed data mining applications. This paper describes how distributed data mining patterns, such as collective learning, ensemble learning, and meta‐learning models, can be implemented as Web Services Resource Framework mining services by exploiting the grid infrastructure. The goal of this work was to design a distributed architectural model that can be exploited for different distributed mining patterns deployed as grid services for the analysis of dispersed data sources. In order to validate such an approach, we presented also the implementation of two clustering algorithms on the developed architecture. In particular, the distributed k‐means and distributed expectation maximization were exploited as pilot examples to show the suitability of the implemented service‐oriented framework. An extensive evaluation of its performance was provided. Copyright © 2011 John Wiley & Sons, Ltd. Eugenio Cesario, Domenico Talia |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | A DHT-based semantic overlay network for service discovery
Giuseppe Pirrò, Domenico Talia, Paolo Trunfio |
Future Gener. Comput. Syst. | 2 |
| 2012 | Energy efficiency in large-scale distributed systems
Tuan Anh Trinh, Helmut Hlavacs, Domenico Talia |
Future Gener. Comput. Syst. | 3 |
| 2012 | P2P-MapReduce: Parallel data processing in dynamic Cloud environments
Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
J. Comput. Syst. Sci. | 2 |
| 2011 | A Sketch-Based Architecture for Mining Frequent Items and Itemsets from Distributed Data StreamsabstractThis paper presents the design and the implementation of an architecture for the analysis of data streams in distributed environments. In particular, data stream analysis has been carried out for the computation of items and item sets that exceed a frequency threshold. The mining approach is hybrid, that is, frequent items are calculated with a single pass, using a sketch algorithm, while frequent item sets are calculated by a further multi-pass analysis. The architecture combines parallel and distributed processing to keep the pace with the rate of distributed data streams. In order to keep computation close to data, miners are distributed among the domains where data streams are generated. The paper also reports the experimental results obtained with a prototype of the architecture, tested on a Grid composed of two domains handling two different data streams. Eugenio Cesario, Antonio Grillo, Carlo Mastroianni, Domenico Talia |
CCGRID | 4 |
| 2011 | A Cloud Framework for Parameter Sweeping Data Mining ApplicationsabstractData mining techniques are used in many application areas to extract useful knowledge from large datasets. Very often, parameter sweeping is used in data mining applications to explore the effects produced on the data analysis result by different values of the algorithm parameters. Parameter sweeping applications can be highly computing demanding, since the number of single tasks to be executed increases with the number of swept parameters and the range of their values. Cloud technologies can be effectively exploited to provide end-users with the computing and storage resources, and the execution mechanisms needed to efficiently run this class of applications. In this paper, we present a Data Mining Cloud App framework that supports the execution of parameter sweeping data mining applications on a Cloud. The framework has been implemented using the Windows Azure platform, and evaluated through a set of parameter sweeping clustering and classification applications. The experimental results demonstrate the effectiveness of the proposed framework, as well as the scalability that can be achieved through the parallel execution of parameter sweeping applications on a pool of virtual servers. Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
CloudCom | 2 |
| 2011 | Energy Efficient Task Allocation over Mobile NetworksabstractIn this paper we present an Energy-Aware Scheduling strategy that assigns computational tasks over a network of mobile devices optimizing the energy usage. The main design principle of our scheduler is to find a task allocation that prolongs network lifetime by balancing the energy load among the devices. We have evaluated the scheduler using a prototype of the system that includes smart phones and Android emulators. Experimental results show that significant energy savings can be achieved by using our energy-aware scheduler compared to classical time-based scheduler, while meeting the specified performance constraints. Carmela Comito, Deborah Falcone, Domenico Talia, Paolo Trunfio |
DASC | 3 |
| 2011 | A Failure Handling Framework for Distributed Data Mining Services on the GridabstractFault tolerance is an important issue in Grid computing, where many and heterogenous machines are used. In this paper we present a flexible failure handling framework which extends a service-oriented architecture for Distributed Data Mining previously proposed, addressing the requirements for fault tolerance in the Grid. The framework allows users to achieve failure recovery whenever a crash can occur on a Grid node involved in the computation. The implemented framework has been evaluated on a real Grid setting to assess its effectiveness and performance. Eugenio Cesario, Domenico Talia |
PDP | 2 |
| 2011 | A Framework for Managing MapReduce Applications in Dynamic Distributed EnvironmentsabstractMapReduce is a programming model widely used in data centers for processing large data sets in a highly parallel way. Current MapReduce systems are based on master-slave architectures that do not cope well with dynamic node participation, since they are mostly designed for conventional parallel computing platforms. On the contrary, in Internet-based computing environments, node churn and failures - including master failures - are likely to happen since nodes join and leave the network at an unpredictable rate. The goal of this work is enabling the use of MapReduce in dynamic distributed environments so as to combine the effectiveness of a well-established programming model with the scalability of a large-scale computing infrastructure. This paper presents an adaptive MapReduce framework, called P2P-MapReduce, which exploits a peer-to-peer model to manage intermittent node participation, master failures and job recovery in a decentralized but effective way, so as to provide a more robust MapReduce middleware that can be effectively exploited in Internet-scale dynamic distributed environments. Fabrizio Marozzo, Domenico Talia, Paolo Trunfio |
PDP | 2 |
| 2011 | P2P schema-mapping over network-bound XML dataabstractAbstract The rise in availability of web‐based data sources has led to new challenges in data integration systems for obtaining decentralized, wide‐scale sharing of data preserving semantics. In this paper, we present a framework for integrating heterogeneous XML data sources distributed over a large‐scale, highly dynamic network of autonomous nodes. We highlight a query reformulation algorithm to combine and query‐distributed XML databases through a decentralized point‐to‐point mediation process among the different data sources by using P2P schema‐mappings. More precisely, our integration model is based on path‐to‐path mappings, using the XPath language. We demonstrate the usefulness and scalability of our ideas and algorithms with a detailed set of experiments. Finally, we discuss our experience implementing the above‐cited query reformulation algorithm as a Web service within the GDIS system, a service‐based Grid architecture. We have evaluated GDIS on several real‐world schemas with promising results. Copyright © 2010 John Wiley & Sons, Ltd. Carmela Comito, Domenico Talia |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Grid-Enabled Virtual Organizations for Next-Generation Learning EnvironmentsabstractNowadays, we are witnesses of a transformation in the e-learning arena. This transformation has different drivers involving all the actors in the learning value chain, from final users to learning institutions through technology providers. All those actors share a common goal: making the learning processes more effective through the information and communication technologies. This is happening through the promotion of a paradigm shift from content-centered to process-centered solutions. In this paper, we present the results from the European Learning Grid Infrastructure project concerning models, processes, and services supported by a service-oriented software architecture for creating dynamic and adaptive virtual organizations for learning using Grid technologies. Matteo Gaeta, Pierluigi Ritrovato, Domenico Talia |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2010 | ERGOT: A Semantic-Based System for Service Discovery in Distributed InfrastructuresabstractThe increasing number of available online services demands distributed architectures to promote scalability as well as semantics to enable their precise and efficient retrieval. Two common approaches toward this goal are Semantic Overlay Networks (SONs) and Distributed Hash Tables (DHTs) with semantic extensions. This paper presents ERGOT, a system that combines DHTs and SONs to enable semantic-based service discovery in distributed infrastructures such as Grids and Clouds. ERGOT takes advantage of semantic annotations that enrich service specifications in two ways: (i) services are advertised in the DHT on the basis of their annotations, thus allowing to establish a SON among service providers, (ii) annotations enable semantic-based service matchmaking, using a novel similarity measure between service requests and descriptions. Experimental evaluations confirmed the efficiency of ERGOT in terms of accuracy of search and network traffic. Giuseppe Pirrò, Paolo Trunfio, Domenico Talia, Paolo Missier, Carole A. Goble |
CCGRID | 3 |
| 2010 | A logic approach to virtual sensor networksabstractThis paper presents a technique that builds a layer of virtual sensors over a sensor network. The virtual sensors are able to infer and provide data for the physical sensors that do not work. The key assumption of our approach is that the physical quantities sensed by the sensors are related. The relations among sensors are unknown, but during a learning phase the layer of virtual sensors infers an approximation of them by means of fuzzy rules. The inferred fuzzy rules capture these relations in a simple way even when the corresponding mathematical models are complex. The set of fuzzy rules inferred for a node can be used to obtain virtual values when the real ones are not available. In order to develop our technique we improved the Tree Routing Protocol in charge to deliver data from the nodes to the base station and used Snlog, a Datalog-like language that supports the implementation of distributed algorithms for Wireless Sensor Network in a declarative way. We developed a system prototype and performed preliminary experiments that prove the validity of our approach. Luciano Caroprese, Carmela Comito, Domenico Talia, Ester Zumpano |
IDEAS | 3 |
| 2010 | Selectivity-based XML query processing in structured peer-to-peer networksabstractDHT-based structured P2P systems have been proposed to index and retrieve many types of contents, including distributed collections of XML documents. During the query processing, a DHT can be used to efficiently identify all nodes storing relevant documents. Carmela Comito, Domenico Talia, Paolo Trunfio |
IDEAS | 2 |
| 2010 | Mining@home: toward a public-resource computing framework for distributed data miningabstractAbstract Several classes of scientific and commercial applications require the execution of a large number of independent tasks. One highly successful and low‐cost mechanism for acquiring the necessary computing power for these applications is the ‘public‐resource computing’, or ‘desktop Grid’ paradigm, which exploits the computational power of private computers. So far, this paradigm has not been applied to data mining applications for two main reasons. First, it is not straightforward to decompose a data mining algorithm into truly independent sub‐tasks. Second, the large volume of the involved data makes it difficult to handle the communication costs of a parallel paradigm. This paper introduces a general framework for distributed data mining applications called Mining@home. In particular, we focus on one of the main data mining problems: the extraction of closed frequent itemsets from transactional databases. We show that it is possible to decompose this problem into independent tasks, which however need to share a large volume of the data. We thus introduce a data‐intensive computing network, which adopts a P2P topology based on super peers with caching capabilities, aiming to support the dissemination of large amounts of information. Finally, we evaluate the execution of a pattern extraction task on such network. Copyright © 2009 John Wiley & Sons, Ltd. Claudio Lucchese, Carlo Mastroianni, Salvatore Orlando 0001, Domenico Talia |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | UFOme: An ontology mapping system with strategy prediction capabilities
Giuseppe Pirrò, Domenico Talia |
Data Knowl. Eng. | 2 |
| 2010 | Special Section: Grid computing, high-performance and distributed applications
Pilar Herrero, Daniel S. Katz, María S. Pérez 0001, Domenico Talia |
Future Gener. Comput. Syst. | 4 |
| 2010 | A framework for distributed knowledge management: Design and implementation
Giuseppe Pirrò, Carlo Mastroianni, Domenico Talia |
Future Gener. Comput. Syst. | 3 |
| 2010 | Perspectives on grid computing
Uwe Schwiegelshohn, Rosa M. Badia, Marian Bubak, Marco Danelutto, Schahram Dustdar, Fabrizio Gagliardi, Alfred Geiger, Ladislav Hluchý, Dieter Kranzlmüller, Erwin Laure, Thierry Priol, Alexander Reinefeld, Michael M. Resch, Andreas Reuter 0001, Otto Rienhoff, Thomas Rüter, Peter M. A. Sloot, Domenico Talia, Klaus Ullmann, Ramin Yahyapour |
Future Gener. Comput. Syst. | 18 |
| 2010 | Enabling Dynamic Querying over Distributed Hash Tables
Domenico Talia, Paolo Trunfio |
J. Parallel Distributed Comput. | 1 |
| 2009 | Introduction
Domenico Talia, Jason Maassen, Fabrice Huet, Shantenu Jha |
Euro-Par | 1 |
| 2009 | A semantic-aware information system for multi-domain applications over service gridsabstractService-oriented Grid frameworks offer resources and facilities to support the design and execution of distributed applications in different domains, ranging from scientific applications and public computing projects to commercial and industrial applications. A critical issue in such a context is the management of the heterogeneity of resources and services offered by a Grid, including computers, data, and software tools provided by different organizations. This paper presents a general architecture of a service-oriented information system, which exploits the characteristics of a multi-domain and semantically enriched metadata model. The main objective of the information system is to uniformly manage service-oriented applications and basic resources by assuring metadata persistence through an XML distributed database, without merely relying on the functionalities of persistent Grid services. The information system has been implemented on the basic services of the WSRF-based Globus Toolkit 4 and its performance has been evaluated in a testbed. Carmela Comito, Carlo Mastroianni, Domenico Talia |
IPDPS | 3 |
| 2009 | Combining DHTs and SONs for Semantic-Based Service DiscoveryabstractThe soaring number of available online services calls for distributed architectures to promote scalability, fault- tolerance and semantics; to provide meaningful descriptions of services; and to support their efficient retrieval. Current approaches exploit either Semantic Overlay Networks (SONs) or Distributed Hash Tables (DHTs) sweetened with some ”semantic sugar.” SONs enable semantic driven query answering but are less scalable than DHTs, which on their turn, feature efficient but semantic-free query answering based on ”exact” match. This paper presents the ERGOT system combining DHTs and SONs to enable distributed and semantic-based service discovery. A preliminary evaluation of the system performance shows the suitability of the approach both in terms of recall and number of messages. Giuseppe Pirrò, Paolo Missier, Paolo Trunfio, Domenico Talia, Gabriele Falace, Carole A. Goble |
ISDA | 4 |
| 2009 | A service-oriented system for distributed data querying and integration on Grids
Carmela Comito, Anastasios Gounaris, Rizos Sakellariou, Domenico Talia |
Future Gener. Comput. Syst. | 4 |
| 2009 | A scalable super-peer approach for public scientific computation
Carlo Mastroianni, Pasquale Cozza, Domenico Talia, Ian Kelley, Ian J. Taylor |
Future Gener. Comput. Syst. | 3 |
| 2008 | Topic 5: Parallel and Distributed Databases
Domenico Talia, Josep Lluís Larriba-Pey, Hillol Kargupta, Esther Pacitti |
Euro-Par | 1 |
| 2008 | An Algorithm for Discovering Ontology Mappings in P2P Systems
Giuseppe Pirrò, Massimo Ruffolo, Domenico Talia |
KES (2) | 3 |
| 2008 | The Weka4WS framework for distributed data mining in service-oriented GridsabstractAbstract The service‐oriented architecture paradigm can be exploited for the implementation of data and knowledge‐based applications in distributed environments. The Web services resource framework (WSRF) has recently emerged as the standard for the implementation of Grid services and applications. WSRF can be exploited for developing high‐level services for distributed data mining applications. This paper describes Weka4WS, a framework that extends the widely used open source Weka toolkit to support distributed data mining on WSRF‐enabled Grids. Weka4WS adopts the WSRF technology for running remote data mining algorithms and managing distributed computations. The Weka4WS user interface supports the execution of both local and remote data mining tasks. On every computing node, a WSRF‐compliant Web service is used to expose all the data mining algorithms provided by the Weka library. The paper describes the design and implementation of Weka4WS using the WSRF libraries and services provided by Globus Toolkit 4. A performance analysis of Weka4WS for executing distributed data mining tasks in different network scenarios is presented. Copyright © 2008 John Wiley & Sons, Ltd. Domenico Talia, Paolo Trunfio, Oreste Verta |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | SIGMCC: A system for sharing meta patient records in a Peer-to-Peer environment
Mario Cannataro, Domenico Talia, Giuseppe Tradigo, Paolo Trunfio, Pierangelo Veltri |
Future Gener. Comput. Syst. | 2 |
| 2008 | Modeling and Supporting Grid Scheduling
Andrea Pugliese 0001, Domenico Talia, Ramin Yahyapour |
J. Grid Comput. | 2 |
| 2008 | Service-oriented middleware for distributed data mining on the grid
Antonio Congiusta, Domenico Talia, Paolo Trunfio |
J. Parallel Distributed Comput. | 2 |
| 2008 | Designing an information system for Grids: Comparing hierarchical, decentralized P2P and super-peer models
Carlo Mastroianni, Domenico Talia, Oreste Verta |
Parallel Comput. | 2 |
| 2007 | A Service-Oriented System to Support Data Integration on Data GridsabstractData Grids provide transparent access to heterogeneous and autonomous data resources. The main contribution of this paper is the presentation of a data sharing system that (i) is tailored to data grids, (ii) supports well established and widely spread relational DBMSs, and (iii) adopts a hybrid architecture by relying on a peer model for query reformulation for retrieving semantically equivalent expressions, and on a wrapper-mediator integration model for accessing and querying distributed data sources. The system builds upon the infrastructure provided by the OGSA-DQP distributed query processor and the XMAP query reformulation algorithm. The paper discusses the implementation methodology, and also presents empirical evaluation results. Anastasios Gounaris, Carmela Comito, Rizos Sakellariou, Domenico Talia |
CCGRID | 4 |
| 2007 | Evaluating Resource Discovery Protocols for Hierarchical and Super-Peer Grid Information SystemsabstractMost currently deployed grids adopt a hierarchical model for their information system. However, nowadays the research and development community is heading towards the use of scalable models of information services based on decentralized approaches such as the peer-to-peer paradigm. This is mainly due to the poor scalability, resiliency and load-balancing features of the hierarchical model. This paper evaluates a resource discovery protocol exploitable in a hierarchical grid and compares it with a super-peer based model which has recently been introduced. Performance analysis, carried out through simulation, shows that the hierarchical model is valuable for small and medium sized grids, while the super-peer model is better suited for very large grids Carlo Mastroianni, Domenico Talia, Oreste Verta |
PDP | 2 |
| 2007 | Using Grids for Exploiting Data Abundance in ScienceabstractDigital data volumes are growing exponentially both in all sciences. To handle this abundance in data availability, scientists must embody data analysis techniques in their scientific practices and solving environments to get the benefits coming from knowledge that can be extracted from large data sources. Domenico Talia |
PDP | 1 |
| 2007 | Distributed data mining services leveraging WSRF
Antonio Congiusta, Domenico Talia, Paolo Trunfio |
Future Gener. Comput. Syst. | 2 |
| 2007 | Peer-to-Peer resource discovery in Grids: Models and systems
Paolo Trunfio, Domenico Talia, Harris Papadakis, Paraskevi Fragopoulou, Matteo Mordacchini, M. Pennanen, Konstantin Popov, Vladimir Vlassov, Seif Haridi |
Future Gener. Comput. Syst. | 2 |
| 2006 | Topic 5: Parallel and Distributed Databases, Data Mining and Knowledge Discovery
Patrick Valduriez, Wolfgang Lehner, Domenico Talia, Paul Watson 0001 |
Euro-Par | 3 |
| 2006 | WSRF Services for Composing Distributed Data Mining Applications on Grids: Functionality and Performance
Domenico Talia, Paolo Trunfio, Oreste Verta |
ICCSA (1) | 1 |
| 2006 | A Semantic Overlay Network for P2P Schema-Based Data IntegrationabstractToday data sources are pervasive and their number is growing tremendously. Current tools are not prepared to exploit this unprecedented amount of information and to cope with this highly heterogeneous, autonomous and dynamic environment. In this paper, we propose a novel semantic overlay network architecture, PARIS, aimed at addressing these issues. In PARIS, the combination of decentralized semantic data integration with gossip-based (unstructured) overlay topology management and (structured) distributed hash tables provides the required level of flexibility, adaptability and scalability, and still allows to perform rich queries on a number of autonomous data sources. We describe the logical model that supports the architecture and show how its original topology is constructed. We present the usage of the system in detail, in particular, the algorithms used to let new peers join the network and to execute queries on top of it and show simulation results that assess the scalability and robustness of the architecture. Carmela Comito, Simon Patarin, Domenico Talia |
ISCC | 3 |
| 2005 | Topic 5 - Parallel and Distributed Databases, Data Mining and Knowledge Discovery
Domenico Talia, Hillol Kargupta, Patrick Valduriez, Rui Camacho |
Euro-Par | 1 |
| 2005 | A Metadata Model and Information System for the Management of Resources in a Grid-Based PSE Toolkit
Carmela Comito, Carlo Mastroianni, Domenico Talia |
HPCC | 3 |
| 2005 | Weka4WS: A WSRF-Enabled Weka Toolkit for Distributed Data Mining on Grids
Domenico Talia, Paolo Trunfio, Oreste Verta |
PKDD | 1 |
| 2005 | P2P computing and interaction with grids
Adriana Iamnitchi, Domenico Talia |
Future Gener. Comput. Syst. | 2 |
| 2005 | A super-peer model for resource discovery services in large-scale Grids
Carlo Mastroianni, Domenico Talia, Oreste Verta |
Future Gener. Comput. Syst. | 2 |
| 2004 | Topic 16: Integrated Problem Solving Environments
Daniela di Serafino, Elias N. Houstis, Peter M. A. Sloot, Domenico Talia |
Euro-Par | 4 |
| 2004 | A P2P Grid Services-Based Protocol: Design and Evaluation
Domenico Talia, Paolo Trunfio |
Euro-Par | 1 |
| 2004 | Application-Oriented Scheduling in the Knowledge Grid: A Model and Architecture
Andrea Pugliese 0001, Domenico Talia |
ICCSA (2) | 2 |
| 2004 | Metadata for Managing Grid Resources in Data Mining Applications
Carlo Mastroianni, Domenico Talia, Paolo Trunfio |
J. Grid Comput. | 2 |
| 2004 | Distributed data mining on grids: services, tools, and applicationsabstractData mining algorithms are widely used today for the analysis of large corporate and scientific datasets stored in databases and data archives. Industry, science, and commerce fields often need to analyze very large datasets maintained over geographically distributed sites by using the computational power of distributed and parallel systems. The grid can play a significant role in providing an effective computational support for distributed knowledge discovery applications. For the development of data mining applications on grids we designed a system called Knowledge Grid. This paper describes the Knowledge Grid framework and presents the toolset provided by the Knowledge Grid for implementing distributed knowledge discovery. The paper discusses how to design and implement data mining applications by using the Knowledge Grid tools starting from searching grid resources, composing software and data components, and executing the resulting data mining process on a grid. Some performance results are also discussed. Mario Cannataro, Antonio Congiusta, Andrea Pugliese 0001, Domenico Talia, Paolo Trunfio |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2003 | Topic Introduction
Bernhard Mitschang, David B. Skillicorn, Philippe Bonnet, Domenico Talia |
Euro-Par | 4 |
| 2003 | Knowledge Discovery Services and Tools on Grids
Domenico Talia |
ISMIS | 1 |
| 2003 | P-AutoClass: Scalable Parallel Clustering for Mining Large Data SetsabstractData clustering is an important task in the area of data mining. Clustering is the unsupervised classification of data items into homogeneous groups called clusters. Clustering methods partition a set of data items into clusters, such that items in the same cluster are more similar to each other than items in different clusters according to some defined criteria. Clustering algorithms are computationally intensive, particularly when they are used to analyze large amounts of data. A possible approach to reduce the processing time is based on the implementation of clustering algorithms on scalable parallel computers. This paper describes the design and implementation of P-AutoClass, a parallel version of the AutoClass system based upon the Bayesian model for determining optimal classes in large data sets. The P-AutoClass implementation divides the clustering task among the processors of a multicomputer so that each processor works on its own partition and exchanges intermediate results with the other processors. The system architecture, its implementation, and experimental performance results on different processor numbers and data sets are presented and compared with theoretical performance. In particular, experimental and predicted scalability and efficiency of P-AutoClass versus the sequential AutoClass system are evaluated and compared. Clara Pizzuti, Domenico Talia |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2003 | Designing Grid services for distributed knowledge discovery
Antonio Congiusta, Andrea Pugliese 0001, Domenico Talia, Paolo Trunfio |
Web Intell. Agent Syst. | 3 |
| 2002 | Eureka! : A Tool for Interactive Knowledge Discovery
Giuseppe Manco 0001, Clara Pizzuti, Domenico Talia |
DEXA | 3 |
| 2002 | Parallel and Distributed Databases, Data Mining and Knowledge Discovery
Harald Kosch, David B. Skillicorn, Domenico Talia |
Euro-Par | 3 |
| 2002 | Distributed data mining on the grid
Mario Cannataro, Domenico Talia, Paolo Trunfio |
Future Gener. Comput. Syst. | 2 |
| 2002 | Parallel data intensive computing in scientific and commercial applications
Mario Cannataro, Domenico Talia, Pradip K. Srimani |
Parallel Comput. | 2 |
| 2002 | Parallel data-intensive algorithms and applications (guest editorial)
Domenico Talia, Pradip K. Srimani |
Parallel Comput. | 1 |
| 2000 | Guest Editor's Introduction: Special Issues on Architecture-Independent Languages and Software tools for Parallel Processing
Domenico Talia, Pradip K. Srimani, Mehdi Jazayeri |
IEEE Trans. Software Eng. | 1 |
| 2000 | Guest Editor's Introduction: Special Issues on Architecture-Independent Languages and Software tools for Parallel Processing
Domenico Talia, Pradip K. Srimani, Mehdi Jazayeri |
IEEE Trans. Software Eng. | 1 |
| 1999 | High Performance Data Mining and Knowledge Discovery - Introduction
David B. Skillicorn, Domenico Talia |
Euro-Par | 2 |
| 1999 | A Divise Initialisation Method for Clustering Algorithms
Clara Pizzuti, Domenico Talia, Giorgio Vonella |
PKDD | 2 |
| 1999 | Programming cellular automata algorithms on parallel computers
Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 2 |
| 1999 | Cellular Automata: Promise and Prospects in Computational Science
Domenico Talia, Peter M. A. Sloot |
Future Gener. Comput. Syst. | 1 |
| 1998 | Language Constructs and Run-Time System for Parallel Cellular Programming
Giandomenico Spezzano, Domenico Talia |
Euro-Par | 2 |
| 1998 | Designing parallel models of soil contamination by the CARPET language
Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 2 |
| 1997 | High performance scientific computing by a parallel cellular environment
Salvatore Di Gregorio, Rocco Rongo, William Spataro, Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 5 |
| 1995 | A Parallel Cellular Automata Environment on Multicomputers for Computational Science
Mario Cannataro, Salvatore Di Gregorio, Rocco Rongo, William Spataro, Giandomenico Spezzano, Domenico Talia |
Parallel Comput. | 6 |
| 1993 | A Survey of Parlog and Concurrent Prolog: The Integration of Logic and Parallelism
Domenico Talia |
Comput. Lang. | 1 |
| 1993 | Distributed termination of concurrent processes in Occam
Domenico Talia |
Comput. Lang. | 1 |
| 1992 | Design, implementation and evaluation of a deadlock-free routing algorithm for concurrent computersabstractAbstract This paper describes the design, the implementation, and the performance results of a routing algorithm which provides deadlock‐free communication in a tightly coupled message‐passing concurrent computer. The algorithm is adaptive, isolated and uses the store‐and‐forward technique. It allows message communication between two processes regardless of where they are physically located on the network. The routing algorithm has many positive characteristics including provable deadlock freedom, guaranteed message arrival, and automatic local congestion reduction. It can be used as a basis for the design of high‐level communication primitives. An Occam implementation on a network of inmos Transputers is discussed. The experimental results show that the routing algorithm is effective to support process to process communication on a concurrent computer. Mario Cannataro, Giandomenico Spezzano, Domenico Talia, E. Gallizzi |
Concurr. Pract. Exp. | 3 |
| 1992 | High level communication mechanisms for distributed parallel computers from an adaptive message routing
Mario Cannataro, Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 3 |
| 1992 | A model of efficient asynchronous parallel algorithms on multicomputer systems
Domenico Conforti, Lucio Grandinetti, Roberto Musmanno, Mario Cannataro, Giandomenico Spezzano, Domenico Talia |
Parallel Comput. | 6 |
| 1991 | A parallel logic system on a multicomputer architecture
Mario Cannataro, Giandomenico Spezzano, Domenico Talia |
Future Gener. Comput. Syst. | 3 |