Giuseppe Psaila

dblp:85/5528 · DBLP profile ↗
← Back
49ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-9228-560XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 29 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Bayesian generation of synthetic datasets for machine-learning tasks: a performance study
abstract
Performing Machine Learning (ML) tasks on large-scale datasets, as well as simply storing them for subsequent analysis or for long-term archival, could require large computational power. The described approach builds on the technique known as “Bayesian Generation” to produce synthetic datasets in such a way that the probability distribution in the source dataset is maintained as much as possible in the new synthetic ones, even if they are much smaller than the original (large) dataset. In fact, this study investigates the impact of generating smaller synthetic datasets for training ML models in place of the original dataset, adopting a twofold perspective. Firstly, the impact on the effectiveness of ML models trained on these smaller synthetic datasets is assessed. Secondly, the amount of computational resources required to generate the synthetic datasets, train ML models on them, and perform the testing phase is measured. Specifically, both execution time and main memory usage are taken into account. Finally, this research work shows that the loss in terms of effectiveness remains consistently limited and stable, and it identifies the scenarios and ML techniques for which incorporating the generation of small synthetic datasets into the ML pipeline can be beneficial for practical deployment in environments with constrained computational resources, such as mobile or industrial devices.
Paolo Fosci, Javier Nieves, Giuseppe Psaila, Jacopo Boffelli, Pablo García Bringas
Neurocomputing3
2025 Detecting Semantic Relationships Among Datasets
Paolo Fosci, Vincenzo Carbone, Matteo Leo, Andrea Marmorato, Giuseppe Psaila, Giampiero Rosa, Mohammadsadegh Torabi
FQAS5
2025 Linguistic Analogies in Word Embeddings: Where Are They?
Riccardo Contessi, Paolo Fosci, Giuseppe Psaila
WEBIST3
2025 Evolving J-CO-QL+ with fuzzy evaluators for flexible queryisng of JSON data sets
abstract
How to introduce soft querying based on fuzzy sets in the novel J-CO-QL + query language (specifically designed to query collections of JSON documents from NoSQL databases) has been investigated by the authors in their past work. Specifically, capabilities for defining fuzzy operators and fuzzy aggregators were introduced through two distinct concepts on which two different language constructs were based. This paper proposes the unified concept of “fuzzy evaluator”, by means of which it is possible to define complex methods for evaluating the membership degrees of JSON documents to fuzzy sets, so as to capture complex semantics while analyzing data in a soft way. The paper both provides a formal meta-model for fuzzy evaluators, and proposes a novel statement for the J-CO-QL + language, so as to further foster soft-querying capabilities.
Paolo Fosci, Giuseppe Psaila
Neurocomputing2
2024 Soft Querying JSON Datasets with Personalized Preferences and Aggregations
abstract
Soft conditions are a powerful and established formal tool to select data on the basis of linguistic predicates. In previous work, the J-CO Framework (and its query language) was used to perform Soft Web Intelligence, i.e., a practical interpretation of the concept of Web Intelligence that exploits soft conditions to search for desired items in JSON datasets acquired from Web sources. However, the effectiveness of soft conditions depends on how elementary conditions are combined: in this sense, a plethora of proposals are available, such as the vector p-norm. This paper shows how a generic concept, named "user-defined fuzzy evaluator", that has been recently introduced in the query language, actually allows users to define their own operators, so as to express advanced operators such as "and possibly". The paper also shows how the AND operator defined as a vector p-norm actually behaves, depending on different configurations of parameters, so as to let the reader understand how to use it in practice.
Paolo Fosci, Giuseppe Psaila
WEBIST2
2024 A unified view of multi-grade fuzzy-set models in J-CO-QL+
abstract
The complexity of reality has driven the evolution of Fuzzy-Set Theory from the initial proposal made by Zadeh in 1965, towards more complex models. Moving from a quick survey of the evolution of Fuzzy-Set Theory, this paper highlights the aspects that are common to many Fuzzy-Set Models, in order to define a meta-model that is capable of providing a unified view to a wide variety of fuzzy-set models. In particular, this work focuses the attention on the family of “Multi-grade Fuzzy Sets”, which are fuzzy sets characterized by more than one degree. The lack of tools capable of querying the large amount of data that are nowadays available in NoSQL databases, has pushed us to devise the J-CO Framework: it is a platform-independent tool that is capable to manage, transform and query collections of JSON documents; the J-CO Framework relies on J-CO-QL+, which is a high-level, general-purpose language with soft-querying capabilities. The latest advancements of J-CO-QL+ allow for defining and exploiting user-defined Multi-grade Fuzzy-Set Models and Operators. In the paper, a case-study demonstrates the effectiveness of the J-CO Framework in performing a non-trivial soft query based on a Multi-grade Fuzzy-Set Model defined by the user.
Paolo Fosci, Giuseppe Psaila
Neurocomputing2
2023 Enhancing Soft Web Intelligence with User-Defined Fuzzy Aggregators
abstract
In our previous work, we proposed Soft Web Intelligence as the interpretation of the general notion of Web Intelligence in the current technological panorama, in such a way JSON data sets are acquired from the Internet, stored within JSON document stores and then processed and queried by means of soft computing and soft querying methods. Specific extensions to the J-CO Framework and to its query language (named J-CO-QL+) made possible to practically implement the concept. However, any “data intelligence” activity does not exclude aggregating data, but J-CO-QL+ did not provide statements for defining “user-defined fuzzy aggregators”. In this paper, we present the novel constructs introduced into J-CO-QL+ to allow users to define and use their own fuzzy aggregators, so as to evaluate membership degrees to fuzzy sets moving from array fields within processed JSON documents. This way, complex soft queries are enabled, so as to enhance Soft Web Intelligence.
Paolo Fosci, Giuseppe Psaila
WEBIST2
2023 Soft querying powered by user-defined functions in J-CO-QL+
Paolo Fosci, Giuseppe Psaila
Neurocomputing2
2022 Soft Spatial Querying on JSON Data Sets
Paolo Fosci, Giuseppe Psaila
ADBIS2
2022 Towards Soft Web Intelligence by Collecting and Processing JSON Data Sets from Web Sources
abstract
Since the last two decades, Web Intelligence has denoted a plethora of approaches to discover useful knowledge from the vast World-Wide Web; however, dealing with the immense variety of the Web is not easy and the challenge is still open. In this paper, we moved from the previous functionalities provided by the J-CO Framework (a research project under development at University of Bergamo Italy), to identify a vision ofWeb Intelligence scopes in which capabilities of soft computing and soft querying provided by a stand-alone tool can actually create novel possibilities of making useful analysis of JSON data sets directly coming from Web sources. The paper identifies some extensions to the J-CO Framework, which we implemented; then it shows an example of soft querying enabled by these extensions.
Paolo Fosci, Giuseppe Psaila
WEBIST2
2021 J-CO, A Framework for Fuzzy Querying Collections of JSON Documents (Demo)
Paolo Fosci, Giuseppe Psaila
FQAS2
2021 Quality assessment methodology based on machine learning with small datasets: Industrial castings defects
Iker Pastor-López, Borja Sanz 0001, Alberto Tellaeche, Giuseppe Psaila, José Gaviria de la Puerta, Pablo García Bringas
Neurocomputing4
2020 Soft Querying GeoJSON Documents within the J-CO Framework
abstract
GeoJSON documents have become important sources of information over the Web, because they describe geographical information layers. Supposing to have such documents stored in some JSON store, the problem of querying them in a flexible and easy way arises. In this paper, we propose a soft-querying model to easily express queries on features (i.e., data items) within GeoJSON documents, based on linguistic predicates. These are fuzzy predicates that evaluate the membership degree to fuzzy sets; this way, imprecise conditions can be expressed and features can be ranked, accordingly. The paper presents a rewriting technique that translates soft queries on GeoJSON documents into fuzzy JCO-QL queries: this is the query language of the J-CO Framework, an Internet-based framework able to get, manipulate and save collections of JSON documents in a way totally independent of the source JSON store.
Giuseppe Psaila, Stefania Marrara, Paolo Fosci
WEBIST1
2019 Can BlockChain Technology Provide Information Systems with Trusted Database? The Case of HyperLedger Fabric
Pablo García Bringas, Iker Pastor-López, Giuseppe Psaila
FQAS3
2017 The Challenge of using Map-reduce to Query Open Data
abstract
For transparency and democracy reasons, a few years ago Public Administrations started publishing data sets concerning public services and territories. These data sets are called open, because they are publicly available through many web sites. Due to the rapid growth of open data corpora, both in terms of number of corpora and in terms of open data sets available in each single corpus, the need for a centralized query engine arises, able to select single data items from within a mess of heterogeneous open data sets. We gave a first answer to this need in (Pelucchi et al., 2017), where we defined a technique for blindly querying a corpus of open data. In this paper, we face the challenge of implementing this technique on top of the Map-Reduce approach, the most famous solution to parallelize computational tasks in the Big Data world.
Mauro Pelucchi, Giuseppe Psaila, Maurizio Toccu
DATA2
2017 A flexible framework to cross-analyze heterogeneous multi-source geo-referenced information: the J-CO-QL proposal and its implementation
abstract
The need for cross-analyzing JSON objects representing heterogeneous geo-referenced information coming from multiple sources, such as open data published on the Web by public administrations and crowd-sourced posts and images from social networks, is becoming common for studying, predicting and planning social dynamics. Nevertheless, although NoSQL databases have emerged as a de facto standard means to store JSON objects, a query language that can be easily used by not-programmers to manipulate and correlate such data is still missing. Furthermore, when the information is geo-referenced, we also need both spatial analysis and mapping facilities.
Gloria Bordogna, Daniele E. Ciriello, Giuseppe Psaila
WI3
2017 Building a Query Engine for a Corpus of Open Data
abstract
Public Administrations openly publish many data sets concerning citizens and territories in order to increase the amount of information made available for people, firms and public administrators. As an effect, Open Data corpora has become so huge that it is impossible to deal with them by hand; as a consequence, it is necessary to use tools that include innovative techniques able to query them. In this paper, we present a technique to select open data sets containing specific pieces of information, and retrieve them in a corpus published by a portal of open data. In particular, users can formulate structured queries blindly submitted to our search engine prototype (i.e., being unaware of the actual structure of data sets). Our approach reinterpret and mixes several known information retrieval approaches, giving at the same time a database view of the problem. We implemented this technique within a prototype, that we tested on a corpus containing more that over 2000 data sets. We noted that our technique provides focused results w.r.t. the baseline experiments performed with Apache Solr.
Mauro Pelucchi, Giuseppe Psaila, Maurizio Toccu
WEBIST2
2016 An Effective and Efficient Similarity-Matrix-Based Algorithm for Clustering Big Mobile Social Data
abstract
Nowadays a great deal of attention is devoted to the issue of supporting big data analytics over big mobile social data. These data are generated by modern emerging social systems like Twitter, Facebook, Instagram, and so forth. Mining big mobile social data has been of great interest, as analyzing such data is critical for a wide spectrum of big data applications (e.g., smart cities). Among several proposals, clustering is a well-known solution for extracting interesting and actionable knowledge from massive amounts of big mobile (geo-located) social data. Inspired by this main thesis, this paper proposes an effective and efficient similarity-matrix-based algorithm for clustering big mobile social data, called TourMiner, which is specifically targeted to clustering trips extracted from tweets, in order to mine most popular tours. The main characteristic of TourMiner consists in applying clustering over a well-suited similarity matrix computed on top of trips. A comprehensive experimental assessment and analysis over Twitter data finally comfirms the benefits coming from our proposal.
Gloria Bordogna, Luca Frigerio, Alfredo Cuzzocrea, Giuseppe Psaila
ICMLA4
2016 An Innovative Framework for Effectively and Efficiently Supporting Big Data Analytics over Geo-Located Mobile Social Media
abstract
Mobile Social Media are gaining momentum in the broader context of Big Data Analytics, where the main issue is represented by the problem of extracting interesting and actionable knowledge from big data repositories. Mobile social media sources like Twitter and Instagram are indeed producing massive amounts of data (namely, posts) that represent a very rich source of knowledge for predictive analytics. In line with this emerging trend, this paper proposes an innovative approach for effectively and efficiently supporting big data analytics over geo-localized mobile social media, with particular emphasis with the context of modern tourist information systems. In this context, the innovative FollowMe suite, which implements the proposed methodology, is also described in details. We complement our analytical contribution with a real-life case study focusing on the EXPO 2015 event in Milan, Italy which clearly shows benefits and potentialities of our proposed big data analytics framework.
Alfredo Cuzzocrea, Giuseppe Psaila, Maurizio Toccu
IDEAS2
2015 Knowledge Discovery from Geo-Located Tweets for Supporting Advanced Big Data Analytics: A Real-Life Experience
Alfredo Cuzzocrea, Giuseppe Psaila, Maurizio Toccu
MEDI2
2013 The Hints from the Crowd Project
Paolo Fosci, Giuseppe Psaila, Marcello Di Stefano
DEXA (1)2
2013 Hints from the Crowd: A Novel NoSQL Database
Paolo Fosci, Giuseppe Psaila, Marcello Di Stefano
MEDI2
2012 Toward a Product Search Engine based on User Reviews
Paolo Fosci, Giuseppe Psaila
DATA2
2012 Distinct Interpretations of Importance Query Weights in the Vector p - norm Database Model
Gloria Bordogna, Alberto Marcellini, Giuseppe Psaila
IPMU (1)3
2012 Finding the Best Source of Information by means of a Socially-enabled Search Engine
abstract
In the era of Web 2.0, users strongly contribute to blogs, providing their opinions and useful information about specific topics; typically, users that are interested in a topic looks for related blogs in pull mode. But social networks have become even more effective in disseminating opinions and information in push mode, because users directly receive tweets and posts.
Paolo Fosci, Giuseppe Psaila
KES2
2012 Geographic information retrieval: Modeling uncertainty of user's context
Gloria Bordogna, Giorgio Ghisalberti, Giuseppe Psaila
Fuzzy Sets Syst.3
2012 Disambiguated query suggestions and personalized content-similarity and novelty ranking of clustered results to optimize web searches
Gloria Bordogna, Alessandro Campi, Giuseppe Psaila, Stefania Ronchi
Inf. Process. Manag.3
2012 Integrating trust management and access control in data-intensive Web applications
abstract
The widespread diffusion of Web-based services provided by public and private organizations emphasizes the need for a flexible solution for protecting the information accessible through Web applications. A promising approach is represented by credential-based access control and trust management. However, although much research has been done and several proposals exist, a clear obstacle to the realization of their benefits in data-intensive Web applications is represented by the lack of adequate support in the DBMSs. As a matter of fact, DBMSs are often responsible for the management of most of the information that is accessed using a Web browser or a Web service invocation. In this article, we aim at eliminating this gap, and present an approach integrating trust management with the access control of the DBMS. We propose a trust model with a SQL syntax and illustrate an algorithm for the efficient verification of a delegation path for certificates. Our solution nicely complements current trust management proposals allowing the efficient realization of the services of an advanced trust management model within current relational DBMSs. An important benefit of our approach lies in its potential for a robust end-to-end design of security for personal data in Web scenario, where vulnerabilities of Web applications cannot be used to violate the protection of the data residing on the database server. We also illustrate the implementation of our approach within an open-source DBMS discussing design choices and performance impact.
Sabrina De Capitani di Vimercati, Sara Foresti, Sushil Jajodia, Stefano Paraboschi, Giuseppe Psaila, Pierangela Samarati
ACM Trans. Web5
2011 Discovering and Analyzing Multi-granular Web Search Results
Gloria Bordogna, Giuseppe Psaila
FQAS2
2009 Query Disambiguation Based on Novelty and Similarity User's Feedback
Gloria Bordogna, Alessandro Campi, Giuseppe Psaila, Stefania Ronchi
FQAS3
2009 Managing uncertainty in location-based queries
Gloria Bordogna, Marco Pagani, Gabriella Pasi, Giuseppe Psaila
Fuzzy Sets Syst.4
2009 Soft Aggregation in Flexible Databases Querying Based on the Vector P-Norm
abstract
In this paper a model of soft aggregation of preferences in flexible database querying is proposed, based on the vector p-norm operator. The model allows aggregating conditions with distinct importance by modeling both veto and favour semantics of the conditions. We outline how the semantics of the compound query varies for increasing values of the parameter p, when the selection conditions are ANDed and ORed.
Gloria Bordogna, Giuseppe Psaila
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2008 A language for manipulating clustered web documents results
abstract
We propose a novel conception language for exploring the results retrieved by several internet search services (like search engines) that cluster retrieved documents. The goal is to offer users a tool to discover relevant hidden relationships between clustered documents.
Gloria Bordogna, Alessandro Campi, Giuseppe Psaila, Stefania Ronchi
CIKM3
2008 An interaction framework for mobile web search
abstract
This paper describes the architecture and the functionalities of a clustering search engine prototypal system, named Matrioshka, whose peculiarity is the flexibility of the interaction framework on which it is based.
Gloria Bordogna, Alessandro Campi, Giuseppe Psaila, Stefania Ronchi
MoMM3
2007 Flexible Location-Based Spatial Queries
Gloria Bordogna, Marco Pagani, Gabriella Pasi, Giuseppe Psaila
IFSA (2)4
2004 Fuzzy-Spatial SQL
Gloria Bordogna, Giuseppe Psaila
FQAS2
2002 Toward XML-Based Knowledge Discovery Systems
abstract
Inductive databases are intended to be general purpose databases in which both source data and mined patterns can be represented, retrieved and manipulated. However, the heterogeneity of models for mined patterns makes difficult to realize them. In this paper, we explore the feasibility of using XML as the unifying framework for inductive databases, introducing a suitable data model called XDM (XML for data mining). XDM is designed to describe source raw data, heterogeneous mined patterns and data mining statements, so that they can be stored inside a unique XML-based inductive database.
Rosa Meo, Giuseppe Psaila
ICDM2
2001 Partitioning of Hierarchical Automation Systems
abstract
The research described concerns the partitioning of large control applications for a multi-computer system in order to meet plant localization requirements and to exploit parallelism. The considered applications have hierarchical structure and are composed by a network of automata. Our application domain is the automation of power stations and electricity distribution. Because of strong EM noise in such environments, the software architecture is organized to be tolerant to transient faults, which could affect the stability of the control system. The hierarchical structure provides a decompositional approach to the design of complex applications. The context for this work is the ASFA platform, originally designed by the Italian board of electricity. The main result is a new partitioning algorithm for hierarchical automata networks, that splits the application into sub-networks which are deadlock-free, compliant with localization constraints, and as parallelizable as possible. The algorithm is also able to satisfy mutual exclusion constraints and to take into account computation/communication weights to achieve balancing of partitions.
Emanuele Ciapessoni, Francesco Maestri, Judit Szanto, Stefano Crespi-Reghizzi, Andrea C. Ornstein, Giuseppe Psaila
ECRTS6
1999 Incremental Refinement of Mining Queries
Elena Baralis, Giuseppe Psaila
DaWaK2
1999 Discovery of Association Rule Meta-Patterns
Giuseppe Psaila
DaWaK1
1998 A Tightly-Coupled Architecture for Data Mining
abstract
Current approaches to data mining are based on the use of a decoupled architecture, where data are first extracted from a database and then processed by a specialized data mining engine. This paper proposes instead a tightly-coupled architecture, where data mining is integrated within a classical SQL server. The premise of this work is a SQL-like operator, called MINE RULE. We show how the various syntactic features of the operator can be managed by either a SQL engine or a classical data mining engine; our main objective is to identify the border between typical relational processing, executed by the relational server, and data mining processing, executed by a specialized component. The resulting architecture exhibits portability at the SQL level and integration of inputs and outputs of the data mining operator with the database, and provides the guidelines for promoting the integration of other data mining techniques and systems with SQL servers.
Rosa Meo, Giuseppe Psaila, Stefano Ceri
ICDE2
1998 Grammar Partitioning and Modular Deterministic Parsing
Stefano Crespi-Reghizzi, Giuseppe Psaila
Comput. Lang.2
1998 An Extension to SQL for Mining Association Rules
Rosa Meo, Giuseppe Psaila, Stefano Ceri
Data Min. Knowl. Discov.2
1997 Designing Templates for Mining Association Rules
Elena Baralis, Giuseppe Psaila
J. Intell. Inf. Syst.2
1996 Composite Events in Chimera
Rosa Meo, Giuseppe Psaila, Stefano Ceri
EDBT2
1996 A New SQL-like Operator for Mining Association Rules
Rosa Meo, Giuseppe Psaila, Stefano Ceri
VLDB2
1995 Active Data Mining
Rakesh Agrawal 0001, Giuseppe Psaila
KDD2
1995 The Algres Testbed of CHIMERA: An Active Object-Oriented Database System
Stefano Ceri, Piero Fraternali, Stefano Paraboschi, Giuseppe Psaila
SIGMOD Conference4
1995 Querying Shapes of Histories
Rakesh Agrawal 0001, Giuseppe Psaila, Edward L. Wimmers, Mohamed Zaït
VLDB2