Alejandro A. Vaisman

dblp:58/6139 · DBLP profile ↗
← Back
47ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-3945-4187ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 44 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 16 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Reconciling tuple and attribute timestamping for temporal data warehouses
Waqas Ahmed 0003, Leticia I. Gómez, Alejandro A. Vaisman, Esteban Zimányi
VLDB J.3
2024 Querying Mobile Pollution Data using MobilityDB
abstract
Air pollution monitoring requires a large number of expensive devices, especially in large cities. To reduce the cost of this process, the use of mobile devices has been proposed. Some proposals promote the use of cheap sensors on board of public buses instead of just a limited number of specialized vehicles. Analyzing the data provided by these mobile devices is computationally costly and complex with traditional database tools. Instead, we propose using MobilityDB, a novel database implemented as an extension to PostgreSQL and PostGIS, which provides support for storing and querying geospatial trajectory data and their time-varying properties (like air-pollution parameters), implementing persistent database types and query operations for such data. We use public data from the city of Delhi to show the viability and advantages of our approach and how analytical queries are expressed in a concise and elegant way using MobilityDB data types and functions.
Leticia I. Gómez, Alejandro A. Vaisman, Esteban Zimányi
MDM2
2021 A model and query language for temporal graph databases
Ariel Debrouvier, Eliseo Parodi, Matías Perazzo, Valeria Soliani, Alejandro A. Vaisman
VLDB J.5
2020 Design and implementation of ETL processes using BPMN and relational algebra
Judith Awiti, Alejandro A. Vaisman, Esteban Zimányi
Data Knowl. Eng.2
2020 Online analytical processsing on graph data
abstract
Online Analytical Processing (OLAP) comprises tools and algorithms that allow querying multidimensional databases. It is based on the multidimensional model, where data can be seen as a cube such that each cell contains one or more measures that can be aggregated along dimensions. In a “Big Data” s cenario, traditional data warehousing and OLAP operations are clearly not sufficient to address current data analysis requirements, for example, social network analysis. Furthermore, OLAP operations and models can expand the possibilities of graph analysis beyond the traditional graph-based computation. Nevertheless, there is not much work on the problem of taking OLAP analysis to the graph data model. This paper proposes a formal multidimensional model for graph analysis, that considers the basic graph data, and also background information in the form of dimension hierarchies. The graphs in this model are node- and edge-labelled directed multi-hypergraphs, called graphoids, which can be defined at several different levels of granularity using the dimensions associated with them. Operations analogous to the ones used in typical OLAP over cubes are defined over graphoids. The paper presents a formal definition of the graphoid model for OLAP, proves that the typical OLAP operations on cubes can be expressed over the graphoid model, and shows that the classic data cube model is a particular case of the graphoid data model. Finally, a case study supports the claim that, for many kinds of OLAP-like analysis on graphs, the graphoid model works better than the typical relational OLAP alternative, and for the classic OLAP queries, it remains competitive.
Leticia I. Gómez, Bart Kuijpers, Alejandro A. Vaisman
Intell. Data Anal.3
2019 From Conceptual to Logical ETL Design Using BPMN and Relational Algebra
Judith Awiti, Alejandro A. Vaisman, Esteban Zimányi
DaWaK2
2018 Data Quality in a Big Data Context
Franco Arolfo, Alejandro A. Vaisman
ADBIS2
2017 An algebra for OLAP
abstract
Online Analytical Processing (OLAP) comprises tools and algorithms that allow querying multidimensional databases. It is based on the multidimensional model, where data can be seen as a cube, where each cell contains one or more measures can be aggregated along dimensions. Despite the extensive cor pus of work in the field, a standard language for OLAP is still needed, since there is no well-defined, accepted semantics, for many of the usual OLAP operations. In this paper, we address this problem, and present a set of operations for manipulating a data cube. We clearly define the semantics of these operations, and prove that they can be composed, yielding a language powerful enough to express complex OLAP queries. We express these operations as a sequence of atomic transformations over a fixed multidimensional matrix, whose cells contain a sequence of measures. Each atomic transformation produces a new measure. When a sequence of transformations defines an OLAP operation, a flag is produced indicating which cells must be considered as input for the next operation. In this way, an elegant algebra is defined. Our main contribution, with respect to other similar efforts in the field is that, for the first time, a formal proof of the correctness of the operations is given, thus providing a clear semantics for them. We believe the present work will serve as a basis to build more solid practical tools for data analysis.
Bart Kuijpers, Alejandro A. Vaisman
Intell. Data Anal.2
2016 Rule-Based Multidimensional Data Quality Assessment Using Contexts
Adriana Marotta, Alejandro A. Vaisman
DaWaK2
2016 QB2OLAP: Enabling OLAP on Statistical Linked Open Data
abstract
Publication and sharing of multidimensional (MD) data on the Semantic Web (SW) opens new opportunities for the use of On-Line Analytical Processing (OLAP). The RDF Data Cube (QB) vocabulary, the current standard for statistical data publishing, however, lacks key MD concepts such as dimension hierarchies and aggregate functions. QB4OLAP was proposed to remedy this. However, QB4OLAP requires extensive manual annotation and users must still write queries in SPARQL, the standard query language for RDF, which typical OLAP users are not familiar with. In this demo, we present QB2OLAP, a tool for enabling OLAP on existing QB data. Without requiring any RDF, QB(4OLAP), or SPARQL skills, it allows semi-automatic transformation of a QB data set into a QB4OLAP one via enrichment with QB4OLAP semantics, exploration of the enriched schema, and querying with the high-level OLAP language QL that exploits the QB4OLAP semantics and is automatically translated to SPARQL.
Jovan Varga, Lorena Etcheverry, Alejandro A. Vaisman, Oscar Romero 0001, Torben Bach Pedersen, Christian Thomsen 0001
ICDE3
2016 Dimensional enrichment of statistical linked open data
Jovan Varga, Alejandro A. Vaisman, Oscar Romero 0001, Lorena Etcheverry, Torben Bach Pedersen, Christian Thomsen 0001
J. Web Semant.2
2015 A Framework for Building OLAP Cubes on Graphs
Amine Ghrab, Oscar Romero 0001, Sabri Skhiri, Alejandro A. Vaisman, Esteban Zimányi
ADBIS4
2015 Efficient repair of dimension hierarchies under inconsistent reclassification
Mónica Caniupán Marileo, Alejandro A. Vaisman, Raúl Arredondo
Data Knowl. Eng.2
2014 Modeling and Querying Data Warehouses on the Semantic Web Using QB4OLAP
Lorena Etcheverry, Alejandro A. Vaisman, Esteban Zimányi
DaWaK2
2013 Mining semantic trajectories
abstract
A typical problem in the field of moving object (MO) databases consists in discovering interesting trajectory patterns. To solve this problem, data mining techniques are commonly used. Due to the huge volume of these trajectory data, some form of com
Leticia I. Gómez, Alejandro A. Vaisman
Intell. Data Anal.2
2012 BPMN-Based Conceptual Modeling of ETL Processes
Zineb El Akkaoui, Jose-Norberto Mazón, Alejandro A. Vaisman, Esteban Zimányi
DaWaK3
2012 A generic data model and query language for spatiotemporal OLAP cube analysis
abstract
Nowadays, organizations need to use OLAP (On Line Analytical Processing) tools together with geographical information. To support this, the notion of SOLAP (Spatial OLAP) arouse, aimed at exploring spatial data in the same way as OLAP operates over tables. SOLAP however, only accounts for discrete spatial data. More sophisticated GIS-based decision support systems are increasingly being needed, to handle more complex types of data, like continuous fields. Fields describe physical phenomena that change continuously in time and/or space (e.g., temperature). Although many models have been proposed for adding spatial information to OLAP tools, no one allows the user to perceive data as a cube, and analyze any type of spatial data, continuous or discrete, together with typical alphanumerical discrete OLAP data, using only the classic OLAP operators (e.g., Roll-up, Drill-down). In this paper we propose an algebra that operates over data cubes, independently of the underlying data types and physical data representation. That means, in our approach, the final user only sees the typical OLAP operators at the query level. At lower abstraction levels we provide discrete and continuous spatial data support as well as different ways of partitioning the space. We also describe a proof-of-concept implementation to illustrate the ideas presented in the paper. As far as we are aware of, this is the first proposal that allows analyzing discrete and continuous spatiotemporal data and OLAP cubes together, using just the traditional OLAP operations, thus providing a very general framework for spatiotemporal data analysis.
Leticia I. Gómez, Silvia A. Gómez, Alejandro A. Vaisman
EDBT3
2012 Enhancing OLAP Analysis with Web Cubes
Lorena Etcheverry, Alejandro A. Vaisman
ESWC2
2011 Analyzing continuous fields with OLAP cubes
abstract
Although raster data is the most popular discrete representation for continuous fields, other forms of space tessellation can be used, like Voronoi diagrams or Triangulated Irregular Networks (TIN). Besides, algebras for field manipulation have been proposed, but they are representation-dependent and not closed. To address this problems, we propose a (closed) generic map algebra over spatio-temporal continuous fields, independent of the underlying field representation. Based on this algebra, we define a language that allows analyzing continuous fields data and OLAP cubes together, using the traditional OLAP operations, providing a very general framework for spatio-temporal data analysis.
Leticia I. Gómez, Silvia A. Gómez, Alejandro A. Vaisman
DOLAP3
2011 RDFS Update: From Theory to Practice
Claudio Gutierrez 0001, Carlos A. Hurtado, Alejandro A. Vaisman
ESWC (2)3
2011 A data model and query language for spatio-temporal decision support
Leticia I. Gómez, Bart Kuijpers, Alejandro A. Vaisman
GeoInformatica3
2010 Physical Design and Implementation of Spatial Data Warehouses Supporting Continuous Fields
Leticia I. Gómez, Alejandro A. Vaisman, Esteban Zimányi
DaWak2
2010 Exploring XML web collections with DescribeX
abstract
As Web applications mature and evolve, the nature of the semistructured data that drives these applications also changes. An important trend is the need for increased flexibility in the structure of Web documents. Hence, applications cannot rely solely on schemas to provide the complex knowledge needed to visualize, use, query and manage documents. Even when XML Web documents are valid with regard to a schema, the actual structure of such documents may exhibit significant variations across collections for several reasons: the schema may be very lax (e.g., RSS feeds), the schema may be large and different subsets of it may be used in different documents (e.g., industry standards like UBL), or open content models may allow arbitrary schemas to be mixed (e.g., RSS extensions like those used for podcasting). For these reasons, many applications that incorporate XPath queries to process a large Web document collection require an understanding of the actual structure present in the collection, and not just the schema. To support modern Web applications, we introduce DescribeX, a powerful framework that is capable of describing complex XML summaries of Web collections. DescribeX supports the construction of heterogenous summaries that can be declaratively defined and refined by means of axis path regular expression (AxPREs). AxPREs provide the flexibility necessary for declaratively defining complex mappings between instance nodes (in the documents) and summary nodes. These mappings are capable of expressing order and cardinality, among other properties, which can significantly help in the understanding of the structure of large collections of XML documents and enhance the performance of Web applications over these collections. DescribeX captures most summary proposals in the literature by providing (for the first time) a common declarative definition for them. Experimental results demonstrate the scalability of DescribeX summary operations (summary creation, as well as refinement and stabilization, two key enablers for tailoring summaries) on multi-gigabyte Web collections.
Mariano P. Consens, Renée J. Miller, Flavio Rizzolo, Alejandro A. Vaisman
ACM Trans. Web4
2009 What Is Spatio-Temporal Data Warehousing?
Alejandro A. Vaisman, Esteban Zimányi
DaWaK1
2009 Efficient constraint evaluation in categorical sequential pattern mining for trajectory databases
abstract
The classic Generalized Sequential Patterns (GSP) algorithm re-turns all frequent sequences present in a database. However, usu-ally a few ones are interesting from a user’s point of view. Thus, post-processing tasks are required in order to discard uninterest-ing sequences. To avoid this drawback, languages based on regular expressions (RE) were proposed to restrict frequent sequences to the ones that satisfy user-specified constraints. In all of these lan-guages, REs are applied over items, which limits their applicability in complex real-world situations. We propose a much powerful language, based on regular expressions, denoted RE-SPaM, where the basic elements are constraints defined over the (temporal and non-temporal) attributes of the items to be mined. Expressions in this language may include attributes, functions over attributes, and variables. We specify the syntax and semantics of RE-SPaM, and present a comprehensive set of examples to illustrate its expressive power. We study in detail how the expressions can be used to prune the resulting sequences in the mining process. In addition, we in-troduce techniques that allow pruning sequences in the early stages of the process, reducing the need to access the database, making use of the categorization of the attributes that compose the items, and of the automaton that accepts the language generated by the RE. Finally, we present experimental results. Although in this paper we focus on trajectory databases, our approach is general enough for being applied to other settings. 1.
Leticia I. Gómez, Alejandro A. Vaisman
EDBT2
2009 Map matching and uncertainty: an algorithm and real-world experiments
abstract
A common problem in moving object databases (MOD) is the reconstruction of a trajectory from a trajectory sample (i.e., a finite sequence of time-space points). A typical solution to this problem is linear interpolation. A more realistic model is based on the notion of uncertainty modelled by space-time prisms, which capture the positions where the object could have been, when it moved from a to b. Often, object positions measured by location-aware devices are not on a road network. Thus, matching the user's position to a location on the digital map is required. This problem is called map matching. In this paper we study the relation between map matching and uncertainty, and propose an algorithm that combines weighted k-shortest paths with space-time prisms. We apply this algorithm to two real-world case studies and we show that accounting for uncertainty leads to obtaining more positive matchings.
Kristof Ghys, Bart Kuijpers, Bart Moelans, Walied Othman, Dries Vangoidsenhoven, Alejandro A. Vaisman
GIS6
2009 A multidimensional model representing continuous fields in spatial data warehouses
abstract
Data warehouses and On-Line Analytical Processing (OLAP) provide an analysis framework supporting the decision making process. In many application domains, complex analysis tasks often require to take geographical information into account. Several proposals exist for integrating OLAP and Geographic Information Systems (GIS). However, there are very few attempts to support continuous fields, i.e., phenomena that are perceived as having a value at each point in space and/or time. Examples of such phenomena include temperature, altitude, or land use. In this paper, we extend a conceptual multidimensional model with continuous fields, showing that this can be achieved by defining an appropriate data type that encapsulates the different operations needed for manipulating such fields. We also define a query language based on relational calculus that allows expressing spatial OLAP queries involving continuous fields, and use this language to formally characterize this class of queries.
Alejandro A. Vaisman, Esteban Zimányi
GIS1
2009 Analyzing Trajectories Using Uncertainty and Background Information
Bart Kuijpers, Bart Moelans, Walied Othman, Alejandro A. Vaisman
SSTD4
2009 Spatial aggregation: Data model and implementation
Leticia I. Gómez, Sofie Haesevoets, Bart Kuijpers, Alejandro A. Vaisman
Inf. Syst.4
2009 P2P OLAP: Data model, implementation and case study
Alejandro A. Vaisman, Mauricio Minuto Espil, Martín Paradela
Inf. Syst.1
2008 Piet-QL: a query language for GIS-OLAP integration
abstract
Commercial Geographic Information Systems (GIS) still fail to provide integration with OLAP (On Line Analytical Processing) tools. In previous work we have introduced Piet, a system aimed at providing this integration. In this paper we describe a new powerful query language developed for the system, denoted Piet-QL. Piet-QL not only extends the original language for Piet, allowing for example to filter geometric queries with information in a data cube, but also has a more natural semantics. In this paper we formally study the syntax and semantics of Piet-QL, provide a comprehensive set of examples, describe the implementation, and discuss experimental results.
Leticia I. Gómez, Alejandro A. Vaisman, Sebastián Zich
GIS2
2008 AxPRE Summaries: Exploring the (Semi-)Structure of XML Web Collections
abstract
This paper introduces AxPRE summaries, a formalism that allows exploring the (semi-)structure of large XML collections. AxPRE summaries are implemented in a tool, DescribeX, that supports visualizing XML collections via summaries that can be interactively refined using a powerful and descriptive axis path regular expression language. Experimental results on gigabyte collections have shown that this flexibility does not come at the expense of efficiency.
Mariano P. Consens, Flavio Rizzolo, Alejandro A. Vaisman
ICDE3
2008 Temporal XML: modeling, indexing, and query processing
Flavio Rizzolo, Alejandro A. Vaisman
VLDB J.2
2007 Piet: a GIS-OLAP implementation
abstract
Data aggregation in Geographic Information Systems (GIS) is a desirable feature, although only marginally present in commercial systems, which also fail to provide integration between GIS and OLAP (On Line Analytical Processing). With this in mind, we have developed Piet, a system that makes use of a novel query processing technique: first, a process called sub-polygonization decomposes each thematic layer in a GIS, into open convex polygons; then, another process computes and stores in a database the overlay of those layers for later use by a query processor. We describe the implementation of Piet, and provide experimental evidence that overlay precomputation can outperform GIS systems that employ indexing schemes based on R-trees.
Ariel Escribano, Leticia I. Gómez, Bart Kuijpers, Alejandro A. Vaisman
DOLAP4
2007 A model for enriching trajectories with semantic geographical information
abstract
The collection of moving object data is becoming more and more common, and therefore there is an increasing need for the efficient analysis and knowledge extraction of these data in different application domains. Trajectory data are normally available as sample points, and do not carry semantic information, which is of fundamental importance for the comprehension of these data. Therefore, the analysis of trajectory data becomes expensive from a computational point of view and complex from a user's perspective. Enriching trajectories with semantic geographical information may simplify queries, analysis, and mining of moving object data. In this paper we propose a data preprocessing model to add semantic information to trajectories in order to facilitate trajectory data analysis in different application domains. The model is generic enough to represent the important parts of trajectories that are relevant to the application, not being restricted to one specific application. We present an algorithm to compute the important parts and show that the query complexity for the semantic analysis of trajectories will be significantly reduced with the proposed model.
Luis Otávio Alvares, Vania Bogorny, Bart Kuijpers, José A. F. de Macêdo, Bart Moelans, Alejandro A. Vaisman
GIS6
2007 Introducing Time into RDF
abstract
The resource description framework (RDF) is a metadata model and language recommended by the W3C. This paper presents a framework to incorporate temporal reasoning into RDF, yielding temporal RDF graphs. We present a semantics for these kinds of graphs which includes the notion of temporal entailment and a syntax to incorporate this framework into standard RDF graphs, using the RDF vocabulary plus temporal labels. We give a characterization of temporal entailment in terms of RDF entailment and show that the former does not yield extra asymptotic complexity with respect to nontemporal RDF graphs. We also discuss temporal RDF graphs with anonymous timestamps, providing a theoretical framework for the study of temporal anonymity. Finally, we sketch a temporal query language for RDF, along with complexity results for query evaluation that show that the time dimension preserves the tractability of answers
Claudio Gutierrez 0001, Carlos A. Hurtado, Alejandro A. Vaisman
IEEE Trans. Knowl. Data Eng.3
2006 The Meaning of Erasing in RDF under the Katsuno-Mendelzon Approach
Claudio Gutierrez 0001, Carlos A. Hurtado, Alejandro A. Vaisman
WebDB3
2005 Temporal RDF
Claudio Gutierrez 0001, Carlos A. Hurtado, Alejandro A. Vaisman
ESWC3
2004 Aggregate queries in peer-to-peer OLAP
abstract
A peer-to-peer (P2P) data management system consists essentially in a network of peer systems, each maintaining full autonomy over its own data resources. Data exchange between peers occurs when one of them, in the role of a local peer, needs data available in other nodes, denoted the acquaintances of the local peer. No global schema is assumed to exist for any data under this computing paradigm. Henceforth, data provided by an acquaintance of a local peer must be adapted, in a manner that answers to queries posed by local peer users conform the view those users have of their data. Because multidimensional data normally consists in a collection of views of aggregated data, a careful translation process is needed in this case, in order to transform any summary concept that appears in a peer acquaintance into a summary concept meaningful to the requesting peer. We present a model for multidimensional data distributed in a P2P network, and a query rewriting technique, that allows a local peer to propagate OLAP queries among its acquaintances, obtaining a meaningful and correct answer.
Mauricio Minuto Espil, Alejandro A. Vaisman
DOLAP2
2004 Indexing Temporal XML Documents
Alberto O. Mendelzon, Flavio Rizzolo, Alejandro A. Vaisman
VLDB3
2004 Supporting dimension updates in an OLAP server
Alejandro A. Vaisman, Alberto O. Mendelzon, Walter Ruaro, Sergio G. Cymerman
Inf. Syst.1
2003 Revising aggregation hierarchies in OLAP: a rule-based approach
Mauricio Minuto Espil, Alejandro A. Vaisman
Data Knowl. Eng.2
2002 Supporting Dimension Updates in an OLAP Server
Alejandro A. Vaisman, Alberto O. Mendelzon, Walter Ruaro, Sergio G. Cymerman
CAiSE1
2001 Efficient Intensional Redefinition of Aggregation Hierarchies in Multidimensional Databases
abstract
Enhancing multidimensional database models with aggregation hierarchies allows viewing data at different levels of aggregation. Usually, hierarchy instances are represented by means of so-called rollup functions. Rollup between adjacent levels in the hierarchy are given extensionally, while rollups between connected non-adjacent levels are obtained by means of function composition. In many real-life cases, this model cannot capture accurately the meaning of common situations, particularly when exceptions arise. Exceptions may appear due to corporate policies, unreliable data or uncertainty, and their presence may turn the notion of rollup composition unsuitable for representing real relationships in the aggregation hierarchies. In this paper we present a language allowing augmenting traditional extensional rollup functions with intensional knowledge. We denote this language IRAH (Intensional Redefinition for Aggregation Hierarchies). Programs in IRAH consist of intensional rules, which can be regarded as patterns for: (a) overriding natural composition between rollup functions on adjacent levels in the concept hierarchy, (b) canceling the effect of rollup functions for specific values. Our proposal is presented as a stratified default theory. We show that a unique model for the underlying theory always exists, and can be computed in a bottom-up fashion. Finally, we present an algorithm that computes the revised dimension in polynomial time, although under more realistic assumptions, complexity becomes linear on the number of paths in the hierarchy of the dimension instance.
Mauricio Minuto Espil, Alejandro A. Vaisman
DOLAP2
2000 Temporal Queries in OLAP
Alberto O. Mendelzon, Alejandro A. Vaisman
VLDB2
1999 Updating OLAP Dimensions
abstract
OLAP systems support data analysis through a multidimensional data model, according to which data facts are viewed as points in a space of application-related “dimensions” , organized into levels which conform a hierarchy. Although the usual assumption is that these points reflect the dynamic aspect of the data warehouse while dimensions are relatively static, in practice it turns out that dimension updates are often necessary to adapt the multidimensional database to changing requirements. These updates can take place either at the structural level (e.g. addition of categories or modification of the hierarchical structure) or at the instance level (elements can be inserted, deleted, merged, etc.). They are poorly supported (or not supported at all) in current commercial systems and have not been addressed in the literature. In a previous paper we introduced a formal model supporting dimension updates. Here, we extend the model, adding a set of semantically meaningful operators which encapsulate common sequences of primitive dimension updates in a more efficient way. We also formally define two mappings (normalized and denormalized) from the multidimensional to the relational model, and compare an implementation of dimension updates using these two approaches.
Carlos A. Hurtado, Alberto O. Mendelzon, Alejandro A. Vaisman
DOLAP3
1999 Maintaining Data Cubes under Dimension Updates
abstract
OLAP systems support data analysis through a multidimensional data model, according to which data facts are viewed as points in a space of application-related "dimensions", organized into levels which conform to a hierarchy. The usual assumption is that the data points reflect the dynamic aspect of the data warehouse, while dimensions are relatively static. However, in practice, dimension updates are often necessary to adapt the multidimensional database to changing requirements. Structural updates can also take place, like addition of categories or modification of the hierarchical structure. When these updates are performed, the materialized aggregate views that are typically stored in OLAP systems must be efficiently maintained. These updates are poorly supported (or not supported at all) in current commercial systems, and have received little attention in the research literature. We present a formal model of dimension updates in a multidimensional model, a collection of primitive operators to perform them, and a study of the effect of these updates on a class of materialized views, giving an algorithm to efficiently maintain them.
Carlos A. Hurtado, Alberto O. Mendelzon, Alejandro A. Vaisman
ICDE3