José Manuel Moreira

dblp:330/3839 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
9since 2021 · last 2024
0000-0003-4633-6944ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 7 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2024 Moving Region Representations on the Spread of a Forest Fire
Henrique Macías da Silva, Tiago F. R. Ribeiro, Rogério Luís C. Costa, José Manuel Moreira
CIKM4
2024 Improving conformalized quantile regression through cluster-based feature relevance
abstract
Conformalized quantile regression, a cutting-edge and model-agnostic algorithm, has emerged as a recent innovation to generate valid prediction intervals on finite samples while addressing heteroscedasticity. It starts by employing quantile regression to estimate conditional quantiles. Subsequently, these estimated conditional quantiles undergo a rectification process using conformal prediction. Under the assumption of exchangeability, a slightly weaker form of independent and identically distributed (i.i.d.) data, the resulting prediction intervals are valid in finite samples. However, a drawback of the proposed conformalization step is identified: it lacks the capacity to adapt to heteroscedasticity due to its independence from the input. To overcome this limitation, we propose an improvement that involves partitioning the covariates space into clusters, assigning higher weights to features with greater predictive power. Following that, within each cluster, a conformal step is applied, leveraging a rectification that is reliant on the input cluster-wise. To demonstrate the superiority of our improved version over the classic version of conformalized quantile regression, we conducted a comprehensive comparison of their respective prediction intervals using synthetic data.
Martim Sousa, Ana Maria Tomé, José Manuel Moreira
Expert Syst. Appl.3
2024 A general framework for multi-step ahead adaptive conformal heteroscedastic time series forecasting
abstract
This paper introduces a novel model-agnostic algorithm called adaptive ensemble batch multi-input multi-output conformalized quantile regression (AEnbMIMOCQR) that enables forecasters to generate multi-step ahead prediction intervals for a fixed pre-specified miscoverage rate α in a distribution-free manner. Our method is grounded on conformal prediction principles, however, it does not require data splitting and provides close to exact coverage even when the data is not exchangeable. Moreover, the resulting prediction intervals, besides being empirically valid along the forecast horizon, do not neglect heteroscedasticity. AEnbMIMOCQR is designed to be robust to distribution shifts, which means that its prediction intervals remain reliable over an unlimited period of time, without entailing retraining or imposing unrealistic strict assumptions on the data-generating process. Through methodically experimentation, we demonstrate that our approach outperforms other competitive methods on both real-world and synthetic datasets. The code used in the experimental part and a tutorial on how to use AEnbMIMOCQR can be found at the following GitHub repository: https://github.com/Quilograma/AEnbMIMOCQR.
Martim Sousa, Ana Maria Tomé, José Manuel Moreira
Neurocomputing3
2023 Logical big data integration and near real-time data analytics
abstract
In the context of decision-making, there is a growing demand for near real-time data that traditional solutions, like data warehousing based on long-running ETL processes, cannot fully meet. On the other hand, existing logical data integration solutions are challenging because users must focus on data location and distribution details rather than on data analytics and decision-making. EasyBDI is an open-source system that provides logical integration of data and high-level business-oriented abstractions. It uses schema matching, integration, and mapping techniques, to automatically identify partitioned data and propose a global schema. Users can then specify star schemas based on global entities and submit analytical queries to retrieve data from distributed data sources without knowing the organization and other technical details of the underlying systems. This work presents the algorithms and methods for global schema creation and query execution. Experimental results show that the overhead imposed by logical integration layers is relatively small compared to the execution times of distributed queries.
José Manuel Moreira, Rogério Luís C. Costa
Data Knowl. Eng.2
2023 Approximating the evolution of rotating moving regions using Bezier curves
abstract
The region interpolation methods proposed in the moving objects databases literature impose restrictions that can have a significant impact on the representation of the evolution of moving regions, in particular, when a rotation occurs between two observations. In this paper, we propose a data model for moving regions that allows moving segments to rotate and change their length during their evolution between two observations and uses quadratic Bezier curves to define the trajectories of their endpoints. This introduces a new class of moving regions called rotating moving regions (rmregions). We present algorithms for operations involving rmregions and we propose a strategy to allow different interpolation methods to be used in the context of moving objects databases by approximating the interpolations they create using rmregions. We demonstrate our strategy using a reference implementation and compare results obtained when using the strategy presented here and the region interpolation methods and the spatiotemporal operations proposed in the state-of-the-art. Experimental results show that our strategy can be used to complement the region interpolation methods proposed in the moving objects databases literature.
José Duarte, Paulo Dias, José Manuel Moreira
Int. J. Geogr. Inf. Sci.3
2022 Why- and How-Provenance in Distributed Environments
Paulo Pintor, Rogério Luís C. Costa, José Manuel Moreira
DEXA (1)3
2022 Provenance in Spatial Queries
abstract
Despite data growth being a known problem for several years, there are more and more people, tools and devices to create and share data, and the need for tools to infer their provenance and quality is even more important than before. Research on data provenance focuses on W3C PROV and databases (where, why, how). However, in the particular case of spatial data, research has mainly focused on handling spatial data provenance from documents and workflows, but there is no literature approaching the topic of spatial data provenance in DBMS and queries.
Paulo Pintor, Rogério Luís C. Costa, José Manuel Moreira
IDEAS3
2021 Automatic Quality Improvement of Data on the Evolution of 2D Regions
Rogério Luís C. Costa, José Manuel Moreira
ADMA2
2021 EasyBDI: Near Real-Time Data Analytics over Heterogeneous Data Sources
abstract
The large volume of currently available data creates several opportunities for sciences and industry, especially with the application of data analytics. But also raises challenges that make unfeasible the use of batch-based ETL processes. Indeed, near real-time data analytics is a requirement in several domains as an alternative to traditional data warehouses. In the last years, big data platforms have been developed to enable query execution over distributed data sources. However, they do not deal with subject-oriented analysis, do not provide data distribution transparency, or do not assist with schema mapping and integration. In this demonstration, we present EasyBDI. It's a near real-time big data analytics prototype that enables users to run queries over heterogeneous data sources based on global logical abstractions created by the system and provides some usual concepts of data warehouse systems, like facts and dimensions. We use two motivating scenarios, one based on three years of real data on photovoltaic energy production and consumption, and the other based on the SSB+ benchmark. We will also present implementation challenges, issues, solutions, and insights.
José Manuel Moreira, Rogério Luís C. Costa
EDBT2
2019 Modeling and Representing Real-World Spatio-Temporal Data in Databases (Vision Paper)
abstract
Research in general-purpose spatio-temporal databases has focused mainly on the development of data models and query languages. However, since spatio-temporal data are captured as snapshots, an important research question is how to compute and represent the spatial evolution of the data between observations in databases. Current methods impose constraints to ensure data integrity, but, in some cases, these constraints do not allow the methods to obtain a natural representation of the evolution of spatio-temporal phenomena over time. This paper discusses a different approach where morphing techniques are used to represent the evolution of spatio-temporal data in databases. First, the methods proposed in the spatio-temporal databases literature are presented and their main limitations are discussed with the help of illustrative examples. Then, the paper discusses the use of morphing techniques to handle spatio-temporal data, and the requirements and the challenges that must be investigated to allow the use of these techniques in databases. Finally, a set of examples is presented to compare the approaches investigated in this work. The need for benchmarking methodologies for spatio-temporal databases is also highlighted.
José Manuel Moreira, José Duarte, Paulo Dias
COSIT1
2019 Towards a qualitative analysis of interpolation methods for deformable moving regions
abstract
Spatio-temporal data on the evolution of real-world phenomena are normally acquired as snapshots in discrete time. The continuous evolution of a phenomenon between observations can be approximated using interpolation methods capable of generating deformable moving regions. Several region interpolation methods have been proposed in the spatio-temporal databases literature, each one with its own characteristics that can be more suited to represent the evolution of specific physical phenomena.
José Duarte, José Manuel Moreira, Paulo Dias, Enrico S. Miranda, Rogério Luís C. Costa
SIGSPATIAL/GIS3
2016 Representation of continuously changing data over time and space: Modeling the shape of spatiotemporal phenomena
abstract
There are numerous technologies and tools to acquire data related to the evolution of spatial phenomena over time. These data are typically organized as sequences of 2D geometric shapes obtained from observations taken at different times. The transformation of such sequences of 2D geometric shapes into spatiotemporal data representations, which can be easily processed and interpreted, has the potential to enable novel applications in fields as diverse as environmental sciences, climate sciences, biology or medicine. This paper focuses on the representation of moving 2D geometric shapes acquired at discrete times using continuous models of time and space. Using morphing techniques based on compatible triangulations, issues regarding the representation of spatiotemporal data in databases, as well as the influence of different design strategies on the fidelity of the approximations with respect to the modelled phenomena, are investigated. An experimental study using synthetic and real data was performed. The findings show that the use of triangulation based interpolation is a promising approach, because it allows creating continuous spatiotemporal representations that are more realistic than those obtained using the solutions proposed in previous work. Open issues regarding the representation of spatiotemporal data in information systems are also highlighted.
José Manuel Moreira, Paulo Dias
eScience1
2001 Oporto: A Realistic Scenario Generator for Moving Objects
Jean-Marc Saglio, José Manuel Moreira
GeoInformatica2