Tiago Brasileiro Araújo

dblp:170/8311 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
3since 2021 · last 2026
0000-0001-6339-9117ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Responsible Entity Resolution Over Streaming Data
Kostas Stefanidis, Vasilis Efthymiou, Tiago Brasileiro Araújo
ICDE3
2026 X-TREATS: Integrating explainability and fairness into streaming entity resolution
abstract
Entity Resolution (ER) is a fundamental task in data integration, particularly in streaming environments where entities arrive continuously and decisions must be made under strict time constraints. Existing approaches primarily optimize efficiency and accuracy, but often overlook the interpretability of matching decisions and may propagate group-level disparities. To address these limitations, this paper introduces X-TREATS, a streaming-oriented ER workflow that integrates pair-level explanations directly into the resolution pipeline, jointly combining similarity, fairness, and explanation signals during ranking step. The proposed approach is evaluated on real-world datasets under incremental processing, assessing effectiveness, fairness, and explainability. Experimental results show that X-TREATS substantially improves interpretability, increasing the Explanation Score by up to 20%, while reducing group-level disparities by 40–70% and preserving high matching precision across all datasets. These findings demonstrate the practical benefits of integrating explanation-aware mechanisms into real-time ER pipelines.
Tiago Brasileiro Araújo, Vasilis Efthymiou, Kostas Stefanidis
Inf. Sci.1
2025 TREATS: Fairness-aware entity resolution over streaming data
abstract
Currently, the growing proliferation of information systems generates large volumes of data continuously, stemming from a variety of sources such as web platforms, social networks, and multiple devices. These data, often lacking a defined schema, require an initial process of consolidation and cleansing before analysis and knowledge extraction can occur. In this context, Entity Resolution (ER) plays a crucial role, facilitating the integration of knowledge bases and identifying similarities among entities from different sources. However, the traditional ER process is computationally expensive, and becomes more complicated in the streaming context where the data arrive continuously. Moreover, there is a lack of studies involving fairness and ER, which is related to the absence of discrimination or bias. In this sense, fairness criteria aim to mitigate the implications of data bias in ER systems, which requires more than just optimizing accuracy, as traditionally done. Considering this context, this work presents TREATS, a schema-agnostic and fairness-aware ER workflow developed for managing streaming data incrementally. The proposed fairness-aware ER framework tackles constraints across various groups of interest, presenting a resilient and equitable solution to the related challenges. Through experimental evaluation, the proposed techniques and heuristics are compared against state-of-the-art approaches over five real-world data source pairs, in which the results demonstrated significant improvements in terms of fairness, without degradation of effectiveness and efficiency measures in the streaming environment. In summary, our contributions aim to propel the ER field forward by providing a workflow that addresses both technical challenges and ethical concerns.
Tiago Brasileiro Araújo, Vasilis Efthymiou, Vassilis Christophides, Evaggelia Pitoura, Kostas Stefanidis
Inf. Syst.1
2020 Estimating record linkage costs in distributed environments
Dimas C. Nascimento, Carlos Eduardo S. Pires, Tiago Brasileiro Araújo, Demetrio Gomes Mestre
J. Parallel Distributed Comput.3
2019 Incremental Blocking for Entity Resolution over Web Streaming Data
abstract
The widespread use of information systems has become a valuable source of semi-structured data. In this context, Entity Resolution (ER) emerges as a fundamental task to integrate multiple knowledge bases or identify similarities between data items (i.e., entities). Since ER is an inherently quadratic task, blocking techniques are often used to improve efficiency. Beyond the challenges related to the data volume and heterogeneity, blocking techniques also face two other challenges: streaming data and incremental processing. To address these challenges, we propose PRIME, a novel incremental schema-agnostic blocking technique that utilizes parallelism to enhance blocking efficiency. The proposed technique deals with streaming and incremental data using a distributed computational infrastructure. To improve efficiency, the technique avoids unnecessary comparisons and applies a time window strategy to prevent excessive memory consumption.
Tiago Brasileiro Araújo, Kostas Stefanidis, Carlos Eduardo S. Pires, Jyrki Nummenmaa, Thiago Pereira da Nóbrega
WI1
2017 Towards Reliable Data Analyses for Smart Cities
abstract
As cities are becoming green and smart, public information systems are being revamped to adopt digital technologies. There are several sources (official or not) that can provide information related to a city. The availability of multiple sources enables the design of advanced analyses for offering valuable services to both citizens and municipalities. However, such analyses would fail if the considered data were affected by errors and uncertainties: Data Quality is one of the main requirements for the successful exploitation of the available information. This paper highlights the importance of the Data Quality evaluation in the context of geographical data sources. Moreover, we describe how the Entity Matching task can provide additional information to refine the quality assessment and, consequently, obtain a better evaluation of the reliability data sources. Data gathered from the public transportation and urban areas of Curitiba, Brazil, are used to show the strengths and effectiveness of the presented approach.
Tiago Brasileiro Araújo, Cinzia Cappiello, Nádia P. Kozievitch, Demetrio Gomes Mestre, Carlos Eduardo S. Pires, Monica Vitali
IDEAS1
2017 Spark-based Streamlined Metablocking
abstract
Blocking techniques are widely applied in Entity Resolution (ER) approaches as preprocessing step in order to avoid the quadratic cost of the ER task. In this context, heterogeneous data and Big Data emerges as the major challenges that are faced by blocking techniques. In this sense, we propose the novel approach Spark-based Streamlined Metablocking (SS-Metablocking). Moreover, this work proposes the Cardinality-based load balancing technique to be applied in SS-Metablocking in order to improve its efficiency. To improve the effectiveness of the SS-Metablocking, the GWNP pruning algorithm is proposed in this work. Based on the experimental results, we can highlight that the proposed approach presents better results regarding efficiency and effectiveness than the state-of-the-art approach.
Tiago Brasileiro Araújo, Carlos Eduardo S. Pires, Thiago Pereira da Nóbrega
ISCC1
2017 An efficient spark-based adaptive windowing for entity matching
Demetrio Gomes Mestre, Carlos Eduardo S. Pires, Dimas C. Nascimento, Andreza Raquel Monteiro de Queiroz, Veruska Borges Santos, Tiago Brasileiro Araújo
J. Syst. Softw.6
2016 A fine-grained load balancing technique for improving partition-parallel-based ontology matching approaches
Tiago Brasileiro Araújo, Carlos Eduardo S. Pires, Thiago Pereira da Nóbrega, Dimas C. Nascimento
Knowl. Based Syst.1