EDBT 2026 Demo / reviewers in the wild / expert
Carlos Eduardo S. Pires
dblp:59/4336 · also Carlos Eduardo Santos Pires
· DBLP profile ↗
22ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-7743-899XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3Computer networks · 2Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A generalized approach to perform unsupervised blocking key selection for entity resolution
Dimas C. Nascimento, Carlos Eduardo S. Pires, Thiago Pereira da Nóbrega |
Inf. Sci. | 2 |
| 2023 | Towards automatic Privacy-Preserving Record Linkage: A Transfer Learning based classification step
Thiago Pereira da Nóbrega, Carlos Eduardo S. Pires, Dimas C. Nascimento, Leandro Balby Marinho |
Data Knowl. Eng. | 2 |
| 2023 | Leveraging BERT for extractive text summarization on federal police documents
Thierry S. Barros, Carlos Eduardo S. Pires, Dimas C. Nascimento |
Knowl. Inf. Syst. | 2 |
| 2022 | A Decision Tree Ensemble Model for Predicting Bus BunchingabstractAbstract Travel delays and bus overcrowding are some of the daily dissatisfactions of public transportation users. These problems may be caused by bus bunching, an event that occurs when two or more buses are running the same route together, i.e. out of schedule. Due to the stochastic nature of the traffic, a static schedule is not effective to avoid the occurrence of these events; thus, preventive actions are necessary to improve the reliability of the public transportation system. In this context, we propose a decision tree ensemble model to predict bus bunching. We use an ensemble of Random Forest, eXtreme Gradient Boosting and Categorical Boosting models applied to Global Positioning System, General Transit Feed Specification, weather and traffic situation data. The efficacy of the proposed model has been demonstrated using real data sets and has been compared with four baselines: Linear Regression, Logistic Regression, Support Vector Machine and Relevance Vector Machine. According to the results, the proposed model can achieve an efficacy between 74 and 80% and can be used to predict bus bunching in real time up to 10 stops before its occurrence. Veruska Borges Santos, Carlos Eduardo S. Pires, Dimas C. Nascimento, Andreza Raquel Monteiro de Queiroz |
Comput. J. | 2 |
| 2022 | Explanation and answers to critiques on: Blockchain-based Privacy-Preserving Record Linkage
Thiago Pereira da Nóbrega, Carlos Eduardo S. Pires, Dimas C. Nascimento |
Inf. Syst. | 2 |
| 2021 | Blockchain-based Privacy-Preserving Record Linkage: enhancing data privacy in an untrusted environment
Thiago Pereira da Nóbrega, Carlos Eduardo S. Pires, Dimas C. Nascimento |
Inf. Syst. | 2 |
| 2020 | Leveraging active learning to reduce human effort in the generation of ground-truth for entity resolutionabstractSummary Several methods of entity resolution (ER) have been developed in academia and industry over the years, with the intention to identify duplicate entities (eg, records) in datasets. To evaluate the efficacy of such methods, it is necessary to compare their results with a ground‐truth, which consists of a document containing all known duplicate record pairs in a dataset. In general, the generation of ground‐truths for real datasets is performed manually from the inspection of all combinations of pairs of records in a dataset. This is subject to error and presents quadratic complexity, with respect to the size(s) of the dataset(s), requiring a long time to be performed. In this context, some works present (semi)automatic approaches for the generation of ground‐truths for the ER task. However, such approaches are either not applicable to several domains or still present a considerable manual effort. In this work, we propose GTGenERAL, a semiautomatic approach that combines results from multiple algorithms of ER together with active learning to generate accurate ground‐truths employing reduced manual effort. Experiments using real datasets show that, with great manual effort reduction, GTGenERAL is able to generate ground‐truths close to those generated by the state‐of‐the‐art approach. Diego Fernandes de Araújo, Carlos Eduardo S. Pires, Dimas C. Nascimento |
Comput. Intell. | 2 |
| 2020 | Configurable assembly of classification rules for enhancing entity resolution results
Dimas C. Nascimento, Carlos Eduardo S. Pires, Thiago Pereira da Nóbrega |
Inf. Process. Manag. | 2 |
| 2020 | Estimating record linkage costs in distributed environments
Dimas C. Nascimento, Carlos Eduardo S. Pires, Tiago Brasileiro Araújo, Demetrio Gomes Mestre |
J. Parallel Distributed Comput. | 2 |
| 2020 | Exploiting block co-occurrence to control block sizes for entity resolution
Dimas C. Nascimento, Carlos Eduardo S. Pires, Demetrio Gomes Mestre |
Knowl. Inf. Syst. | 2 |
| 2019 | Incremental Blocking for Entity Resolution over Web Streaming DataabstractThe widespread use of information systems has become a valuable source of semi-structured data. In this context, Entity Resolution (ER) emerges as a fundamental task to integrate multiple knowledge bases or identify similarities between data items (i.e., entities). Since ER is an inherently quadratic task, blocking techniques are often used to improve efficiency. Beyond the challenges related to the data volume and heterogeneity, blocking techniques also face two other challenges: streaming data and incremental processing. To address these challenges, we propose PRIME, a novel incremental schema-agnostic blocking technique that utilizes parallelism to enhance blocking efficiency. The proposed technique deals with streaming and incremental data using a distributed computational infrastructure. To improve efficiency, the technique avoids unnecessary comparisons and applies a time window strategy to prevent excessive memory consumption. Tiago Brasileiro Araújo, Kostas Stefanidis, Carlos Eduardo S. Pires, Jyrki Nummenmaa, Thiago Pereira da Nóbrega |
WI | 3 |
| 2019 | BIGSEA: A Big Data analytics platform for public transportation information
Andy S. Alic, Jussara M. Almeida, Giovanni Aloisio, Nazareno Andrade, Nuno Antunes, Danilo Ardagna, Rosa M. Badia, Tânia Basso, Ignacio Blanquer, Tarciso Braz, Andrey Brito, Donatello Elia, Sandro Fiore, Dorgival O. Guedes, Marco Lattuada 0001, Daniele Lezzi, Matheus Maciel, Wagner Meira Jr., Demetrio Gomes Mestre, Regina Lúcia de Oliveira Moraes, Fábio Morais 0001, Carlos Eduardo S. Pires, Nádia P. Kozievitch, Walter Santos, Paulo Silva 0002, Marco Vieira |
Future Gener. Comput. Syst. | 22 |
| 2018 | Heuristic-based approaches for speeding up incremental record linkage
Dimas C. Nascimento, Carlos Eduardo S. Pires, Demetrio Gomes Mestre |
J. Syst. Softw. | 2 |
| 2018 | Estimating Inefficiency in Bus Trip Choices From a User Perspective With Schedule, Positioning, and Ticketing DataabstractThe availability of historical data on the global positioning systems' trajectories of vehicles and passenger boarding information for public bus fleets of large municipalities has given researchers and practitioners the opportunity to explore new challenges regarding the analysis of public transportation systems. This paper performs one such analysis as a case study examining the margin of improvement that passengers of a 1.8M people Brazilian city have when choosing their daily bus trips. In doing so, we document a number of not readily apparent challenges that must be overcome to leverage public transportation big data to policymakers, transportation systems operators, and citizens. Solutions are devised to each of these challenges and demonstrated on the analysis of the aforementioned 1.8M people city. Tarciso Braz, Matheus Maciel, Demetrio Gomes Mestre, Nazareno Andrade, Carlos Eduardo S. Pires, Andreza Raquel Monteiro de Queiroz, Veruska Borges Santos |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2017 | Towards Reliable Data Analyses for Smart CitiesabstractAs cities are becoming green and smart, public information systems are being revamped to adopt digital technologies. There are several sources (official or not) that can provide information related to a city. The availability of multiple sources enables the design of advanced analyses for offering valuable services to both citizens and municipalities. However, such analyses would fail if the considered data were affected by errors and uncertainties: Data Quality is one of the main requirements for the successful exploitation of the available information. This paper highlights the importance of the Data Quality evaluation in the context of geographical data sources. Moreover, we describe how the Entity Matching task can provide additional information to refine the quality assessment and, consequently, obtain a better evaluation of the reliability data sources. Data gathered from the public transportation and urban areas of Curitiba, Brazil, are used to show the strengths and effectiveness of the presented approach. Tiago Brasileiro Araújo, Cinzia Cappiello, Nádia P. Kozievitch, Demetrio Gomes Mestre, Carlos Eduardo S. Pires, Monica Vitali |
IDEAS | 5 |
| 2017 | Spark-based Streamlined MetablockingabstractBlocking techniques are widely applied in Entity Resolution (ER) approaches as preprocessing step in order to avoid the quadratic cost of the ER task. In this context, heterogeneous data and Big Data emerges as the major challenges that are faced by blocking techniques. In this sense, we propose the novel approach Spark-based Streamlined Metablocking (SS-Metablocking). Moreover, this work proposes the Cardinality-based load balancing technique to be applied in SS-Metablocking in order to improve its efficiency. To improve the effectiveness of the SS-Metablocking, the GWNP pruning algorithm is proposed in this work. Based on the experimental results, we can highlight that the proposed approach presents better results regarding efficiency and effectiveness than the state-of-the-art approach. Tiago Brasileiro Araújo, Carlos Eduardo S. Pires, Thiago Pereira da Nóbrega |
ISCC | 2 |
| 2017 | Towards the efficient parallelization of multi-pass adaptive blocking for entity matching
Demetrio Gomes Mestre, Carlos Eduardo S. Pires, Dimas C. Nascimento |
J. Parallel Distributed Comput. | 2 |
| 2017 | An efficient spark-based adaptive windowing for entity matching
Demetrio Gomes Mestre, Carlos Eduardo S. Pires, Dimas C. Nascimento, Andreza Raquel Monteiro de Queiroz, Veruska Borges Santos, Tiago Brasileiro Araújo |
J. Syst. Softw. | 2 |
| 2016 | Applying machine learning techniques for scaling out data quality algorithms in cloud computing environments
Dimas C. Nascimento, Carlos Eduardo S. Pires, Demetrio Gomes Mestre |
Appl. Intell. | 2 |
| 2016 | A fine-grained load balancing technique for improving partition-parallel-based ontology matching approaches
Tiago Brasileiro Araújo, Carlos Eduardo S. Pires, Thiago Pereira da Nóbrega, Dimas C. Nascimento |
Knowl. Based Syst. | 2 |
| 2013 | Improving load balancing for MapReduce-based entity matchingabstractThe effectiveness and scalability of MapReduce-based implementations for data-intensive tasks depends on the data assignment made from map to reduce tasks. The robustness of this assignment strategy is crucial to achieve skewed data handling and balanced workload distribution among all reduce tasks. For the entity matching problem in the Big Data context, we propose BlockSlicer, a MapReduce-based approach that supports blocking techniques to reduce the entity matching search space. The approach utilizes a preprocessing MapReduce job to analyze the data distribution and provides an improved load balancing by applying an efficient block slice strategy as well as a well-known optimization algorithm to assign the generated match tasks. We evaluate the approach against an existing one that addresses the same problem on a real cloud infrastructure. The results show that our approach increases significantly the performance of distributed entity matching task by reducing the amount of data generated from the map phase and diminishing the overall execution time. Demetrio Gomes Mestre, Carlos Eduardo S. Pires |
ISCC | 2 |
| 2011 | Generating Synthetic Database Schemas for Simulation Purposes
Carlos Eduardo S. Pires, Priscilla Vieira, Márcio Saraiva, Denilson Barbosa 0001 |
DEXA (2) | 1 |