EDBT 2026 Demo / reviewers in the wild / expert
Renata Galante
dblp:37/2418 · also Renata de Matos Galante
· DBLP profile ↗
33ranked-venue papers
3as first author
2since 2021 · last 2026
0000-0003-3589-1619ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 17 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 15 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 43% Data mining · 29% Machine learning and data management · 27% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
sampling |
0.5 | 2 | 2016 | A practical and effective sampling selection strategy for large scale deduplication · ICDE 2016 A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015 |
Data integration and cleaning › entity resolution
record deduplication |
0.2 | 1 | 2016 | A practical and effective sampling selection strategy for large scale deduplication · ICDE 2016 |
Machine learning and data management
active learning |
0.2 | 1 | 2015 | A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015 |
Machine learning and data management
data selection |
0.2 | 1 | 2015 | A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015 |
Data integration and cleaning › entity resolution
deduplication |
0.2 | 1 | 2015 | A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015 |
Data integration and cleaning
entity resolution |
0.2 | 1 | 2015 | A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015 |
Methods — techniques the papers use, named apart from their topics
sampling · 0.2two-stage sampling · 0.2active selection · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic Governance Paradigm of Heterogeneous World Models for Agentic AI Systems
Silvio Fernando Angonese, Renata Galante |
COMPSAC | 2 |
| 2021 | FIP-SHA - Finding Individual Profiles Through SHared Accounts
Carolina Nery, Renata Galante, Weverton Luis da Costa Cordeiro |
DEXA (2) | 2 |
| 2020 | Drink2Vec: Improving the classification of alcohol-related tweets using distributional semantics and external contextual enrichment
Marcos A. Grzeça, Karin Becker, Renata Galante |
Inf. Process. Manag. | 3 |
| 2019 | Broker-Insights: An Interactive and Visual Recommendation System for Insurance Brokerage
Paul Dany Flores Atauchi, Luciana Porcher Nedel, Renata Galante |
CGI | 3 |
| 2019 | FFT-2PCA: A New Feature Extraction Method for Data-Based Fault Detection
Matheus Maia de Souza, João Cesar Netto, Renata Galante |
DEXA (1) | 3 |
| 2019 | Bag of textual graphs (BoTG): A general graph-based text representation modelabstractText representation models are the fundamental basis for information retrieval and text mining tasks. Although different text models have been proposed, they typically target specific task aspects in isolation, such as time efficiency, accuracy, or applicability for different scenarios. Here we present Bag of Textual Graphs (BoTG), a general text representation model that addresses these three requirements at the same time. The proposed textual representation is based on a graph‐based scheme that encodes term proximity and term ordering, and represents text documents into an efficient vector space that addresses all these aspects as well as provides discriminative textual patterns. Extensive experiments are conducted in two experimental scenarios—classification and retrieval—considering multiple well‐known text collections. We also compare our model against several methods from the literature. Experimental results demonstrate that our model is generic enough to handle different tasks and collections. It is also more efficient than the widely used state‐of‐the‐art methods in textual classification and retrieval tasks, with a competitive effectiveness, sometimes with gains by large margins. Ícaro C. Dourado, Renata Galante, Marcos André Gonçalves, Ricardo da Silva Torres |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2019 | Combining URL and HTML Features for Entity Discovery in the WebabstractThe web is a large repository of entity-pages. An entity-page is a page that publishes data representing an entity of a particular type, for example, a page that describes a driver on a website about a car racing championship. The attribute values published in the entity-pages can be used for many data-driven companies, such as insurers, retailers, and search engines. In this article, we define a novel method, called SSUP , which discovers the entity-pages on the websites. The novelty of our method is that it combines URL and HTML features in a way that allows the URL terms to have different weights depending on their capacity to distinguish entity-pages from other pages, and thus the efficacy of the entity-page discovery task is increased. SSUP determines the similarity thresholds on each website without human intervention. We carried out experiments on a dataset with different real-world websites and a wide range of entity types. SSUP achieved a 95% rate of precision and 85% recall rate. Our method was compared with two state-of-the-art methods and outperformed them with a precision gain between 51% and 66%. Edimar Manica, Carina F. Dorneles, Renata Galante |
ACM Trans. Web | 3 |
| 2018 | Improving the Classification of Drunk Texting in Tweets Using Semantic EnrichmentabstractExcessive alcohol consumption is a worldwide problem, and social networks such as Twitter can provide valuable data that help understanding factors related to alcoholism, particularly among youngsters. The identification of drunk tweets (i.e. posted under the influence of alcohol) is complex because tweets are short, sparse and written with diverse and internet specific vocabulary, possibly with errors due to alcohol influence. In this paper, we propose an enriching framework that integrates conceptual and semantic features that expand and generalize the vocabulary, providing context to tweet terms. It also handles misspellings and the selection of discriminative features resulting from contextual enrichment. We outperformed the baseline, achieving improvements of 13.79 percentage points in recall, with no significant harm to precision. We illustrate the value of drunk tweets classification by developing an exploratory analysis that reveals drunk tweeters demographics and tweet properties. Marcos A. Grzeça, Karin Becker, Renata Galante |
WI | 3 |
| 2017 | R-Extractor: A Method for Data Extraction from Template-Based Entity-PagesabstractThe challenges in Big Data start during the data acquisition, where it is necessary to transform non-structured data into a structured format. The Big Data Era brought a new challenge to data acquisition: the need for eliminating user intervention. In this context, the domain-centric data extraction (DCDE) methods arose through replacing user intervention with content redundancy. The DCDE methods extract the attribute values of entities of the web that are restricted to a specific application domain. A research gap was identified from various analyzed methods, i.e., the DCDE methods are not effective in extracting attribute values that are differently presented within the same website. We have responded to this research gap by proposing the R-Extractor method, which extends a state-of-the-art DCDE method by adding a reinforcement stage for the treatment of attribute values that are differently presented within the same website. The R-Extractor method re-analyzes the extraction rules to identify those that can be combined to extract the values of a given attribute from all the pages that describe entities on a website. This identification is based on a novel score function that takes into account different features of the extraction rules. We carried out experiments on a dataset with more than 50k web pages from different real-world websites of a wide range of application domains. The R-Extractor method reached 98% of precision. Our method was compared with two baselines (an XPath-based method and a tree-based method) and outperformed them with an increase in precision up to 14%. Edimar Manica, Carina F. Dorneles, Renata Galante |
COMPSAC (1) | 3 |
| 2017 | KANDOR - Knowledge Analysis of Neighborhood Dynamics and Online RelationshipsabstractWith the emergence of smartphones and location-based social networks, a large amount of user-generated data has become available to better understand city dynamics and help urban planning. While most of the related works choose to focus on specific dimensions of the data, the proposed models aim to benefit from the extent of the information existing in social platforms. Thus, this paper explores the full potential of social media data and proposes novel clustering models for retrieving information about city dynamics and urban characterization by incorporating new dimensions to state of the art algorithms. Aspects such as venue rating, entropy and popularity may lead to new and more complete understanding of activities and trends in a city. Preliminary experiments show that it is possible to aggregate a large diversity of information from different social networks, and generate different and complementary visualizations of the city. Moreover, by applying these methods to urban environments, governments and citizens can better understand and build better sustainable cities together. Ricardo Chagas Rapacki, Leandro Krug Wives, Renata Galante |
COMPSAC (1) | 3 |
| 2017 | Orion: A Cypher-Based Web Data Extractor
Edimar Manica, Carina F. Dorneles, Renata Galante |
DEXA (1) | 3 |
| 2017 | The effects of classifiers diversity on the accuracy of stackingabstractIn recent years several data classification techniques have been proposed.However, it is not a trivial task to choose the most appropriate classifier for deal with a particular problem and set it up properly.In addition, there is no optimal algorithm to solve all prediction problems.In order to improve the result of the classification process, the stacking strategy combines the knowledge acquired by individual learning algorithms aiming to discover new patterns not yet identified.Stacking combines the outputs of base classifiers, induced by several learning algorithms using the same dataset, by means of a meta-classifier.The main goal of this paper is to evaluate the effects of classifier diversity on the accuracy of stacking.We have performed a lot of experiments which results show the impact of multiple diversity measures on the gain of stacking, considering many real datasets extracted from UCI machine learning repository and three synthetic twodimensional datasets.The results revealed connections between some measures and the gain of stacking, but they imply a weak or moderate relationship that suggest predicting the improvement on the best base classifier accuracy using diversity measures is inappropriate. Mariele Lanes, Eduardo N. Borges, Renata Galante |
SEKE | 3 |
| 2017 | Twenty years of object-relational mapping: A survey on patterns, solutions, and their implications on application design
Alexandre Torres, Renata Galante, Marcelo Soares Pimenta, Alexandre Jonatan B. Martins |
Inf. Softw. Technol. | 2 |
| 2016 | A practical and effective sampling selection strategy for large scale deduplicationabstractRecord deduplication aims at identifying entities that are potentially the same in a data repository. A set of pairs that is manually labeled is generally used to tune the deduplication process, as each dataset has a particular dirtiness pattern. However, producing an informative set of pairs is a very costly task, especially in very large datasets (even for expert users). We propose a new sampling strategy that is able to select a very small and informative set of pairs from large datasets. Our results show that our approach reduces user effort substantially while achieving a competitive or superior matching quality. Guilherme Dal Bianco, Renata Galante, Carlos Alberto Heuser, Marcos André Gonçalves, Sérgio D. Canuto |
ICDE | 2 |
| 2015 | Improving Financial Time Series Prediction Through Output Classification by a Neural Network Ensemble
Felipe Giacomel, Adriano C. M. Pereira, Renata Galante |
DEXA (2) | 3 |
| 2015 | A Practical and Effective Sampling Selection Strategy for Large Scale DeduplicationabstractThe data deduplication task has attracted a considerable amount of attention from the research community in order to provide effective and efficient solutions. The information provided by the user to tune the deduplication process is usually represented by a set of manually labeled pairs. In very large datasets, producing this kind of labeled set is a daunting task since it requires an expert to select and label a large number of informative pairs. In this article, we propose a two-stage sampling selection strategy (T3S) that selects a reduced set of pairs to tune the deduplication process in large datasets. T3S selects the most representative pairs by following two stages. In the first stage, we propose a strategy to produce balanced subsets of candidate pairs for labeling. In the second stage, an active selection is incrementally invoked to remove the redundant pairs in the subsets created in the first stage in order to produce an even smaller and more informative training set. This training set is effectively used both to identify where the most ambiguous pairs lie and to configure the classification approaches. Our evaluation shows that T3S is able to reduce the labeling effort substantially while achieving a competitive or superior matching quality when compared with state-of-the-art deduplication methods in large datasets. Guilherme Dal Bianco, Renata Galante, Marcos André Gonçalves, Sérgio D. Canuto, Carlos Alberto Heuser |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | SSUP - A URL-Based Method to Entity-Page Discovery
Edimar Manica, Renata Galante, Carina F. Dorneles |
ICWE | 2 |
| 2013 | Tuning large scale deduplication with reduced effortabstractDeduplication is the task of identifying which objects are potentially the same in a data repository. It usually demands user intervention in several steps of the process, mainly to identify some pairs representing matchings and non-matchings. This information is then used to help in identifying other potentially duplicated records. When deduplication is applied to very large datasets, the performance and matching quality depends on expert users to configure the most important steps of the process (e.g., blocking and classification). In this paper, we propose a new framework called FS-Dedup able to help tuning the deduplication process on large datasets with a reduced effort from the user, who is only required to label a small, automatically selected, subset of pairs. FS-Dedup exploits Signature-Based Deduplication (Sig-Dedup) algorithms in its deduplication core. Sig-Dedup is characterized by high efficiency and scalability in large datasets but requires an expert user to tune several parameters. FS-Dedup helps in solving this drawback by providing a framework that does not demand specialized user knowledge about the dataset or thresholds to produce high effectiveness. Our evaluation over large real and synthetic datasets (containing millions of records) shows that FS-Dedup is able to reach or even surpass the maximal matching quality obtained by Sig-Dedup techniques with a reduced manual effort from the user. Guilherme Dal Bianco, Renata Galante, Carlos Alberto Heuser, Marcos André Gonçalves |
SSDBM | 2 |
| 2012 | Measuring Change Impact Based on Usage ProfilesabstractService evolution is a critical issue because even small changes, if not compatible, can potentially affect a huge number of client applications. However, particularly in the context of large scale service usage, changes have different impact on clients according to its use. This paper proposes a change management framework that supports service providers to scope and quantify the impact of changes based on usage analysis. The framework adopts a finer-grained versioning model in order to easily locate and assess the compatibility of changes in service descriptions. The framework also clusters client applications based on similar patterns of usage, summarizing them in usage profiles. A usage profile quantifies the functionality of the service used by the corresponding applications, enabling to assess the impact of incompatible changes against the profile. Marcelo Yamashita, Bruno Vollino, Karin Becker, Renata Galante |
ICWS | 4 |
| 2011 | Service Evolution Management Based on Usage ProfileabstractServices have been increasingly used as the building blocks for decoupled and flexible applications. Service evolution is a critical issue because even small changes, if not compatible, can potentially affect a huge number of client applications. However, particularly in the context of large scale usage of a service, changes cause different impact on client applications according to its use. This paper proposes to focus on compatibility from the point of view of usage patterns in order to deal with service evolution issues in more flexible and less costly way. The idea is to summarize the behavior of client applications into usage profiles, from which metrics that represent the impact of changes can be derived. This valuable information may support service providers on decisions about service lifecycle. The paper discusses the adoption of usage profiles and presents a framework for the automatic evaluation of service changes impact during its lifecycle. Marcelo Yamashita, Karin Becker, Renata Galante |
ICWS | 3 |
| 2011 | An unsupervised heuristic-based approach for bibliographic metadata deduplication
Eduardo N. Borges, Moisés G. de Carvalho, Renata Galante, Marcos André Gonçalves, Alberto H. F. Laender |
Inf. Process. Manag. | 3 |
| 2011 | A synergistic model-driven approach for persistence modeling with UML
Alexandre Torres, Renata Galante, Marcelo Soares Pimenta |
J. Syst. Softw. | 2 |
| 2009 | TRIple Content-based OnTology (TRICOt) for XML Dissemination
Mirella M. Moro, Deise de Brum Saccol, Renata Galante |
SEKE | 3 |
| 2009 | MD-JPA Profile: A Model Driven Language for Java Persistence
Alexandre Torres, Renata Galante, Marcelo Soares Pimenta |
SEKE | 2 |
| 2008 | A Metadata Model for Managing and Querying XML Resources in Peer-to-peer Systems
Deise de Brum Saccol, Nina Edelweiss, Renata Galante |
SEKE | 3 |
| 2007 | Detecting, Managing and Querying Replicas and Versions in a Peer-to-Peer Environmentabstract(P2P) systems provide sharing of resources, which may be duplicated or versioned in several peers. However, traditional P2P systems are not aware of such replicas and versions, which arises an inefficiency and ineffectiveness problem. To solve this issue, our work proposes an automatic mechanism for replica and version detection. This functionality is presented as part of DetVX, an environment for managing replicas and versions of XML documents in a P2P context. Deise de Brum Saccol, Nina Edelweiss, Renata Galante |
CCGRID | 3 |
| 2007 | XML version detectionabstractThe problem of version detection is critical in many important application scenarios, including software clone identification, Web page ranking, plagiarism detection, and peer-to-peer searching. A natural and commonly used approach to version detection relies on analyzing the similarity between files. Most of the techniques proposed so far rely on the use of hard thresholds for similarity measures. However, defining a threshold value is problematic for several reasons: in particular (i) the threshold value is not the same when considering different similarity functions, and (ii) it is not semantically meaningful for the user. To overcome this problem, our work proposes a version detection mechanism for XML documents based on Naïve Bayesian classifiers. Thus, our approach turns the detection problem into a classification problem. In this paper, we present the results of various experiments on synthetic data that show that our approach produces very good results, both in terms of recall and precision measures. Deise de Brum Saccol, Nina Edelweiss, Renata Galante, Carlo Zaniolo |
ACM Symposium on Document Engineering | 3 |
| 2007 | A Deep Classification of Temporal Versioned Integrity Constraints for Designing Database Applications
Robson L. F. Cordeiro, Renata Galante, Nina Edelweiss, Clesio Saraiva dos Santos |
SEKE | 2 |
| 2007 | CXPath: a Query Language for Conceptual Models of Integrated XML Data
Diego de Vargas Feijó, Cláudio Naoto Fuzitaki, Álvaro F. Moreira, Renata Galante, Carlos Alberto Heuser |
SEKE | 4 |
| 2007 | Managing XML Versions and Replicas in a P2P Context
Deise de Brum Saccol, Nina Edelweiss, Renata Galante, Carlo Zaniolo |
SEKE | 3 |
| 2005 | Temporal and versioning model for schema evolution in object-oriented databases
Renata Galante, Clesio Saraiva dos Santos, Nina Edelweiss, Álvaro F. Moreira |
Data Knowl. Eng. | 1 |
| 2003 | TVL_SE - Temporal and Versioning Language for Schema Evolution in Object-Oriented Databases
Renata Galante, Nina Edelweiss, Clesio Saraiva dos Santos |
DEXA | 1 |
| 2002 | Dynamic Schema Evolution Management Using Version in Temporal Object-Oriented Databases
Renata Galante, Adriana Bueno da Silva Roma, Anelise Jantsch, Nina Edelweiss, Clesio Saraiva dos Santos |
DEXA | 1 |