Cassio P. de Campos

dblp:05/2010 · also Cassio Polpo de Campos, Cassio de Campos · DBLP profile ↗
← Back
6ranked-venue papers in the field
0as first author
4since 2021 · last 2025
0000-0001-9130-1287ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4Database Systems & Data Management · 2
YearPublicationVenuePosition
2025 Imposing Constraints in Probabilistic Circuits via Gradient Optimization
Soroush Ghandi, Benjamin Quost, Cassio P. de Campos
IDA3
2024 Probabilistic Circuits with Constraints via Convex Optimization
Soroush Ghandi, Benjamin Quost, Cassio P. de Campos
ECML/PKDD (3)3
2022 High-Value Token-Blocking: Efficient Blocking Method for Record Linkage
abstract
Data integration is an important component of Big Data analytics. One of the key challenges in data integration is record linkage, that is, matching records that represent the same real-world entity. Because of computational costs, methods referred to as blocking are employed as a part of the record linkage pipeline in order to reduce the number of comparisons among records. In the past decade, a range of blocking techniques have been proposed. Real-world applications require approaches that can handle heterogeneous data sources and do not rely on labelled data. We propose high-value token-blocking (HVTB), a simple and efficient approach for blocking that is unsupervised and schema-agnostic, based on a crafted use of Term Frequency-Inverse Document Frequency. We compare HVTB with multiple methods and over a range of datasets, including a novel unstructured dataset composed of titles and abstracts of scientific papers. We thoroughly discuss results in terms of accuracy, use of computational resources, and different characteristics of datasets and records. The simplicity of HVTB yields fast computations and does not harm its accuracy when compared with existing approaches. It is shown to be significantly superior to other methods, suggesting that simpler methods for blocking should be considered before resorting to more sophisticated methods.
Kevin O'Hare, Anna Jurek-Loughrey, Cassio P. de Campos
ACM Trans. Knowl. Discov. Data3
2021 Bayesian Independence Test with Mixed-type Variables
abstract
A fundamental task in AI is to assess (in)dependence between mixed-type variables (text, image, sound). We propose a Bayesian kernelised correlation test of (in)dependence using a Dirichlet process model. The new measure of (in)dependence allows us to answer some fundamental questions: Based on data, are (mixed-type) variables independent? How likely is dependence/independence to hold? How high is the probability that two mixed-type variables are more than just weakly dependent? We theoretically show the properties of the approach, as well as algorithms for fast computation with it. We empirically demonstrate the effectiveness of the proposed method by analysing its performance and by comparing it with other frequentist and Bayesian approaches on a range of datasets and tasks with mixed-type variables.
Alessio Benavoli, Cassio P. de Campos
DSAA2
2019 An unsupervised blocking technique for more efficient record linkage
Kevin O'Hare, Anna Jurek-Loughrey, Cassio P. de Campos
Data Knowl. Eng.3
2018 A new technique of selecting an optimal blocking method for better record linkage
Kevin O'Hare, Anna Jurek-Loughrey, Cassio P. de Campos
Inf. Syst.3