VLDB 2026 Research / reviewers in the wild / expert
André Freitas
dblp:47/9409
· DBLP profile ↗
24ranked-venue papers in the field
8as first author
7since 2021 · last 2025
0000-0002-4430-4837ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 11 (5 first)Database Systems & Data Management · 9 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gem: Gaussian Mixture Model Embeddings for Numerical Feature Distributions
Hafiz Tayyab Rauf, Alex Teodor Bogatu, Norman W. Paton, André Freitas |
EDBT | 4 |
| 2025 | TableDC: Deep Clustering for Tabular DataabstractDeep clustering (DC), a fusion of deep representation learning and clustering, has recently demonstrated positive results in data science, particularly text processing and computer vision. However, joint optimization of feature learning and data distribution in the multi-dimensional space is domain-specific, so existing DC methods struggle to generalize to other application domains (such as data integration). In data management tasks, where high-density embeddings and overlapping clusters dominate, a data management-specific DC algorithm should be able to interact better with the data properties to support data integration tasks. This paper presents a deep clustering algorithm for tabular data (TableDC) that reflects the properties of data management applications that cluster tables (schema inference), rows (entity resolution) and columns (domain discovery). To address overlapping clusters, TableDC integrates Mahalanobis distance, which considers variance and correlation within the data, offering a similarity method suitable for tabular data in high-dimensional latent spaces. TableDC also shows higher tolerance to outliers through its heavy-tailed Cauchy distribution as the similarity kernel. The proposed similarity measure is particularly beneficial where the embeddings of raw data are densely packed and exhibit high degrees of overlap. Data integration tasks may also involve large numbers of clusters, which challenges the scalability of existing DC methods. TableDC learns data embeddings with a large number of clusters more efficiently than baseline DC methods, which scale in quadratic time. We evaluated TableDC with several existing DC, Standard Clustering (SC), and state-of-the-art bespoke methods over benchmark datasets. TableDC consistently outperforms existing DC, SC and bespoke methods. Hafiz Tayyab Rauf, André Freitas, Norman W. Paton |
Proc. ACM Manag. Data | 2 |
| 2024 | Deep Clustering for Data Cleaning and Integration
Hafiz Tayyab Rauf, André Freitas, Norman W. Paton |
EDBT | 2 |
| 2022 | Voyager: Data Discovery and Integration for Onboarding in Data Science
Alex Teodor Bogatu, Norman W. Paton, Mark Douthwaite, André Freitas |
EDBT | 4 |
| 2021 | Natural Language Inference over Tables: Enabling Explainable Data Exploration on Data Lakes
Mario Ramirez, Alex Teodor Bogatu, Norman W. Paton, André Freitas |
ESWC | 4 |
| 2021 | Cost-effective Variational Active Entity ResolutionabstractAccurately identifying different representations of the same real-world entity is an integral part of data cleaning and many methods have been proposed to accomplish it. The challenges of this entity resolution task that demand so much research attention are often rooted in the task-specificity and user-dependence of the process. Adopting deep learning techniques has the potential to lessen these challenges. In this paper, we set out to devise an entity resolution method that builds on the robustness conferred by deep autoencoders to reduce human-involvement costs. Specifically, we reduce the cost of training deep entity resolution models by performing unsupervised representation learning. This unveils a transferability property of the resulting model that can further reduce the cost of applying the approach to new datasets by means of transfer learning. Finally, we reduce the cost of labeling training data through an active learning approach that builds on the properties conferred by the use of deep autoencoders. Empirical evaluation confirms the accomplishment of our cost-reduction desideratum, while achieving comparable effectiveness with state-of-the-art alternatives. Alex Teodor Bogatu, Norman W. Paton, Mark Douthwaite, Stuart Davie, André Freitas |
ICDE | 5 |
| 2021 | What's Your Value of Travel Time? Collecting Traveler-Centered Mobility Data via Crowdsourcing
Cristian Consonni, Silvia Basile, Matteo Manca, Ludovico Boratto, André Freitas, Tatiana Kovacikova, Ghadir Pourhashem, Yannick Cornet |
ICWSM | 5 |
| 2020 | SKG4J 2020: 1st International Workshop on Semantic and Knowledge Graph Advances for JournalismabstractSKG4J targeted contributions at the interface between Artificial Intelligence, Data Management and its implications for journalistic practice. The first version of the workshop accepted three submissions with topics emphasising the complementary requirements for delivering realistic journalistic knowledge extraction/management platforms. Tareq Al-Moslmi, Raphaël Troncy, André Freitas, Davide Ceolin, Abdullatif Abolohom |
CIKM | 3 |
| 2020 | A User-centred Analysis of Explanations for a Multi-component Semantic Parser
Juliano Sales, André Freitas, Siegfried Handschuh |
NLDB | 2 |
| 2018 | Classification of composite semantic relations by a distributional-relational modelabstractDifferent semantic interpretation tasks such as text entailment and question answering require the classification of semantic relations between terms or entities within text. However, in most cases, it is not possible to assign a direct semantic relation between entities/terms. This paper proposes an approach for composite semantic relation classification using one or more relations between entities/term mentions, extending the traditional semantic relation classification task . The proposed model is different from existing approaches which typically use machine learning models built over lexical and distributional word vector features in that is uses a combination of a large commonsense knowledge base of binary relations , a distributional navigational algorithm and sequence classification to provide a solution for the composite semantic relation classification problem. The proposed approach outperformed existing baselines with regard to F1-score, Accuracy, Precision and Recall. Siamak Barzegar, Brian Davis 0001, Siegfried Handschuh, André Freitas |
Data Knowl. Eng. | 4 |
| 2017 | Composite Semantic Relation Classification
Siamak Barzegar, André Freitas, Siegfried Handschuh, Brian Davis 0001 |
NLDB | 2 |
| 2016 | Semantic Relatedness for All (Languages): A Comparative Analysis of Multilingual Semantic Relatedness Using Machine Translation
André Freitas, Siamak Barzegar, Juliano Sales, Siegfried Handschuh, Brian Davis 0001 |
EKAW | 1 |
| 2016 | Word Tagging with Foundational Ontology Classes: Extending the WordNet-DOLCE Mapping to Verbs
Vivian Dos Santos Silva, André Freitas, Siegfried Handschuh |
EKAW | 2 |
| 2016 | Determining Data Relevance Using Semantic Types and Graphical Interpretation Cues
Eduardo Haruo Kamioka, André Freitas, Frederico Tommasi Caroli, Siegfried Handschuh |
IDA | 2 |
| 2016 | Preface
Chris Biemann, André Freitas, Siegfried Handschuh, Elisabeth Métais, Farid Meziane |
Data Knowl. Eng. | 2 |
| 2015 | DINFRA: A One Stop Shop for Computing Multilingual Semantic RelatednessabstractThis demonstration presents an infrastructure for computing multilingual semantic relatedness and correlation for twelve natural languages by using three distributional semantic models (DSMs). Our demonsrator - DInfra (Distributional Infrastructure) provides researchers and developers with a highly useful platform for processing large-scale corpora and conducting experiments with distributional semantics. We integrate several multilingual DSMs in our webservice so end user can obtain a result without worrying about the complexities involved in building DSMs. Our webservice allows the users to have easy access to a wide range of comparisons of DSMs with different parameters. In addition, users can configure and access DSM parameters using a easy to use API. Siamak Barzegar, Juliano Sales, André Freitas, Siegfried Handschuh, Brian Davis 0001 |
SIGIR | 3 |
| 2015 | Linse: A Distributional Semantics Entity Search EngineabstractEntering 'Football Players from United States' when searching for 'American Footballers' is an example of vocabulary mismatch, which occurs when different words are used to express the same concepts. In order to address this phenomenon for entity search targeting descriptors for complex categories, we propose a compositional-distributional semantics entity search engine, which extracts semantic and commonsense knowledge from large-scale corpora to address the vocabulary gap between query and data. Juliano Sales, André Freitas, Siegfried Handschuh, Brian Davis 0001 |
SIGIR | 2 |
| 2015 | Approximate and selective reasoning on knowledge graphs: A distributional semantics approach
André Freitas, João C. P. da Silva, Edward Curry, Paul Buitelaar |
Data Knowl. Eng. | 1 |
| 2014 | A Distributional Semantics Approach for Selective Reasoning on Commonsense Graph Knowledge Bases
André Freitas, João C. P. da Silva, Edward Curry, Paul Buitelaar |
NLDB | 1 |
| 2014 | On the Semantic Representation and Extraction of Complex Category Descriptors
André Freitas, Rafael Vieira, Edward Curry, Danilo S. Carvalho, João C. P. da Silva |
NLDB | 1 |
| 2013 | Answering natural language queries over linked data graphs: a distributional semantics approachabstractThis paper demonstrates Treo, a natural language query mechanism for Linked Data graphs. The approach uses a distributional semantic vector space model to semantically match user query terms with data, supporting vocabulary-independent (or schema-agnostic) queries over structured data. André Freitas, Fabrício Firmino de Faria, Seán O'Riain, Edward Curry |
SIGIR | 1 |
| 2013 | Querying linked data graphs using semantic relatedness: A vocabulary independent approach
André Freitas, João Gabriel Oliveira, Seán O'Riain, João C. P. da Silva, Edward Curry |
Data Knowl. Eng. | 1 |
| 2011 | Querying Linked Data Using Semantic Relatedness: A Vocabulary Independent Approach
André Freitas, João Gabriel Oliveira, Seán O'Riain, Edward Curry, João C. P. da Silva |
NLDB | 1 |
| 2011 | Treo: Best-Effort Natural Language Queries over Linked Data
André Freitas, João Gabriel Oliveira, Seán O'Riain, Edward Curry, João C. P. da Silva |
NLDB | 1 |