EDBT 2026 Demo / reviewers in the wild / expert
Alex Teodor Bogatu
dblp:202/0373 · also Alex Bogatu
· DBLP profile ↗
10ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0003-1604-5097ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gem: Gaussian Mixture Model Embeddings for Numerical Feature Distributions
Hafiz Tayyab Rauf, Alex Teodor Bogatu, Norman W. Paton, André Freitas |
EDBT | 2 |
| 2023 | Meta-analysis informed machine learning: Supporting cytokine storm detection during CAR-T cell TherapyabstractCytokine release syndrome (CRS), also known as cytokine storm, is one of the most consequential adverse effects of chimeric antigen receptor therapies that have shown otherwise promising results in cancer treatment. When emerging, CRS could be identified by the analysis of specific cytokine and chemokine profiles that tend to exhibit similarities across patients. In this paper, we exploit these similarities using machine learning algorithms and set out to pioneer a meta-review informed method for the identification of CRS based on specific cytokine peak concentrations and evidence from previous clinical studies. To this end we also address a widespread challenge of the applicability of machine learning in general: reduced training data availability. We do so by augmenting available (but often insufficient) patient cytokine concentrations with statistical knowledge extracted from domain literature. We argue that such methods could support clinicians in analyzing suspect cytokine profiles by matching them against the said CRS knowledge from past clinical studies, with the ultimate aim of swift CRS diagnosis. We evaluate our proposed methods under several design choices, achieving performance of more than 90% in terms of CRS identification accuracy, and showing that many of our choices outperform a purely data-driven alternative. During evaluation with real-world CRS clinical data, we emphasize the potential of our proposed method of producing interpretable results, in addition to being effective in identifying the onset of cytokine storm. Alex Teodor Bogatu, Magdalena Wysocka, Oskar Wysocki, Holly Butterworth, Manon Pillai, Jennifer Allison, Donal Landers, Elaine Kilgour, Fiona Thistlethwaite, André Freitas |
J. Biomed. Informatics | 1 |
| 2022 | Voyager: Data Discovery and Integration for Onboarding in Data Science
Alex Teodor Bogatu, Norman W. Paton, Mark Douthwaite, André Freitas |
EDBT | 1 |
| 2021 | Natural Language Inference over Tables: Enabling Explainable Data Exploration on Data Lakes
Mario Ramirez, Alex Teodor Bogatu, Norman W. Paton, André Freitas |
ESWC | 2 |
| 2021 | Cost-effective Variational Active Entity ResolutionabstractAccurately identifying different representations of the same real-world entity is an integral part of data cleaning and many methods have been proposed to accomplish it. The challenges of this entity resolution task that demand so much research attention are often rooted in the task-specificity and user-dependence of the process. Adopting deep learning techniques has the potential to lessen these challenges. In this paper, we set out to devise an entity resolution method that builds on the robustness conferred by deep autoencoders to reduce human-involvement costs. Specifically, we reduce the cost of training deep entity resolution models by performing unsupervised representation learning. This unveils a transferability property of the resulting model that can further reduce the cost of applying the approach to new datasets by means of transfer learning. Finally, we reduce the cost of labeling training data through an active learning approach that builds on the properties conferred by the use of deep autoencoders. Empirical evaluation confirms the accomplishment of our cost-reduction desideratum, while achieving comparable effectiveness with state-of-the-art alternatives. Alex Teodor Bogatu, Norman W. Paton, Mark Douthwaite, Stuart Davie, André Freitas |
ICDE | 1 |
| 2021 | Incorporating Data Context to Cost-Effectively Automate End-to-End Data WranglingabstractThe process of preparing potentially large and complex data sets for further analysis or manual examination is often called data wrangling. In classical warehousing environments, the steps in such a process are carried out using Extract-Transform-Load platforms, with significant manual involvement in specifying, configuring or tuning many of them. In typical big data applications, we need to ensure that all wrangling steps, including web extraction, selection, integration and cleaning, benefit from automation wherever possible. Towards this goal, in the paper we: (i) introduce a notion of data context, which associates portions of a target schema with extensional data of types that are commonly available; (ii) define a scalable methodology to bootstrap an end-to-end data wrangling process based on data profiling; (iii) describe how data context is used to inform automation in several steps within wrangling, specifically, matching, value format transformation, data repair, and mapping generation and selection to optimise the accuracy, consistency and relevance of the result; and (iv) we evaluate the approach with real estate data and financial data, showing substantial improvements in the results of automated wrangling. Martin Koehler, Edward Abel, Alex Teodor Bogatu, Cristina Civili, Lacramioara Mazilu, Nikolaos Konstantinou 0001, Alvaro A. A. Fernandes, John A. Keane, Leonid Libkin, Norman W. Paton |
IEEE Trans. Big Data | 3 |
| 2020 | Dataset Discovery in Data LakesabstractData analytics stands to benefit from the increasing availability of datasets that are held without their conceptual relationships being explicitly known. When collected, these datasets form a data lake from which, by processes like data wrangling, specific target datasets can be constructed that enable value- adding analytics. Given the potential vastness of such data lakes, the issue arises of how to pull out of the lake those datasets that might contribute to wrangling out a given target. We refer to this as the problem of dataset discovery in data lakes and this paper contributes an effective and efficient solution to it. Our approach uses features of the values in a dataset to construct hash- based indexes that map those features into a uniform distance space. This makes it possible to define similarity distances between features and to take those distances as measurements of relatedness w.r.t. a target table. Given the latter (and exemplar tuples), our approach returns the most related tables in the lake. We provide a detailed description of the approach and report on empirical results for two forms of relatedness (unionability and joinability) comparing them with prior work, where pertinent, and showing significant improvements in all of precision, recall, target coverage, indexing and discovery times. Alex Teodor Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos Konstantinou 0001 |
ICDE | 1 |
| 2019 | SynthEdit: Format transformations by example using edit operationsabstractFormat transformation is one of the most labor intensive tasks of a data wrangling process. Recent advances in programming by example proposed synthesis algorithms that showed promising results on spreadsheet data. However, when employed on repositories consisting of multiple sources and large number of examples, such algorithms manifest scalability issues. This paper introduces a new transformation synthesis technique based on edit operations that enables efficient learning of transformation programs. Empirical results show comparable effectiveness and dramatic improvements in efficiency over the state-of-the art. Alex Teodor Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, Nikolaos Konstantinou 0001 |
EDBT | 1 |
| 2019 | Towards Automatic Data Format Transformations: Data Wrangling at ScaleabstractAbstract Data wrangling is the process whereby data are cleaned and integrated for analysis. Data wrangling, even with tool support, is typically a labour intensive process. One aspect of data wrangling involves carrying out format transformations on attribute values, for example so that names or phone numbers are represented consistently. Recent research has developed techniques for synthesizing format transformation programs from examples of the source and target representations. This is valuable, but still requires a user to provide suitable examples, something that may be challenging in applications in which there are huge datasets or numerous data sources. In this paper, we investigate the automatic discovery of examples that can be used to synthesize format transformation programs. In particular, we propose two approaches to identifying candidate data examples and validating the transformations that are synthesized from them. The approaches are evaluated empirically using datasets from open government data. Alex Teodor Bogatu, Norman W. Paton, Alvaro A. A. Fernandes, Martin Koehler |
Comput. J. | 1 |
| 2017 | Data context informed data wranglingabstractThe process of preparing potentially large and complex data sets for further analysis or manual examination is often called data wrangling. In classical warehousing environments, the steps in such a process have been carried out using Extract-Transform-Load platforms, with significant manual involvement in specifying, configuring or tuning many of them. Cost-effective data wrangling processes need to ensure that data wrangling steps benefit from automation wherever possible. In this paper, we define a methodology to fully automate an end-to-end data wrangling process incorporating data context, which associates portions of a target schema with potentially spurious extensional data of types that are commonly available. Instance-based evidence together with data profiling paves the way to inform automation in several steps within the wrangling process, specifically, matching, mapping validation, value format transformation, and data repair. The approach is evaluated with real estate data showing substantial improvements in the results of automated wrangling. Martin Koehler, Alex Teodor Bogatu, Cristina Civili, Nikolaos Konstantinou 0001, Edward Abel, Alvaro A. A. Fernandes, John A. Keane, Leonid Libkin, Norman W. Paton |
IEEE BigData | 2 |