VLDB 2026 Research / reviewers in the wild / expert
Giorgos Borboudakis
dblp:39/8355
· DBLP profile ↗
10ranked-venue papers
7as first author
1since 2021 · last 2021
0000-0001-7355-8871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Probabilistic and Bayesian machine learning · 64% Knowledge representation and reasoning · 21% Representation and self-supervised learning · 16% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.4 | 1 | 2019 | Forward-Backward Selection with Early Dropping · J. Mach. Learn. Res. 2019 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.4 | 1 | 2019 | Forward-Backward Selection with Early Dropping · J. Mach. Learn. Res. 2019 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
markov blanket discovery |
0.4 | 1 | 2019 | Forward-Backward Selection with Early Dropping · J. Mach. Learn. Res. 2019 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery |
0.2 | 1 | 2016 | Towards Robust and Versatile Causal Discovery for Business Applications · KDD 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.2 | 1 | 2016 | Towards Robust and Versatile Causal Discovery for Business Applications · KDD 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling › preference reasoning
CP-nets |
0.2 | 1 | 2016 | Towards Robust and Versatile Causal Discovery for Business Applications · KDD 2016 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders |
0.2 | 1 | 2016 | Towards Robust and Versatile Causal Discovery for Business Applications · KDD 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network |
0.1 | 1 | 2012 | Incorporating Causal Prior Knowledge as Path-Constraints in Bayesian Networks and Maximal Ancestral Graphs · ICML 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2012 | Incorporating Causal Prior Knowledge as Path-Constraints in Bayesian Networks and Maximal Ancestral Graphs · ICML 2012 |
Methods — techniques the papers use, named apart from their topics
lasso · 0.4path constraints · 0.1causal prior knowledge · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Extending greedy feature selection algorithms to multiple solutionsabstractMost feature selection methods identify only a single solution. This is acceptable for predictive purposes, but is not sufficient for knowledge discovery if multiple solutions exist. We propose a strategy to extend a class of greedy methods to efficiently identify multiple solutions, and show under which conditions it identifies all solutions. We also introduce a taxonomy of features that takes the existence of multiple solutions into account. Furthermore, we explore different definitions of statistical equivalence of solutions, as well as methods for testing equivalence. A novel algorithm for compactly representing and visualizing multiple solutions is also introduced. In experiments we show that (a) the proposed algorithm is significantly more computationally efficient than the TIE* algorithm, the only alternative approach with similar theoretical guarantees, while identifying similar solutions to it, and (b) that the identified solutions have similar predictive performance. Giorgos Borboudakis, Ioannis Tsamardinos |
Data Min. Knowl. Discov. | 1 |
| 2019 | Forward-Backward Selection with Early DroppingabstractForward-backward selection is one of the most basic and commonly-used feature selection algorithms available. It is also general and conceptually applicable to many different types of data. In this paper, we propose a heuristic that significantly improves its running time, while preserving predictive performance. The idea is to temporarily discard the variables that are conditionally independent with the outcome given the selected variable set. Depending on how those variables are reconsidered and reintroduced, this heuristic gives rise to a family of algorithms with increasingly stronger theoretical guarantees. In distributions that can be faithfully represented by Bayesian networks or maximal ancestral graphs, members of this algorithmic family are able to correctly identify the Markov blanket in the sample limit. In experiments we show that the proposed heuristic increases computational efficiency by about 1-2 orders of magnitude, while selecting fewer or the same number of variables and retaining predictive performance. Furthermore, we show that the proposed algorithm and feature selection with LASSO perform similarly when restricted to select the same number of variables, making the proposed algorithm an attractive alternative for problems where no (efficient) algorithm for LASSO exists. Giorgos Borboudakis, Ioannis Tsamardinos |
J. Mach. Learn. Res. | 1 |
| 2019 | A greedy feature selection algorithm for Big Data of high dimensionalityabstractWe present the Parallel, Forward–Backward with Pruning (PFBP) algorithm for feature selection (FS) for Big Data of high dimensionality. PFBP partitions the data matrix both in terms of rows as well as columns. By employing the concepts of p -values of conditional independence tests and meta-analysis techniques, PFBP relies only on computations local to a partition while minimizing communication costs, thus massively parallelizing computations. Similar techniques for combining local computations are also employed to create the final predictive model. PFBP employs asymptotically sound heuristics to make early, approximate decisions, such as Early Dropping of features from consideration in subsequent iterations, Early Stopping of consideration of features within the same iteration, or Early Return of the winner in each iteration. PFBP provides asymptotic guarantees of optimality for data distributions faithfully representable by a causal network (Bayesian network or maximal ancestral graph). Empirical analysis confirms a super-linear speedup of the algorithm with increasing sample size, linear scalability with respect to the number of features and processing cores. An extensive comparative evaluation also demonstrates the effectiveness of PFBP against other algorithms in its class. The heuristics presented are general and could potentially be employed to other greedy-type of FS algorithms. An application on simulated Single Nucleotide Polymorphism (SNP) data with 500K samples is provided as a use case. Ioannis Tsamardinos, Giorgos Borboudakis, Pavlos Katsogridakis, Polyvios Pratikakis, Vassilis Christophides |
Mach. Learn. | 2 |
| 2018 | Bootstrapping the out-of-sample predictions for efficient and accurate cross-validationabstractCross-Validation (CV), and out-of-sample performance-estimation protocols in general, are often employed both for (a) selecting the optimal combination of algorithms and values of hyper-parameters (called a configuration) for producing the final predictive model, and (b) estimating the predictive performance of the final model. However, the cross-validated performance of the best configuration is optimistically biased. We present an efficient bootstrap method that corrects for the bias, called Bootstrap Bias Corrected CV (BBC-CV). BBC-CV's main idea is to bootstrap the whole process of selecting the best-performing configuration on the out-of-sample predictions of each configuration, without additional training of models. In comparison to the alternatives, namely the nested cross-validation (Varma and Simon in BMC Bioinform 7(1):91, 2006) and a method by Tibshirani and Tibshirani (Ann Appl Stat 822-829, 2009), BBC-CV is computationally more efficient, has smaller variance and bias, and is applicable to any metric of performance (accuracy, AUC, concordance index, mean squared error). Subsequently, we employ again the idea of bootstrapping the out-of-sample predictions to speed up the CV process. Specifically, using a bootstrap-based statistical criterion we stop training of models on new folds of inferior (with high probability) configurations. We name the method Bootstrap Bias Corrected with Dropping CV (BBCD-CV) that is both efficient and provides accurate performance estimates. Ioannis Tsamardinos, Elissavet Greasidou, Giorgos Borboudakis |
Mach. Learn. | 3 |
| 2016 | Towards Robust and Versatile Causal Discovery for Business ApplicationsabstractCausal discovery algorithms can induce some of the causal relations from the data, commonly in the form of a causal network such as a causal Bayesian network. Arguably however, all such algorithms lack far behind what is necessary for a true business application. We develop an initial version of a new, general causal discovery algorithm called ETIO with many features suitable for business applications. These include (a) ability to accept prior causal knowledge (e.g., taking senior driving courses improves driving skills), (b) admitting the presence of latent confounding factors, (c) admitting the possibility of (a certain type of) selection bias in the data (e.g., clients sampled mostly from a given region), (d) ability to analyze data with missing-by-design (i.e., not planned to measure) values (e.g., if two companies merge and their databases measure different attributes), and (e) ability to analyze data from different interventions (e.g., prior and posterior to an advertisement campaign). ETIO is an instance of the logical approach to integrative causal discovery that has been relatively recently introduced and enables the solution of complex reverse-engineering problems in causal discovery. ETIO is compared against the state-of-the-art and is shown to be more effective in terms of speed, with only a slight degradation in terms of learning accuracy, while incorporating all the features above. The code is available on the mensxmachina.org website. Giorgos Borboudakis, Ioannis Tsamardinos |
KDD | 1 |
| 2015 | Bayesian Network Learning with Discrete Case-Control Data
Giorgos Borboudakis, Ioannis Tsamardinos |
UAI | 1 |
| 2013 | Scoring and Searching over Bayesian Networks with Causal and Associative Priors
Giorgos Borboudakis, Ioannis Tsamardinos |
UAI | 1 |
| 2012 | Incorporating Causal Prior Knowledge as Path-Constraints in Bayesian Networks and Maximal Ancestral Graphs
Giorgos Borboudakis, Ioannis Tsamardinos |
ICML | 1 |
| 2011 | A constraint-based approach to incorporate prior knowledge in causal models
Giorgos Borboudakis, Sofia Triantafyllou, Vincenzo Lagani, Ioannis Tsamardinos |
ESANN | 1 |
| 2010 | Permutation Testing Improves Bayesian Network Learning
Ioannis Tsamardinos, Giorgos Borboudakis |
ECML/PKDD (3) | 2 |