Panagiotis Mandros

dblp:186/1481 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
1since 2021 · last 2024
0009-0008-9638-9722ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 5 first-authorArtificial intelligence and machine learning · 6 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Data integration and cleaning · 48% Data mining · 45% Database theory · 7%
Theoretical computer science
4 papers
Computational complexity · 49% Information theory · 38% Mathematical optimization · 14%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › dependency discovery
functional dependency discovery
1.232020
Discovering Approximate Functional Dependencies using Smoothed Mutual Information · KDD 2020
Discovering Functional Dependencies from Mixed-Type Data · KDD 2020
Discovering Reliable Approximate Functional Dependencies · KDD 2017
Data mining
pattern mining
1.132019
Discovering Reliable Dependencies from Data: Hardness and Improved Algorithms · IJCAI 2019
Discovering Reliable Correlations in Categorical Data · ICDM 2019
Discovering Reliable Dependencies from Data: Hardness and Improved Algorithms · ICDM 2018
Bioinformatics and computational biology › biological network › network biology
network inference
0.812024
BONOBO: Bayesian Optimized Sample-Specific Networks Obtained by Omics Data · RECOMB 2024
Data integration and cleaning › dependency discovery › functional dependency discovery
approximate functional dependency discovery
0.722020
Discovering Approximate Functional Dependencies using Smoothed Mutual Information · KDD 2020
Discovering Reliable Approximate Functional Dependencies · KDD 2017
Data mining › statistical analysis › statistical estimation
mutual information estimation
0.412020
Discovering Functional Dependencies from Mixed-Type Data · KDD 2020
Information theory › information measures
mutual information
0.412020
Discovering Approximate Functional Dependencies using Smoothed Mutual Information · KDD 2020
Data mining › multivariate data analysis
correlation analysis
0.412019
Discovering Reliable Correlations in Categorical Data · ICDM 2019
Database theory › dependency theory
functional dependency
0.312017
Discovering Reliable Approximate Functional Dependencies · KDD 2017
Data integration and cleaning
data profiling
0.112020
Discovering Approximate Functional Dependencies using Smoothed Mutual Information · KDD 2020
Information theory › information measures › mutual information
total correlation
0.112019
Discovering Reliable Correlations in Categorical Data · ICDM 2019
Mathematical optimization › integer programming
branch-and-bound
0.112018
Discovering Reliable Dependencies from Data: Hardness and Improved Algorithms · ICDM 2018
Mathematical optimization
combinatorial optimization
0.112018
Discovering Reliable Dependencies from Data: Hardness and Improved Algorithms · ICDM 2018

Methods — techniques the papers use, named apart from their topics

branch-and-bound · 2.3greedy algorithm · 1.4bounding function · 1.4uniform priors · 0.9smoothed mutual information · 0.9heuristic search · 0.8exact search · 0.8bayesian optimization · 0.8approximate search · 0.8mutual information estimator · 0.4discretization · 0.4admissible bounding function · 0.4
YearPublicationVenuePosition
2024 BONOBO: Bayesian Optimized Sample-Specific Networks Obtained by Omics Data
Enakshi Saha, Viola Fanfani, Panagiotis Mandros, Marouen Ben Guebila, Jonas Fischer, Katherine H. Shutta, Kimberly Glass, Dawn L. DeMeo, Camila Miranda Lopes-Ramos, John Quackenbush
RECOMB3
2020 Discovering Functional Dependencies from Mixed-Type Data
abstract
Given complex data collections, practitioners can perform non-parametric functional dependency discovery (FDD) to uncover relationships between variables that were previously unknown. However, known FDD methods are applicable to nominal data, and in practice non-nominal variables are discretized, e.g., in a pre-processing step. This is problematic because, as soon as a mix of discrete and continuous variables is involved, the interaction of discretization with the various dependency measures from the literature is poorly understood. In particular, it is unclear whether a given discretization method even leads to a consistent dependency estimate. In this paper, we analyze these fundamental questions and derive formal criteria as to when a discretization process applied to a mixed set of random variables leads to consistent estimates of mutual information. With these insights, we derive an estimator framework applicable to any task that involves estimating mutual information from multivariate and mixed-type data. Last, we extend with this framework a previously proposed FDD approach for reliable dependencies. Experimental evaluation shows that the derived reliable estimator is both computationally and statistically efficient, and leads to effective FDD algorithms for mixed-type data.
Panagiotis Mandros, David Kaltenpoth, Mario Boley, Jilles Vreeken
KDD1
2020 Discovering Approximate Functional Dependencies using Smoothed Mutual Information
abstract
We consider the task of discovering the top-K reliable approximate functional dependencies X -> Y from high dimensional data. While naively maximizing mutual information involving high dimensional entropies over empirical data is subject to false discoveries, correcting the empirical estimator against data sparsity can lead to efficient exact algorithms for robust dependency discovery. Previous approaches focused on correcting by subtracting expected values of different null hypothesis models. In this paper, we consider a different correction strategy and counter data sparsity using uniform priors and smoothing techniques, that leads to an efficient and robust estimating process. In addition, we derive an admissible and tight bounding function for the smoothed estimator that allows us to efficiently solve via branch-and-bound the hard search problem for the top-K dependencies. Our experiments show that our approach is much faster than previous proposals, and leads to the discovery of sparse and informative functional dependencies.
Frédéric Pennerath, Panagiotis Mandros, Jilles Vreeken
KDD2
2020 Discovering dependencies with reliable mutual information
abstract
Abstract We consider the task of discovering functional dependencies in data for target attributes of interest. To solve it, we have to answer two questions: How do we quantify the dependency in a model-agnostic and interpretable way as well as reliably against sample size and dimensionality biases? How can we efficiently discover the exact or $$\alpha $$ α -approximate top-kdependencies? We address the first question by adopting information-theoretic notions. Specifically, we consider the mutual information score, for which we propose a reliable estimator that enables robust optimization in high-dimensional data. To address the second question, we then systematically explore the algorithmic implications of using this measure for optimization. We show the problem is NP-hard and justify worst-case exponential-time as well as heuristic search methods. We propose two bounding functions for the estimator, which we use as pruning criteria in branch-and-bound search to efficiently mine dependencies with approximation guarantees. Empirical evaluation shows that the derived estimator has desirable statistical properties, the bounding functions lead to effective exact and greedy search algorithms, and when combined, qualitative experiments show the framework indeed discovers highly informative dependencies.
Panagiotis Mandros, Mario Boley, Jilles Vreeken
Knowl. Inf. Syst.1
2019 Discovering Reliable Correlations in Categorical Data
abstract
In many scientific tasks we are interested in finding correlations in our data. This raises many questions, such as how to reliably and interpretably measure correlation between a multivariate set of attributes, how to do so without having to make assumptions on data distribution or the type of correlation, and, how to search efficiently for the most correlated attribute sets. We answer these questions for discovery tasks with categorical data. In particular, we propose a corrected-for-chance, consistent, and efficient estimator for normalized total correlation, in order to obtain a reliable, interpretable, and non-parametric measure for correlation over multivariate sets. For the discovery of the top-k correlated sets, we derive an effective algorithmic framework based on a tight bounding function. This framework offers exact, approximate, and heuristic search. Empirical evaluation shows that already for small sample sizes the estimator leads to low-regret optimization outcomes, while the algorithms are shown to be highly effective for both large and high-dimensional data. Through a case study we confirm that our discovery framework identifies interesting and meaningful correlations.
Panagiotis Mandros, Mario Boley, Jilles Vreeken
ICDM1
2019 Discovering Reliable Dependencies from Data: Hardness and Improved Algorithms
abstract
The reliable fraction of information is an attractive score for quantifying (functional) dependencies in high-dimensional data. In this paper, we systematically explore the algorithmic implications of using this measure for optimization. We show that the problem is NP-hard, justifying worst-case exponential-time as well as heuristic search methods. We then substantially improve the practical performance for both optimization styles by deriving a novel admissible bounding function that has an unbounded potential for additional pruning over the previously proposed one. Finally, we empirically investigate the approximation ratio of the greedy algorithm and show that it produces highly competitive results in a fraction of time needed for complete branch-and-bound style search.
Panagiotis Mandros, Mario Boley, Jilles Vreeken
IJCAI1
2018 Discovering Reliable Dependencies from Data: Hardness and Improved Algorithms
abstract
The reliable fraction of information is an attractive score for quantifying (functional) dependencies in high-dimensional data. In this paper, we systematically explore the algorithmic implications of using this measure for optimization. We show that the problem is NP-hard, which justifies the usage of worst-case exponential-time as well as heuristic search methods. We then substantially improve the practical performance for both optimization styles by deriving a novel admissible bounding function that has an unbounded potential for additional pruning over the previously proposed one. Finally, we empirically investigate the approximation ratio of the greedy algorithm and show that it produces highly competitive results in a fraction of time needed for complete branch-and-bound style search.
Panagiotis Mandros, Mario Boley, Jilles Vreeken
ICDM1
2017 Discovering Reliable Approximate Functional Dependencies
abstract
Given a database and a target attribute of interest, how can we tell whether there exists a functional, or approximately functional dependence of the target on any set of other attributes in the data? How can we reliably, without bias to sample size or dimensionality, measure the strength of such a dependence? And, how can we efficiently discover the optimal or α-approximate top-k dependencies? These are exactly the questions we answer in this paper.
Panagiotis Mandros, Mario Boley, Jilles Vreeken
KDD1
2016 Universal Dependency Analysis
abstract
Most data is multi-dimensional. Discovering whether any subset of dimensions, or subspaces, shows dependence is a core task in data mining. To do so, we require a measure that quantifies how dependent a subspace is. For practical use, such a measure should be universal in the sense that it captures correlation in subspaces of any dimensionality and allows to meaningfully compare scores across different subspaces, regardless how many dimensions they have and what specific statistical properties their dimensions possess. Further, it would be nice if the measure can non-parametrically and efficiently capture both linear and non-linear correlations. In this paper, we propose UDS, a multivariate dependence measure that fulfils all of these desiderata. In short, we define UDS based on cumulative entropy and propose a principled normalisation scheme to bring its scores across different subspaces to the same domain, enabling universal dependence assessment. UDS is purely non-parametric as we make no assumption on data distributions nor types of correlation. To compute it on empirical data, we introduce an efficient and non-parametric method. Extensive experiments show that UDS outperforms state of the art.
Hoang Vu Nguyen, Panagiotis Mandros, Jilles Vreeken
SDM2