EDBT 2026 Demo / reviewers in the wild / expert
Ata Kabán
dblp:k/AtaKaban
· DBLP profile ↗
17ranked-venue papers in the field
6as first author
3since 2021 · last 2024
0000-0003-3733-7064ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 14 (5 first)Database Systems & Data Management · 1Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Self-certified Tuple-Wise Deep Learning
Yunwen Lei, Ata Kabán |
ECML/PKDD (2) | 3 |
| 2024 | Efficient learning with projected histogramsabstractAbstract High dimensional learning is a perennial problem due to challenges posed by the “curse of dimensionality”; learning typically demands more computing resources as well as more training data. In differentially private (DP) settings, this is further exacerbated by noise that needs adding to each dimension to achieve the required privacy. In this paper, we present a surprisingly simple approach to address all of these concerns at once, based on histograms constructed on a low-dimensional random projection (RP) of the data. Our approach exploits RP to take advantage of hidden low-dimensional structures in the data, yielding both computational efficiency, and improved error convergence with respect to the sample size—whereby less training data suffice for learning. We also propose a variant for efficient differentially private (DP) classification that further exploits the data-oblivious nature of both the histogram construction and the RP based dimensionality reduction, resulting in an efficient management of the privacy budget. We present a detailed and rigorous theoretical analysis of generalisation of our algorithms in several settings, showing that our approach is able to exploit low-dimensional structure of the data, ameliorates the ill-effects of noise required for privacy, and has good generalisation under minimal conditions. We also corroborate our findings experimentally, and demonstrate that our algorithms achieve competitive classification accuracy in both non-private and private settings. Zhanliang Huang, Ata Kabán, Henry W. J. Reeve |
Data Min. Knowl. Discov. | 2 |
| 2022 | Noise-Efficient Learning of Differentially Private Partitioning Machine Ensembles
Zhanliang Huang, Yunwen Lei, Ata Kabán |
ECML/PKDD (4) | 3 |
| 2015 | Improved Bounds on the Dot Product under Random Projection and Random Sign ProjectionabstractDot product is a key building block in a number of data mining algorithms from classification, regression, correlation clustering, to information retrieval and many others. When data is high dimensional, the use of random projections may serve as a universal dimensionality reduction method that provides both low distortion guarantees and computational savings. Yet, contrary to the optimal guarantees that are known on the preservation of the Euclidean distance cf. the Johnson-Lindenstrauss lemma, the existing guarantees on the dot product under random projection are loose and incomplete in the current data mining and machine learning literature. Some recent literature even suggested that the dot product may not be preserved when the angle between the original vectors is obtuse. Ata Kabán |
KDD | 1 |
| 2012 | Label-Noise Robust Logistic Regression and Its Applications
Jakramate Bootkrajang, Ata Kabán |
ECML/PKDD (1) | 2 |
| 2010 | Compressed fisher linear discriminant analysis: classification of randomly projected dataabstractWe consider random projections in conjunction with classification, specifically the analysis of Fisher's Linear Discriminant (FLD) classifier in randomly projected data spaces. Robert J. Durrant, Ata Kabán |
KDD | 2 |
| 2008 | Learning with Lq<1 vs L1-Norm Regularisation with Exponentially Many Irrelevant Features
Ata Kabán, Robert J. Durrant |
ECML/PKDD (1) | 1 |
| 2008 | A dynamic bibliometric model for identifying online communities
Xin Wang 0006, Ata Kabán |
Data Min. Knowl. Discov. | 2 |
| 2007 | Robust Visual Mining of Data with Error Information
Jianyong Sun, Ata Kabán, Somak Raychaudhury |
PKDD | 2 |
| 2006 | Deconvolutive Clustering of Markov States
Ata Kabán, Xin Wang 0006 |
ECML | 1 |
| 2005 | Finding Young Stellar Populations in Elliptical Galaxies from Independent Components of Optical SpectraabstractElliptical galaxies are believed to consist of a single population of old stars formed together at an early epoch in the Universe, yet recent analyses of galaxy spectra seem to indicate the presence of significant younger populations of stars in them. The detailed physical modelling of such populations is computationally expensive, inhibiting the detailed analysis of the several million galaxy spectra becoming available over the next few years. Here we present a data mining application aimed at decomposing the spectra of galaxies into several coeval stellar populations, without the use of detailed physical models. This is achieved by performing a linear independent basis transformation that essentially decouples the initial problem of joint processing of a set of correlated spectral measurements into that of the independent processing of a small set of prototypical spectra. Two methods are investigated: (1) A fast projection approach is derived by exploiting the correlation structure of neighboring wavelength bins within the spectral data. (2) A factorisation method that takes advantage of the positivity of the spectra is also investigated. The preliminary results show that typical features observed in stellar population spectra of different evolutionary histories can be convincingly disentangled by these methods, despite the absence of input physics. The success of this basis transformation analysis in recovering physically interpretable representations indicates that this technique is a potentially powerful tool for astronomical data mining. Ata Kabán, Louisa Nolan, Somak Raychaudhury |
SDM | 1 |
| 2005 | Sequential Activity Profiling: Latent Dirichlet Allocation of Markov Chains
Mark A. Girolami, Ata Kabán |
Data Min. Knowl. Discov. | 2 |
| 2005 | Semisupervised Learning of Hierarchical Latent Trait Models for Data VisualizationabstractRecently, we have developed the hierarchical generative topographic mapping (HGTM), an interactive method for visualization of large high-dimensional real-valued data sets. We propose a more general visualization system by extending HGTM in three ways, which allows the user to visualize a wider range of data sets and better support the model development process. 1) We integrate HGTM with noise models from the exponential family of distributions. The basic building block is the latent trait model (LTM). This enables us to visualize data of inherently discrete nature, e.g., collections of documents, in a hierarchical manner. 2) We give the user a choice of initializing the child plots of the current plot in either interactive, or automatic mode. In the interactive mode, the user selects "regions of interest", whereas in the automatic mode, an unsupervised minimum message length (MML)-inspired construction of a mixture of LTMs is employed. The unsupervised construction is particularly useful when high-level plots are covered with dense clusters of highly overlapping data projections, making it difficult to use the interactive mode. Such a situation often arises when visualizing large data sets. 3) We derive general formulas for magnification factors in latent trait models. Magnification factors are a useful tool to improve our understanding of the visualization plots, since they can highlight the boundaries between data clusters. We illustrate our approach on a toy example and evaluate it on three more complex real data sets. Ian T. Nabney, Yi Sun 0001, Peter Tiño, Ata Kabán |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2004 | A generative probabilistic approach to visualizing sets of symbolic sequencesabstractThere is a notable interest in extending probabilistic generative modeling principles to accommodate for more complex structured data types. In this paper we develop a generative probabilistic model for visualizing sets of discrete symbolic sequences. The model, a constrained mixture of discrete hidden Markov models, is a generalization of density-based visualization methods previously developed for static data sets. We illustrate our approach on sequences representing web-log data and chorals by J.S. Bach. Peter Tiño, Ata Kabán, Yi Sun 0001 |
KDD | 2 |
| 2004 | Learning to Read Between the Lines: The Aspect Bernoulli ModelabstractWe present a novel probabilistic multiple cause model for binary observations. In contrast to other approaches, the model is linear and it infers reasons behind both observed and unobserved attributes with the aid of an explanatory variable. We exploit this distinctive feature of the method to automatically distinguish between attributes that are ‘off’ by content and those that are missing. Results on artificially corrupted binary images as well as the expansion of short text documents are given by way of demonstration. Ata Kabán, Ella Bingham, T. Hirsimäki |
SDM | 1 |
| 2003 | On an equivalence between PLSI and LDAabstractLatent Dirichlet Allocation (LDA) is a fully generative approach to language modelling which overcomes the inconsistent generative semantics of Probabilistic Latent Semantic Indexing (PLSI). This paper shows that PLSI is a maximum a posteriori estimated LDA model under a uniform Dirichlet prior, therefore the perceived shortcomings of PLSI can be resolved and elucidated within the LDA framework. Mark A. Girolami, Ata Kabán |
SIGIR | 2 |
| 2002 | A Dynamic Probabilistic Model to Visualise Topic Evolution in Text Streams
Ata Kabán, Mark A. Girolami |
J. Intell. Inf. Syst. | 1 |