EDBT 2026 Demo / reviewers in the wild / expert
Adam Kowalczyk
dblp:48/3093
· DBLP profile ↗
37ranked-venue papers
17as first author
0since 2021 · last 2019
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 15 first-authorApplied, interdisciplinary, general and emerging computing · 13 · 1 first-authorDatabases, data management, data science and information retrieval · 5 · 3 first-authorComputer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 92% Computational science and engineering · 8% | |
| Artificial intelligence
11 papers |
Learning theory · 42% Kernel, tree and ensemble methods · 29% Learning paradigms · 17% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction |
0.1 | 2 | 2011 | is-rSNP: a novel technique for in silico regulatory SNP detection · Bioinform. 2010 Genome annotation test with validation on transcription start site and ChIP-Seq for Pol-II binding data · Bioinform. 2011 |
Bioinformatics and computational biology › cancer genomics › cancer subtype analysis
cancer subtype classification |
0.1 | 1 | 2012 | FSR: feature set reduction for scalable and accurate multi-class cancer subtype classification based on copy number · Bioinform. 2012 |
Bioinformatics and computational biology
genome annotation |
0.1 | 1 | 2011 | Genome annotation test with validation on transcription start site and ChIP-Seq for Pol-II binding data · Bioinform. 2011 |
Bioinformatics and computational biology › gene regulation › promoter analysis
transcription start site prediction |
0.1 | 1 | 2011 | Genome annotation test with validation on transcription start site and ChIP-Seq for Pol-II binding data · Bioinform. 2011 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection |
0.1 | 1 | 2010 | Exploiting sequence similarity to validate the sensitivity of SNP arrays in detecting fine-scaled copy number variations · Bioinform. 2010 |
Bioinformatics and computational biology › gene expression analysis
differential expression analysis |
0.1 | 1 | 2010 | The Poisson Margin Test for Normalisation Free Significance Analysis of NGS Data · RECOMB 2010 |
Bioinformatics and computational biology › sequence analysis
sequencing data analysis |
0.1 | 1 | 2010 | The Poisson Margin Test for Normalisation Free Significance Analysis of NGS Data · RECOMB 2010 |
Computational science and engineering
statistical significance testing |
0.1 | 1 | 2010 | The Poisson Margin Test for Normalisation Free Significance Analysis of NGS Data · RECOMB 2010 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel machines |
0.1 | 2 | 2001 | Kernel Machines and Boolean Functions · NIPS 2001 Sparsity of Data Representation of Optimal Kernel Machine and Leave-one-out Estimator · NIPS 2000 |
Data mining
clustering |
0.0 | 2 | 2002 | Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002 Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002 |
Machine learning › Learning theory
generalization bounds |
0.0 | 2 | 2000 | Sparsity of Data Representation of Optimal Kernel Machine and Leave-one-out Estimator · NIPS 2000 MLP Can Provably Generalize Much Better than VC-bounds Indicate · NIPS 1996 |
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis |
0.0 | 1 | 2011 | Genome annotation test with validation on transcription start site and ChIP-Seq for Pol-II binding data · Bioinform. 2011 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.0 | 1 | 2002 | Multi-Instance Kernels · ICML 2002 |
Machine learning › Learning paradigms
multiple instance learning |
0.0 | 1 | 2002 | Multi-Instance Kernels · ICML 2002 |
Machine learning › Learning paradigms
semi-supervised learning |
0.0 | 1 | 2002 | Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002 |
Natural language and speech › Information extraction and text analysis
text classification |
0.0 | 1 | 2002 | Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002 |
Data mining › predictive modeling › classification
clustering-based classification |
0.0 | 1 | 2002 | Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002 |
Data mining › semi-supervised learning
co-training |
0.0 | 1 | 2002 | Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002 |
Data mining
semi-supervised learning |
0.0 | 1 | 2002 | Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002 |
Data mining › text mining
text classification |
0.0 | 1 | 2002 | Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002 |
Bioinformatics and computational biology
next-generation sequencing |
0.0 | 1 | 2010 | The Poisson Margin Test for Normalisation Free Significance Analysis of NGS Data · RECOMB 2010 |
Machine learning › Learning theory
computational learning theory |
0.0 | 1 | 2001 | Kernel Machines and Boolean Functions · NIPS 2001 |
Machine learning › Kernel, tree and ensemble methods › support vector machine
support vector bounds |
0.0 | 1 | 2000 | Sparsity of Data Representation of Optimal Kernel Machine and Leave-one-out Estimator · NIPS 2000 |
Machine learning › Learning theory › computational learning theory › machine teaching
teaching dimension |
0.0 | 1 | 1997 | Dense Shattering and Teaching Dimensions for Differentiable Families (Extended Abstract) · COLT 1997 |
Network management and operations
network control |
0.0 | 1 | 1997 | Experiments with Simple Neural Networks for Real-Time Control · IEEE J. Sel. Areas Commun. 1997 |
Network optimization and economics
resource allocation |
0.0 | 1 | 1997 | Experiments with Simple Neural Networks for Real-Time Control · IEEE J. Sel. Areas Commun. 1997 |
Machine learning › Learning theory
learning curves |
0.0 | 1 | 1995 | Examples of learning curves from a modified VC-formalism · NIPS 1995 |
Machine learning › Learning theory › computational learning theory
VC theory |
0.0 | 1 | 1995 | Examples of learning curves from a modified VC-formalism · NIPS 1995 |
Hardware accelerators and domain-specific architectures
neural network control |
0.0 | 1 | 1995 | Experiments with Neural Networks for Real Time Implementation of Control · NIPS 1995 |
Embedded and real-time systems
real-time control |
0.0 | 1 | 1995 | Experiments with Neural Networks for Real Time Implementation of Control · NIPS 1995 |
Methods — techniques the papers use, named apart from their topics
feature selection · 0.1dimensionality reduction · 0.1supervised machine learning · 0.1position weight matrix · 0.1calibrated precision metric · 0.1rank-order statistics · 0.1position weight matrix scoring · 0.1poisson margin test · 0.1convolution · 0.1DRECS · 0.1cluster parameters · 0.1support vector machine · 0.0kernel methods · 0.0co-training · 0.0clustering · 0.0regularized risk · 0.0perceptron · 0.0maximal perceptron learning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Exploring effective approaches for haplotype block phasingabstractBACKGROUND: Knowledge of phase, the specific allele sequence on each copy of homologous chromosomes, is increasingly recognized as critical for detecting certain classes of disease-associated mutations. One approach for detecting such mutations is through phased haplotype association analysis. While the accuracy of methods for phasing genotype data has been widely explored, there has been little attention given to phasing accuracy at haplotype block scale. Understanding the combined impact of the accuracy of phasing tool and the method used to determine haplotype blocks on the error rate within the determined blocks is essential to conduct accurate haplotype analyses. RESULTS: We present a systematic study exploring the relationship between seven widely used phasing methods and two common methods for determining haplotype blocks. The evaluation focuses on the number of haplotype blocks that are incorrectly phased. Insights from these results are used to develop a haplotype estimator based on a consensus of three tools. The consensus estimator achieved the most accurate phasing in all applied tests. Individually, EAGLE2, BEAGLE and SHAPEIT2 alternate in being the best performing tool in different scenarios. Determining haplotype blocks based on linkage disequilibrium leads to more correctly phased blocks compared to a sliding window approach. We find that there is little difference between phasing sections of a genome (e.g. a gene) compared to phasing entire chromosomes. Finally, we show that the location of phasing error varies when the tools are applied to the same data several times, a finding that could be important for downstream analyses. CONCLUSIONS: The choice of phasing and block determination algorithms and their interaction impacts the accuracy of phased haplotype blocks. This work provides guidance and evidence for the different design choices needed for analyses using haplotype blocks. The study highlights a number of issues that may have limited the replicability of previous haplotype analysis. Ziad Al Bkhetan, Justin Zobel, Adam Kowalczyk, Karin Verspoor, Benjamin Goudey |
BMC Bioinform. | 3 |
| 2014 | GWISFI: A universal GPU interface for exhaustive search of pairwise interactions in case-control GWAS in minutesabstractEpistatic interactions between genes are believed to be a critical component in the genetic architecture of complex diseases. Genome Wide Association Studies (GWAS) may be able to detect such genetic interactions indirectly, via the identification of associated SNP markers. Major obstacles to progress in this area are: the unknown nature of epistatic interactions, little understanding of the capabilities of different filtering methods, and the computational difficulties for exhaustive analysis. A common platform enabling various detection methods is needed to avoid practical issues such as software compatibility and portability, incompatible input and output formats and varying demands on computational resources. We developed a highly optimised GPU system capable of exhaustively analysing all SNP-pairs in typical GWAS data (0.5M SNPs, 5K samples) in a few minutes on a standard desktop computer. A number of programming elements provided by a functional interface can be used to construct user-defined statistical tests to efficiently score every SNP pair. As a proof of principle, we have implemented 8 methods from the literature via our interface. We have applied all of them using a single GPU to exhaustively scan the 7 popular WTCCC case-control GWAS datasets. We present timing results for these methods, both in their original software implementations and using our platform. Significant improvements in timing are observed, up to 10000 times for CPU implementations of the popular FastEpistasis in PLINK and up to 2 orders of magnitude for some GPU implementations in the literature. As an initial discovery we show plots for overlaps of list of selected pairs by 8 algorithms for Type 2 Diabetes, WTCCC data. Andrew Kowalczyk, Richard M. Campbell, Benjamin Goudey, David Rawlinson 0001, Aaron Harwood, Herman L. Ferrá, Adam Kowalczyk |
BIBM | 9 |
| 2012 | FSR: feature set reduction for scalable and accurate multi-class cancer subtype classification based on copy numberabstractMOTIVATION: Feature selection is a key concept in machine learning for microarray datasets, where features represented by probesets are typically several orders of magnitude larger than the available sample size. Computational tractability is a key challenge for feature selection algorithms in handling very high-dimensional datasets beyond a hundred thousand features, such as in datasets produced on single nucleotide polymorphism microarrays. In this article, we present a novel feature set reduction approach that enables scalable feature selection on datasets with hundreds of thousands of features and beyond. Our approach enables more efficient handling of higher resolution datasets to achieve better disease subtype classification of samples for potentially more accurate diagnosis and prognosis, which allows clinicians to make more informed decisions in regards to patient treatment options. RESULTS: We applied our feature set reduction approach to several publicly available cancer single nucleotide polymorphism (SNP) array datasets and evaluated its performance in terms of its multiclass predictive classification accuracy over different cancer subtypes, its speedup in execution as well as its scalability with respect to sample size and array resolution. Feature Set Reduction (FSR) was able to reduce the dimensions of an SNP array dataset by more than two orders of magnitude while achieving at least equal, and in most cases superior predictive classification performance over that achieved on features selected by existing feature selection methods alone. An examination of the biological relevance of frequently selected features from FSR-reduced feature sets revealed strong enrichment in association with cancer. AVAILABILITY: FSR was implemented in MATLAB R2010b and is available at http://ww2.cs.mu.oz.au/~gwong/FSR. Gerard Wong, Christopher Leckie, Adam Kowalczyk |
Bioinform. | 3 |
| 2012 | SparSNP: Fast and memory-efficient analysis of all SNPs for phenotype predictionabstractBACKGROUND: A central goal of genomics is to predict phenotypic variation from genetic variation. Fitting predictive models to genome-wide and whole genome single nucleotide polymorphism (SNP) profiles allows us to estimate the predictive power of the SNPs and potentially develop diagnostic models for disease. However, many current datasets cannot be analysed with standard tools due to their large size. RESULTS: We introduce SparSNP, a tool for fitting lasso linear models for massive SNP datasets quickly and with very low memory requirements. In analysis on a large celiac disease case/control dataset, we show that SparSNP runs substantially faster than four other state-of-the-art tools for fitting large scale penalised models. SparSNP was one of only two tools that could successfully fit models to the entire celiac disease dataset, and it did so with superior performance. Compared with the other tools, the models generated by SparSNP had better than or equal to predictive performance in cross-validation. CONCLUSIONS: Genomic datasets are rapidly increasing in size, rendering existing approaches to model fitting impractical due to their prohibitive time or memory requirements. This study shows that SparSNP is an essential addition to the genomic analysis toolkit.SparSNP is available at http://www.genomics.csse.unimelb.edu.au/SparSNP. Gad Abraham, Adam Kowalczyk, Justin Zobel, Michael Inouye |
BMC Bioinform. | 2 |
| 2011 | Genome annotation test with validation on transcription start site and ChIP-Seq for Pol-II binding dataabstractMOTIVATION: Many ChIP-Seq experiments are aimed at developing gold standards for determining the locations of various genomic features such as transcription start or transcription factor binding sites on the whole genome. Many such pioneering experiments lack rigorous testing methods and adequate 'gold standard' annotations to compare against as they themselves are the most reliable source of empirical data available. To overcome this problem, we propose a self-consistency test whereby a dataset is tested against itself. It relies on a supervised machine learning style protocol for in silico annotation of a genome and accuracy estimation to guarantee, at least, self-consistency. RESULTS: The main results use a novel performance metric (a calibrated precision) in order to assess and compare the robustness of the proposed supervised learning method across different test sets. As a proof of principle, we applied the whole protocol to two recent ChIP-Seq ENCODE datasets of STAT1 and Pol-II binding sites. STAT1 is benchmarked against in silicodetection of binding sites using available position weight matrices. Pol-II, the main focus of this paper, is benchmarked against 17 algorithms for the closely related and well-studied problem of in silico transcription start site (TSS) prediction. Our results also demonstrate the feasibility of in silico genome annotation extension with encouraging results from a small portion of annotated genome to the remainder. AVAILABILITY: Available fromhttp://www.genomics.csse.unimelb.edu.au/gat. Justin Bedo, Adam Kowalczyk |
Bioinform. | 2 |
| 2011 | Replication of epistatic DNA loci in two case-control GWAS studies using OPE algorithmabstractBackground One of the limiting factors of current genome-wide association studies (GWAS) is the inability of current methods to comprehensively examine SNP interactions for a reasonable sized dataset. It is hypothesised that this limitation is one of the reasons that GWAS studies have not been able to have a greater impact [1,2]. Many current methods for handling interactions are computationally expensive and do not scale to entire studies. Those methods that do scale often achieve this by pruning their datasets in some manner. This is commonly done by considering only those SNPs that show strong marginal effects, despite the fact that a strongly interacting pair may consist of SNPs with low effects individually. Benjamin Goudey, David Rawlinson 0001, Armita Zarnegar, Eder Kikianty, John Markham, Geoff MacIntyre, Gad Abraham, Linda Stern, Michael Inouye, Izhak Haviv, Adam Kowalczyk |
BMC Bioinform. | 12 |
| 2011 | Meta-analysis of gene expression microarrays with missing replicatesabstractBACKGROUND: Many different microarray experiments are publicly available today. It is natural to ask whether different experiments for the same phenotypic conditions can be combined using meta-analysis, in order to increase the overall sample size. However, some genes are not measured in all experiments, hence they cannot be included or their statistical significance cannot be appropriately estimated in traditional meta-analysis. Nonetheless, these genes, which we refer to as incomplete genes, may also be informative and useful. RESULTS: We propose a meta-analysis framework, called "Incomplete Gene Meta-analysis", which can include incomplete genes by imputing the significance of missing replicates, and computing a meta-score for every gene across all datasets. We demonstrate that the incomplete genes are worthy of being included and our method is able to appropriately estimate their significance in two groups of experiments. We first apply the Incomplete Gene Meta-analysis and several comparable methods to five breast cancer datasets with an identical set of probes. We simulate incomplete genes by randomly removing a subset of probes from each dataset and demonstrate that our method consistently outperforms two other methods in terms of their false discovery rate. We also apply the methods to three gastric cancer datasets for the purpose of discriminating diffuse and intestinal subtypes. CONCLUSIONS: Meta-analysis is an effective approach that identifies more robust sets of differentially expressed genes from multiple studies. The incomplete genes that mainly arise from the use of different platforms may also have statistical and biological importance but are ignored or are not appropriately involved by previous studies. Our Incomplete Gene Meta-analysis is able to incorporate the incomplete genes by estimating their significance. The results on both breast and gastric cancer datasets suggest that the highly ranked genes and associated GO terms produced by our method are more significant and biologically meaningful according to the previous literature. Gad Abraham, Christopher Leckie, Izhak Haviv, Adam Kowalczyk |
BMC Bioinform. | 5 |
| 2010 | The Poisson Margin Test for Normalisation Free Significance Analysis of NGS Data
Adam Kowalczyk, Justin Bedo, Thomas C. Conway, Bryan Beresford-Smith |
RECOMB | 1 |
| 2010 | is-rSNP: a novel technique for in silico regulatory SNP detectionabstractMOTIVATION: Determining the functional impact of non-coding disease-associated single nucleotide polymorphisms (SNPs) identified by genome-wide association studies (GWAS) is challenging. Many of these SNPs are likely to be regulatory SNPs (rSNPs): variations which affect the ability of a transcription factor (TF) to bind to DNA. However, experimental procedures for identifying rSNPs are expensive and labour intensive. Therefore, in silico methods are required for rSNP prediction. By scoring two alleles with a TF position weight matrix (PWM), it can be determined which SNPs are likely rSNPs. However, predictions in this manner are noisy and no method exists that determines the statistical significance of a nucleotide variation on a PWM score. RESULTS: We have designed an algorithm for in silico rSNP detection called is-rSNP. We employ novel convolution methods to determine the complete distributions of PWM scores and ratios between allele scores, facilitating assignment of statistical significance to rSNP effects. We have tested our method on 41 experimentally verified rSNPs, correctly predicting the disrupted TF in 28 cases. We also analysed 146 disease-associated SNPs with no known functional impact in an attempt to identify candidate rSNPs. Of the 11 significantly predicted disrupted TFs, 9 had previous evidence of being associated with the disease in the literature. These results demonstrate that is-rSNP is suitable for high-throughput screening of SNPs for potential regulatory function. This is a useful and important tool in the interpretation of GWAS. AVAILABILITY: is-rSNP software is available for use at: www.genomics.csse.unimelb.edu.au/is-rSNP. Geoff MacIntyre, James Bailey 0001, Izhak Haviv, Adam Kowalczyk |
Bioinform. | 4 |
| 2010 | Exploiting sequence similarity to validate the sensitivity of SNP arrays in detecting fine-scaled copy number variationsabstractMOTIVATION: High-density single nucleotide polymorphism (SNP) genotyping arrays are efficient and cost effective platforms for the detection of copy number variation (CNV). To ensure accuracy in probe synthesis and to minimize production costs, short oligonucleotide probe sequences are used. The use of short probe sequences limits the specificity of binding targets in the human genome. The specificity of these short probeset sequences has yet to be fully analysed against a normal reference human genome. Sequence similarity can artificially elevate or suppress copy number measurements, and hence reduce the reliability of affected probe readings. For the purpose of detecting narrow CNVs reliably down to the width of a single probeset, sequence similarity is an important issue that needs to be addressed. RESULTS: We surveyed the Affymetrix Human Mapping SNP arrays for probeset sequence similarity against the reference human genome. Utilizing sequence similarity results, we identified a collection of fine-scaled putative CNVs between gender from autosomal probesets whose sequence matches various loci on the sex chromosomes. To detect these variations, we utilized our statistical approach, Detecting REcurrent Copy number change using rank-order Statistics (DRECS), and showed that its performance was superior and more stable than the t-test in detecting CNVs. Through the application of DRECS on the HapMap population datasets with multi-matching probesets filtered, we identified biologically relevant SNPs in aberrant regions across populations with known association to physical traits, such as height, covered by the span of a single probe. This provided empirical confirmation of the existence of naturally occurring narrow CNVs as well as the sensitivity of the Affymetrix SNP array technology in detecting them. AVAILABILITY: The MATLAB implementation of DRECS is available at http://ww2.cs.mu.oz.au/ approximately gwong/DRECS/index.html. Gerard Wong, Christopher Leckie, Kylie L. Gorringe, Izhak Haviv, Ian G. Campbell, Adam Kowalczyk |
Bioinform. | 6 |
| 2010 | Prediction of breast cancer prognosis using gene set statistics provides signature stability and biological contextabstractBACKGROUND: Different microarray studies have compiled gene lists for predicting outcomes of a range of treatments and diseases. These have produced gene lists that have little overlap, indicating that the results from any one study are unstable. It has been suggested that the underlying pathways are essentially identical, and that the expression of gene sets, rather than that of individual genes, may be more informative with respect to prognosis and understanding of the underlying biological process. RESULTS: We sought to examine the stability of prognostic signatures based on gene sets rather than individual genes. We classified breast cancer cases from five microarray studies according to the risk of metastasis, using features derived from predefined gene sets. The expression levels of genes in the sets are aggregated, using what we call a set statistic. The resulting prognostic gene sets were as predictive as the lists of individual genes, but displayed more consistent rankings via bootstrap replications within datasets, produced more stable classifiers across different datasets, and are potentially more interpretable in the biological context since they examine gene expression in the context of their neighbouring genes in the pathway. In addition, we performed this analysis in each breast cancer molecular subtype, based on ER/HER2 status. The prognostic gene sets found in each subtype were consistent with the biology based on previous analysis of individual genes. CONCLUSIONS: To date, most analyses of gene expression data have focused at the level of the individual genes. We show that a complementary approach of examining the data using predefined gene sets can reduce the noise and could provide increased insight into the underlying biological pathways. Gad Abraham, Adam Kowalczyk, Sherene Loi, Izhak Haviv, Justin Zobel |
BMC Bioinform. | 2 |
| 2010 | is-rSNP: a novel technique for in silico regulatory SNP detection
Geoff MacIntyre, James Bailey 0001, Izhak Haviv, Adam Kowalczyk |
BMC Bioinform. | 4 |
| 2010 | A bi-ordering approach to linking gene expression with clinical annotations in gastric cancerabstractBACKGROUND: In the study of cancer genomics, gene expression microarrays, which measure thousands of genes in a single assay, provide abundant information for the investigation of interesting genes or biological pathways. However, in order to analyze the large number of noisy measurements in microarrays, effective and efficient bioinformatics techniques are needed to identify the associations between genes and relevant phenotypes. Moreover, systematic tests are needed to validate the statistical and biological significance of those discoveries. RESULTS: In this paper, we develop a robust and efficient method for exploratory analysis of microarray data, which produces a number of different orderings (rankings) of both genes and samples (reflecting correlation among those genes and samples). The core algorithm is closely related to biclustering, and so we first compare its performance with several existing biclustering algorithms on two real datasets - gastric cancer and lymphoma datasets. We then show on the gastric cancer data that the sample orderings generated by our method are highly statistically significant with respect to the histological classification of samples by using the Jonckheere trend test, while the gene modules are biologically significant with respect to biological processes (from the Gene Ontology). In particular, some of the gene modules associated with biclusters are closely linked to gastric cancer tumorigenesis reported in previous literature, while others are potentially novel discoveries. CONCLUSION: In conclusion, we have developed an effective and efficient method, Bi-Ordering Analysis, to detect informative patterns in gene expression microarrays by ranking genes and samples. In addition, a number of evaluation metrics were applied to assess both the statistical and biological significance of the resulting bi-orderings. The methodology was validated on gastric cancer and lymphoma datasets. Christopher Leckie, Geoff MacIntyre, Izhak Haviv, Alex Boussioutas, Adam Kowalczyk |
BMC Bioinform. | 6 |
| 2010 | Using Gene Ontology annotations in exploratory microarray clustering to understand cancer etiology
Geoff MacIntyre, James Bailey 0001, Daniel Gustafsson, Izhak Haviv, Adam Kowalczyk |
Pattern Recognit. Lett. | 5 |
| 2007 | Continuity of Performance Metrics for Thin Feature Maps
Adam Kowalczyk |
ALT | 1 |
| 2007 | Classification of Anti-learnable Biological and Synthetic Data
Adam Kowalczyk |
PKDD | 1 |
| 2005 | An Analysis of the Anti-learning Phenomenon for the Class Symmetric Polyhedron
Adam Kowalczyk, Olivier Chapelle |
ALT | 1 |
| 2004 | Exploring Potential of Leave-One-Out Estimator for Calibration of SVM in Text Mining
Adam Kowalczyk, Bhavani Raskutti, Herman L. Ferrá |
PAKDD | 1 |
| 2003 | Exploring Fringe Settings of SVMs for Classification
Adam Kowalczyk, Bhavani Raskutti |
PKDD | 1 |
| 2002 | Multi-Instance Kernels
Thomas Gärtner 0001, Peter A. Flach, Adam Kowalczyk, Alexander J. Smola |
ICML | 3 |
| 2002 | Using Unlabelled Data for Text Classification through Addition of Cluster Parameters
Bhavani Raskutti, Herman L. Ferrá, Adam Kowalczyk |
ICML | 3 |
| 2002 | Combining clustering and co-training to enhance text classification using unlabelled dataabstractIn this paper, we present a new co-training strategy that makes use of unlabelled data. It trains two predictors in parallel, with each predictor labelling the unlabelled data for training the other predictor in the next round. Both predictors are support vector machines, one trained using data from the original feature space, the other trained with new features that are derived by clustering both the labelled and unlabelled data. Hence, unlike standard co-training methods, our method does not require a priori the existence of two redundant views either of which can be used for classification, nor is it dependent on the availability of two different supervised learning algorithms that complement each other.We evaluated our method with two classifiers and three text benchmarks: WebKB, Reuters newswire articles and 20 NewsGroups. Our evaluation shows that our co-training technique improves text classification accuracy especially when the number of labelled examples are very few. Bhavani Raskutti, Herman L. Ferrá, Adam Kowalczyk |
KDD | 3 |
| 2001 | Second Order Features for Maximising Text Classification Performance
Bhavani Raskutti, Herman L. Ferrá, Adam Kowalczyk |
ECML | 3 |
| 2001 | Kernel Machines and Boolean FunctionsabstractWe give results about the learnability and required complexity of logical formulae to solve classification problems. These results are obtained by linking propositional logic with kernel machines. In particular we show that decision trees and disjunctive normal forms (DNF) can be repre- sented by the help of a special kernel, linking regularized risk to separa- tion margin. Subsequently we derive a number of lower bounds on the required complexity of logic formulae using properties of algorithms for generation of linear estimators, such as perceptron and maximal percep- tron learning. Adam Kowalczyk, Alexander J. Smola, Robert C. Williamson |
NIPS | 1 |
| 2000 | Sparsity of Data Representation of Optimal Kernel Machine and Leave-one-out EstimatorabstractVapnik's result that the expectation of the generalisation error ofthe opti(cid:173) mal hyperplane is bounded by the expectation of the ratio of the number of support vectors to the number of training examples is extended to a broad class of kernel machines. The class includes Support Vector Ma(cid:173) chines for soft margin classification and regression, and Regularization Networks with a variety of kernels and cost functions. We show that key inequalities in Vapnik's result become equalities once "the classification error" is replaced by "the margin error", with the latter defined as an in(cid:173) stance with positive cost. In particular we show that expectations of the true margin error and the empirical margin error are equal, and that the sparse solutions for kernel machines are possible only if the cost function is "partially" insensitive. Adam Kowalczyk |
NIPS | 1 |
| 1997 | Dense Shattering and Teaching Dimensions for Differentiable Families (Extended Abstract)
Adam Kowalczyk |
COLT | 1 |
| 1997 | Experiments with Simple Neural Networks for Real-Time ControlabstractWe demonstrate the practical ability of neural networks (NNs) trained in a supervised mode to extract useful control "knowledge" from a large, high-dimensional empirical database, and then to deliver almost optimal control in "real time". In particular, this paper describes experiments with NN-based controllers for allocating bandwidth capacity in a telecommunications network (SDH). This system was proposed in order to overcome a "real time" response constraint. Two basic architectures, each consisting of a combination of two methods, are evaluated: (1) a feedforward network-heuristic combination and (2) a feedforward network-recurrent network combination. These architectures are compared against a linear programming (LP) optimizer as a benchmark. This LP optimizer was also used as a teacher to label the data samples for the feedforward NN training algorithm. NN-based solutions are very accurate (/spl sim/98% of optimal throughput) and, in contrast to the algorithmic approach, can be delivered in "real time". It is found that while the "human" generated heuristics (greedy search optimization) fail to find a solution in approximately 30% of cases, the best NN fails only in 4.9% of cases. Moreover, it has been found that in spite of the very high dimensionality of the problem (55 inputs and 126 outputs), the solution can be delivered by surprisingly compact NNs, with as little as around 1000 synaptic weights. This proves that on this occasion the NNs were able to extract simple but powerful "heuristics" hidden in the complex sets of numerical data. Peter K. Campbell, Alan Christiansen, Michael Dale, Herman L. Ferrá, Adam Kowalczyk, Jacek Szymanski |
IEEE J. Sel. Areas Commun. | 5 |
| 1997 | Estimates of Storage Capacity of Multilayer Perceptron with Threshold Logic Hidden Units
Adam Kowalczyk |
Neural Networks | 1 |
| 1996 | MLP Can Provably Generalize Much Better than VC-bounds Indicate
Adam Kowalczyk, Herman L. Ferrá |
NIPS | 1 |
| 1995 | Experiments with Neural Networks for Real Time Implementation of Control
Peter K. Campbell, Michael Dale, Herman L. Ferrá, Adam Kowalczyk |
NIPS | 4 |
| 1995 | Examples of learning curves from a modified VC-formalism
Adam Kowalczyk, Jacek Szymanski, Peter L. Bartlett, Robert C. Williamson |
NIPS | 1 |
| 1994 | Generalisation in Feedforward NetworksabstractWe discuss a model of consistent learning with an additional re(cid:173) striction on the probability distribution of training samples, the target concept and hypothesis class. We show that the model pro(cid:173) vides a significant improvement on the upper bounds of sample complexity, i.e. the minimal number of random training samples allowing a selection of the hypothesis with a predefined accuracy and confidence. Further, we show that the model has the poten(cid:173) tial for providing a finite sample complexity even in the case of infinite VC-dimension as well as for a sample complexity below VC-dimension. This is achieved by linking sample complexity to an "average" number of implement able dichotomies of a training sample rather than the maximal size of a shattered sample, i.e. VC-dimension. Adam Kowalczyk, Herman L. Ferrá |
NIPS | 1 |
| 1994 | Developing higher-order networks with empirically selected unitsabstractIntroduces a class of simple polynomial neural network classifiers, called mask perceptrons. A series of algorithms for practical development of such structures is outlined. It relies on ordering of input attributes with respect to their potential usefulness and heuristic driven generation and selection of hidden units (monomial terms) in order to combat the exponential explosion in the number of higher-order monomial terms to choose from. Results of tests for two popular machine learning benchmarking domains (mushroom classification and faulty LED-display), and for two nonstandard domains (spoken digit recognition and article category determination) are given. All results are compared against a number of other classifiers. A procedure for converting a mask perceptron to a classical logic production rule is outlined and shown to produce a number of 100% percent accurate simple rules after training on 6-20% of a database. Adam Kowalczyk, Herman L. Ferrá |
IEEE Trans. Neural Networks | 1 |
| 1993 | Counting Function Theorem for Multi-Layer Networks
Adam Kowalczyk |
NIPS | 1 |
| 1993 | Constructive higher-order network that is polynomial time
Nicholas J. Redding, Adam Kowalczyk, Tom Downs |
Neural Networks | 2 |
| 1992 | Some Estimates on the Number of Connections and Hidden Units for Feed-Forward Networks
Adam Kowalczyk |
NIPS | 1 |
| 1991 | Discovering Production Rules with Higher Order Neural Networks
Adam Kowalczyk, Herman L. Ferrá, Ken Gardiner |
ML | 1 |