Herman L. Ferrá

dblp:50/18 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-authorDatabases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 100%
Artificial intelligence
4 papers
Learning theory · 31% Information extraction and text analysis · 28% Learning paradigms · 28%
Computer networks
1 paper
Network optimization and economics · 50% Network management and operations · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Embedded and real-time systems · 50% Hardware accelerators and domain-specific architectures · 50%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.022002
Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002
Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002
Machine learning › Learning paradigms
semi-supervised learning
0.012002
Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002
Natural language and speech › Information extraction and text analysis
text classification
0.012002
Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002
Data mining › predictive modeling › classification
clustering-based classification
0.012002
Using Unlabelled Data for Text Classification through Addition of Cluster Parameters · ICML 2002
Data mining › semi-supervised learning
co-training
0.012002
Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002
Data mining
semi-supervised learning
0.012002
Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002
Data mining › text mining
text classification
0.012002
Combining clustering and co-training to enhance text classification using unlabelled data · KDD 2002
Network management and operations
network control
0.011997
Experiments with Simple Neural Networks for Real-Time Control · IEEE J. Sel. Areas Commun. 1997
Network optimization and economics
resource allocation
0.011997
Experiments with Simple Neural Networks for Real-Time Control · IEEE J. Sel. Areas Commun. 1997
Machine learning › Learning theory
generalization bounds
0.011996
MLP Can Provably Generalize Much Better than VC-bounds Indicate · NIPS 1996
Hardware accelerators and domain-specific architectures
neural network control
0.011995
Experiments with Neural Networks for Real Time Implementation of Control · NIPS 1995
Embedded and real-time systems
real-time control
0.011995
Experiments with Neural Networks for Real Time Implementation of Control · NIPS 1995
Machine learning › Learning theory
sample complexity
0.011994
Generalisation in Feedforward Networks · NIPS 1994
Machine learning › Learning theory › computational learning theory › VC theory
VC dimension
0.011994
Generalisation in Feedforward Networks · NIPS 1994
Machine learning › Deep learning architectures and training › feedforward neural network
higher-order neural network
0.011991
Discovering Production Rules with Higher Order Neural Networks · ML 1991
Knowledge, reasoning and agents › Knowledge representation and reasoning
rule learning
0.011991
Discovering Production Rules with Higher Order Neural Networks · ML 1991

Methods — techniques the papers use, named apart from their topics

cluster parameters · 0.1support vector machine · 0.0co-training · 0.0clustering · 0.0recurrent neural network · 0.0linear programming · 0.0greedy search heuristic · 0.0feedforward neural network · 0.0generalization bound analysis · 0.0neural network · 0.0average dichotomies analysis · 0.0higher-order neural networks · 0.0
YearPublicationVenuePosition
2014 GWISFI: A universal GPU interface for exhaustive search of pairwise interactions in case-control GWAS in minutes
abstract
Epistatic interactions between genes are believed to be a critical component in the genetic architecture of complex diseases. Genome Wide Association Studies (GWAS) may be able to detect such genetic interactions indirectly, via the identification of associated SNP markers. Major obstacles to progress in this area are: the unknown nature of epistatic interactions, little understanding of the capabilities of different filtering methods, and the computational difficulties for exhaustive analysis. A common platform enabling various detection methods is needed to avoid practical issues such as software compatibility and portability, incompatible input and output formats and varying demands on computational resources. We developed a highly optimised GPU system capable of exhaustively analysing all SNP-pairs in typical GWAS data (0.5M SNPs, 5K samples) in a few minutes on a standard desktop computer. A number of programming elements provided by a functional interface can be used to construct user-defined statistical tests to efficiently score every SNP pair. As a proof of principle, we have implemented 8 methods from the literature via our interface. We have applied all of them using a single GPU to exhaustively scan the 7 popular WTCCC case-control GWAS datasets. We present timing results for these methods, both in their original software implementations and using our platform. Significant improvements in timing are observed, up to 10000 times for CPU implementations of the popular FastEpistasis in PLINK and up to 2 orders of magnitude for some GPU implementations in the literature. As an initial discovery we show plots for overlaps of list of selected pairs by 8 algorithms for Type 2 Diabetes, WTCCC data.
Andrew Kowalczyk, Richard M. Campbell, Benjamin Goudey, David Rawlinson 0001, Aaron Harwood, Herman L. Ferrá, Adam Kowalczyk
BIBM8
2004 Exploring Potential of Leave-One-Out Estimator for Calibration of SVM in Text Mining
Adam Kowalczyk, Bhavani Raskutti, Herman L. Ferrá
PAKDD3
2003 Applying Reinforcement Learning to Packet Scheduling in Routers
Herman L. Ferrá, Ken Lau, Christopher Leckie, Anderson Tang
IAAI1
2002 Using Unlabelled Data for Text Classification through Addition of Cluster Parameters
Bhavani Raskutti, Herman L. Ferrá, Adam Kowalczyk
ICML2
2002 Combining clustering and co-training to enhance text classification using unlabelled data
abstract
In this paper, we present a new co-training strategy that makes use of unlabelled data. It trains two predictors in parallel, with each predictor labelling the unlabelled data for training the other predictor in the next round. Both predictors are support vector machines, one trained using data from the original feature space, the other trained with new features that are derived by clustering both the labelled and unlabelled data. Hence, unlike standard co-training methods, our method does not require a priori the existence of two redundant views either of which can be used for classification, nor is it dependent on the availability of two different supervised learning algorithms that complement each other.We evaluated our method with two classifiers and three text benchmarks: WebKB, Reuters newswire articles and 20 NewsGroups. Our evaluation shows that our co-training technique improves text classification accuracy especially when the number of labelled examples are very few.
Bhavani Raskutti, Herman L. Ferrá, Adam Kowalczyk
KDD2
2001 Second Order Features for Maximising Text Classification Performance
Bhavani Raskutti, Herman L. Ferrá, Adam Kowalczyk
ECML2
1997 Experiments with Simple Neural Networks for Real-Time Control
abstract
We demonstrate the practical ability of neural networks (NNs) trained in a supervised mode to extract useful control "knowledge" from a large, high-dimensional empirical database, and then to deliver almost optimal control in "real time". In particular, this paper describes experiments with NN-based controllers for allocating bandwidth capacity in a telecommunications network (SDH). This system was proposed in order to overcome a "real time" response constraint. Two basic architectures, each consisting of a combination of two methods, are evaluated: (1) a feedforward network-heuristic combination and (2) a feedforward network-recurrent network combination. These architectures are compared against a linear programming (LP) optimizer as a benchmark. This LP optimizer was also used as a teacher to label the data samples for the feedforward NN training algorithm. NN-based solutions are very accurate (/spl sim/98% of optimal throughput) and, in contrast to the algorithmic approach, can be delivered in "real time". It is found that while the "human" generated heuristics (greedy search optimization) fail to find a solution in approximately 30% of cases, the best NN fails only in 4.9% of cases. Moreover, it has been found that in spite of the very high dimensionality of the problem (55 inputs and 126 outputs), the solution can be delivered by surprisingly compact NNs, with as little as around 1000 synaptic weights. This proves that on this occasion the NNs were able to extract simple but powerful "heuristics" hidden in the complex sets of numerical data.
Peter K. Campbell, Alan Christiansen, Michael Dale, Herman L. Ferrá, Adam Kowalczyk, Jacek Szymanski
IEEE J. Sel. Areas Commun.4
1996 MLP Can Provably Generalize Much Better than VC-bounds Indicate
Adam Kowalczyk, Herman L. Ferrá
NIPS2
1995 Experiments with Neural Networks for Real Time Implementation of Control
Peter K. Campbell, Michael Dale, Herman L. Ferrá, Adam Kowalczyk
NIPS3
1994 Generalisation in Feedforward Networks
abstract
We discuss a model of consistent learning with an additional re(cid:173) striction on the probability distribution of training samples, the target concept and hypothesis class. We show that the model pro(cid:173) vides a significant improvement on the upper bounds of sample complexity, i.e. the minimal number of random training samples allowing a selection of the hypothesis with a predefined accuracy and confidence. Further, we show that the model has the poten(cid:173) tial for providing a finite sample complexity even in the case of infinite VC-dimension as well as for a sample complexity below VC-dimension. This is achieved by linking sample complexity to an "average" number of implement able dichotomies of a training sample rather than the maximal size of a shattered sample, i.e. VC-dimension.
Adam Kowalczyk, Herman L. Ferrá
NIPS2
1994 Developing higher-order networks with empirically selected units
abstract
Introduces a class of simple polynomial neural network classifiers, called mask perceptrons. A series of algorithms for practical development of such structures is outlined. It relies on ordering of input attributes with respect to their potential usefulness and heuristic driven generation and selection of hidden units (monomial terms) in order to combat the exponential explosion in the number of higher-order monomial terms to choose from. Results of tests for two popular machine learning benchmarking domains (mushroom classification and faulty LED-display), and for two nonstandard domains (spoken digit recognition and article category determination) are given. All results are compared against a number of other classifiers. A procedure for converting a mask perceptron to a classical logic production rule is outlined and shown to produce a number of 100% percent accurate simple rules after training on 6-20% of a database.
Adam Kowalczyk, Herman L. Ferrá
IEEE Trans. Neural Networks2
1991 Discovering Production Rules with Higher Order Neural Networks
Adam Kowalczyk, Herman L. Ferrá, Ken Gardiner
ML2