Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nello Cristianini

dblp:11/4380 · DBLP profile ↗
← Back
75ranked-venue papers
14as first author
3since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 10 first-author · 2 since 2021Databases, data management, data science and information retrieval · 23 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Theory of computation · 2Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Kernel, tree and ensemble methods · 40% Learning theory · 32% Information extraction and text analysis · 7%
Databases, data mining, and information retrieval
7 papers
Web and social media mining · 32% Information retrieval · 31% Data mining · 24%
Theoretical computer science
6 papers
Mathematical optimization · 73% Graph algorithms and graph theory · 14% Automata and formal languages · 12%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 85% Computational social science and digital humanities · 15%

Topics — the 30 heaviest of 66, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.392005
Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004
Learning the Kernel Matrix with Semi-Definite Programming · ICML 2002
Spectral Kernel Methods for Clustering · NIPS 2001
Data mining
causal inference
0.112011
Refining causality: who copied from whom? · KDD 2011
Web and social media mining
information diffusion
0.112011
Refining causality: who copied from whom? · KDD 2011
Web and social media mining
news analysis
0.112011
NOAM: news outlets analysis and monitoring system · SIGMOD Conference 2011
Mathematical optimization
semidefinite programming
0.132004
Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004
Convex Methods for Transduction · NIPS 2003
Learning the Kernel Matrix with Semi-Definite Programming · ICML 2002
Machine learning › Learning theory
generalization bounds
0.132005
On the eigenspectrum of the gram matrix and the generalization error of kernel-PCA · IEEE Trans. Inf. Theory 2005
On the generalization of soft margin algorithms · IEEE Trans. Inf. Theory 2002
Further Results on the Margin Distribution · COLT 1999
Mathematical optimization › continuous optimization
convex optimization
0.122006
Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006
Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004
Information retrieval
retrieval models
0.122002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Learning Semantic Similarity · NIPS 2002
Natural language and speech › Information extraction and text analysis
text classification
0.122002
Text Classification using String Kernels · J. Mach. Learn. Res. 2002
Text Classification using String Kernels · NIPS 2000
Machine learning and data management
kernel methods
0.122002
Learning Semantic Similarity · NIPS 2002
Text Classification using String Kernels · NIPS 2000
Machine learning › Representation and self-supervised learning › similarity measure
centered kernel alignment
0.122001
Spectral Kernel Methods for Clustering · NIPS 2001
On Kernel-Target Alignment · NIPS 2001
Bioinformatics and computational biology › molecular evolution
gene family evolution
0.112006
CAFE: a computational tool for the study of gene family evolution · Bioinform. 2006
Bioinformatics and computational biology
phylogenetics
0.112006
CAFE: a computational tool for the study of gene family evolution · Bioinform. 2006
Mathematical optimization
combinatorial optimization
0.112006
Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006
Graph algorithms and graph theory
graph partitioning
0.112006
Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006
Mathematical optimization › convex relaxation
semidefinite relaxation
0.112006
Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006
Automata and formal languages
transductions
0.112006
Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006
Machine learning › Learning theory
spectral analysis
0.112005
On the eigenspectrum of the gram matrix and the generalization error of kernel-PCA · IEEE Trans. Inf. Theory 2005
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning
0.122004
Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004
Dynamically Adapting Kernels in Support Vector Machines · NIPS 1998
Machine learning › Learning paradigms › semi-supervised learning
transductive learning
0.022003
Convex Methods for Transduction · NIPS 2003
Large Margin Trees for Induction and Transduction · ICML 1999
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
multiple kernel learning
0.012004
Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004
Machine learning › Optimization for machine learning
convex relaxation
0.012003
Convex Methods for Transduction · NIPS 2003
Machine learning › Kernel, tree and ensemble methods
support vector machine
0.021998
Dynamically Adapting Kernels in Support Vector Machines · NIPS 1998
The Kernel-Adatron Algorithm: A Fast and Simple Learning Procedure for Support Vector Machines · ICML 1998
Computational social science and digital humanities
social network analysis
0.012011
Refining causality: who copied from whom? · KDD 2011
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
kernel matrix learning
0.012002
Learning the Kernel Matrix with Semi-Definite Programming · ICML 2002
Machine learning › Learning theory › generalization bounds
margin bounds
0.012002
On the generalization of soft margin algorithms · IEEE Trans. Inf. Theory 2002
Information retrieval
cross-language information retrieval
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Information retrieval › retrieval models › latent semantic models
latent semantic indexing
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Information retrieval › retrieval models
semantic representation
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Information retrieval › similarity measure
semantic similarity
0.012002
Learning Semantic Similarity · NIPS 2002

Methods — techniques the papers use, named apart from their topics

semidefinite programming · 0.4temporal feature analysis · 0.2natural language processing · 0.2machine learning · 0.2data mining · 0.2support vector machine · 0.1kernel matrix optimization · 0.1eigenproblem approximation · 0.1bag-of-words · 0.1stochastic birth and death process · 0.1maximum likelihood · 0.1convex relaxation · 0.1kernel PCA · 0.1eigenvalue analysis · 0.1large margin · 0.0kernel methods · 0.0convex optimization · 0.0kernel function · 0.0
YearPublicationVenuePosition
2023 QBERT: Generalist Model for Processing Questions
Zhaozhen Xu, Nello Cristianini
IDA2
2023 On Compositionality in Data Embedding
Zhaozhen Xu, Zhijin Guo, Nello Cristianini
IDA3
2022 Self-attention Networks for Non-recurrent Handwritten Text Recognition
Rafael d'Arce, Terence J. T. Norton, Sion L. Hannuna, Nello Cristianini
ICFHR4
2018 Fact Checking from Natural Text with Probabilistic Soft Logic
Nouf Bindris, Saatviga Sudhahar, Nello Cristianini
IDA3
2018 Right for the Right Reason: Training Agnostic Networks
Sen Jia 0002, Thomas Lansdall-Welfare, Nello Cristianini
IDA3
2018 Detecting Shifts in Public Opinion: A Big Data Study of Global News Content
Saatviga Sudhahar, Nello Cristianini
IDA2
2018 Biased Embeddings from Wild Data: Measuring, Understanding and Removing
Adam Sutton, Thomas Lansdall-Welfare, Nello Cristianini
IDA3
2017 Seasonal Variation in Collective Mood via Twitter Content and Medical Purchases
Fabon Dzogang, James Goulding, Stafford Lightman, Nello Cristianini
IDA4
2017 Freudian Slips: Analysing the Internal Representations of a Neural Network from Its Mistakes
Sen Jia 0002, Thomas Lansdall-Welfare, Nello Cristianini
IDA3
2017 The Actors of History: Narrative Network Analysis Reveals the Institutions of Power in British Society Between 1800-1950
Thomas Lansdall-Welfare, Saatviga Sudhahar, Nello Cristianini
IDA4
2015 ThinkBIG - Understanding the Impact of Big Data on Science and Society
Nello Cristianini
ICPRAM (1)1
2015 Network analysis of narrative content in large corpora
abstract
Abstract We present a methodology for the extraction of narrative information from a large corpus. The key idea is to transform the corpus into a network, formed by linking the key actors and objects of the narration, and then to analyse this network to extract information about their relations. By representing information into a single network it is possible to infer relations between these entities, including when they have never been mentioned together. We discuss various types of information that can be extracted by our method, various ways to validate the information extracted and two different application scenarios. Our methodology is very scalable, and addresses specific research needs in social sciences.
Saatviga Sudhahar, Gianluca de Fazio, Roberto Franzosi, Nello Cristianini
Nat. Lang. Eng.4
2015 Learning to classify gender from four million images
Sen Jia 0002, Nello Cristianini
Pattern Recognit. Lett.2
2014 On the coverage of science in the media: A big data study on the impact of the Fukushima disaster
abstract
The contents of English-language online-news over 5 years have been analyzed to explore the impact of the Fukushima disaster on the media coverage of nuclear power. This big data study, based on millions of news articles, involves the extraction of narrative networks, association networks, and sentiment time series. The key finding is that media attitude towards nuclear power has significantly changed in the wake of the Fukushima disaster, in terms of sentiment and in terms of framing, showing a long lasting effect that does not appear to recover before the end of the period covered by this study. In particular, we find that the media discourse has shifted from one of public debate about nuclear power as a viable option for energy supply needs to a re-emergence of the public views of nuclear power and the risks associated with it. The methodology used presents an opportunity to leverage big data for corpus analysis and opens up new possibilities in social scientific research.
Thomas Lansdall-Welfare, Saatviga Sudhahar, Giuseppe A. Veltri, Nello Cristianini
IEEE BigData4
2014 Efficient classification of multi-labeled text streams by clashing
Ricardo Ñanculef, Ilias N. Flaounas, Nello Cristianini
Expert Syst. Appl.3
2013 Finite-Time Analysis of Kernelised Contextual Bandits
Michal Valko, Nathaniel Korda, Rémi Munos, Ilias N. Flaounas, Nello Cristianini
UAI5
2013 Modelling and predicting news popularity
Elena Hensinger, Ilias N. Flaounas, Nello Cristianini
Pattern Anal. Appl.3
2012 ElectionWatch: Detecting Patterns in News Coverage of US Elections
Saatviga Sudhahar, Thomas Lansdall-Welfare, Ilias N. Flaounas, Nello Cristianini
EACL4
2012 Learning Machine Translation from In-domain and Out-of-domain Data
Marco Turchi, Cyril Goutte, Nello Cristianini
EAMT3
2012 An Empirical Comparison of Label Prediction Algorithms on Automatically Inferred Networks
Omar Ali, Giovanni Zappella, Tijl De Bie, Nello Cristianini
ICPRAM (2)4
2012 What Makes Us Click? - Modelling and Predicting the Appeal of News Articles
Elena Hensinger, Ilias N. Flaounas, Nello Cristianini
ICPRAM (2)3
2012 Scalable Corpus Annotation by Graph Construction and Label Propagation
Thomas Lansdall-Welfare, Ilias N. Flaounas, Nello Cristianini
ICPRAM (1)3
2012 The NetCover algorithm for the reconstruction of causal networks
Nick Fyson, Tijl De Bie, Nello Cristianini
Neurocomputing3
2012 Nowcasting Events from the Social Web with Statistical Learning
abstract
We present a general methodology for inferring the occurrence and magnitude of an event or phenomenon by exploring the rich amount of unstructured textual information on the social part of the Web. Having geo-tagged user posts on the microblogging service of Twitter as our input data, we investigate two case studies. The first consists of a benchmark problem, where actual levels of rainfall in a given location and time are inferred from the content of tweets . The second one is a real-life task, where we infer regional Influenza-like Illness rates in the effort of detecting timely an emerging epidemic disease. Our analysis builds on a statistical learning framework, which performs sparse learning via the bootstrapped version of LASSO to select a consistent subset of textual features from a large amount of candidates. In both case studies, selected features indicate close semantic correlation with the target topics and inference, conducted by regression, has a significant performance, especially given the short length --approximately one year-- of Twitter’s data time series.
Vasileios Lampos, Nello Cristianini
ACM Trans. Intell. Syst. Technol.2
2012 An intelligent Web agent that autonomously learns how to translate
abstract
We describe the design of an autonomous agent that can teach itself how to translate from a foreign language, by first assembling its own training set, then using it to improve its vocabulary and language model. The key idea is that a Statistical Mac
Marco Turchi, Tijl De Bie, Nello Cristianini
Web Intell. Agent Syst.3
2011 Automatic Discovery of Patterns in Media Content
Nello Cristianini
CPM1
2011 Refining causality: who copied from whom?
abstract
Inferring causal networks behind observed data is an active area of research with wide applicability to areas such as epidemiology, microbiology and social science. In particular recent research has focused on identifying how information propagates through the Internet. This research has so far only used temporal features of observations, and while reasonable results have been achieved, there is often further information which can be used.
Tristan Snowsill, Nick Fyson, Tijl De Bie, Nello Cristianini
KDD4
2011 Celebrity Watch: Browsing News Content by Exploiting Social Intelligence
Omar Ali, Ilias N. Flaounas, Tijl De Bie, Nello Cristianini
ECML/PKDD (3)4
2011 NOAM: news outlets analysis and monitoring system
abstract
We present NOAM, an integrated platform for the monitoring and analysis of news media content. NOAM is the data management system behind various applications and scientific studies aiming at modelling the mediasphere. The system is also intended to address the need in the AI community for platforms where various AI technologies are integrated and deployed in the real world. It combines a relational database (DB) with state of the art AI technologies, including data mining, machine learning and natural language processing. These technologies are organised in a robust, distributed architecture of collaborating modules, that are used to populate and annotate the DB. NOAM manages tens of millions of news items in multiple languages, automatically annotating them in order to enable queries based on their semantic properties. The system also includes a unified user interface for interacting with its various modules.
Ilias N. Flaounas, Omar Ali, Marco Turchi, Tristan Snowsill, Florent Nicart, Tijl De Bie, Nello Cristianini
SIGMOD Conference7
2010 Flu Detector - Tracking Epidemics on Twitter
Vasileios Lampos, Tijl De Bie, Nello Cristianini
ECML/PKDD (3)3
2010 Detecting Events in a Million New York Times Articles
Tristan Snowsill, Ilias N. Flaounas, Tijl De Bie, Nello Cristianini
ECML/PKDD (3)4
2010 Are we there yet?
Nello Cristianini
Neural Networks1
2009 Estimating the Sentence-Level Quality of Machine Translation Systems
Lucia Specia, Marco Turchi, Nicola Cancedda, Nello Cristianini, Marc Dymetman
EAMT4
2009 Are We There Yet?
Nello Cristianini
ECML/PKDD (1)1
2009 Inference and Validation of Networks
Ilias N. Flaounas, Marco Turchi, Tijl De Bie, Nello Cristianini
ECML/PKDD (1)4
2009 Found in Translation
Marco Turchi, Ilias N. Flaounas, Omar Ali, Tijl De Bie, Tristan Snowsill, Nello Cristianini
ECML/PKDD (2)6
2007 Terminator Detection by Support Vector Machine Utilizing a Stochastic Context-Free Grammar
abstract
A 2-stage detector was designed to find rho-independent transcription terminators in the Escherichia coli genome. The detector includes a stochastic context free grammar (SCFG) component and a support vector machine (SVM) component. To find terminators, the SCFG searches the intergenic regions of nucleotide sequence for local matches to a terminator grammar that was designed and trained utilizing examples of known terminators. The grammar selects sequences that are the best candidates for terminators and assigns them a prefix, stem-loop, suffix structure using the Cocke-Younger-Kasaami (CYK) algorithm, modified to incorporate energy effects of base pairing. The parameters from this inferred structure are passed to the SVM classifier, which distinguishes terminators from non-terminators that score high according to the terminator grammar. The SVM was trained with negative examples drawn from intergenic sequences that include both featureless and RNA gene regions (which were assigned prefix, stem-loop, suffix structure by the SCFG), so that it successfully distinguishes terminators from either of these. The classifier was found to be 96.4% successful during testing
Patricia Francis-Lyon, Nello Cristianini, Stephen R. Holbrook
CIBCB2
2007 Discriminative Sequence Labeling by Z-Score Optimization
Elisa Ricci 0001, Tijl De Bie, Nello Cristianini
ECML3
2007 Learning to Align: A Statistical Approach
Elisa Ricci 0001, Tijl De Bie, Nello Cristianini
IDA3
2007 MINI: Mining Informative Non-redundant Itemsets
Arianna Gallo, Tijl De Bie, Nello Cristianini
PKDD3
2006 CAFE: a computational tool for the study of gene family evolution
abstract
SUMMARY: We present CAFE (Computational Analysis of gene Family Evolution), a tool for the statistical analysis of the evolution of the size of gene families. It uses a stochastic birth and death process to model the evolution of gene family sizes over a phylogeny. For a specified phylogenetic tree, and given the gene family sizes in the extant species, CAFE can estimate the global birth and death rate of gene families, infer the most likely gene family size at all internal nodes, identify gene families that have accelerated rates of gain and loss (quantified by a p-value) and identify which branches cause the p-value to be small for significant families. AVAILABILITY: Software is available from http://www.bio.indiana.edu/~hahnlab/Software.html
Tijl De Bie, Nello Cristianini, Jeffery P. Demuth, Matthew W. Hahn
Bioinform.2
2006 Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem
Tijl De Bie, Nello Cristianini
J. Mach. Learn. Res.2
2005 On the eigenspectrum of the gram matrix and the generalization error of kernel-PCA
abstract
In this paper, the relationships between the eigenvalues of the m/spl times/m Gram matrix K for a kernel /spl kappa/(/spl middot/,/spl middot/) corresponding to a sample x/sub 1/,...,x/sub m/ drawn from a density p(x) and the eigenvalues of the corresponding continuous eigenproblem is analyzed. The differences between the two spectra are bounded and a performance bound on kernel principal component analysis (PCA) is provided showing that good performance can be expected even in very-high-dimensional feature spaces provided the sample eigenvalues fall sufficiently quickly.
John Shawe-Taylor, Christopher K. I. Williams, Nello Cristianini, Jaz S. Kandola
IEEE Trans. Inf. Theory3
2004 A statistical framework for genomic data fusion
abstract
MOTIVATION: During the past decade, the new focus on genomics has highlighted a particular challenge: to integrate the different views of the genome that are provided by various types of experimental data. RESULTS: This paper describes a computational framework for integrating and drawing inferences from a collection of genome-wide measurements. Each dataset is represented via a kernel function, which defines generalized similarity relationships between pairs of entities, such as genes or proteins. The kernel representation is both flexible and efficient, and can be applied to many different types of data. Furthermore, kernel functions derived from different types of data can be combined in a straightforward fashion. Recent advances in the theory of kernel methods have provided efficient algorithms to perform such combinations in a way that minimizes a statistical loss function. These methods exploit semidefinite programming techniques to reduce the problem of finding optimizing kernel combinations to a convex optimization problem. Computational experiments performed using yeast genome-wide datasets, including amino acid sequences, hydropathy profiles, gene expression data and known protein-protein interactions, demonstrate the utility of this approach. A statistical learning algorithm trained from all of these data to recognize particular classes of proteins--membrane proteins and ribosomal proteins--performs significantly better than the same algorithm trained on any single type of data. AVAILABILITY: Supplementary data at http://noble.gs.washington.edu/proj/sdp-svm
Gert R. G. Lanckriet, Tijl De Bie, Nello Cristianini, Michael I. Jordan, William Stafford Noble
Bioinform.3
2004 Learning the Kernel Matrix with Semidefinite Programming
Gert R. G. Lanckriet, Nello Cristianini, Peter L. Bartlett, Laurent El Ghaoui, Michael I. Jordan
J. Mach. Learn. Res.2
2003 Efficiently Learning the Metric with Side-Information
Tijl De Bie, Michinari Momma, Nello Cristianini
ALT3
2003 Kernel Methods for Pattern Analysis
Nello Cristianini
ICTAI1
2003 Convex Methods for Transduction
abstract
The 2-class transduction problem, as formulated by Vapnik [1], involves finding a separating hyperplane for a labelled data set that is also maximally distant from a given set of unlabelled test points. In this form, the problem has exponential computational complexity in the size of the working set. So far it has been attacked by means of integer programming techniques [2] that do not scale to reasonable problem sizes, or by local search procedures [3]. In this paper we present a relaxation of this task based on semi- definite programming (SDP), resulting in a convex optimization problem that has polynomial complexity in the size of the data set. The results are very encouraging for mid sized data sets, however the cost is still too high for large scale problems, due to the high di- mensional search space. To this end, we restrict the feasible region by introducing an approximation based on solving an eigenproblem. With this approximation, the computational cost of the algorithm is such that problems with more than 1000 points can be treated.
Tijl De Bie, Nello Cristianini
NIPS2
2002 On the Eigenspectrum of the Gram Matrix and Its Relationship to the Operator Eigenspectrum
John Shawe-Taylor, Christopher K. I. Williams, Nello Cristianini, Jaz S. Kandola
ALT3
2002 On the Eigenspectrum of the Gram Matrix and Its Relationship to the Operator Eigenspectrum
John Shawe-Taylor, Christopher K. I. Williams, Nello Cristianini, Jaz S. Kandola
Discovery Science3
2002 Learning the Kernel Matrix with Semi-Definite Programming
Gert R. G. Lanckriet, Nello Cristianini, Peter L. Bartlett, Laurent El Ghaoui, Michael I. Jordan
ICML2
2002 Learning Semantic Similarity
abstract
The standard representation of text documents as bags of words suffers from well known limitations, mostly due to its inability to exploit semantic similarity between terms. Attempts to incorpo(cid:173) rate some notion of term similarity include latent semantic index(cid:173) ing [8], the use of semantic networks [9], and probabilistic methods [5]. In this paper we propose two methods for inferring such sim(cid:173) ilarity from a corpus. The first one defines word-similarity based on document-similarity and viceversa, giving rise to a system of equations whose equilibrium point we use to obtain a semantic similarity measure. The second method models semantic relations by means of a diffusion process on a graph defined by lexicon and co-occurrence information. Both approaches produce valid kernel functions parametrised by a real number. The paper shows how the alignment measure can be used to successfully perform model selection over this parameter. Combined with the use of support vector machines we obtain positive results.
Jaz S. Kandola, John Shawe-Taylor, Nello Cristianini
NIPS3
2002 Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis
abstract
The problem of learning a semantic representation of a text document from data is addressed, in the situation where a corpus of unlabeled paired documents is available, each pair being formed by a short En- glish document and its French translation. This representation can then be used for any retrieval, categorization or clustering task, both in a stan- dard and in a cross-lingual setting. By using kernel functions, in this case simple bag-of-words inner products, each part of the corpus is mapped to a high-dimensional space. The correlations between the two spaces are then learnt by using kernel Canonical Correlation Analysis. A set of directions is found in the first and in the second space that are max- imally correlated. Since we assume the two representations are com- pletely independent apart from the semantic content, any correlation be- tween them should reflect some semantic similarity. Certain patterns of English words that relate to a specific meaning should correlate with cer- tain patterns of French words corresponding to the same meaning, across the corpus. Using the semantic representation obtained in this way we first demonstrate that the correlations detected between the two versions of the corpus are significantly higher than random, and hence that a rep- resentation based on such features does capture statistical patterns that should reflect semantic information. Then we use such representation both in cross-language and in single-language retrieval tasks, observing performance that is consistently and significantly superior to LSI on the same data.
Alexei Vinokourov, John Shawe-Taylor, Nello Cristianini
NIPS3
2002 Latent Semantic Kernels
Nello Cristianini, John Shawe-Taylor, Huma Lodhi
J. Intell. Inf. Syst.1
2002 Text Classification using String Kernels
Huma Lodhi, Craig Saunders, John Shawe-Taylor, Nello Cristianini, Christopher J. C. H. Watkins
J. Mach. Learn. Res.4
2002 Editorial: Kernel Methods: Current Research and Future Directions
Nello Cristianini, Colin Campbell, Christopher J. C. Burges
Mach. Learn.1
2002 On the generalization of soft margin algorithms
abstract
Generalization bounds depending on the margin of a classifier are a relatively new development. They provide an explanation of the performance of state-of-the-art learning systems such as support vector machines (SVMs) and Adaboost. The difficulty with these bounds has been either their lack of robustness or their looseness. The question of whether the generalization of a classifier can be more tightly bounded in terms of a robust measure of the distribution of margin values has remained open for some time. The paper answers this open question in the affirmative and, furthermore, the analysis leads to bounds that motivate the previously heuristic soft margin SVM algorithms as well as justifying the use of the quadratic loss in neural network training algorithms. The results are extended to give bounds for the probability of failing to achieve a target accuracy in regression prediction, with a statistical analysis of ridge regression and Gaussian processes as a special case. The analysis presented in the paper has also lead to new boosting algorithms described elsewhere.
John Shawe-Taylor, Nello Cristianini
IEEE Trans. Inf. Theory2
2001 Latent Semantic Kernels
Nello Cristianini, John Shawe-Taylor, Huma Lodhi
ICML1
2001 Composite Kernels for Hypertext Categorisation
Thorsten Joachims, Nello Cristianini, John Shawe-Taylor
ICML2
2001 On Kernel-Target Alignment
abstract
We introduce the notion of kernel-alignment, a measure of similar(cid:173) ity between two kernel functions or between a kernel and a target function. This quantity captures the degree of agreement between a kernel and a given learning task, and has very natural interpre(cid:173) tations in machine learning, leading also to simple algorithms for model selection and learning. We analyse its theoretical properties, proving that it is sharply concentrated around its expected value, and we discuss its relation with other standard measures of per(cid:173) formance. Finally we describe some of the algorithms that can be obtained within this framework, giving experimental results show(cid:173) ing that adapting the kernel to improve alignment on the labelled data significantly increases the alignment on the test set, giving improved classification accuracy. Hence, the approach provides a principled method of performing transduction. Keywords: Kernels, alignment, eigenvectors, eigenvalues, transduction
Nello Cristianini, John Shawe-Taylor, André Elisseeff, Jaz S. Kandola
NIPS1
2001 Spectral Kernel Methods for Clustering
abstract
In this paper we introduce new algorithms for unsupervised learn(cid:173) ing based on the use of a kernel matrix. All the information re(cid:173) quired by such algorithms is contained in the eigenvectors of the matrix or of closely related matrices. We use two different but re(cid:173) lated cost functions, the Alignment and the 'cut cost'. The first one is discussed in a companion paper [3], the second one is based on graph theoretic concepts. Both functions measure the level of clustering of a labeled dataset, or the correlation between data clus(cid:173) ters and labels. We state the problem of unsupervised learning as assigning labels so as to optimize these cost functions. We show how the optimal solution can be approximated by slightly relaxing the corresponding optimization problem, and how this corresponds to using eigenvector information. The resulting simple algorithms are tested on real world data with positive results.
Nello Cristianini, John Shawe-Taylor, Jaz S. Kandola
NIPS1
2001 On the Concentration of Spectral Properties
abstract
We consider the problem of measuring the eigenvalues of a ran(cid:173) domly drawn sample of points. We show that these values can be reliably estimated as can the sum of the tail of eigenvalues. Fur(cid:173) thermore, the residuals when data is projected into a subspace is shown to be reliably estimated on a random sample. Experiments are presented that confirm the theoretical results.
John Shawe-Taylor, Nello Cristianini, Jaz S. Kandola
NIPS2
2000 Query Learning with Large Margin Classifiers
Colin Campbell, Nello Cristianini, Alexander J. Smola
ICML2
2000 Large margin strategies in machine learning
abstract
Controlling the capacity of a learning system in a way that does not depend on the dimensionality of the hypothesis space provides the key for effectively using large neural networks and decision trees, ensemble methods and kernel-induced feature spaces. This extended abstract will provide an overview of recent work in this direction, based on the concepts of margin and margin distribution.
Nello Cristianini
ISCAS1
2000 Text Classification using String Kernels
abstract
We introduce a novel kernel for comparing two text documents. The kernel is an inner product in the feature space consisting of all subsequences of length k. A subsequence is any ordered se(cid:173) quence of k characters occurring in the text though not necessarily contiguously. The subsequences are weighted by an exponentially decaying factor of their full length in the text, hence emphasising those occurrences which are close to contiguous. A direct compu(cid:173) tation of this feature vector would involve a prohibitive amount of computation even for modest values of k, since the dimension of the feature space grows exponentially with k. The paper describes how despite this fact the inner product can be efficiently evaluated by a dynamic programming technique. A preliminary experimental comparison of the performance of the kernel compared with a stan(cid:173) dard word feature space kernel results. [6] is made showing encouraging
Huma Lodhi, John Shawe-Taylor, Nello Cristianini, Christopher J. C. H. Watkins
NIPS3
2000 Support vector machine classification and validation of cancer tissue samples using microarray expression data
abstract
MOTIVATION: DNA microarray experiments generating thousands of gene expression measurements, are being used to gather information from tissue and cell samples regarding gene expression differences that will be useful in diagnosing disease. We have developed a new method to analyse this kind of data using support vector machines (SVMs). This analysis consists of both classification of the tissue samples, and an exploration of the data for mis-labeled or questionable tissue results. RESULTS: We demonstrate the method in detail on samples consisting of ovarian cancer tissues, normal ovarian tissues, and other normal tissues. The dataset consists of expression experiment results for 97,802 cDNAs for each tissue. As a result of computational analysis, a tissue sample is discovered and confirmed to be wrongly labeled. Upon correction of this mistake and the removal of an outlier, perfect classification of tissues is achieved, but not with high confidence. We identify and analyse a subset of genes from the ovarian dataset whose expression is highly differentiated between the types of tissues. To show robustness of the SVM method, two previously published datasets from other types of tissues or cells are analysed. The results are comparable to those previously obtained. We show that other machine learning methods also perform comparably to the SVM on many of those datasets. AVAILABILITY: The SVM software is available at http://www.cs. columbia.edu/ approximately bgrundy/svm.
Terrence S. Furey, Nello Cristianini, Nigel Duffy, David W. Bednarski, Michèl Schummer, David Haussler
Bioinform.2
2000 Enlarging the Margins in Perceptron Decision Trees
Kristin P. Bennett, Nello Cristianini, John Shawe-Taylor, Donghui Wu
Mach. Learn.2
1999 Further Results on the Margin Distribution
abstract
A number of results have bounded generalization error of a classifier in terms of its margin on the training points.There has been some debate about whether the minimum margin is the best measure of the distribution of training set margin values with which to estimate the generalization error.Freund and Schapire [7] have shown how a different function of the margin distribution can be used to bound the number of mistakes of an on-line learning algorithm for a perceptron, as well as an expected error bound.Shawe-Taylor and Cristianini [ 131 showed that a slight generalization of their construction can be used to give a pat style bound on the tail of the distribution of the generalization errors that arise from a given sample size when using threshold linear classifiers.We show that in the linear case the approach can be viewed as a change of kernel and that the algorithms arising from the approach are exactly those originally proposed by Cortes and Vapnik [4].We generalise the basic result to function classes with bounded fat-shattering dimension and the Ii measure for slack variables which gives rise to Vapnik's box constraint algorithm.Finally, application to regression is considered, which includes standard least squares as a special case.Permission to make digital or hard
John Shawe-Taylor, Nello Cristianini
COLT2
1999 A multiplicative updating algorithm for training support vector machine
Nello Cristianini, Colin Campbell, John Shawe-Taylor
ESANN1
1999 Large Margin Trees for Induction and Transduction
Donghui Wu, Kristin P. Bennett, Nello Cristianini, John Shawe-Taylor
ICML3
1999 Large Margin DAGs for Multiclass Classification
John C. Platt, Nello Cristianini, John Shawe-Taylor
NIPS2
1998 Bayesian Classifiers Are Large Margin Hyperplanes in a Hilbert Space
Nello Cristianini, John Shawe-Taylor, Peter Sykacek
ICML1
1998 The Kernel-Adatron Algorithm: A Fast and Simple Learning Procedure for Support Vector Machines
Thilo-Thomas Frieß, Nello Cristianini, Colin Campbell
ICML2
1998 Dynamically Adapting Kernels in Support Vector Machines
Nello Cristianini, Colin Campbell, John Shawe-Taylor
NIPS1
1997 Data-Dependent Structural Risk Minimization for Perceptron Decision Trees
John Shawe-Taylor, Nello Cristianini
NIPS2