EDBT 2026 Demo / reviewers in the wild / expert
Nello Cristianini
dblp:11/4380
· DBLP profile ↗
75ranked-venue papers
14as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 58 · 10 first-author · 2 since 2021Databases, data management, data science and information retrieval · 23 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6Theory of computation · 2Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
21 papers |
Kernel, tree and ensemble methods · 40% Learning theory · 32% Information extraction and text analysis · 7% | |
| Databases, data mining, and information retrieval
7 papers |
Web and social media mining · 32% Information retrieval · 31% Data mining · 24% | |
| Theoretical computer science
6 papers |
Mathematical optimization · 73% Graph algorithms and graph theory · 14% Automata and formal languages · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 85% Computational social science and digital humanities · 15% |
Topics — the 30 heaviest of 66, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.3 | 9 | 2005 | Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004 Learning the Kernel Matrix with Semi-Definite Programming · ICML 2002 Spectral Kernel Methods for Clustering · NIPS 2001 |
Data mining
causal inference |
0.1 | 1 | 2011 | Refining causality: who copied from whom? · KDD 2011 |
Web and social media mining
information diffusion |
0.1 | 1 | 2011 | Refining causality: who copied from whom? · KDD 2011 |
Web and social media mining
news analysis |
0.1 | 1 | 2011 | NOAM: news outlets analysis and monitoring system · SIGMOD Conference 2011 |
Mathematical optimization
semidefinite programming |
0.1 | 3 | 2004 | Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004 Convex Methods for Transduction · NIPS 2003 Learning the Kernel Matrix with Semi-Definite Programming · ICML 2002 |
Machine learning › Learning theory
generalization bounds |
0.1 | 3 | 2005 | On the eigenspectrum of the gram matrix and the generalization error of kernel-PCA · IEEE Trans. Inf. Theory 2005 On the generalization of soft margin algorithms · IEEE Trans. Inf. Theory 2002 Further Results on the Margin Distribution · COLT 1999 |
Mathematical optimization › continuous optimization
convex optimization |
0.1 | 2 | 2006 | Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006 Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004 |
Information retrieval
retrieval models |
0.1 | 2 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 Learning Semantic Similarity · NIPS 2002 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 2 | 2002 | Text Classification using String Kernels · J. Mach. Learn. Res. 2002 Text Classification using String Kernels · NIPS 2000 |
Machine learning and data management
kernel methods |
0.1 | 2 | 2002 | Learning Semantic Similarity · NIPS 2002 Text Classification using String Kernels · NIPS 2000 |
Machine learning › Representation and self-supervised learning › similarity measure
centered kernel alignment |
0.1 | 2 | 2001 | Spectral Kernel Methods for Clustering · NIPS 2001 On Kernel-Target Alignment · NIPS 2001 |
Bioinformatics and computational biology › molecular evolution
gene family evolution |
0.1 | 1 | 2006 | CAFE: a computational tool for the study of gene family evolution · Bioinform. 2006 |
Bioinformatics and computational biology
phylogenetics |
0.1 | 1 | 2006 | CAFE: a computational tool for the study of gene family evolution · Bioinform. 2006 |
Mathematical optimization
combinatorial optimization |
0.1 | 1 | 2006 | Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006 |
Graph algorithms and graph theory
graph partitioning |
0.1 | 1 | 2006 | Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006 |
Mathematical optimization › convex relaxation
semidefinite relaxation |
0.1 | 1 | 2006 | Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006 |
Automata and formal languages
transductions |
0.1 | 1 | 2006 | Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem · J. Mach. Learn. Res. 2006 |
Machine learning › Learning theory
spectral analysis |
0.1 | 1 | 2005 | On the eigenspectrum of the gram matrix and the generalization error of kernel-PCA · IEEE Trans. Inf. Theory 2005 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning |
0.1 | 2 | 2004 | Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004 Dynamically Adapting Kernels in Support Vector Machines · NIPS 1998 |
Machine learning › Learning paradigms › semi-supervised learning
transductive learning |
0.0 | 2 | 2003 | Convex Methods for Transduction · NIPS 2003 Large Margin Trees for Induction and Transduction · ICML 1999 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
multiple kernel learning |
0.0 | 1 | 2004 | Learning the Kernel Matrix with Semidefinite Programming · J. Mach. Learn. Res. 2004 |
Machine learning › Optimization for machine learning
convex relaxation |
0.0 | 1 | 2003 | Convex Methods for Transduction · NIPS 2003 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.0 | 2 | 1998 | Dynamically Adapting Kernels in Support Vector Machines · NIPS 1998 The Kernel-Adatron Algorithm: A Fast and Simple Learning Procedure for Support Vector Machines · ICML 1998 |
Computational social science and digital humanities
social network analysis |
0.0 | 1 | 2011 | Refining causality: who copied from whom? · KDD 2011 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
kernel matrix learning |
0.0 | 1 | 2002 | Learning the Kernel Matrix with Semi-Definite Programming · ICML 2002 |
Machine learning › Learning theory › generalization bounds
margin bounds |
0.0 | 1 | 2002 | On the generalization of soft margin algorithms · IEEE Trans. Inf. Theory 2002 |
Information retrieval
cross-language information retrieval |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Information retrieval › retrieval models › latent semantic models
latent semantic indexing |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Information retrieval › retrieval models
semantic representation |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Information retrieval › similarity measure
semantic similarity |
0.0 | 1 | 2002 | Learning Semantic Similarity · NIPS 2002 |
Methods — techniques the papers use, named apart from their topics
semidefinite programming · 0.4temporal feature analysis · 0.2natural language processing · 0.2machine learning · 0.2data mining · 0.2support vector machine · 0.1kernel matrix optimization · 0.1eigenproblem approximation · 0.1bag-of-words · 0.1stochastic birth and death process · 0.1maximum likelihood · 0.1convex relaxation · 0.1kernel PCA · 0.1eigenvalue analysis · 0.1large margin · 0.0kernel methods · 0.0convex optimization · 0.0kernel function · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | QBERT: Generalist Model for Processing Questions
Zhaozhen Xu, Nello Cristianini |
IDA | 2 |
| 2023 | On Compositionality in Data Embedding
Zhaozhen Xu, Zhijin Guo, Nello Cristianini |
IDA | 3 |
| 2022 | Self-attention Networks for Non-recurrent Handwritten Text Recognition
Rafael d'Arce, Terence J. T. Norton, Sion L. Hannuna, Nello Cristianini |
ICFHR | 4 |
| 2018 | Fact Checking from Natural Text with Probabilistic Soft Logic
Nouf Bindris, Saatviga Sudhahar, Nello Cristianini |
IDA | 3 |
| 2018 | Right for the Right Reason: Training Agnostic Networks
Sen Jia 0002, Thomas Lansdall-Welfare, Nello Cristianini |
IDA | 3 |
| 2018 | Detecting Shifts in Public Opinion: A Big Data Study of Global News Content
Saatviga Sudhahar, Nello Cristianini |
IDA | 2 |
| 2018 | Biased Embeddings from Wild Data: Measuring, Understanding and Removing
Adam Sutton, Thomas Lansdall-Welfare, Nello Cristianini |
IDA | 3 |
| 2017 | Seasonal Variation in Collective Mood via Twitter Content and Medical Purchases
Fabon Dzogang, James Goulding, Stafford Lightman, Nello Cristianini |
IDA | 4 |
| 2017 | Freudian Slips: Analysing the Internal Representations of a Neural Network from Its Mistakes
Sen Jia 0002, Thomas Lansdall-Welfare, Nello Cristianini |
IDA | 3 |
| 2017 | The Actors of History: Narrative Network Analysis Reveals the Institutions of Power in British Society Between 1800-1950
Thomas Lansdall-Welfare, Saatviga Sudhahar, Nello Cristianini |
IDA | 4 |
| 2015 | ThinkBIG - Understanding the Impact of Big Data on Science and Society
Nello Cristianini |
ICPRAM (1) | 1 |
| 2015 | Network analysis of narrative content in large corporaabstractAbstract We present a methodology for the extraction of narrative information from a large corpus. The key idea is to transform the corpus into a network, formed by linking the key actors and objects of the narration, and then to analyse this network to extract information about their relations. By representing information into a single network it is possible to infer relations between these entities, including when they have never been mentioned together. We discuss various types of information that can be extracted by our method, various ways to validate the information extracted and two different application scenarios. Our methodology is very scalable, and addresses specific research needs in social sciences. Saatviga Sudhahar, Gianluca de Fazio, Roberto Franzosi, Nello Cristianini |
Nat. Lang. Eng. | 4 |
| 2015 | Learning to classify gender from four million images
Sen Jia 0002, Nello Cristianini |
Pattern Recognit. Lett. | 2 |
| 2014 | On the coverage of science in the media: A big data study on the impact of the Fukushima disasterabstractThe contents of English-language online-news over 5 years have been analyzed to explore the impact of the Fukushima disaster on the media coverage of nuclear power. This big data study, based on millions of news articles, involves the extraction of narrative networks, association networks, and sentiment time series. The key finding is that media attitude towards nuclear power has significantly changed in the wake of the Fukushima disaster, in terms of sentiment and in terms of framing, showing a long lasting effect that does not appear to recover before the end of the period covered by this study. In particular, we find that the media discourse has shifted from one of public debate about nuclear power as a viable option for energy supply needs to a re-emergence of the public views of nuclear power and the risks associated with it. The methodology used presents an opportunity to leverage big data for corpus analysis and opens up new possibilities in social scientific research. Thomas Lansdall-Welfare, Saatviga Sudhahar, Giuseppe A. Veltri, Nello Cristianini |
IEEE BigData | 4 |
| 2014 | Efficient classification of multi-labeled text streams by clashing
Ricardo Ñanculef, Ilias N. Flaounas, Nello Cristianini |
Expert Syst. Appl. | 3 |
| 2013 | Finite-Time Analysis of Kernelised Contextual Bandits
Michal Valko, Nathaniel Korda, Rémi Munos, Ilias N. Flaounas, Nello Cristianini |
UAI | 5 |
| 2013 | Modelling and predicting news popularity
Elena Hensinger, Ilias N. Flaounas, Nello Cristianini |
Pattern Anal. Appl. | 3 |
| 2012 | ElectionWatch: Detecting Patterns in News Coverage of US Elections
Saatviga Sudhahar, Thomas Lansdall-Welfare, Ilias N. Flaounas, Nello Cristianini |
EACL | 4 |
| 2012 | Learning Machine Translation from In-domain and Out-of-domain Data
Marco Turchi, Cyril Goutte, Nello Cristianini |
EAMT | 3 |
| 2012 | An Empirical Comparison of Label Prediction Algorithms on Automatically Inferred Networks
Omar Ali, Giovanni Zappella, Tijl De Bie, Nello Cristianini |
ICPRAM (2) | 4 |
| 2012 | What Makes Us Click? - Modelling and Predicting the Appeal of News Articles
Elena Hensinger, Ilias N. Flaounas, Nello Cristianini |
ICPRAM (2) | 3 |
| 2012 | Scalable Corpus Annotation by Graph Construction and Label Propagation
Thomas Lansdall-Welfare, Ilias N. Flaounas, Nello Cristianini |
ICPRAM (1) | 3 |
| 2012 | The NetCover algorithm for the reconstruction of causal networks
Nick Fyson, Tijl De Bie, Nello Cristianini |
Neurocomputing | 3 |
| 2012 | Nowcasting Events from the Social Web with Statistical LearningabstractWe present a general methodology for inferring the occurrence and magnitude of an event or phenomenon by exploring the rich amount of unstructured textual information on the social part of the Web. Having geo-tagged user posts on the microblogging service of Twitter as our input data, we investigate two case studies. The first consists of a benchmark problem, where actual levels of rainfall in a given location and time are inferred from the content of tweets . The second one is a real-life task, where we infer regional Influenza-like Illness rates in the effort of detecting timely an emerging epidemic disease. Our analysis builds on a statistical learning framework, which performs sparse learning via the bootstrapped version of LASSO to select a consistent subset of textual features from a large amount of candidates. In both case studies, selected features indicate close semantic correlation with the target topics and inference, conducted by regression, has a significant performance, especially given the short length --approximately one year-- of Twitter’s data time series. Vasileios Lampos, Nello Cristianini |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | An intelligent Web agent that autonomously learns how to translateabstractWe describe the design of an autonomous agent that can teach itself how to translate from a foreign language, by first assembling its own training set, then using it to improve its vocabulary and language model. The key idea is that a Statistical Mac Marco Turchi, Tijl De Bie, Nello Cristianini |
Web Intell. Agent Syst. | 3 |
| 2011 | Automatic Discovery of Patterns in Media Content
Nello Cristianini |
CPM | 1 |
| 2011 | Refining causality: who copied from whom?abstractInferring causal networks behind observed data is an active area of research with wide applicability to areas such as epidemiology, microbiology and social science. In particular recent research has focused on identifying how information propagates through the Internet. This research has so far only used temporal features of observations, and while reasonable results have been achieved, there is often further information which can be used. Tristan Snowsill, Nick Fyson, Tijl De Bie, Nello Cristianini |
KDD | 4 |
| 2011 | Celebrity Watch: Browsing News Content by Exploiting Social Intelligence
Omar Ali, Ilias N. Flaounas, Tijl De Bie, Nello Cristianini |
ECML/PKDD (3) | 4 |
| 2011 | NOAM: news outlets analysis and monitoring systemabstractWe present NOAM, an integrated platform for the monitoring and analysis of news media content. NOAM is the data management system behind various applications and scientific studies aiming at modelling the mediasphere. The system is also intended to address the need in the AI community for platforms where various AI technologies are integrated and deployed in the real world. It combines a relational database (DB) with state of the art AI technologies, including data mining, machine learning and natural language processing. These technologies are organised in a robust, distributed architecture of collaborating modules, that are used to populate and annotate the DB. NOAM manages tens of millions of news items in multiple languages, automatically annotating them in order to enable queries based on their semantic properties. The system also includes a unified user interface for interacting with its various modules. Ilias N. Flaounas, Omar Ali, Marco Turchi, Tristan Snowsill, Florent Nicart, Tijl De Bie, Nello Cristianini |
SIGMOD Conference | 7 |
| 2010 | Flu Detector - Tracking Epidemics on Twitter
Vasileios Lampos, Tijl De Bie, Nello Cristianini |
ECML/PKDD (3) | 3 |
| 2010 | Detecting Events in a Million New York Times Articles
Tristan Snowsill, Ilias N. Flaounas, Tijl De Bie, Nello Cristianini |
ECML/PKDD (3) | 4 |
| 2010 | Are we there yet?
Nello Cristianini |
Neural Networks | 1 |
| 2009 | Estimating the Sentence-Level Quality of Machine Translation Systems
Lucia Specia, Marco Turchi, Nicola Cancedda, Nello Cristianini, Marc Dymetman |
EAMT | 4 |
| 2009 | Are We There Yet?
Nello Cristianini |
ECML/PKDD (1) | 1 |
| 2009 | Inference and Validation of Networks
Ilias N. Flaounas, Marco Turchi, Tijl De Bie, Nello Cristianini |
ECML/PKDD (1) | 4 |
| 2009 | Found in Translation
Marco Turchi, Ilias N. Flaounas, Omar Ali, Tijl De Bie, Tristan Snowsill, Nello Cristianini |
ECML/PKDD (2) | 6 |
| 2007 | Terminator Detection by Support Vector Machine Utilizing a Stochastic Context-Free GrammarabstractA 2-stage detector was designed to find rho-independent transcription terminators in the Escherichia coli genome. The detector includes a stochastic context free grammar (SCFG) component and a support vector machine (SVM) component. To find terminators, the SCFG searches the intergenic regions of nucleotide sequence for local matches to a terminator grammar that was designed and trained utilizing examples of known terminators. The grammar selects sequences that are the best candidates for terminators and assigns them a prefix, stem-loop, suffix structure using the Cocke-Younger-Kasaami (CYK) algorithm, modified to incorporate energy effects of base pairing. The parameters from this inferred structure are passed to the SVM classifier, which distinguishes terminators from non-terminators that score high according to the terminator grammar. The SVM was trained with negative examples drawn from intergenic sequences that include both featureless and RNA gene regions (which were assigned prefix, stem-loop, suffix structure by the SCFG), so that it successfully distinguishes terminators from either of these. The classifier was found to be 96.4% successful during testing Patricia Francis-Lyon, Nello Cristianini, Stephen R. Holbrook |
CIBCB | 2 |
| 2007 | Discriminative Sequence Labeling by Z-Score Optimization
Elisa Ricci 0001, Tijl De Bie, Nello Cristianini |
ECML | 3 |
| 2007 | Learning to Align: A Statistical Approach
Elisa Ricci 0001, Tijl De Bie, Nello Cristianini |
IDA | 3 |
| 2007 | MINI: Mining Informative Non-redundant Itemsets
Arianna Gallo, Tijl De Bie, Nello Cristianini |
PKDD | 3 |
| 2006 | CAFE: a computational tool for the study of gene family evolutionabstractSUMMARY: We present CAFE (Computational Analysis of gene Family Evolution), a tool for the statistical analysis of the evolution of the size of gene families. It uses a stochastic birth and death process to model the evolution of gene family sizes over a phylogeny. For a specified phylogenetic tree, and given the gene family sizes in the extant species, CAFE can estimate the global birth and death rate of gene families, infer the most likely gene family size at all internal nodes, identify gene families that have accelerated rates of gain and loss (quantified by a p-value) and identify which branches cause the p-value to be small for significant families. AVAILABILITY: Software is available from http://www.bio.indiana.edu/~hahnlab/Software.html Tijl De Bie, Nello Cristianini, Jeffery P. Demuth, Matthew W. Hahn |
Bioinform. | 2 |
| 2006 | Fast SDP Relaxations of Graph Cut Clustering, Transduction, and Other Combinatorial Problem
Tijl De Bie, Nello Cristianini |
J. Mach. Learn. Res. | 2 |
| 2005 | On the eigenspectrum of the gram matrix and the generalization error of kernel-PCAabstractIn this paper, the relationships between the eigenvalues of the m/spl times/m Gram matrix K for a kernel /spl kappa/(/spl middot/,/spl middot/) corresponding to a sample x/sub 1/,...,x/sub m/ drawn from a density p(x) and the eigenvalues of the corresponding continuous eigenproblem is analyzed. The differences between the two spectra are bounded and a performance bound on kernel principal component analysis (PCA) is provided showing that good performance can be expected even in very-high-dimensional feature spaces provided the sample eigenvalues fall sufficiently quickly. John Shawe-Taylor, Christopher K. I. Williams, Nello Cristianini, Jaz S. Kandola |
IEEE Trans. Inf. Theory | 3 |
| 2004 | A statistical framework for genomic data fusionabstractMOTIVATION: During the past decade, the new focus on genomics has highlighted a particular challenge: to integrate the different views of the genome that are provided by various types of experimental data. RESULTS: This paper describes a computational framework for integrating and drawing inferences from a collection of genome-wide measurements. Each dataset is represented via a kernel function, which defines generalized similarity relationships between pairs of entities, such as genes or proteins. The kernel representation is both flexible and efficient, and can be applied to many different types of data. Furthermore, kernel functions derived from different types of data can be combined in a straightforward fashion. Recent advances in the theory of kernel methods have provided efficient algorithms to perform such combinations in a way that minimizes a statistical loss function. These methods exploit semidefinite programming techniques to reduce the problem of finding optimizing kernel combinations to a convex optimization problem. Computational experiments performed using yeast genome-wide datasets, including amino acid sequences, hydropathy profiles, gene expression data and known protein-protein interactions, demonstrate the utility of this approach. A statistical learning algorithm trained from all of these data to recognize particular classes of proteins--membrane proteins and ribosomal proteins--performs significantly better than the same algorithm trained on any single type of data. AVAILABILITY: Supplementary data at http://noble.gs.washington.edu/proj/sdp-svm Gert R. G. Lanckriet, Tijl De Bie, Nello Cristianini, Michael I. Jordan, William Stafford Noble |
Bioinform. | 3 |
| 2004 | Learning the Kernel Matrix with Semidefinite Programming
Gert R. G. Lanckriet, Nello Cristianini, Peter L. Bartlett, Laurent El Ghaoui, Michael I. Jordan |
J. Mach. Learn. Res. | 2 |
| 2003 | Efficiently Learning the Metric with Side-Information
Tijl De Bie, Michinari Momma, Nello Cristianini |
ALT | 3 |
| 2003 | Kernel Methods for Pattern Analysis
Nello Cristianini |
ICTAI | 1 |
| 2003 | Convex Methods for TransductionabstractThe 2-class transduction problem, as formulated by Vapnik [1], involves finding a separating hyperplane for a labelled data set that is also maximally distant from a given set of unlabelled test points. In this form, the problem has exponential computational complexity in the size of the working set. So far it has been attacked by means of integer programming techniques [2] that do not scale to reasonable problem sizes, or by local search procedures [3]. In this paper we present a relaxation of this task based on semi- definite programming (SDP), resulting in a convex optimization problem that has polynomial complexity in the size of the data set. The results are very encouraging for mid sized data sets, however the cost is still too high for large scale problems, due to the high di- mensional search space. To this end, we restrict the feasible region by introducing an approximation based on solving an eigenproblem. With this approximation, the computational cost of the algorithm is such that problems with more than 1000 points can be treated. Tijl De Bie, Nello Cristianini |
NIPS | 2 |
| 2002 | On the Eigenspectrum of the Gram Matrix and Its Relationship to the Operator Eigenspectrum
John Shawe-Taylor, Christopher K. I. Williams, Nello Cristianini, Jaz S. Kandola |
ALT | 3 |
| 2002 | On the Eigenspectrum of the Gram Matrix and Its Relationship to the Operator Eigenspectrum
John Shawe-Taylor, Christopher K. I. Williams, Nello Cristianini, Jaz S. Kandola |
Discovery Science | 3 |
| 2002 | Learning the Kernel Matrix with Semi-Definite Programming
Gert R. G. Lanckriet, Nello Cristianini, Peter L. Bartlett, Laurent El Ghaoui, Michael I. Jordan |
ICML | 2 |
| 2002 | Learning Semantic SimilarityabstractThe standard representation of text documents as bags of words suffers from well known limitations, mostly due to its inability to exploit semantic similarity between terms. Attempts to incorpo(cid:173) rate some notion of term similarity include latent semantic index(cid:173) ing [8], the use of semantic networks [9], and probabilistic methods [5]. In this paper we propose two methods for inferring such sim(cid:173) ilarity from a corpus. The first one defines word-similarity based on document-similarity and viceversa, giving rise to a system of equations whose equilibrium point we use to obtain a semantic similarity measure. The second method models semantic relations by means of a diffusion process on a graph defined by lexicon and co-occurrence information. Both approaches produce valid kernel functions parametrised by a real number. The paper shows how the alignment measure can be used to successfully perform model selection over this parameter. Combined with the use of support vector machines we obtain positive results. Jaz S. Kandola, John Shawe-Taylor, Nello Cristianini |
NIPS | 3 |
| 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation AnalysisabstractThe problem of learning a semantic representation of a text document from data is addressed, in the situation where a corpus of unlabeled paired documents is available, each pair being formed by a short En- glish document and its French translation. This representation can then be used for any retrieval, categorization or clustering task, both in a stan- dard and in a cross-lingual setting. By using kernel functions, in this case simple bag-of-words inner products, each part of the corpus is mapped to a high-dimensional space. The correlations between the two spaces are then learnt by using kernel Canonical Correlation Analysis. A set of directions is found in the first and in the second space that are max- imally correlated. Since we assume the two representations are com- pletely independent apart from the semantic content, any correlation be- tween them should reflect some semantic similarity. Certain patterns of English words that relate to a specific meaning should correlate with cer- tain patterns of French words corresponding to the same meaning, across the corpus. Using the semantic representation obtained in this way we first demonstrate that the correlations detected between the two versions of the corpus are significantly higher than random, and hence that a rep- resentation based on such features does capture statistical patterns that should reflect semantic information. Then we use such representation both in cross-language and in single-language retrieval tasks, observing performance that is consistently and significantly superior to LSI on the same data. Alexei Vinokourov, John Shawe-Taylor, Nello Cristianini |
NIPS | 3 |
| 2002 | Latent Semantic Kernels
Nello Cristianini, John Shawe-Taylor, Huma Lodhi |
J. Intell. Inf. Syst. | 1 |
| 2002 | Text Classification using String Kernels
Huma Lodhi, Craig Saunders, John Shawe-Taylor, Nello Cristianini, Christopher J. C. H. Watkins |
J. Mach. Learn. Res. | 4 |
| 2002 | Editorial: Kernel Methods: Current Research and Future Directions
Nello Cristianini, Colin Campbell, Christopher J. C. Burges |
Mach. Learn. | 1 |
| 2002 | On the generalization of soft margin algorithmsabstractGeneralization bounds depending on the margin of a classifier are a relatively new development. They provide an explanation of the performance of state-of-the-art learning systems such as support vector machines (SVMs) and Adaboost. The difficulty with these bounds has been either their lack of robustness or their looseness. The question of whether the generalization of a classifier can be more tightly bounded in terms of a robust measure of the distribution of margin values has remained open for some time. The paper answers this open question in the affirmative and, furthermore, the analysis leads to bounds that motivate the previously heuristic soft margin SVM algorithms as well as justifying the use of the quadratic loss in neural network training algorithms. The results are extended to give bounds for the probability of failing to achieve a target accuracy in regression prediction, with a statistical analysis of ridge regression and Gaussian processes as a special case. The analysis presented in the paper has also lead to new boosting algorithms described elsewhere. John Shawe-Taylor, Nello Cristianini |
IEEE Trans. Inf. Theory | 2 |
| 2001 | Latent Semantic Kernels
Nello Cristianini, John Shawe-Taylor, Huma Lodhi |
ICML | 1 |
| 2001 | Composite Kernels for Hypertext Categorisation
Thorsten Joachims, Nello Cristianini, John Shawe-Taylor |
ICML | 2 |
| 2001 | On Kernel-Target AlignmentabstractWe introduce the notion of kernel-alignment, a measure of similar(cid:173) ity between two kernel functions or between a kernel and a target function. This quantity captures the degree of agreement between a kernel and a given learning task, and has very natural interpre(cid:173) tations in machine learning, leading also to simple algorithms for model selection and learning. We analyse its theoretical properties, proving that it is sharply concentrated around its expected value, and we discuss its relation with other standard measures of per(cid:173) formance. Finally we describe some of the algorithms that can be obtained within this framework, giving experimental results show(cid:173) ing that adapting the kernel to improve alignment on the labelled data significantly increases the alignment on the test set, giving improved classification accuracy. Hence, the approach provides a principled method of performing transduction. Keywords: Kernels, alignment, eigenvectors, eigenvalues, transduction Nello Cristianini, John Shawe-Taylor, André Elisseeff, Jaz S. Kandola |
NIPS | 1 |
| 2001 | Spectral Kernel Methods for ClusteringabstractIn this paper we introduce new algorithms for unsupervised learn(cid:173) ing based on the use of a kernel matrix. All the information re(cid:173) quired by such algorithms is contained in the eigenvectors of the matrix or of closely related matrices. We use two different but re(cid:173) lated cost functions, the Alignment and the 'cut cost'. The first one is discussed in a companion paper [3], the second one is based on graph theoretic concepts. Both functions measure the level of clustering of a labeled dataset, or the correlation between data clus(cid:173) ters and labels. We state the problem of unsupervised learning as assigning labels so as to optimize these cost functions. We show how the optimal solution can be approximated by slightly relaxing the corresponding optimization problem, and how this corresponds to using eigenvector information. The resulting simple algorithms are tested on real world data with positive results. Nello Cristianini, John Shawe-Taylor, Jaz S. Kandola |
NIPS | 1 |
| 2001 | On the Concentration of Spectral PropertiesabstractWe consider the problem of measuring the eigenvalues of a ran(cid:173) domly drawn sample of points. We show that these values can be reliably estimated as can the sum of the tail of eigenvalues. Fur(cid:173) thermore, the residuals when data is projected into a subspace is shown to be reliably estimated on a random sample. Experiments are presented that confirm the theoretical results. John Shawe-Taylor, Nello Cristianini, Jaz S. Kandola |
NIPS | 2 |
| 2000 | Query Learning with Large Margin Classifiers
Colin Campbell, Nello Cristianini, Alexander J. Smola |
ICML | 2 |
| 2000 | Large margin strategies in machine learningabstractControlling the capacity of a learning system in a way that does not depend on the dimensionality of the hypothesis space provides the key for effectively using large neural networks and decision trees, ensemble methods and kernel-induced feature spaces. This extended abstract will provide an overview of recent work in this direction, based on the concepts of margin and margin distribution. Nello Cristianini |
ISCAS | 1 |
| 2000 | Text Classification using String KernelsabstractWe introduce a novel kernel for comparing two text documents. The kernel is an inner product in the feature space consisting of all subsequences of length k. A subsequence is any ordered se(cid:173) quence of k characters occurring in the text though not necessarily contiguously. The subsequences are weighted by an exponentially decaying factor of their full length in the text, hence emphasising those occurrences which are close to contiguous. A direct compu(cid:173) tation of this feature vector would involve a prohibitive amount of computation even for modest values of k, since the dimension of the feature space grows exponentially with k. The paper describes how despite this fact the inner product can be efficiently evaluated by a dynamic programming technique. A preliminary experimental comparison of the performance of the kernel compared with a stan(cid:173) dard word feature space kernel results. [6] is made showing encouraging Huma Lodhi, John Shawe-Taylor, Nello Cristianini, Christopher J. C. H. Watkins |
NIPS | 3 |
| 2000 | Support vector machine classification and validation of cancer tissue samples using microarray expression dataabstractMOTIVATION: DNA microarray experiments generating thousands of gene expression measurements, are being used to gather information from tissue and cell samples regarding gene expression differences that will be useful in diagnosing disease. We have developed a new method to analyse this kind of data using support vector machines (SVMs). This analysis consists of both classification of the tissue samples, and an exploration of the data for mis-labeled or questionable tissue results. RESULTS: We demonstrate the method in detail on samples consisting of ovarian cancer tissues, normal ovarian tissues, and other normal tissues. The dataset consists of expression experiment results for 97,802 cDNAs for each tissue. As a result of computational analysis, a tissue sample is discovered and confirmed to be wrongly labeled. Upon correction of this mistake and the removal of an outlier, perfect classification of tissues is achieved, but not with high confidence. We identify and analyse a subset of genes from the ovarian dataset whose expression is highly differentiated between the types of tissues. To show robustness of the SVM method, two previously published datasets from other types of tissues or cells are analysed. The results are comparable to those previously obtained. We show that other machine learning methods also perform comparably to the SVM on many of those datasets. AVAILABILITY: The SVM software is available at http://www.cs. columbia.edu/ approximately bgrundy/svm. Terrence S. Furey, Nello Cristianini, Nigel Duffy, David W. Bednarski, Michèl Schummer, David Haussler |
Bioinform. | 2 |
| 2000 | Enlarging the Margins in Perceptron Decision Trees
Kristin P. Bennett, Nello Cristianini, John Shawe-Taylor, Donghui Wu |
Mach. Learn. | 2 |
| 1999 | Further Results on the Margin DistributionabstractA number of results have bounded generalization error of a classifier in terms of its margin on the training points.There has been some debate about whether the minimum margin is the best measure of the distribution of training set margin values with which to estimate the generalization error.Freund and Schapire [7] have shown how a different function of the margin distribution can be used to bound the number of mistakes of an on-line learning algorithm for a perceptron, as well as an expected error bound.Shawe-Taylor and Cristianini [ 131 showed that a slight generalization of their construction can be used to give a pat style bound on the tail of the distribution of the generalization errors that arise from a given sample size when using threshold linear classifiers.We show that in the linear case the approach can be viewed as a change of kernel and that the algorithms arising from the approach are exactly those originally proposed by Cortes and Vapnik [4].We generalise the basic result to function classes with bounded fat-shattering dimension and the Ii measure for slack variables which gives rise to Vapnik's box constraint algorithm.Finally, application to regression is considered, which includes standard least squares as a special case.Permission to make digital or hard John Shawe-Taylor, Nello Cristianini |
COLT | 2 |
| 1999 | A multiplicative updating algorithm for training support vector machine
Nello Cristianini, Colin Campbell, John Shawe-Taylor |
ESANN | 1 |
| 1999 | Large Margin Trees for Induction and Transduction
Donghui Wu, Kristin P. Bennett, Nello Cristianini, John Shawe-Taylor |
ICML | 3 |
| 1999 | Large Margin DAGs for Multiclass Classification
John C. Platt, Nello Cristianini, John Shawe-Taylor |
NIPS | 2 |
| 1998 | Bayesian Classifiers Are Large Margin Hyperplanes in a Hilbert Space
Nello Cristianini, John Shawe-Taylor, Peter Sykacek |
ICML | 1 |
| 1998 | The Kernel-Adatron Algorithm: A Fast and Simple Learning Procedure for Support Vector Machines
Thilo-Thomas Frieß, Nello Cristianini, Colin Campbell |
ICML | 2 |
| 1998 | Dynamically Adapting Kernels in Support Vector Machines
Nello Cristianini, Colin Campbell, John Shawe-Taylor |
NIPS | 1 |
| 1997 | Data-Dependent Structural Risk Minimization for Perceptron Decision Trees
John Shawe-Taylor, Nello Cristianini |
NIPS | 2 |