Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Alexei Vinokourov

dblp:49/3456 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2005
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
2 papers
Kernel, tree and ensemble methods · 87% Machine translation · 13%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods › kernel function
fisher kernel
0.012002
String Kernels, Fisher Kernels and Finite State Automata · NIPS 2002
Machine learning › Kernel, tree and ensemble methods › kernel methods › structured kernel
string kernel
0.012002
String Kernels, Fisher Kernels and Finite State Automata · NIPS 2002
Information retrieval
cross-language information retrieval
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Information retrieval › retrieval models › latent semantic models
latent semantic indexing
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Information retrieval
retrieval models
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Information retrieval › retrieval models
semantic representation
0.012002
Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002
Automata and formal languages
finite automata
0.012002
String Kernels, Fisher Kernels and Finite State Automata · NIPS 2002

Methods — techniques the papers use, named apart from their topics

markov process · 0.1kernel methods · 0.1kernel canonical correlation analysis · 0.1bag-of-words · 0.1
YearPublicationVenuePosition
2005 A probabilistic framework for mismatch and profile string kernels
Alexei Vinokourov, Andrei N. Soklakov, Craig Saunders
ESANN1
2002 String Kernels, Fisher Kernels and Finite State Automata
abstract
In this paper we show how the generation of documents can be thought of as a k-stage Markov process, which leads to a Fisher ker(cid:173) nel from which the n-gram and string kernels can be re-constructed. The Fisher kernel view gives a more flexible insight into the string kernel and suggests how it can be parametrised in a way that re(cid:173) flects the statistics of the training corpus. Furthermore, the prob(cid:173) abilistic modelling approach suggests extending the Markov pro(cid:173) cess to consider sub-sequences of varying length, rather than the standard fixed-length approach used in the string kernel. We give a procedure for determining which sub-sequences are informative features and hence generate a Finite State Machine model, which can again be used to obtain a Fisher kernel. By adjusting the parametrisation we can also influence the weighting received by the features . In this way we are able to obtain a logarithmic weighting in a Fisher kernel. Finally, experiments are reported comparing the different kernels using the standard Bag of Words kernel as a baseline.
Craig Saunders, John Shawe-Taylor, Alexei Vinokourov
NIPS3
2002 Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis
abstract
The problem of learning a semantic representation of a text document from data is addressed, in the situation where a corpus of unlabeled paired documents is available, each pair being formed by a short En- glish document and its French translation. This representation can then be used for any retrieval, categorization or clustering task, both in a stan- dard and in a cross-lingual setting. By using kernel functions, in this case simple bag-of-words inner products, each part of the corpus is mapped to a high-dimensional space. The correlations between the two spaces are then learnt by using kernel Canonical Correlation Analysis. A set of directions is found in the first and in the second space that are max- imally correlated. Since we assume the two representations are com- pletely independent apart from the semantic content, any correlation be- tween them should reflect some semantic similarity. Certain patterns of English words that relate to a specific meaning should correlate with cer- tain patterns of French words corresponding to the same meaning, across the corpus. Using the semantic representation obtained in this way we first demonstrate that the correlations detected between the two versions of the corpus are significantly higher than random, and hence that a rep- resentation based on such features does capture statistical patterns that should reflect semantic information. Then we use such representation both in cross-language and in single-language retrieval tasks, observing performance that is consistently and significantly superior to LSI on the same data.
Alexei Vinokourov, John Shawe-Taylor, Nello Cristianini
NIPS1
2002 A Probabilistic Framework for the Hierarchic Organisation and Classification of Document Collections
Alexei Vinokourov, Mark A. Girolami
J. Intell. Inf. Syst.1
2000 Probabilistic Hierarchical Clustering Method for Organizing Collections of Text Documents
abstract
A generic probabilistic framework for the unsupervised hierarchical clustering of large-scale sparse high-dimensional data collections is proposed. The framework is based on a hierarchical probabilistic mixture methodology. Two classes of models emerge from the analysis and these have been called symmetric and asymmetric models. For text data specifically both asymmetric and symmetric models based on multinomial and binomial distributions are most appropriate. An expectation maximisation parameter estimation method is provided for all of these models. An experimental comparison of the models is obtained for two extensive online document collections.
Alexei Vinokourov, Mark A. Girolami
ICPR1