VLDB 2026 Research / reviewers in the wild / expert
Alexei Vinokourov
dblp:49/3456
· DBLP profile ↗
5ranked-venue papers
4as first author
0since 2021 · last 2005
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
2 papers |
Kernel, tree and ensemble methods · 87% Machine translation · 13% | |
| Theoretical computer science
1 paper |
Automata and formal languages · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods › kernel function
fisher kernel |
0.0 | 1 | 2002 | String Kernels, Fisher Kernels and Finite State Automata · NIPS 2002 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › structured kernel
string kernel |
0.0 | 1 | 2002 | String Kernels, Fisher Kernels and Finite State Automata · NIPS 2002 |
Information retrieval
cross-language information retrieval |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Information retrieval › retrieval models › latent semantic models
latent semantic indexing |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Information retrieval
retrieval models |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Information retrieval › retrieval models
semantic representation |
0.0 | 1 | 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation Analysis · NIPS 2002 |
Automata and formal languages
finite automata |
0.0 | 1 | 2002 | String Kernels, Fisher Kernels and Finite State Automata · NIPS 2002 |
Methods — techniques the papers use, named apart from their topics
markov process · 0.1kernel methods · 0.1kernel canonical correlation analysis · 0.1bag-of-words · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2005 | A probabilistic framework for mismatch and profile string kernels
Alexei Vinokourov, Andrei N. Soklakov, Craig Saunders |
ESANN | 1 |
| 2002 | String Kernels, Fisher Kernels and Finite State AutomataabstractIn this paper we show how the generation of documents can be thought of as a k-stage Markov process, which leads to a Fisher ker(cid:173) nel from which the n-gram and string kernels can be re-constructed. The Fisher kernel view gives a more flexible insight into the string kernel and suggests how it can be parametrised in a way that re(cid:173) flects the statistics of the training corpus. Furthermore, the prob(cid:173) abilistic modelling approach suggests extending the Markov pro(cid:173) cess to consider sub-sequences of varying length, rather than the standard fixed-length approach used in the string kernel. We give a procedure for determining which sub-sequences are informative features and hence generate a Finite State Machine model, which can again be used to obtain a Fisher kernel. By adjusting the parametrisation we can also influence the weighting received by the features . In this way we are able to obtain a logarithmic weighting in a Fisher kernel. Finally, experiments are reported comparing the different kernels using the standard Bag of Words kernel as a baseline. Craig Saunders, John Shawe-Taylor, Alexei Vinokourov |
NIPS | 3 |
| 2002 | Inferring a Semantic Representation of Text via Cross-Language Correlation AnalysisabstractThe problem of learning a semantic representation of a text document from data is addressed, in the situation where a corpus of unlabeled paired documents is available, each pair being formed by a short En- glish document and its French translation. This representation can then be used for any retrieval, categorization or clustering task, both in a stan- dard and in a cross-lingual setting. By using kernel functions, in this case simple bag-of-words inner products, each part of the corpus is mapped to a high-dimensional space. The correlations between the two spaces are then learnt by using kernel Canonical Correlation Analysis. A set of directions is found in the first and in the second space that are max- imally correlated. Since we assume the two representations are com- pletely independent apart from the semantic content, any correlation be- tween them should reflect some semantic similarity. Certain patterns of English words that relate to a specific meaning should correlate with cer- tain patterns of French words corresponding to the same meaning, across the corpus. Using the semantic representation obtained in this way we first demonstrate that the correlations detected between the two versions of the corpus are significantly higher than random, and hence that a rep- resentation based on such features does capture statistical patterns that should reflect semantic information. Then we use such representation both in cross-language and in single-language retrieval tasks, observing performance that is consistently and significantly superior to LSI on the same data. Alexei Vinokourov, John Shawe-Taylor, Nello Cristianini |
NIPS | 1 |
| 2002 | A Probabilistic Framework for the Hierarchic Organisation and Classification of Document Collections
Alexei Vinokourov, Mark A. Girolami |
J. Intell. Inf. Syst. | 1 |
| 2000 | Probabilistic Hierarchical Clustering Method for Organizing Collections of Text DocumentsabstractA generic probabilistic framework for the unsupervised hierarchical clustering of large-scale sparse high-dimensional data collections is proposed. The framework is based on a hierarchical probabilistic mixture methodology. Two classes of models emerge from the analysis and these have been called symmetric and asymmetric models. For text data specifically both asymmetric and symmetric models based on multinomial and binomial distributions are most appropriate. An expectation maximisation parameter estimation method is provided for all of these models. An experimental comparison of the models is obtained for two extensive online document collections. Alexei Vinokourov, Mark A. Girolami |
ICPR | 1 |