Yury Ustinovskiy

dblp:38/10440 · also Yury Mikhailovich Ustinovskiy, Yury Ustinovsky · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
0since 2021 · last 2018
0000-0002-2510-4098ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 4 first-authorArtificial intelligence and machine learning · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 42% Kernel, tree and ensemble methods · 42% Optimization for machine learning · 16%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 63% Data integration and cleaning · 16% Data mining · 16%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking
learning to rank
0.522016
An Optimization Framework for Remapping and Reweighting Noisy Relevance Labels · SIGIR 2016
Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosted decision trees
0.312018
Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018
Machine learning › Trustworthy machine learning
interpretability
0.312018
Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018
Machine learning › Trustworthy machine learning › interpretability
training data attribution
0.312018
Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018
Machine learning › Kernel, tree and ensemble methods › ensemble learning
tree ensembles
0.312018
Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018
Machine learning › Optimization for machine learning
hyperparameter optimization
0.212016
Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016
Data mining
crowdsourcing
0.212016
An Optimization Framework for Remapping and Reweighting Noisy Relevance Labels · SIGIR 2016
Data integration and cleaning › data quality
label noise
0.212016
Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016
Information retrieval
retrieval models
0.212016
Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016
Information retrieval
personalized search
0.212015
An Optimization Framework for Weighting Implicit Relevance Labels for Personalized Web Search · WWW 2015
Recommender systems
personalized ranking
0.112015
An Optimization Framework for Weighting Implicit Relevance Labels for Personalized Web Search · WWW 2015

Methods — techniques the papers use, named apart from their topics

gradient-based hyperparameter optimization · 0.5decision tree learning · 0.5leave-one-out retraining · 0.3influence functions · 0.3sample reweighting · 0.2majority voting · 0.2label remapping · 0.2optimization framework · 0.2click-based feedback · 0.2
YearPublicationVenuePosition
2018 Finding Influential Training Samples for Gradient Boosted Decision Trees
abstract
We address the problem of finding influential training samples for a particular case of tree ensemble-based models, e.g., Random Forest (RF) or Gradient Boosted Decision Trees (GBDT). A natural way of formalizing this problem is studying how the model’s predictions change upon leave-one-out retraining, leaving out each individual training sample. Recent work has shown that, for parametric models, this analysis can be conducted in a computationally efficient way. We propose several ways of extending this framework to non-parametric GBDT ensembles under the assumption that tree structures remain fixed. Furthermore, we introduce a general scheme of obtaining further approximations to our method that balance the trade-off between performance and computational complexity. We evaluate our approaches on various experimental setups and use-case scenarios and demonstrate both the quality of our approach to finding influential training samples in comparison to the baselines and its computational efficiency.
Boris Sharchilev, Yury Ustinovskiy, Pavel Serdyukov, Maarten de Rijke
ICML2
2018 On Face Numbers of Flag Simplicial Complexes
Yury Ustinovskiy
Discret. Comput. Geom.1
2016 Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning
abstract
Labeled training data is an essential part of any supervised machine learning framework. In practice, there is a trade-off between the quality of a label and its cost. In this paper, we consider a problem of learning to rank on a large-scale dataset with low-quality relevance labels aiming at maximizing the quality of a trained ranker on a small validation dataset with high-quality ground truth relevance labels. Motivated by the classical Gauss-Markov theorem for the linear regression problem, we formulate the problems of (1) reweighting training instances and (2) remapping learning targets. We propose meta–gradient decision tree learning framework for optimizing weight and target functions by applying gradient-based hyperparameter optimization. Experiments on a large-scale real-world dataset demonstrate that we can significantly improve state-of-the-art machine-learning algorithms by incorporating our framework.
Yury Ustinovskiy, Valentina Fedorova, Gleb Gusev, Pavel Serdyukov
ICML1
2016 An Optimization Framework for Remapping and Reweighting Noisy Relevance Labels
abstract
Relevance labels is the essential part of any learning to rank framework. The rapid development of crowdsourcing platforms led to a significant reduction of the cost of manual labeling. This makes it possible to collect very large sets of labeled documents to train a ranking algorithm. However, relevance labels acquired via crowdsourcing are typically coarse and noisy, so certain consensus models are used to measure the quality of labels and to reduce the noise. This noise is likely to affect a ranker trained on such labels, and, since none of the existing consensus models directly optimizes ranking quality, one has to apply some heuristics to utilize the output of a consensus model in a ranking algorithm, e.g., to use majority voting among workers to get consensus labels. The major goal of this paper is to unify existing approaches to consensus modeling and noise reduction within a learning to rank framework. Namely, we present a machine learning algorithm aimed at improving the performance of a ranker trained on a crowdsourced dataset by proper remapping of labels and reweighting of samples. In the experimental part, we use several characteristics of workers/labels extracted via various consensus models in order to learn the remapping and reweighting functions. Our experiments on a large-scale dataset demonstrate that we can significantly improve state-of-the-art machine-learning algorithms by incorporating our framework.
Yury Ustinovskiy, Valentina Fedorova, Gleb Gusev, Pavel Serdyukov
SIGIR1
2015 Adaptive Caching of Fresh Web Search Results
Liudmila Ostroumova, Yury Ustinovskiy, Egor Samosvat, Damien Lefortier, Pavel Serdyukov
ECIR2
2015 An Optimization Framework for Weighting Implicit Relevance Labels for Personalized Web Search
abstract
Implicit feedback from users of a web search engine is an essential source providing consistent personal relevance labels from the actual population of users. However, previous studies on personalized search employ this source in a rather straightforward manner. Basically, documents that were clicked on get maximal gain, and the rest of the documents are assigned the zero gain. As we demonstrate in our paper, a ranking algorithm trained using these gains directly as the ground truth relevance labels leads to a suboptimal personalized ranking.
Yury Ustinovskiy, Gleb Gusev, Pavel Serdyukov
WWW1
2013 Predicting the impact of expansion terms using semantic and user interaction features
abstract
Query expansion for Information Retrieval is a challenging task. On the one hand, low quality expansion may hurt either recall, due to vocabulary mismatch, or precision, due to topic drift, and therefore reduce user satisfaction. On the other hand, utilizing a large number of expansion terms for a query may easily lead to resource consumption overhead. As web search engines apply strict constraints on response time, it is essential to estimate the impact of each expansion term on query performance at the pre-retrieval time. Our experimental results confirm that a significant part of expansions do not improve query performance, and it is possible to detect such expansions at the pre-retrieval time.
Anton Bakhtin, Yury Ustinovskiy, Pavel Serdyukov
CIKM2
2013 Personalization of web-search using short-term browsing context
abstract
Search and browsing activity is known to be a valuable source of information about user's search intent. It is extensively utilized by most of modern search engines to improve ranking by constructing certain ranking features as well as by personalizing search. Personalization aims at two major goals: extraction of stable preferences of a user and specification and disambiguation of the current query. The common way to approach these problems is to extract information from user's search and browsing long-term history and to utilize short-term history to determine the context of a given query. Personalization of the web search for the first queries in new search sessions of new users is more difficult due to the lack of both long- and short-term data.
Yury Ustinovskiy, Pavel Serdyukov
CIKM1
2013 Intent-Based Browse Activity Segmentation
Yury Ustinovskiy, Anna Mazur, Pavel Serdyukov
ECIR1
2012 Session-based query performance prediction
abstract
Search sessions are known to be a rich source of diverse valuable information for individual query analysis. In this paper, we address the problem of query performance prediction by utilizing the entire logical search sessions containing the given query. Guided by the intuitions based on the observations made after the analysis of the search sessions' properties and performance of the queries they contain, we propose a number of features that significantly advance the existing query performance prediction models. Some of them specifically allow to focus on tail queries with sparse click-through statistics.
Andrey Kustarev, Yury Ustinovskiy, Anna Mazur, Pavel Serdyukov
CIKM2
2011 Smoothing NDCG metrics using tied scores
abstract
One of promising directions in research on learning to rank concerns the problem of appropriate choice of the objective function to maximize by means of machine learning algorithms. We describe a novel technique of smoothing an arbitrary ranking metric and demonstrate how to utilize it to maximize the retrieval quality in terms of the $NDCG$ metric. The idea behind our listwise ranking model called TieRank is artificial probabilistic tying of predicted relevance scores at each iteration of learning process, which defines a distribution on the set of all permutations of retrieved documents. Such distribution provides a desired smoothed version of the target retrieval quality metric. This smooth function is possible to maximize using a gradient descent method. Experiments on LETOR collections show that TieRank outperforms most of the existing learning to rank algorithms.
Andrey Kustarev, Yury Ustinovskiy, Yury Logachev, Evgeny Grechnikov, Ilya Segalovich, Pavel Serdyukov
CIKM2