EDBT 2026 Demo / reviewers in the wild / expert
Yury Ustinovskiy
dblp:38/10440 · also Yury Mikhailovich Ustinovskiy, Yury Ustinovsky
· DBLP profile ↗
11ranked-venue papers
6as first author
0since 2021 · last 2018
0000-0002-2510-4098ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-authorArtificial intelligence and machine learning · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 42% Kernel, tree and ensemble methods · 42% Optimization for machine learning · 16% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 63% Data integration and cleaning · 16% Data mining · 16% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › ranking
learning to rank |
0.5 | 2 | 2016 | An Optimization Framework for Remapping and Reweighting Noisy Relevance Labels · SIGIR 2016 Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016 |
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosted decision trees |
0.3 | 1 | 2018 | Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2018 | Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018 |
Machine learning › Trustworthy machine learning › interpretability
training data attribution |
0.3 | 1 | 2018 | Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
tree ensembles |
0.3 | 1 | 2018 | Finding Influential Training Samples for Gradient Boosted Decision Trees · ICML 2018 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.2 | 1 | 2016 | Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016 |
Data mining
crowdsourcing |
0.2 | 1 | 2016 | An Optimization Framework for Remapping and Reweighting Noisy Relevance Labels · SIGIR 2016 |
Data integration and cleaning › data quality
label noise |
0.2 | 1 | 2016 | Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016 |
Information retrieval
retrieval models |
0.2 | 1 | 2016 | Meta-Gradient Boosted Decision Tree Model for Weight and Target Learning · ICML 2016 |
Information retrieval
personalized search |
0.2 | 1 | 2015 | An Optimization Framework for Weighting Implicit Relevance Labels for Personalized Web Search · WWW 2015 |
Recommender systems
personalized ranking |
0.1 | 1 | 2015 | An Optimization Framework for Weighting Implicit Relevance Labels for Personalized Web Search · WWW 2015 |
Methods — techniques the papers use, named apart from their topics
gradient-based hyperparameter optimization · 0.5decision tree learning · 0.5leave-one-out retraining · 0.3influence functions · 0.3sample reweighting · 0.2majority voting · 0.2label remapping · 0.2optimization framework · 0.2click-based feedback · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Finding Influential Training Samples for Gradient Boosted Decision TreesabstractWe address the problem of finding influential training samples for a particular case of tree ensemble-based models, e.g., Random Forest (RF) or Gradient Boosted Decision Trees (GBDT). A natural way of formalizing this problem is studying how the model’s predictions change upon leave-one-out retraining, leaving out each individual training sample. Recent work has shown that, for parametric models, this analysis can be conducted in a computationally efficient way. We propose several ways of extending this framework to non-parametric GBDT ensembles under the assumption that tree structures remain fixed. Furthermore, we introduce a general scheme of obtaining further approximations to our method that balance the trade-off between performance and computational complexity. We evaluate our approaches on various experimental setups and use-case scenarios and demonstrate both the quality of our approach to finding influential training samples in comparison to the baselines and its computational efficiency. Boris Sharchilev, Yury Ustinovskiy, Pavel Serdyukov, Maarten de Rijke |
ICML | 2 |
| 2018 | On Face Numbers of Flag Simplicial Complexes
Yury Ustinovskiy |
Discret. Comput. Geom. | 1 |
| 2016 | Meta-Gradient Boosted Decision Tree Model for Weight and Target LearningabstractLabeled training data is an essential part of any supervised machine learning framework. In practice, there is a trade-off between the quality of a label and its cost. In this paper, we consider a problem of learning to rank on a large-scale dataset with low-quality relevance labels aiming at maximizing the quality of a trained ranker on a small validation dataset with high-quality ground truth relevance labels. Motivated by the classical Gauss-Markov theorem for the linear regression problem, we formulate the problems of (1) reweighting training instances and (2) remapping learning targets. We propose meta–gradient decision tree learning framework for optimizing weight and target functions by applying gradient-based hyperparameter optimization. Experiments on a large-scale real-world dataset demonstrate that we can significantly improve state-of-the-art machine-learning algorithms by incorporating our framework. Yury Ustinovskiy, Valentina Fedorova, Gleb Gusev, Pavel Serdyukov |
ICML | 1 |
| 2016 | An Optimization Framework for Remapping and Reweighting Noisy Relevance LabelsabstractRelevance labels is the essential part of any learning to rank framework. The rapid development of crowdsourcing platforms led to a significant reduction of the cost of manual labeling. This makes it possible to collect very large sets of labeled documents to train a ranking algorithm. However, relevance labels acquired via crowdsourcing are typically coarse and noisy, so certain consensus models are used to measure the quality of labels and to reduce the noise. This noise is likely to affect a ranker trained on such labels, and, since none of the existing consensus models directly optimizes ranking quality, one has to apply some heuristics to utilize the output of a consensus model in a ranking algorithm, e.g., to use majority voting among workers to get consensus labels. The major goal of this paper is to unify existing approaches to consensus modeling and noise reduction within a learning to rank framework. Namely, we present a machine learning algorithm aimed at improving the performance of a ranker trained on a crowdsourced dataset by proper remapping of labels and reweighting of samples. In the experimental part, we use several characteristics of workers/labels extracted via various consensus models in order to learn the remapping and reweighting functions. Our experiments on a large-scale dataset demonstrate that we can significantly improve state-of-the-art machine-learning algorithms by incorporating our framework. Yury Ustinovskiy, Valentina Fedorova, Gleb Gusev, Pavel Serdyukov |
SIGIR | 1 |
| 2015 | Adaptive Caching of Fresh Web Search Results
Liudmila Ostroumova, Yury Ustinovskiy, Egor Samosvat, Damien Lefortier, Pavel Serdyukov |
ECIR | 2 |
| 2015 | An Optimization Framework for Weighting Implicit Relevance Labels for Personalized Web SearchabstractImplicit feedback from users of a web search engine is an essential source providing consistent personal relevance labels from the actual population of users. However, previous studies on personalized search employ this source in a rather straightforward manner. Basically, documents that were clicked on get maximal gain, and the rest of the documents are assigned the zero gain. As we demonstrate in our paper, a ranking algorithm trained using these gains directly as the ground truth relevance labels leads to a suboptimal personalized ranking. Yury Ustinovskiy, Gleb Gusev, Pavel Serdyukov |
WWW | 1 |
| 2013 | Predicting the impact of expansion terms using semantic and user interaction featuresabstractQuery expansion for Information Retrieval is a challenging task. On the one hand, low quality expansion may hurt either recall, due to vocabulary mismatch, or precision, due to topic drift, and therefore reduce user satisfaction. On the other hand, utilizing a large number of expansion terms for a query may easily lead to resource consumption overhead. As web search engines apply strict constraints on response time, it is essential to estimate the impact of each expansion term on query performance at the pre-retrieval time. Our experimental results confirm that a significant part of expansions do not improve query performance, and it is possible to detect such expansions at the pre-retrieval time. Anton Bakhtin, Yury Ustinovskiy, Pavel Serdyukov |
CIKM | 2 |
| 2013 | Personalization of web-search using short-term browsing contextabstractSearch and browsing activity is known to be a valuable source of information about user's search intent. It is extensively utilized by most of modern search engines to improve ranking by constructing certain ranking features as well as by personalizing search. Personalization aims at two major goals: extraction of stable preferences of a user and specification and disambiguation of the current query. The common way to approach these problems is to extract information from user's search and browsing long-term history and to utilize short-term history to determine the context of a given query. Personalization of the web search for the first queries in new search sessions of new users is more difficult due to the lack of both long- and short-term data. Yury Ustinovskiy, Pavel Serdyukov |
CIKM | 1 |
| 2013 | Intent-Based Browse Activity Segmentation
Yury Ustinovskiy, Anna Mazur, Pavel Serdyukov |
ECIR | 1 |
| 2012 | Session-based query performance predictionabstractSearch sessions are known to be a rich source of diverse valuable information for individual query analysis. In this paper, we address the problem of query performance prediction by utilizing the entire logical search sessions containing the given query. Guided by the intuitions based on the observations made after the analysis of the search sessions' properties and performance of the queries they contain, we propose a number of features that significantly advance the existing query performance prediction models. Some of them specifically allow to focus on tail queries with sparse click-through statistics. Andrey Kustarev, Yury Ustinovskiy, Anna Mazur, Pavel Serdyukov |
CIKM | 2 |
| 2011 | Smoothing NDCG metrics using tied scoresabstractOne of promising directions in research on learning to rank concerns the problem of appropriate choice of the objective function to maximize by means of machine learning algorithms. We describe a novel technique of smoothing an arbitrary ranking metric and demonstrate how to utilize it to maximize the retrieval quality in terms of the $NDCG$ metric. The idea behind our listwise ranking model called TieRank is artificial probabilistic tying of predicted relevance scores at each iteration of learning process, which defines a distribution on the set of all permutations of retrieved documents. Such distribution provides a desired smoothed version of the target retrieval quality metric. This smooth function is possible to maximize using a gradient descent method. Experiments on LETOR collections show that TieRank outperforms most of the existing learning to rank algorithms. Andrey Kustarev, Yury Ustinovskiy, Yury Logachev, Evgeny Grechnikov, Ilya Segalovich, Pavel Serdyukov |
CIKM | 2 |