VLDB 2026 Research / reviewers in the wild / expert
Dmitry Ustalov
dblp:148/4508
· DBLP profile ↗
12ranked-venue papers in the field
5as first author
8since 2021 · last 2024
0000-0002-9979-2188ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 6 (3 first)Data Mining & Knowledge Discovery · 3 (2 first)Other / Interdisciplinary · 2Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Overview of PAN 2024: Multi-author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification - Extended Abstract
Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashraf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korencic, Maximilian Mayerl, Animesh Mukherjee 0001, Alexander Panchenko, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Alisa Smirnova, Efstathios Stamatatos, Benno Stein 0001, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, Eva Zangerle |
ECIR (6) | 19 |
| 2023 | Clustering Without Knowing How To: Application and Evaluation
Daniil Likhobaba, Daniil Fedulov, Dmitry Ustalov |
ECIR (3) | 3 |
| 2023 | Crowdsourcing for Information Retrieval
Dmitry Ustalov, Alisa Smirnova, Natalia Fedorova, Nikita Pavlichenko |
ECIR (3) | 1 |
| 2023 | Best Prompts for Text-to-Image Models and How to Find ThemabstractAdvancements in text-guided diffusion models have allowed for the creation of visually appealing images similar to those created by professional artists. The effectiveness of these models depends on the composition of the textual description, known as the prompt, and its accompanying keywords. Evaluating aesthetics computationally is difficult, so human input is necessary to determine the ideal prompt formulation and keyword combination. In this study, we propose a human-in-the-loop method for discovering the most effective combination of prompt keywords using a genetic algorithm. Our approach demonstrates how this can lead to an improvement in the visual appeal of images generated from the same description. Nikita Pavlichenko, Dmitry Ustalov |
SIGIR | 2 |
| 2023 | 4th Crowd Science Workshop - CANDLE: Collaboration of Humans and Learning Algorithms for Data LabelingabstractCrowdsourcing has been used to produce impactful and large-scale datasets for Machine Learning and Artificial Intelligence (AI), such as ImageNET, SuperGLUE, etc. Since the rise of crowdsourcing in early 2000s, the AI community has been studying its computational, system design, and data-centric aspects at various angles. We welcome the studies on developing and enhancing of crowdworker-centric tools, that offer task matching, requester assessment, instruction validation, among other topics. We are also interested in exploring methods that leverage the integration of crowdworkers to improve the recognition and performance of the machine learning models. Thus, we invite studies that focus on shipping active learning techniques, methods for joint learning from noisy data and from crowds, novel approaches for crowd-computer interaction, repetitive task automation, and role separation between humans and machines. Moreover, we invite works on designing and applying such techniques in various domains, including e-commerce and medicine. Dmitry Ustalov, Saiph Savage, Niels van Berkel, Yang Liu 0018 |
WSDM | 1 |
| 2022 | Web Engineering with Human-in-the-Loop
Dmitry Ustalov, Nikita Pavlichenko, Boris Tseytlin, Daria Baidakova, Alexey Drutsa |
ICWE | 1 |
| 2022 | Improving Recommender Systems with Human-in-the-LoopabstractToday, most recommender systems employ Machine Learning to recommend posts, products, and other items, usually produced by the users. Although the impressive progress in Deep Learning and Reinforcement Learning, we observe that recommendations made by such systems still do not correlate with actual human preferences. In our tutorial, we will bridge the gap between crowdsourcing and recommender systems communities by showing how one can incorporate human-in-the-loop into their recommender system to gather the real human feedback on the ranked recommendations. We will discuss the ranking data lifecycle and run through it step-by-step. A significant portion of tutorial time is devoted to a hands-on practice, when the attendees will, under our guidance, sample recommendations and build the ground truth dataset using crowdsourced data, and compute the offline evaluation scores. Dmitry Ustalov, Natalia Fedorova, Nikita Pavlichenko |
RecSys | 1 |
| 2022 | Challenges in Data Production for AI with Human-in-the-LoopabstractToday, successful Artificial Intelligence applications rely on three pillars: machine learning algorithms, hardware for running them, and data for training and evaluating models. Although algorithms and hardware have already become commodities, obtaining up-to-date and high-quality data at scale is still challenging-but possible by building hybrid human-computer pipelines called human-in-the-loop. This talk will show how to make a significant business impact using human-in-the-loop pipelines that combine machine learning with crowdsourcing. We will share the experience of one of the world's largest search engines, Yandex. Dmitry Ustalov |
WSDM | 1 |
| 2020 | Crowdsourcing Practice for Efficient Data Labeling: Aggregation, Incremental Relabeling, and PricingabstractIn this tutorial, we present a portion of unique industry experience in efficient data labeling via crowdsourcing shared by both leading researchers and engineers from Yandex. We will make an introduction to data labeling via public crowdsourcing marketplaces and will present the key components of efficient label collection. This will be followed by a practice session, where participants will choose one of the real label collection tasks, experiment with selecting settings for the labeling process, and launch their label collection project on one of the largest crowdsourcing marketplaces. The projects will be run on real crowds within the tutorial session. While the crowd performers are annotating the project set up by the attendees, we will present the major theoretical results in efficient aggregation, incremental relabeling, and dynamic pricing. We will also discuss their strengths and weaknesses as well as applicability to real-world tasks, summarizing our five year-long research and industrial expertise in crowdsourcing. Finally, participants will receive a feedback about their projects and practical advice on how to make them more efficient. We invite beginners, advanced specialists, and researchers to learn how to collect high quality labeled data and do it efficiently. Alexey Drutsa, Valentina Fedorova, Dmitry Ustalov, Olga Megorskaya, Evfrosiniya Zerminova, Daria Baidakova |
SIGMOD Conference | 3 |
| 2020 | Practice of Efficient Data Collection via Crowdsourcing: Aggregation, Incremental Relabelling, and PricingabstractIn this tutorial, we present a portion of unique industry experience in efficient data labelling via crowdsourcing shared by both leading researchers and engineers from Yandex. We will make an introduction to data labelling via public crowdsourcing marketplaces and will present key components of efficient label collection. This will be followed by a practice session, where participants will choose one of the real label collection tasks, experiment with selecting settings for the labelling process, and launch their label collection project on Yandex.Toloka, one of the largest crowdsourcing marketplaces. The projects will be run on real crowds within the tutorial session. Finally, participants will receive a feedback about their projects and practical advice to make them more efficient. We expect that our tutorial will address an audience with a wide range of background and interests. We do not require specific prerequisite knowledge or skills. We invite beginners, advanced specialists, and researchers to learn how to efficiently collect labelled data. Alexey Drutsa, Valentina Fedorova, Dmitry Ustalov, Olga Megorskaya, Evfrosiniya Zerminova, Daria Baidakova |
WSDM | 3 |
| 2016 | YARN: Spinning-in-ProgressabstractYARN (Yet Another RussNet), a project started in 2013, aims at creating a large open WordNet-like thesaurus for Russian by means of crowdsourcing.The first stage of the project was to create noun synsets.Currently, the resource comprises 48K+ word entries and 44K+ synsets.More than 200 people have taken part in assembling synsets throughout the project.The paper describes the linguistic, technical, and organizational principles of the project, as well as the evaluation results, lessons learned, and the future plans. Pavel Braslavski 0001, Dmitry Ustalov, Mikhail Mukhin, Yuri Kiselev |
GWC | 2 |
| 2016 | Eliminating Fuzzy Duplicates in Crowdsourced Lexical ResourcesabstractCollaboratively created lexical resources is a trending approach to creating high quality thesauri in a short time span at a remarkably low price.The key idea is to invite non-expert participants to express and share their knowledge with the aim of constructing a resource.However, this approach tends to be noisy and error-prone, thus making data cleansing a highly topical task to perform.In this paper, we study different techniques for synset deduplication including machineand crowd-based ones.Eventually, we put forward an approach that can solve the deduplication problem fully automatically, with the quality comparable to the expertbased approach. Yuri Kiselev, Dmitry Ustalov, Sergey Porshnev |
GWC | 2 |