EDBT 2026 Demo / reviewers in the wild / expert
Daria Baidakova
dblp:256/9284
· DBLP profile ↗
4ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 67% Machine learning and data management · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 50% Computational science and engineering · 50% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management
data annotation |
0.9 | 2 | 2020 | Practice of Efficient Data Collection via Crowdsourcing: Aggregation, Incremental Relabelling, and Pricing · WSDM 2020 Crowdsourcing Practice for Efficient Data Labeling: Aggregation, Incremental Relabeling, and Pricing · SIGMOD Conference 2020 |
Data mining › crowdsourcing
label aggregation |
0.9 | 2 | 2020 | Practice of Efficient Data Collection via Crowdsourcing: Aggregation, Incremental Relabelling, and Pricing · WSDM 2020 Crowdsourcing Practice for Efficient Data Labeling: Aggregation, Incremental Relabeling, and Pricing · SIGMOD Conference 2020 |
Computing education › AI education
machine learning education |
0.7 | 1 | 2023 | Data Labeling for Machine Learning Engineers: Project-Based Curriculum and Data-Centric Competitions · AAAI 2023 |
Computational science and engineering › active learning
project-based learning |
0.7 | 1 | 2023 | Data Labeling for Machine Learning Engineers: Project-Based Curriculum and Data-Centric Competitions · AAAI 2023 |
Data mining › crowdsourcing
crowdsourced annotation |
0.4 | 1 | 2020 | Crowdsourcing Practice for Efficient Data Labeling: Aggregation, Incremental Relabeling, and Pricing · SIGMOD Conference 2020 |
Data mining › crowdsourcing
crowdsourced data |
0.4 | 1 | 2020 | Practice of Efficient Data Collection via Crowdsourcing: Aggregation, Incremental Relabelling, and Pricing · WSDM 2020 |
Methods — techniques the papers use, named apart from their topics
project-based curriculum · 0.7inverse ML competitions · 0.7crowdsourcing · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Data Labeling for Machine Learning Engineers: Project-Based Curriculum and Data-Centric CompetitionsabstractThe process of training and evaluating machine learning (ML) models relies on high-quality and timely annotated datasets. While a significant portion of academic and industrial research is focused on creating new ML methods, these communities rely on open datasets and benchmarks. However, practitioners often face issues with unlabeled and unavailable data specific to their domain. We believe that building scalable and sustainable processes for collecting data of high quality for ML is a complex skill that needs focused development. To fill the need for this competency, we created a semester course on Data Collection and Labeling for Machine Learning, integrated into a bachelor program that trains data analysts and ML engineers. The course design and delivery illustrate how to overcome the challenge of putting university students with a theoretical background in mathematics, computer science, and physics through a program that is substantially different from their educational habits. Our goal was to motivate students to focus on practicing and mastering a skill that was considered unnecessary to their work. We created a system of inverse ML competitions that showed the students how high-quality and relevant data affect their work with ML models, and their mindset changed completely in the end. Project-based learning with increasing complexity of conditions at each stage helped to raise the satisfaction index of students accustomed to difficult challenges. During the course, our invited industry practitioners drew on their first-hand experience with data, which helped us avoid overtheorizing and made the course highly applicable to the students’ future career paths. Anastasia Zhdanovskaya, Daria Baidakova, Dmitry Ustalov |
AAAI | 2 |
| 2022 | Web Engineering with Human-in-the-Loop
Dmitry Ustalov, Nikita Pavlichenko, Boris Tseytlin, Daria Baidakova, Alexey Drutsa |
ICWE | 4 |
| 2020 | Crowdsourcing Practice for Efficient Data Labeling: Aggregation, Incremental Relabeling, and PricingabstractIn this tutorial, we present a portion of unique industry experience in efficient data labeling via crowdsourcing shared by both leading researchers and engineers from Yandex. We will make an introduction to data labeling via public crowdsourcing marketplaces and will present the key components of efficient label collection. This will be followed by a practice session, where participants will choose one of the real label collection tasks, experiment with selecting settings for the labeling process, and launch their label collection project on one of the largest crowdsourcing marketplaces. The projects will be run on real crowds within the tutorial session. While the crowd performers are annotating the project set up by the attendees, we will present the major theoretical results in efficient aggregation, incremental relabeling, and dynamic pricing. We will also discuss their strengths and weaknesses as well as applicability to real-world tasks, summarizing our five year-long research and industrial expertise in crowdsourcing. Finally, participants will receive a feedback about their projects and practical advice on how to make them more efficient. We invite beginners, advanced specialists, and researchers to learn how to collect high quality labeled data and do it efficiently. Alexey Drutsa, Valentina Fedorova, Dmitry Ustalov, Olga Megorskaya, Evfrosiniya Zerminova, Daria Baidakova |
SIGMOD Conference | 6 |
| 2020 | Practice of Efficient Data Collection via Crowdsourcing: Aggregation, Incremental Relabelling, and PricingabstractIn this tutorial, we present a portion of unique industry experience in efficient data labelling via crowdsourcing shared by both leading researchers and engineers from Yandex. We will make an introduction to data labelling via public crowdsourcing marketplaces and will present key components of efficient label collection. This will be followed by a practice session, where participants will choose one of the real label collection tasks, experiment with selecting settings for the labelling process, and launch their label collection project on Yandex.Toloka, one of the largest crowdsourcing marketplaces. The projects will be run on real crowds within the tutorial session. Finally, participants will receive a feedback about their projects and practical advice to make them more efficient. We expect that our tutorial will address an audience with a wide range of background and interests. We do not require specific prerequisite knowledge or skills. We invite beginners, advanced specialists, and researchers to learn how to efficiently collect labelled data. Alexey Drutsa, Valentina Fedorova, Dmitry Ustalov, Olga Megorskaya, Evfrosiniya Zerminova, Daria Baidakova |
WSDM | 6 |