EDBT 2026 Demo / reviewers in the wild / expert
Boya Zeng
dblp:395/4030
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 67% Vision and language · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 2 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
dataset bias |
0.8 | 1 | 2024 | Understanding Bias in Large-Scale Visual Datasets · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › dataset bias
visual dataset bias |
0.8 | 1 | 2024 | Understanding Bias in Large-Scale Visual Datasets · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
semantic transformation · 1.5natural language description generation · 1.5frequency analysis · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Understanding Bias in Large-Scale Visual DatasetsabstractA recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a framework to identify the unique visual attributes distinguishing these datasets. Our approach applies various transformations to extract semantic, structural, boundary, color, and frequency information from datasets, and assess how much each type of information reflects their bias. We further decompose their semantic bias with object-level analysis, and leverage natural language methods to generate detailed, open-ended descriptions of each dataset's characteristics. Our work aims to help researchers understand the bias in existing large-scale pre-training datasets, and build more diverse and representative ones in the future. Our project page and code are available at boyazeng.github.io/understand_bias. Boya Zeng, Yida Yin, Zhuang Liu 0003 |
NeurIPS | 1 |