Boya Zeng

dblp:395/4030 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 67% Vision and language · 33%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 2 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
dataset bias
0.812024
Understanding Bias in Large-Scale Visual Datasets · NeurIPS 2024
Machine learning › Trustworthy machine learning › dataset bias
visual dataset bias
0.812024
Understanding Bias in Large-Scale Visual Datasets · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

semantic transformation · 1.5natural language description generation · 1.5frequency analysis · 1.5
YearPublicationVenuePosition
2024 Understanding Bias in Large-Scale Visual Datasets
abstract
A recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a framework to identify the unique visual attributes distinguishing these datasets. Our approach applies various transformations to extract semantic, structural, boundary, color, and frequency information from datasets, and assess how much each type of information reflects their bias. We further decompose their semantic bias with object-level analysis, and leverage natural language methods to generate detailed, open-ended descriptions of each dataset's characteristics. Our work aims to help researchers understand the bias in existing large-scale pre-training datasets, and build more diverse and representative ones in the future. Our project page and code are available at boyazeng.github.io/understand_bias.
Boya Zeng, Yida Yin, Zhuang Liu 0003
NeurIPS1