VLDB 2026 Research / reviewers in the wild / expert
Tomás Horych
dblp:345/8769
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2024
0009-0003-6456-2977ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset Collection · SIGIR 2023 |
Machine learning › Trustworthy machine learning › fairness › bias evaluation › bias detection
media bias detection |
0.7 | 1 | 2023 | Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset Collection · SIGIR 2023 |
Information retrieval › evaluation
benchmark |
0.7 | 1 | 2023 | Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset Collection · SIGIR 2023 |
Information retrieval
evaluation |
0.7 | 1 | 2023 | Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset Collection · SIGIR 2023 |
Methods — techniques the papers use, named apart from their topics
transformer models · 1.3t5 · 1.3BART · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MAGPIE: Multi-Task Analysis of Media-Bias Generalization with Pre-Trained Identification of ExpressionsabstractMedia bias detection poses a complex, multifaceted problem traditionally tackled using single-task models and small in-domain datasets, consequently lacking generalizability. To address this, we introduce MAGPIE, a large-scale multi-task pre-training approach explicitly tailored for media bias detection. To enable large-scale pre-training, we construct Large Bias Mixture (LBM), a compilation of 59 bias-related tasks. MAGPIE outperforms previous approaches in media bias detection on the Bias Annotation By Experts (BABE) dataset, with a relative improvement of 3.3% F1-score. Furthermore, using a RoBERTa encoder, we show that MAGPIE needs only 15% of fine-tuning steps compared to single-task approaches. We provide insight into task learning interference and show that sentiment analysis and emotion detection help learning of all other tasks, and scaling the number of tasks leads to the best results. MAGPIE confirms that MTL is a promising approach for addressing media bias detection, enhancing the accuracy and efficiency of existing models. Furthermore, LBM is the first available resource collection focused on media bias MTL. Tomás Horych, Martin Wessel, Jan Philip Wahle, Terry Ruas, Jerome Waßmuth, André Greiner-Petter, Akiko Aizawa, Bela Gipp, Timo Spinde |
LREC/COLING | 1 |
| 2023 | Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset CollectionabstractAlthough media bias detection is a complex multi-task problem, there is, to date, no unified benchmark grouping these evaluation tasks. We introduce the Media Bias Identification Benchmark (MBIB), a comprehensive benchmark that groups different types of media bias (e.g., linguistic, cognitive, political) under a common framework to test how prospective detection techniques generalize. After reviewing 115 datasets, we select nine tasks and carefully propose 22 associated datasets for evaluating media bias detection techniques. We evaluate MBIB using state-of-the-art Transformer techniques (e.g., T5, BART). Our results suggest that while hate speech, racial bias, and gender bias are easier to detect, models struggle to handle certain bias types, e.g., cognitive and political bias. However, our results show that no single technique can outperform all the others significantly.We also find an uneven distribution of research interest and resource allocation to the individual tasks in media bias. A unified benchmark encourages the development of more robust systems and shifts the current paradigm in media bias detection evaluation towards solutions that tackle not one but multiple media bias types simultaneously. Martin Wessel, Tomás Horych, Terry Ruas, Akiko Aizawa, Bela Gipp, Timo Spinde |
SIGIR | 2 |