VLDB 2026 Research / reviewers in the wild / expert
David Marecek
dblp:23/7248
· DBLP profile ↗
19ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-5327-488XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 31% Information extraction and text analysis · 30% Language models and text generation · 19% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Debiasing Algorithm through Model Adaptation · ICLR 2024 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation |
0.8 | 1 | 2024 | Debiasing Algorithm through Model Adaptation · ICLR 2024 |
Natural language and speech › Language models and text generation
knowledge editing |
0.8 | 1 | 2024 | Debiasing Algorithm through Model Adaptation · ICLR 2024 |
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual parsing |
0.5 | 1 | 2021 | Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning |
0.5 | 1 | 2021 | Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training › regularization
orthogonal constraint |
0.5 | 1 | 2021 | Introducing Orthogonal Constraint in Structural Probes · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.5 | 1 | 2021 | Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text representation › syntactic representation
coordination structure |
0.2 | 1 | 2013 | Coordination Structures in Dependency Treebanks · ACL (1) 2013 |
Natural language and speech › Information extraction and text analysis › data annotation › corpus annotation
syntactic annotation |
0.2 | 1 | 2013 | Coordination Structures in Dependency Treebanks · ACL (1) 2013 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing |
0.1 | 1 | 2012 | Exploiting Reducibility in Unsupervised Dependency Parsing · EMNLP-CoNLL 2012 |
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
unsupervised dependency parsing |
0.1 | 1 | 2012 | Exploiting Reducibility in Unsupervised Dependency Parsing · EMNLP-CoNLL 2012 |
Methods — techniques the papers use, named apart from their topics
model adaptation · 0.8linear projection · 0.8causal analysis · 0.8structural probing · 0.5orthogonal transformation · 0.5orthogonal structural probe · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and TranslationabstractMitigation of biases, such as language models’ reliance on gender stereotypes, is a crucial endeavor required for the creation of reliable and useful language technology. The crucial aspect of debiasing is to ensure that the models preserve their versatile capabilities, including their ability to solve language tasks and equitably represent various genders. To address these issues, we introduce Dual Dabiasing Algorithm through Model Adaptation (2DAMA). Novel Dual Debiasing enables robust reduction of stereotypical bias while preserving desired factual gender information encoded by language models. We show that 2DAMA effectively reduces gender bias in language models for English and is one of the first approaches facilitating the mitigation of their stereotypical tendencies in translation. The proposed method’s key advantage is the preservation of factual gender cues, which are useful in a wide range of natural language processing tasks. Tomasz Limisiewicz, David Marecek, Tomás Musil |
INLG | 2 |
| 2024 | Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder TestabstractIndependent Component Analysis (ICA) is an algorithm originally developed for finding separate sources in a mixed signal, such as a recording of multiple people in the same room speaking at the same time. Unlike Principal Component Analysis (PCA), ICA permits the representation of a word as an unstructured set of features, without any particular feature being deemed more significant than the others. In this paper, we used ICA to analyze word embeddings. We have found that ICA can be used to find semantic features of the words and these features can easily be combined to search for words that satisfy the combination. We show that most of the independent components represent such features. To quantify the interpretability of the components, we use the word intruder test, performed both by humans and by large language models. We propose to use the automated version of the word intruder test as a fast and inexpensive way of quantifying vector interpretability without the need for human effort. Tomás Musil, David Marecek |
LREC/COLING | 2 |
| 2024 | Debiasing Algorithm through Model AdaptationabstractLarge language models are becoming the go-to solution for the ever-growing number of tasks.
However, with growing capacity, models are prone to rely on spurious correlations stemming from biases and stereotypes present in the training data.
This work proposes a novel method for detecting and mitigating gender bias in language models.
We perform causal analysis to identify problematic model components and discover that mid-upper feed-forward layers are most prone to convey bias.
Based on the analysis results, we intervene in the model by applying a linear projection to the weight matrices of these layers.
Our titular method DAMA, significantly decreases bias as measured by diverse metrics while maintaining the model's performance on downstream tasks.
We release code for our method and models, which retrain LLaMA's state-of-the-art performance while being significantly less biased. Tomasz Limisiewicz, David Marecek, Tomás Musil |
ICLR | 2 |
| 2023 | The Functional Relevance of Probed Information: A Case StudyabstractRecent studies have shown that transformer models like BERT rely on number information encoded in their representations of sentences' subjects and head verbs when performing subject-verb agreement.However, probing experiments suggest that subject number is also encoded in the representations of all words in such sentences.In this paper, we use causal interventions to show that BERT only uses the subject plurality information encoded in its representations of the subject and words that agree with it in number.We also demonstrate that current probing metrics are unable to determine which words' representations contain functionally relevant information.This both provides a revised view of subject-verb agreement in language models, and suggests potential pitfalls for current probe usage and evaluation. Michael Hanna 0001, Roberto Zamparelli, David Marecek |
EACL | 3 |
| 2023 | Exploring the Impact of Training Data Distribution and Subword Tokenization on Gender Bias in Machine TranslationabstractBar Iluz, Tomasz Limisiewicz, Gabriel Stanovsky, David Mareček. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Bar Iluz, Tomasz Limisiewicz, Gabriel Stanovsky, David Marecek |
IJCNLP (1) | 4 |
| 2021 | Introducing Orthogonal Constraint in Structural ProbesabstractTomasz Limisiewicz, David Mareček. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tomasz Limisiewicz, David Marecek |
ACL/IJCNLP (1) | 2 |
| 2021 | Examining Cross-lingual Contextual Embeddings with Orthogonal Structural ProbesabstractState-of-the-art contextual embeddings are obtained from large language models available only for a few languages.For others, we need to learn representations using a multilingual model.There is an ongoing debate on whether multilingual embeddings can be aligned in a space shared across many languages.The novel Orthogonal Structural Probe (Limisiewicz and Mareček, 2021) allows us to answer this question for specific linguistic features and learn a projection based only on mono-lingual annotated datasets.We evaluate syntactic (UD) and lexical (WordNet) structural information encoded in MBERT's contextual representations for nine diverse languages.1 We observe that for languages closely related to English, no transformation is needed.The evaluated information is encoded in a shared cross-lingual embedding space.For other languages, it is beneficial to apply orthogonal transformation learned separately for each language.We successfully apply our findings to zero-shot and few-shot cross-lingual parsing. Tomasz Limisiewicz, David Marecek |
EMNLP (1) | 2 |
| 2016 | If You Even Don't Have a Bit of Bible: Learning Delexicalized POS Taggers
David Marecek, Zdenek Zabokrtský, Daniel Zeman |
LREC | 2 |
| 2016 | Planting Trees in the Desert: Delexicalized Tagging and Parsing Combined
Daniel Zeman, David Marecek, Zdenek Zabokrtský |
PACLIC | 2 |
| 2014 | Dealing with Function Words in Unsupervised Dependency Parsing
David Marecek, Zdenek Zabokrtský |
CICLing (1) | 1 |
| 2014 | HamleDT 2.0: Thirty Dependency Treebanks Stanfordized
Rudolf Rosa, Jan Masek 0003, David Marecek, Martin Popel, Daniel Zeman, Zdenek Zabokrtský |
LREC | 3 |
| 2014 | Adaptation of machine translation for multilingual information retrieval in the medical domain
Pavel Pecina, Ondrej Dusek, Lorraine Goeuriot, Jan Hajic 0001, Jaroslava Hlavácová, Gareth J. F. Jones, Liadh Kelly, Johannes Leveling, David Marecek, Michal Novák 0001, Martin Popel, Rudolf Rosa, Ales Tamchyna, Zdenka Uresová |
Artif. Intell. Medicine | 9 |
| 2013 | Stop-probability estimates computed on a large corpus improve Unsupervised Dependency Parsing
David Marecek, Milan Straka |
ACL (1) | 1 |
| 2013 | Coordination Structures in Dependency Treebanks
Martin Popel, David Marecek, Jan Stepánek, Daniel Zeman, Zdenek Zabokrtský |
ACL (1) | 2 |
| 2012 | Exploiting Reducibility in Unsupervised Dependency Parsing
David Marecek, Zdenek Zabokrtský |
EMNLP-CoNLL | 1 |
| 2012 | The Joy of Parallelism with CzEng 1.0
Ondrej Bojar, Zdenek Zabokrtský, Ondrej Dusek, Petra Galuscáková, Martin Majlis, David Marecek, Jirka Marsík, Michal Novák 0001, Martin Popel, Ales Tamchyna |
LREC | 6 |
| 2012 | HamleDT: To Parse or Not to Parse?
Daniel Zeman, David Marecek, Martin Popel, Loganathan Ramasamy, Jan Stepánek, Zdenek Zabokrtský, Jan Hajic 0001 |
LREC | 2 |
| 2011 | Combining Diverse Word-Alignment Symmetrizations Improves Dependency Tree Projection
David Marecek |
CICLing (1) | 1 |
| 2008 | Automatic alignment of Czech and English deep syntactic dependency trees
David Marecek, Zdenek Zabokrtský, Václav Novák |
EAMT | 1 |