Tomás Musil

dblp:205/9019 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 67% Language models and text generation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.812024
Debiasing Algorithm through Model Adaptation · ICLR 2024
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation
0.812024
Debiasing Algorithm through Model Adaptation · ICLR 2024
Natural language and speech › Language models and text generation
knowledge editing
0.812024
Debiasing Algorithm through Model Adaptation · ICLR 2024

Methods — techniques the papers use, named apart from their topics

model adaptation · 0.8linear projection · 0.8causal analysis · 0.8
YearPublicationVenuePosition
2026 SEEM-CZ: Annotation and Classification of Epistemic Markers in Czech
Barbora Stepánková, Michal Novák 0001, Tomás Musil, Lucie Poláková
LREC3
2025 Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
abstract
Mitigation of biases, such as language models’ reliance on gender stereotypes, is a crucial endeavor required for the creation of reliable and useful language technology. The crucial aspect of debiasing is to ensure that the models preserve their versatile capabilities, including their ability to solve language tasks and equitably represent various genders. To address these issues, we introduce Dual Dabiasing Algorithm through Model Adaptation (2DAMA). Novel Dual Debiasing enables robust reduction of stereotypical bias while preserving desired factual gender information encoded by language models. We show that 2DAMA effectively reduces gender bias in language models for English and is one of the first approaches facilitating the mitigation of their stereotypical tendencies in translation. The proposed method’s key advantage is the preservation of factual gender cues, which are useful in a wide range of natural language processing tasks.
Tomasz Limisiewicz, David Marecek, Tomás Musil
INLG3
2024 Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test
abstract
Independent Component Analysis (ICA) is an algorithm originally developed for finding separate sources in a mixed signal, such as a recording of multiple people in the same room speaking at the same time. Unlike Principal Component Analysis (PCA), ICA permits the representation of a word as an unstructured set of features, without any particular feature being deemed more significant than the others. In this paper, we used ICA to analyze word embeddings. We have found that ICA can be used to find semantic features of the words and these features can easily be combined to search for words that satisfy the combination. We show that most of the independent components represent such features. To quantify the interpretability of the components, we use the word intruder test, performed both by humans and by large language models. We propose to use the automated version of the word intruder test as a fast and inexpensive way of quantifying vector interpretability without the need for human effort.
Tomás Musil, David Marecek
LREC/COLING1
2024 Debiasing Algorithm through Model Adaptation
abstract
Large language models are becoming the go-to solution for the ever-growing number of tasks. However, with growing capacity, models are prone to rely on spurious correlations stemming from biases and stereotypes present in the training data. This work proposes a novel method for detecting and mitigating gender bias in language models. We perform causal analysis to identify problematic model components and discover that mid-upper feed-forward layers are most prone to convey bias. Based on the analysis results, we intervene in the model by applying a linear projection to the weight matrices of these layers. Our titular method DAMA, significantly decreases bias as measured by diverse metrics while maintaining the model's performance on downstream tasks. We release code for our method and models, which retrain LLaMA's state-of-the-art performance while being significantly less biased.
Tomasz Limisiewicz, David Marecek, Tomás Musil
ICLR3