Tomasz Danel

dblp:248/8081 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-6053-0028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 72% Generative modeling · 24% Deep learning architectures and training · 4%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
1.012026
Enhancing Chemical Explainability Through Counterfactual Masking · AAAI 2026
Machine learning › Trustworthy machine learning › interpretability
explainable AI
1.012026
Enhancing Chemical Explainability Through Counterfactual Masking · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
Enhancing Chemical Explainability Through Counterfactual Masking · AAAI 2026
Machine learning › Generative modeling › molecular generation
molecular graph generation
1.012026
Enhancing Chemical Explainability Through Counterfactual Masking · AAAI 2026
Bioinformatics and computational biology
molecular property prediction
0.922026
HuggingMolecules: An Open-Source Library for Transformer-Based Molecular Property Prediction (Student Abstract) · AAAI 2022
Enhancing Chemical Explainability Through Counterfactual Masking · AAAI 2026
Bioinformatics and computational biology
drug discovery
0.912025
KinDEL: DNA-Encoded Library Dataset for Kinase Inhibitors · ICML 2025
Bioinformatics and computational biology › drug discovery
virtual screening
0.912025
KinDEL: DNA-Encoded Library Dataset for Kinase Inhibitors · ICML 2025
Information retrieval › evaluation
benchmark dataset
0.912025
KinDEL: DNA-Encoded Library Dataset for Kinase Inhibitors · ICML 2025
Bioinformatics and computational biology › molecular informatics
cheminformatics
0.612022
HuggingMolecules: An Open-Source Library for Transformer-Based Molecular Property Prediction (Student Abstract) · AAAI 2022
Machine learning › Deep learning architectures and training
transformer
0.212022
HuggingMolecules: An Open-Source Library for Transformer-Based Molecular Property Prediction (Student Abstract) · AAAI 2022

Methods — techniques the papers use, named apart from their topics

generative model · 2.0counterfactual masking · 2.0probabilistic modeling · 1.7molecular docking · 1.72d and 3d structure representation · 1.7transformer models · 1.1
YearPublicationVenuePosition
2026 Enhancing Chemical Explainability Through Counterfactual Masking
abstract
Molecular property prediction is a crucial task that guides the design of new compounds, including drugs and materials. While explainable artificial intelligence methods aim to scrutinize model predictions by identifying influential molecular substructures, many existing approaches rely on masking strategies that remove either atoms or atom-level features to assess importance via fidelity metrics. These methods, however, often fail to adhere to the underlying molecular distribution and thus yield unintuitive explanations. In this work, we propose counterfactual masking, a novel framework that replaces masked substructures with chemically reasonable fragments sampled from generative models trained to complete molecular graphs. Rather than evaluating masked predictions against implausible zeroed-out baselines, we assess them relative to counterfactual molecules drawn from the data distribution. Our method offers two key benefits: (1) molecular realism that underpins robust and distribution-consistent explanations, and (2) meaningful counterfactuals that directly indicate how structural modifications may affect predicted properties. We demonstrate that counterfactual masking is well-suited for benchmarking model explainers and yields more actionable insights across multiple datasets and property prediction tasks. Our approach bridges the gap between explainability and molecular design, offering a principled and generative path toward explainable machine learning in chemistry.
Lukasz Janisiów, Marek Kochanczyk, Bartosz Zielinski 0001, Tomasz Danel
AAAI4
2025 KinDEL: DNA-Encoded Library Dataset for Kinase Inhibitors
abstract
DNA-Encoded Libraries (DELs) represent a transformative technology in drug discovery, facilitating the high-throughput exploration of vast chemical spaces. Despite their potential, the scarcity of publicly available DEL datasets presents a bottleneck for the advancement of machine learning methodologies in this domain. To address this gap, we introduce KinDEL, one of the largest publicly accessible DEL datasets and the first one that includes binding poses from molecular docking experiments. Focused on two kinases, Mitogen-Activated Protein Kinase 14 (MAPK14) and Discoidin Domain Receptor Tyrosine Kinase 1 (DDR1), KinDEL includes 81 million compounds, offering a rich resource for computational exploration. Additionally, we provide comprehensive biophysical assay validation data, encompassing both on-DNA and off-DNA measurements, which we use to evaluate a suite of machine learning techniques, including novel structure-based probabilistic models. We hope that our benchmark, encompassing both 2D and 3D structures, will help advance the development of machine learning models for data-driven hit identification using DELs.
Benson Chen, Tomasz Danel, Gabriel H. S. Dreiman, Patrick J. McEnaney, Kirill Novikov, Spurti Umesh Akki, Joshua L. Turnbull, Virja Atul Pandya, Boris P. Belotserkovskii, Jared Bryce Weaver, Ankita Biswas, Kent Gorday, Mohammad Sultan, Nathaniel Stanley, Daniel M. Whalen, Divya Kanichar, Christoph Klein 0006, Emily Fox, R. Edward Watts
ICML2
2024 Feature-Based Interpolation and Geodesics in the Latent Spaces of Generative Models
abstract
Interpolating between points is a problem connected simultaneously with finding geodesics and study of generative models. In the case of geodesics, we search for the curves with the shortest length, while in the case of generative models, we typically apply linear interpolation in the latent space. However, this interpolation uses implicitly the fact that Gaussian is unimodal. Thus, the problem of interpolating in the case when the latent density is non-Gaussian is an open problem. In this article, we present a general and unified approach to interpolation, which simultaneously allows us to search for geodesics and interpolating curves in latent space in the case of arbitrary density. Our results have a strong theoretical background based on the introduced quality measure of an interpolating curve. In particular, we show that maximizing the quality measure of the curve can be equivalently understood as a search of geodesic for a certain redefinition of the Riemannian metric on the space. We provide examples in three important cases. First, we show that our approach can be easily applied to finding geodesics on manifolds. Next, we focus our attention in finding interpolations in pretrained generative models. We show that our model effectively works in the case of arbitrary density. Moreover, we can interpolate in the subset of the space consisting of data possessing a given feature. The last case is focused on finding interpolation in the space of chemical compounds.
Lukasz Struski, Michal Sadowski, Tomasz Danel, Jacek Tabor, Igor T. Podolak
IEEE Trans. Neural Networks Learn. Syst.3
2023 ProGReST: Prototypical Graph Regression Soft Trees for Molecular Property Prediction
abstract
In this work, we propose the novel Prototypical Graph Regression Self-explainable Trees (ProGReST) model, which combines prototype learning, soft decision trees, and Graph Neural Networks. In contrast to other works, our model can be used to address various challenging tasks, including compound property prediction. In ProGReST, the rationale is obtained along with prediction due to the model's built-in interpretability. Additionally, we introduce a new graph prototype projection to accelerate model training. Finally, we evaluate PRoGReST on a wide range of chemical datasets for molecular property prediction and perform in-depth analysis with chemical experts to evaluate obtained interpretations. Our method achieves competitive results against state-of- the-art methods.
Dawid Rymarczyk, Daniel Dobrowolski, Tomasz Danel
SDM3
2023 SONGs: Self-Organizing Neural Graphs
abstract
Recent years have seen a surge in research on combining deep neural networks with other methods, including decision trees and graphs. There are at least three advantages of incorporating decision trees and graphs: they are easy to interpret since they are based on sequential decisions, they can make decisions faster, and they provide a hierarchy of classes. However, one of the well-known drawbacks of decision trees, as compared to decision graphs, is that decision trees cannot reuse the decision nodes. Nevertheless, decision graphs were not commonly used in deep learning due to the lack of efficient gradient-based training techniques. In this paper, we fill this gap and provide a general paradigm based on Markov processes, which allows for efficient training of the special type of decision graphs, which we call Self-Organizing Neural Graphs (SONG). We provide a theoretical study on SONG, complemented by experiments conducted on Letter, Connect4, MNIST, CIFAR, and TinyImageNet datasets, showing that our method performs on par or better than existing decision models.
Lukasz Struski, Tomasz Danel, Marek Smieja, Jacek Tabor, Bartosz Zielinski 0001
WACV2
2022 HuggingMolecules: An Open-Source Library for Transformer-Based Molecular Property Prediction (Student Abstract)
abstract
Large-scale transformer-based methods are gaining popularity as a tool for predicting the properties of chemical compounds, which is of central importance to the drug discovery process. To accelerate their development and dissemination among the community, we are releasing HuggingMolecules -- an open-source library, with a simple and unified API, that provides the implementation of several state-of-the-art transformers for molecular property prediction. In addition, we add a comparison of these methods on several regression and classification datasets. HuggingMolecules package is available at: github.com/gmum/huggingmolecules.
Piotr Gainski, Lukasz Maziarka, Tomasz Danel, Stanislaw Jastrzebski
AAAI3
2021 Multitask Learning Using BERT with Task-Embedded Attention
abstract
Multitask learning helps to obtain a meaningful representation of the data, retaining a small number of parameters needed to train the model. In natural language processing, models often reach as many as a few hundred million trainable parameters, which makes adaptations to new tasks computationally infeasible. Creating shareable layers for multiple tasks allows to spare resources and often leads to better data representation. In this work, we propose a new approach to train a BERT model in a multitask setup, which we call EmBERT. To introduce information about the task, we inject task-specific embeddings to the multi-head attention layers. Our modified architecture requires a minimal number of additional parameters relative to the original BERT model (+0.025% per task) while achieving state-of-the-art results in the GLUE benchmark.
Lukasz Maziarka, Tomasz Danel
IJCNN2
2021 Comparison of Atom Representations in Graph Neural Networks for Molecular Property Prediction
abstract
Graph neural networks have recently become a standard method for analysing chemical compounds. In the field of molecular property prediction, the emphasis is now put on designing new model architectures, and the importance of atom featurisation is oftentimes belittled. When contrasting two graph neural networks, the use of different atom features possibly leads to the incorrect attribution of the results to the network architecture. To provide a better understanding of this issue, we compare multiple atom representations for graph models and evaluate them on the prediction of free energy, solubility, and metabolic stability. To the best of our knowledge, this is the first methodological study that focuses on the relevance of atom representation to the predictive performance of graph neural networks.
Agnieszka Wojtuch, Tomasz Danel, Sabina Podlewska, Jacek Tabor, Lukasz Maziarka
IJCNN2
2020 Processing of Incomplete Images by (Graph) Convolutional Neural Networks
Tomasz Danel, Marek Smieja, Lukasz Struski, Przemyslaw Spurek, Lukasz Maziarka
ICONIP (2)1
2020 Spatial Graph Convolutional Networks
Tomasz Danel, Przemyslaw Spurek, Jacek Tabor, Marek Smieja, Lukasz Struski, Agnieszka Slowik, Lukasz Maziarka
ICONIP (5)1