Petr Smirnov

dblp:178/7481 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Theory of computation · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2022 Tight Bounds for Tseitin Formulas
abstract
Bottom-up knowledge compilation is a paradigm for generating representations of functions by iteratively conjoining constraints using a so-called apply function. When the input is not efficiently compilable into a language - generally a class of circuits - because optimal compiled representations are provably large, the problem is not the compilation algorithm as much as the choice of a language too restrictive for the input. In contrast, in this paper, we look at CNF formulas for which very small circuits exists and look at the efficiency of their bottom-up compilation in one of the most general languages, namely that of structured decomposable negation normal forms (str-DNNF). We prove that, while the inputs have constant size representations as str-DNNF, any bottom-up compilation in the general setting where conjunction and structure modification are allowed takes exponential time and space, since large intermediate results have to be produced. This unconditionally proves that the inefficiency of bottom-up compilation resides in the bottom-up paradigm itself.
Dmitry Itsykson, Artur Riazanov, Petr Smirnov
SAT3
2022 Evaluation of statistical approaches for association testing in noisy drug screening data
abstract
BACKGROUND: Identifying associations among biological variables is a major challenge in modern quantitative biological research, particularly given the systemic and statistical noise endemic to biological systems. Drug sensitivity data has proven to be a particularly challenging field for identifying associations to inform patient treatment. RESULTS: To address this, we introduce two semi-parametric variations on the commonly used concordance index: the robust concordance index and the kernelized concordance index (rCI, kCI), which incorporate measurements about the noise distribution from the data. We demonstrate that common statistical tests applied to the concordance index and its variations fail to control for false positives, and introduce efficient implementations to compute p-values using adaptive permutation testing. We then evaluate the statistical power of these coefficients under simulation and compare with Pearson and Spearman correlation coefficients. Finally, we evaluate the various statistics in matching drugs across pharmacogenomic datasets. CONCLUSIONS: We observe that the rCI and kCI are better powered than the concordance index in simulation and show some improvement on real data. Surprisingly, we observe that the Pearson correlation was the most robust to measurement noise among the different metrics.
Petr Smirnov, Ian Smith, Zhaleh Safikhani, Wail Ba-alawi, Farnoosh Khodakarami, Eva Lin, Yihong Yu, Scott Martin, Janosch Ortmann, Tero Aittokallio, Marc Hafner, Benjamin Haibe-Kains
BMC Bioinform.1
2021 Drug sensitivity prediction from cell line-based pharmacogenomics data: guidelines for developing machine learning models
abstract
The goal of precision oncology is to tailor treatment for patients individually using the genomic profile of their tumors. Pharmacogenomics datasets such as cancer cell lines are among the most valuable resources for drug sensitivity prediction, a crucial task of precision oncology. Machine learning methods have been employed to predict drug sensitivity based on the multiple omics data available for large panels of cancer cell lines. However, there are no comprehensive guidelines on how to properly train and validate such machine learning models for drug sensitivity prediction. In this paper, we introduce a set of guidelines for different aspects of training gene expression-based predictors using cell line datasets. These guidelines provide extensive analysis of the generalization of drug sensitivity predictors and challenge many current practices in the community including the choice of training dataset and measure of drug sensitivity. The application of these guidelines in future studies will enable the development of more robust preclinical biomarkers.
Hossein Sharifi-Noghabi, Soheil Jahangiri-Tazehkand, Petr Smirnov, Casey Hon, Anthony Mammoliti, Sisira Kadambat Nair, Arvind Singh Mer, Martin Ester, Benjamin Haibe-Kains
Briefings Bioinform.3
2021 Near-Optimal Lower Bounds on Regular Resolution Refutations of Tseitin Formulas for All Constant-Degree Graphs
Dmitry Itsykson, Artur Riazanov, Danil Sagunov, Petr Smirnov
Comput. Complex.4
2021 Correction to: Near-Optimal Lower Bounds on Regular Resolution Refutations of Tseitin Formulas for All Constant-Degree Graphs
Dmitry Itsykson, Artur Riazanov, Danil Sagunov, Petr Smirnov
Comput. Complex.4
2019 Dr.VAE: improving drug response prediction via modeling of drug perturbation effects
abstract
MOTIVATION: Individualized drug response prediction is a fundamental part of personalized medicine for cancer. Great effort has been made to discover biomarkers or to develop machine learning methods for accurate drug response prediction in cancers. Incorporating prior knowledge of biological systems into these methods is a promising avenue to improve prediction performance. High-throughput cell line assays of drug-induced transcriptomic perturbation effects are a prior knowledge that has not been fully incorporated into a drug response prediction model yet. RESULTS: We introduce a unified probabilistic approach, Drug Response Variational Autoencoder (Dr.VAE), that simultaneously models both drug response in terms of viability and transcriptomic perturbations. Dr.VAE is a deep generative model based on variational autoencoders. Our experimental results showed Dr.VAE to do as well or outperform standard classification methods for 23 out of 26 tested Food and Drug Administration-approved drugs. In a series of ablation experiments we showed that the observed improvement of Dr.VAE can be credited to the incorporation of drug-induced perturbation effects with joint modeling of treatment sensitivity. AVAILABILITY AND IMPLEMENTATION: Processed data and software implementation using PyTorch (Paszke et al., 2017) are available at: https://github.com/rampasek/DrVAE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ladislav Rampásek, Daniel Hidru, Petr Smirnov, Benjamin Haibe-Kains, Anna Goldenberg
Bioinform.3
2018 Tissue specificity of in vitro drug sensitivity
abstract
Objectives: We sought to investigate the tissue specificity of drug sensitivities in large-scale pharmacological studies and compare these associations to those found in drug clinical indications. Materials and Methods: We leveraged the curated cell line response data from PharmacoGx and applied an enrichment algorithm on drug sensitivity values' area under the drug dose-response curves (AUCs) with and without adjustment for general level of drug sensitivity. Results: We observed tissue specificity in 63% of tested drugs, with 8% of total interactions deemed significant (false discovery rate <0.05). By restricting the drug-tissue interactions to those with AUC > 0.2, we found that in 52% of interactions, the tissue was predictive of drug sensitivity (concordance index > 0.65). When compared with clinical indications, the observed overlap was weak (Matthew correlation coefficient, MCC = 0.0003, P > .10). Discussion: While drugs exhibit significant tissue specificity in vitro, there is little overlap with clinical indications. This can be attributed to factors such as underlying biological differences between in vitro models and patient tumors, or the inability of tissue-specific drugs to bring additional benefits beyond gold standard treatments during clinical trials. Conclusion: Our meta-analysis of pan-cancer drug screening datasets indicates that most tested drugs exhibit tissue-specific sensitivities in a large panel of cancer cell lines. However, the observed preclinical results do not translate to the clinical setting. Our results suggest that additional research into showing parallels between preclinical and clinical data is required to increase the translational potential of in vitro drug screening.
Fupan Yao, Seyed Ali Madani Tonekaboni, Zhaleh Safikhani, Petr Smirnov, Nehme Hachem, Mark Freeman 0002, Venkata Satya Kumar Manem, Benjamin Haibe-Kains
J. Am. Medical Informatics Assoc.4
2016 PharmacoGx: an R package for analysis of large pharmacogenomic datasets
abstract
UNLABELLED: Pharmacogenomics holds great promise for the development of biomarkers of drug response and the design of new therapeutic options, which are key challenges in precision medicine. However, such data are scattered and lack standards for efficient access and analysis, consequently preventing the realization of the full potential of pharmacogenomics. To address these issues, we implemented PharmacoGx, an easy-to-use, open source package for integrative analysis of multiple pharmacogenomic datasets. We demonstrate the utility of our package in comparing large drug sensitivity datasets, such as the Genomics of Drug Sensitivity in Cancer and the Cancer Cell Line Encyclopedia. Moreover, we show how to use our package to easily perform Connectivity Map analysis. With increasing availability of drug-related data, our package will open new avenues of research for meta-analysis of pharmacogenomic data. AVAILABILITY AND IMPLEMENTATION: PharmacoGx is implemented in R and can be easily installed on any system. The package is available from CRAN and its source code is available from GitHub. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Petr Smirnov, Zhaleh Safikhani, Nehme Hachem, Adrian She, Catharina Olsen, Mark Freeman 0002, Heather Marie Selby, Deena M. A. Gendoo, Patrick Grossmann, Andrew H. Beck, Hugo J. W. L. Aerts, Mathieu Lupien, Anna Goldenberg, Benjamin Haibe-Kains
Bioinform.1