Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Artem Cherkasov

dblp:88/2110 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
2since 2021 · last 2025
0000-0002-1599-1439ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 96% Medical and health informatics · 4%
Artificial intelligence
1 paper
Generative modeling · 67% Representation and self-supervised learning · 33%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative flow networks
0.912025
Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation · ICML 2025
Machine learning › Generative modeling › molecular generation
molecular graph generation
0.912025
Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation · ICML 2025
Machine learning › Representation and self-supervised learning
pre-training
0.912025
Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation · ICML 2025
Bioinformatics and computational biology
drug discovery
0.612022
DD-GUI: a graphical user interface for deep learning-accelerated virtual screening of large chemical libraries (Deep Docking) · Bioinform. 2022
Bioinformatics and computational biology › drug discovery
virtual screening
0.612022
DD-GUI: a graphical user interface for deep learning-accelerated virtual screening of large chemical libraries (Deep Docking) · Bioinform. 2022
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemogenomics
0.412020
DeepCOP: deep learning-based approach to predict gene regulating effects of small molecules · Bioinform. 2020
Bioinformatics and computational biology › drug discovery
computational drug discovery
0.412020
DeepCOP: deep learning-based approach to predict gene regulating effects of small molecules · Bioinform. 2020
Bioinformatics and computational biology › drug discovery
drug screening
0.412020
DeepCOP: deep learning-based approach to predict gene regulating effects of small molecules · Bioinform. 2020
Bioinformatics and computational biology › drug discovery
drug design
0.312025
Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation · ICML 2025
Medical and health informatics › precision medicine
precision oncology
0.112020
DeepCOP: deep learning-based approach to predict gene regulating effects of small molecules · Bioinform. 2020
Bioinformatics and computational biology › drug discovery › peptide drug discovery
antimicrobial peptide discovery
0.112007
AMPer: a database and an automated discovery tool for antimicrobial peptides · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
profile hidden markov model
0.112007
AMPer: a database and an automated discovery tool for antimicrobial peptides · Bioinform. 2007
Bioinformatics and computational biology
protein function prediction
0.112007
AMPer: a database and an automated discovery tool for antimicrobial peptides · Bioinform. 2007
Bioinformatics and computational biology
sequence analysis
0.112007
AMPer: a database and an automated discovery tool for antimicrobial peptides · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

proxy rewards · 1.7molecular descriptors · 1.7molecular docking · 0.6deep learning · 0.6molecular fingerprint · 0.4gene ontology descriptors · 0.4deep neural network · 0.4sequence clustering · 0.1hidden markov model · 0.1
YearPublicationVenuePosition
2025 Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation
abstract
Generative Flow Networks (GFlowNets) have recently emerged as a suitable framework for generating diverse and high-quality molecular structures by learning from rewards treated as unnormalized distributions. Previous works in this framework often restrict exploration by using predefined molecular fragments as building blocks, limiting the chemical space that can be accessed. In this work, we introduce Atomic GFlowNets (A-GFNs), a foundational generative model leveraging individual atoms as building blocks to explore drug-like chemical space more comprehensively. We propose an unsupervised pre-training approach using drug-like molecule datasets, which teaches A-GFNs about inexpensive yet informative molecular descriptors such as drug-likeliness, topological polar surface area, and synthetic accessibility scores. These properties serve as proxy rewards, guiding A-GFNs towards regions of chemical space that exhibit desirable pharmacological properties. We further implement a goal-conditioned finetuning process, which adapts A-GFNs to optimize for specific target properties. In this work, we pretrain A-GFN on a subset of ZINC dataset, and by employing robust evaluation metrics we show the effectiveness of our approach when compared to other relevant baseline methods for a wide range of drug design tasks. The code is accessible at https://github.com/diamondspark/AGFN.
Mohit Pandey, Gopeshh Subbaraj, Artem Cherkasov, Martin Ester, Emmanuel Bengio
ICML3
2022 DD-GUI: a graphical user interface for deep learning-accelerated virtual screening of large chemical libraries (Deep Docking)
abstract
SUMMARY: Deep learning (DL) can significantly accelerate virtual screening of ultra-large chemical libraries, enabling the evaluation of billions of compounds at a fraction of the computational cost and time required by conventional docking. Here, we introduce DD-GUI, the graphical user interface for such DL approach we have previously developed, termed Deep Docking (DD). The DD-GUI allows for quick setups of large-scale virtual screens in an intuitive way, and provides convenient tools to track the progress and analyze the outcomes of a drug discovery project. AVAILABILITY AND IMPLEMENTATION: DD-GUI is freely available with an MIT license on GitHub at https://github.com/jamesgleave/DeepDockingGUI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jean Charle Yaacoub, James Gleave, Francesco Gentile, Abraham C. Stern, Artem Cherkasov
Bioinform.5
2020 DeepCOP: deep learning-based approach to predict gene regulating effects of small molecules
abstract
MOTIVATION: Recent advances in the areas of bioinformatics and chemogenomics are poised to accelerate the discovery of small molecule regulators of cell development. Combining large genomics and molecular data sources with powerful deep learning techniques has the potential to revolutionize predictive biology. In this study, we present Deep gene COmpound Profiler (DeepCOP), a deep learning based model that can predict gene regulating effects of low-molecular weight compounds. This model can be used for direct identification of a drug candidate causing a desired gene expression response, without utilizing any information on its interactions with protein target(s). RESULTS: In this study, we successfully combined molecular fingerprint descriptors and gene descriptors (derived from gene ontology terms) to train deep neural networks that predict differential gene regulation endpoints collected in LINCS database. We achieved 10-fold cross-validation RAUC scores of and above 0.80, as well as enrichment factors of >5. We validated our models using an external RNA-Seq dataset generated in-house that described the effect of three potent antiandrogens (with different modes of action) on gene expression in LNCaP prostate cancer cell line. The results of this pilot study demonstrate that deep learning models can effectively synergize molecular and genomic descriptors and can be used to screen for novel drug candidates with the desired effect on gene expression. We anticipate that such models can find a broad use in developing novel cancer therapeutics and can facilitate precision oncology efforts. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Godwin Woo, Michael Fernández, Michael Hsing, Nathan A. Lack, Ayse Derya Cavga, Artem Cherkasov
Bioinform.6
2008 The Relation between Indel Length and Functional Divergence: A Formal Study
Raheleh Salari, Alexander Schönhuth, Fereydoun Hormozdiari, Artem Cherkasov, Süleyman Cenk Sahinalp
WABI4
2008 Indel PDB: A database of structural insertions and deletions derived from sequence alignments of closely related proteins
abstract
BACKGROUND: Insertions and deletions (indels) represent a common type of sequence variations, which are less studied and pose many important biological questions. Recent research has shown that the presence of sizable indels in protein sequences may be indicative of protein essentiality and their role in protein interaction networks. Examples of utilization of indels for structure-based drug design have also been recently demonstrated. Nonetheless many structural and functional characteristics of indels remain less researched or unknown. DESCRIPTION: We have created a web-based resource, Indel PDB, representing a structural database of insertions/deletions identified from the sequence alignments of highly similar proteins found in the Protein Data Bank (PDB). Indel PDB utilized large amounts of available structural information to characterize 1-, 2- and 3-dimensional features of indel sites. Indel PDB contains 117,266 non-redundant indel sites extracted from 11,294 indel-containing proteins. Unlike loop databases, Indel PDB features more indel sequences with secondary structures including alpha-helices and beta-sheets in addition to loops. The insertion fragments have been characterized by their sequences, lengths, locations, secondary structure composition, solvent accessibility, protein domain association and three dimensional structures. CONCLUSION: By utilizing the data available in Indel PDB, we have studied and presented here several sequence and structural features of indels. We anticipate that Indel PDB will not only enable future functional studies of indels, but will also assist protein modeling efforts and identification of indel-directed drug binding sites.
Michael Hsing, Artem Cherkasov
BMC Bioinform.2
2007 AMPer: a database and an automated discovery tool for antimicrobial peptides
abstract
MOTIVATION: Increasing antibiotics resistance in human pathogens represents a pressing public health issue worldwide for which novel antibiotic therapies based on antimicrobial peptides (AMPs) may offer one possible solution. In the current study, we utilized publicly available data on AMPs to construct hidden Markov models (HMMs) that enable recognition of individual classes of antimicrobials peptides (such as defensins, cathelicidins, cecropins, etc.) with up to 99% accuracy and can be used for discovering novel AMP candidates. RESULTS: HMM models for both mature peptides and propeptides were constructed. A total of 146 models for mature peptides and 40 for propeptides have been developed for individual AMP classes. These were created by clustering and analyzing AMP sequences available in the public sources and by consequent iterative scanning of the Swiss-Prot database for previously unknown gene-coded AMPs. As a result, an additional 229 additional AMPs have been identified from Swiss-Prot, and all but 34 could be associated with known antimicrobial activities according to the literature. The final set of 1045 mature peptides and 253 propeptides have been organized into the open-source AMPer database. AVAILABILITY: The developed HMM-based tools and AMP sequences can be accessed through the AMPer resource at http://www.cnbi2.com/cgi-bin/amp.pl
Christopher D. Fjell, Robert E. W. Hancock, Artem Cherkasov
Bioinform.3
2007 Relationship between insertion/deletion (indel) frequency of proteins and essentiality
abstract
BACKGROUND: In a previous study, we demonstrated that some essential proteins from pathogenic organisms contained sizable insertions/deletions (indels) when aligned to human proteins of high sequence similarity. Such indels may provide sufficient spatial differences between the pathogenic protein and human proteins to allow for selective targeting. In one example, an indel difference was targeted via large scale in-silico screening. This resulted in selective antibodies and small compounds which were capable of binding to the deletion-bearing essential pathogen protein without any cross-reactivity to the highly similar human protein. The objective of the current study was to investigate whether indels were found more frequently in essential than non-essential proteins. RESULTS: We have investigated three species, Bacillus subtilis, Escherichia coli, and Saccharomyces cerevisiae, for which high-quality protein essentiality data is available. Using these data, we demonstrated with t-test calculations that the mean indel frequencies in essential proteins were greater than that of non-essential proteins in the three proteomes. The abundance of indels in both types of proteins was also shown to be accurately modeled by the Weibull distribution. However, Receiver Operator Characteristic (ROC) curves showed that indel frequencies alone could not be used as a marker to accurately discriminate between essential and non-essential proteins in the three proteomes. Finally, we analyzed the protein interaction data available for S. cerevisiae and observed that indel-bearing proteins were involved in more interactions and had greater betweenness values within Protein Interaction Networks (PINs). CONCLUSION: Overall, our findings demonstrated that indels were not randomly distributed across the studied proteomes and were likely to occur more often in essential proteins and those that were highly connected, indicating a possible role of sequence insertions and deletions in the regulation and modification of protein-protein interactions. Such observations will provide new insights into indel-based drug design using bioinformatics and cheminformatics tools.
Simon K. Chan, Michael Hsing, Fereydoun Hormozdiari, Artem Cherkasov
BMC Bioinform.4
2004 Structural characterization of genomes by large scale sequence-structure threading
abstract
BACKGROUND: Using sequence-structure threading we have conducted structural characterization of complete proteomes of 37 archaeal, bacterial and eukaryotic organisms (including worm, fly, mouse and human) totaling 167,888 genes. RESULTS: The reported data represent first rather general evaluation of performance of full sequence-structure threading on multiple genomes providing opportunity to evaluate its general applicability for large scale studies. According to the estimated results the sequence-structure threading has assigned protein folds to more then 60% of eukaryotic, 68% of archaeal and 70% of bacterial proteomes.The repertoires of protein classes, architectures, topologies and homologous superfamilies (according to the CATH 2.4 classification) have been established for distant organisms and superkingdoms. It has been found that the average abundance of CATH classes decreases from "alpha and beta" to "mainly beta", followed by "mainly alpha" and "few secondary structures".3-Layer (aba) Sandwich has been characterized as the most abundant protein architecture and Rossman fold as the most common topology. CONCLUSION: The analysis of genomic occurrences of CATH 2.4 protein homologous superfamilies and topologies has revealed the power-law character of their distributions. The corresponding double logarithmic "frequency - genomic occurrence" dependences characteristic of scale-free systems have been established for individual organisms and for three superkingdoms.
Artem Cherkasov, Steven J. M. Jones
BMC Bioinform.1
2004 An approach to large scale identification of non-obvious structural similarities between proteins
abstract
BACKGROUND: A new sequence independent bioinformatics approach allowing genome-wide search for proteins with similar three dimensional structures has been developed. By utilizing the numerical output of the sequence threading it establishes putative non-obvious structural similarities between proteins. When applied to the testing set of proteins with known three dimensional structures the developed approach was able to recognize structurally similar proteins with high accuracy. RESULTS: The method has been developed to identify pathogenic proteins with low sequence identity and high structural similarity to host analogues. Such protein structure relationships would be hypothesized to arise through convergent evolution or through ancient horizontal gene transfer events, now undetectable using current sequence alignment techniques. The pathogen proteins, which could mimic or interfere with host activities, would represent candidate virulence factors. The developed approach utilizes the numerical outputs from the sequence-structure threading. It identifies the potential structural similarity between a pair of proteins by correlating the threading scores of the corresponding two primary sequences against the library of the standard folds. This approach allowed up to 64% sensitivity and 99.9% specificity in distinguishing protein pairs with high structural similarity. CONCLUSION: Preliminary results obtained by comparison of the genomes of Homo sapiens and several strains of Chlamydia trachomatis have demonstrated the potential usefulness of the method in the identification of bacterial proteins with known or potential roles in virulence.
Artem Cherkasov, Steven J. M. Jones
BMC Bioinform.1
2004 Structural characterization of genomes by large scale sequence-structure threading: application of reliability analysis in structural genomics
abstract
BACKGROUND: We establish that the occurrence of protein folds among genomes can be accurately described with a Weibull function. Systems which exhibit Weibull character can be interpreted with reliability theory commonly used in engineering analysis. For instance, Weibull distributions are widely used in reliability, maintainability and safety work to model time-to-failure of mechanical devices, mechanisms, building constructions and equipment. RESULTS: We have found that the Weibull function describes protein fold distribution within and among genomes more accurately than conventional power functions which have been used in a number of structural genomic studies reported to date. It has also been found that the Weibull reliability parameter beta for protein fold distributions varies between genomes and may reflect differences in rates of gene duplication in evolutionary history of organisms. CONCLUSIONS: The results of this work demonstrate that reliability analysis can provide useful insights and testable predictions in the fields of comparative and structural genomics.
Artem Cherkasov, Shannan J. Ho Sui, Robert C. Brunham, Steven J. M. Jones
BMC Bioinform.1
2004 Modeling of cell signaling pathways in macrophages by semantic networks
abstract
BACKGROUND: Substantial amounts of data on cell signaling, metabolic, gene regulatory and other biological pathways have been accumulated in literature and electronic databases. Conventionally, this information is stored in the form of pathway diagrams and can be characterized as highly "compartmental" (i.e. individual pathways are not connected into more general networks). Current approaches for representing pathways are limited in their capacity to model molecular interactions in their spatial and temporal context. Moreover, the critical knowledge of cause-effect relationships among signaling events is not reflected by most conventional approaches for manipulating pathways. RESULTS: We have applied a semantic network (SN) approach to develop and implement a model for cell signaling pathways. The semantic model has mapped biological concepts to a set of semantic agents and relationships, and characterized cell signaling events and their participants in the hierarchical and spatial context. In particular, the available information on the behaviors and interactions of the PI3K enzyme family has been integrated into the SN environment and a cell signaling network in human macrophages has been constructed. A SN-application has been developed to manipulate the locations and the states of molecules and to observe their actions under different biological scenarios. The approach allowed qualitative simulation of cell signaling events involving PI3Ks and identified pathways of molecular interactions that led to known cellular responses as well as other potential responses during bacterial invasions in macrophages. CONCLUSIONS: We concluded from our results that the semantic network is an effective method to model cell signaling pathways. The semantic model allows proper representation and integration of information on biological structures and their interactions at different levels. The reconstruction of the cell signaling network in the macrophage allowed detailed investigation of connections among various essential molecules and reflected the cause-effect relationships among signaling events. The simulation demonstrated the dynamics of the semantic network, where a change of states on a molecule can alter its function and potentially cause a chain-reaction effect in the system.
Michael Hsing, Joel L. Bellenson, Conor Shankey, Artem Cherkasov
BMC Bioinform.4