VLDB 2026 Research / reviewers in the wild / expert
Mikhail S. Gelfand
dblp:11/6691
· DBLP profile ↗
34ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0003-4181-0846ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On "Bioinformatics in Russia: history and present-day landscape" by M.A. Nawaz, I.E. Pamirsky, and K.S. GolokhvastabstractDear Editors As an author of several papers in ‘Briefings in Bioinformatics’ and a reviewer for the journal, I’ve been surprised by the paper “Bioinformatics in Russia: history and present-day landscape” by M.A. Nawaz, I.E. Pamirsky, and K.S. Golokhvast (vol. 25, no. 6, p. bbae513). While the topic is clearly important, the paper does a disservice both to readers of Briefings in Bioinformatics and to Russian bioinformaticians. It provides random information (e.g. the entire chapter about the market state of IT, pharmaceutical, agrotechnology, and biotechnology industries), while giving no real insight into the history and problems of Russian bioinformatics, both applied (e.g. related to medical genetics) and fundamental (that is, molecular evolution, about which nothing is said at all). The choice of scientific topics to discuss seems to follow lists of publications at Internet sites of several universities and research institutes with no comprehensive overview. As I’ve been involved in several of the activities mentioned by Nawaz et al., and as I’m a co-author of more than dozen papers they reference (12, 39–43, 60–62, 144, 184, 186, 192–193), with even more referenced papers by my former students, I think I’m in a good position to note specific errors and omissions. This is not an exhaustive review, as I’ll mainly restrict myself to the facts, events, and results about which I have a first-hand knowledge. The history narrated by Nawaz et al. is rather incomplete. They mention (correctly) the Institute of Mathematical Biology in Pushchino (where the group of Alexei Kondrashov worked) and the Institute of Cytology and Genetics in Novosibirsk (where bioinformatics was pioneered by Vadim Ratner), but neglect early contributions by groups that ceased to exist or changed location, in particular, labs of Andrei Mironov at the State Research Institute of Genetics and Selection of Industrial Microorganisms (now at the Lomonosov Moscow State University) and Aleksandr Aleksandrov at the Institute of Molecular Genetics. Nothing is said about the Bioinformatics section of the Russian Human Genome Program whose conferences in the 90’s collected contributions of almost all active groups. A crucial role in integrating the bioinformatics community was played by international conferences, ‘Bioinformatics of Genome Structure and Regulation” (even years since 1998, Novosibirsk) and ‘Moscow Conference on Computational Molecular Biology’ (odd years since 2003), and the regular ‘Moscow Bioinformatics Seminar’ (biweekly since 1993 until early 10’s, then less regular; I believe this seminar had been a unique phenomenon not only in Russian, but in the World bioinformatics). Important but not mentioned were contacts with the Russian bioinformatics diaspora, in particular (but by no means limited to) Eugene Koonin (National Center for Biotechnology Information), Leonid Mirny (Massachasetts Institute of Technology), Shamil Sunyaev, Vadim Gladyshev, and Peter Kharchenko (Harvard Medical School), Pavel Pevzner (University of Calfornia, San Diego), Oleg Gusev (Juntendo); the latter four at some point had led labs in Russia. I’m flattered that the entire ‘Information Box II’ is dedicated to the work of the Research and Training Center ‘Bioinformatics’ (RTCB) of the Institute for Information Transmission Problems that I have created in 2003 and led until this year, but the choice of papers and results to mention is random: too many minor papers are mentioned while really important ones (e.g. on riboswitches) are missing; the listed Internet resources have been mainly developed by my former students in the US, not in Russia; and Ref. 190 is not related to RTCB, nor even to Russia, as all its authors work in France and none is of the Russian descent. The Genomics Core Facility of the Skolkovo Institute of Science and Technology (SkolTech) performs sequencing, but does not provide bioinformatics service, contrary to what is claimed by Nawaz et al. As in the previous case, the choice of results by Skoltech researchers is capricious, with entire successful research directions completely omitted (e.g. comparative systems biology of the brain or large-scale metagenomics). On the other hand, Skoltech is not involved in the project to sequence 100 000 human genomes aside of the fact that the latter is led by a former SkolTech professor Konstantin Severinov. “Institute of Protein of the USSR Academy of Sciences” and “Protein Institute of the RAS” is commonly known as the “Institute of Protein Research”. Hence it is not surprising that, as reported by Nawaz et al., their ‘literature search produced limited results’; in particular they have not noticed fundamental contributions by the protein structure school created by Oleg Ptitsyn and maintained by Alexei Finkelstein. Nawaz et al. mention bioinformatics programs in a number of universities, but merely list the flagship Faculty of Bioengineering and Bioinformatics at the Lomonosov Moscow State University (established in 2002) and Department of Information Biology at the Novosibirsk State University in the Appendix. To that I add a number of other observations. The most interesting one is “the founding of the Academy of Sciences of the USSR (abbreviated as AH CCCP in Russian) by Pei Kib Breat in 1725”. This mysterious person is the Russian tsar Peter the Great, he would had been surprised to learn back in 1725 about the coming of the Soviet Union in about 200 years (It has been brought to my attention by colleagues that the strange phrase originates from a text recognition error in the electronic copy of a ‘Nature’ note published in 1946 [1]) The authors wrote, “George Gamow optimistically proposed a fairly precise genetic code in a letter to Watson and Crick [33]” — in fact, he did not; Gamow’s major contribution had been understanding that a code linking DNA and proteins should exist, while the code he had suggested was very far from reality, as explicitly described in the cited Ref. 33. Ref. 8 has an interesting author “Union IT” and an editor “Center GI”. Ref. 125 completely and Ref. 155 partially are given in Russian Cyrillic (it is not clear how they have survived the editorial process). Similarly, “the Scientific Research Computing Centre of the AH CCCP” is, of course, “… of the USSR Academy of Sciences”. Three out of four first references (refs. 1, 2, and 4) are not relevant to the papers’ topic; this is just one small example. “Information Box – 1” does not deal with bioinformatics at all, while important papers on the SARS-CoV-2 evolution by Georgii Bazykin’s group are not mentioned. The authors mention that large-scale projects on sequencing chickpea, soybean, and rice varieties are not realized yet — but how does that relate to their topic, Russian bioinformatics? — the references are given to the international projects. Nawaz et al. notice an increase in the number of publications since 2014 and assign it to the establishment of the Russian Science Foundation. A more likely explanation is one of the 2012 decrees of then just re-elected President Vladimir Putin calling for an increase of the Russian share in the international bibliometrics. That percolated downstream and, in many universities, led to establishment of so-called ‘publication bonuses’ of considerable value, hence inflating the publication rates. Similarly, Nawaz et al. write that “it is somewhat understandable from the literature survey that relatively more focus has been given to prokaryotic genome sequencing rather than eukaryotes”. Indeed, the former is orders of magnitude cheaper than the latter, which is an important, yet unmentioned factor in Russia. It is surprising that this paper has passed a strict reviewing process in a respected, high-impact journal. None declared. Mikhail S. Gelfand |
Briefings Bioinform. | 1 |
| 2023 | HiConfidence: a novel approach uncovering the biological signal in Hi-C data affected by technical biasesabstractThe chromatin interaction assays, particularly Hi-C, enable detailed studies of genome architecture in multiple organisms and model systems, resulting in a deeper understanding of gene expression regulation mechanisms mediated by epigenetics. However, the analysis and interpretation of Hi-C data remain challenging due to technical biases, limiting direct comparisons of datasets obtained in different experiments and laboratories. As a result, removing biases from Hi-C-generated chromatin contact matrices is a critical data analysis step. Our novel approach, HiConfidence, eliminates biases from the Hi-C data by weighing chromatin contacts according to their consistency between replicates so that low-quality replicates do not substantially influence the result. The algorithm is effective for the analysis of global changes in chromatin structures such as compartments and topologically associating domains. We apply the HiConfidence approach to several Hi-C datasets with significant technical biases, that could not be analyzed effectively using existing methods, and obtain meaningful biological conclusions. In particular, HiConfidence aids in the study of how changes in histone acetylation pattern affect chromatin organization in Drosophila melanogaster S2 cells. The method is freely available at GitHub: https://github.com/victorykobets/HiConfidence. Victoria A Kobets, Sergey V. Ulianov, Aleksandra A. Galitsyna, Semen A. Doronin, Elena A. Mikhaleva, Mikhail S. Gelfand, Yuri Y. Shevelyov, Sergey V. Razin, Ekaterina Khrameeva |
Briefings Bioinform. | 6 |
| 2021 | Single-cell Hi-C data analysis: safety in numbersabstractOver the past decade, genome-wide assays for chromatin interactions in single cells have enabled the study of individual nuclei at unprecedented resolution and throughput. Current chromosome conformation capture techniques survey contacts for up to tens of thousands of individual cells, improving our understanding of genome function in 3D. However, these methods recover a small fraction of all contacts in single cells, requiring specialised processing of sparse interactome data. In this review, we highlight recent advances in methods for the interpretation of single-cell genomic contacts. After discussing the strengths and limitations of these methods, we outline frontiers for future development in this rapidly moving field. Aleksandra A. Galitsyna, Mikhail S. Gelfand |
Briefings Bioinform. | 2 |
| 2021 | Perspectives for the reconstruction of 3D chromatin conformation using single cell Hi-C dataabstractConstruction of chromosomes 3D models based on single cell Hi-C data constitute an important challenge. We present a reconstruction approach, DPDchrom, that incorporates basic knowledge whether the reconstructed conformation should be coil-like or globular and spring relaxation at contact sites. In contrast to previously published protocols, DPDchrom can naturally form globular conformation due to the presence of explicit solvent. Benchmarking of this and several other methods on artificial polymer models reveals similar reconstruction accuracy at high contact density and DPDchrom advantage at low contact density. To compare 3D structures insensitively to spatial orientation and scale, we propose the Modified Jaccard Index. We analyzed two sources of the contact dropout: contact radius change and random contact sampling. We found that the reconstruction accuracy exponentially depends on the number of contacts per genomic bin allowing to estimate the reconstruction accuracy in advance. We applied DPDchrom to model chromosome configurations based on single-cell Hi-C data of mouse oocytes and found that these configurations differ significantly from a random one, that is consistent with other studies. Pavel Kos, Aleksandra A. Galitsyna, Sergey V. Ulianov, Mikhail S. Gelfand, Sergey V. Razin, Alexander Chertovich |
PLoS Comput. Biol. | 4 |
| 2020 | HiChew: a Tool for TAD Clustering in Embryogenesis
Nikolai S. Bykov, Olga M. Sigalova, Mikhail S. Gelfand, Aleksandra A. Galitsyna |
ISBRA | 3 |
| 2018 | Reconstruction of the chromatin 3D conformation from single cell Hi-C data
Pavel Kos, Aleksandra A. Galitsyna, Sergey V. Ulianov, Mikhail S. Gelfand, Sergey V. Razin, Alexander Chertovich |
BIBM | 4 |
| 2018 | Clustering and Comparison of Hierarchies in the Spatial Organization of Chromatin
Olga Pushkareva, Alexander Kurashenko, Uliana Moskvina, Anatoly R. Rubinov, Mikhail S. Gelfand |
BIBM | 5 |
| 2018 | Prediction of 3D Chromatin Structure Using Recurrent Neural Networks
Michal Rozenwald, Ekaterina Khrameeva, Grigory V. Sapunov, Mikhail S. Gelfand |
BIBM | 4 |
| 2018 | Prediction of chromatin spatial structure characteristics using machine learning methods
Sergei Starikov, Ekaterina Khrameeva, Mikhail S. Gelfand |
BIBM | 3 |
| 2018 | The chromatin structure of Dictyostelium discoideum
Olga Tsoy, Aleksandra A. Galitsyna, Ekaterina Khrameeva, Sergey V. Ulianov, Mikhail S. Gelfand, Sergey V. Razin |
BIBM | 5 |
| 2018 | Nuclear lamina maintains global spatial organization of chromatin in Drosophila cultured cells
Sergey V. Ulianov, Semen A. Doronin, Ekaterina Khrameeva, Pavel Kos, Sergey S. Starikov, Aleksandra A. Galitsyna, Artem Luzhin, Mikhail S. Gelfand, Alexander Chertovich, Sergey V. Razin, Yuri Y. Shevelyov |
BIBM | 8 |
| 2016 | Selectoscope: A Modern Web-App for Positive Selection Analysis of Genomic Data
Andrey V. Zaika, Iakov I. Davydov, Mikhail S. Gelfand |
ISBRA | 3 |
| 2014 | ANA HEp-2 cells image classification using number, size, shape and localization of targeted cell regions
Gennady V. Ponomarev, Vladimir L. Arlazarov, Mikhail S. Gelfand, Marat D. Kazanov |
Pattern Recognit. | 3 |
| 2014 | Evaluation and Comparison of Current Fetal Ultrasound Image Segmentation Methods for Biometric Measurements: A Grand ChallengeabstractThis paper presents the evaluation results of the methods submitted to Challenge US: Biometric Measurements from Fetal Ultrasound Images, a segmentation challenge held at the IEEE International Symposium on Biomedical Imaging 2012. The challenge was set to compare and evaluate current fetal ultrasound image segmentation methods. It consisted of automatically segmenting fetal anatomical structures to measure standard obstetric biometric parameters, from 2D fetal ultrasound images taken on fetuses at different gestational ages (21 weeks, 28 weeks, and 33 weeks) and with varying image quality to reflect data encountered in real clinical environments. Four independent sub-challenges were proposed, according to the objects of interest measured in clinical practice: abdomen, head, femur, and whole fetus. Five teams participated in the head sub-challenge and two teams in the femur sub-challenge, including one team who tackled both. Nobody attempted the abdomen and whole fetus sub-challenges. The challenge goals were two-fold and the participants were asked to submit the segmentation results as well as the measurements derived from the segmented objects. Extensive quantitative (region-based, distance-based, and Bland-Altman measurements) and qualitative evaluation was performed to compare the results from a representative selection of current methods submitted to the challenge. Several experts (three for the head sub-challenge and two for the femur sub-challenge), with different degrees of expertise, manually delineated the objects of interest to define the ground truth used within the evaluation framework. For the head sub-challenge, several groups produced results that could be potentially used in clinical settings, with comparable performance to manual delineations. The femur sub-challenge had inferior performance to the head sub-challenge due to the fact that it is a harder segmentation problem and that the techniques presented relied more on the femur's appearance. Sylvia Rueda, Sana Fathima, Caroline L. Knight, Mohammad Yaqub, Aris T. Papageorghiou, Bahbibi Rahmatullah, Alessandro Foi, Matteo Maggioni, Antonietta Pepe, Jussi Tohka, Richard V. Stebbing, John McManigle, Anca Ciurte, Xavier Bresson, Meritxell Bach Cuadra, Changming Sun, Gennady V. Ponomarev, Mikhail S. Gelfand, Marat D. Kazanov, Ching-Wei Wang, Hsiang-Chou Chen, Chun-Wei Peng, Chu-Mei Hung, J. Alison Noble |
IEEE Trans. Medical Imaging | 18 |
| 2012 | Evidence for Widespread Association of Mammalian Splicing and Conserved Long-Range RNA Structures
Dmitri D. Pervouchine, Ekaterina Khrameeva, Marina Pichugina, Olexii Nikolaienko, Mikhail S. Gelfand, Petr Rubtsov, Andrey A. Mironov |
RECOMB | 5 |
| 2012 | Biases in read coverage demonstrated by interlaboratory and interplatform comparison of 117 mRNA and genome sequencing experimentsabstractHigh-throughput sequencing of whole genomes and transcriptomes allows one to generate large amounts of sequence data very rapidly and at a low cost. The goal of most mRNA sequencing studies is to perform the comparison of the expression level between different samples. However, given a broad variety of modern sequencing protocols, platforms and versions thereof, it is not clear to what extent the obtained results are consistent across platforms and laboratories. The comparison of 117 human mRNA and genome high-throughput sequencing experiments performed on the Illumina and SOLiD platforms at 26 institutions all over the world demonstrated high dependency of the gene coverage profiles on the producing laboratory. Gene coverage profiles showed laboratory-specific non-uniformity that survived the 3'-bias correction and mappability normalization, suggesting that there are other yet unknown mRNA-associated biases. Ekaterina Khrameeva, Mikhail S. Gelfand |
BMC Bioinform. | 2 |
| 2009 | Evolution of Regulatory Systems in Bacteria (Invited Keynote Talk)
Mikhail S. Gelfand, Alexei E. Kazakov, Yuri D. Korostelev, Olga N. Laikova, Andrey A. Mironov, Aleksandra B. Rakhmaninova, Dmitry A. Ravcheev, Dmitry A. Rodionov, Alexei G. Vitreschak |
ISBRA | 1 |
| 2009 | Combining specificity determining and conserved residues improves functional site predictionabstractBACKGROUND: Predicting the location of functionally important sites from protein sequence and/or structure is a long-standing problem in computational biology. Most current approaches make use of sequence conservation, assuming that amino acid residues conserved within a protein family are most likely to be functionally important. Most often these approaches do not consider many residues that act to define specific sub-functions within a family, or they make no distinction between residues important for function and those more relevant for maintaining structure (e.g. in the hydrophobic core). Many protein families bind and/or act on a variety of ligands, meaning that conserved residues often only bind a common ligand sub-structure or perform general catalytic activities. RESULTS: Here we present a novel method for functional site prediction based on identification of conserved positions, as well as those responsible for determining ligand specificity. We define Specificity-Determining Positions (SDPs), as those occupied by conserved residues within sub-groups of proteins in a family having a common specificity, but differ between groups, and are thus likely to account for specific recognition events. We benchmark the approach on enzyme families of known 3D structure with bound substrates, and find that in nearly all families residues predicted by SDPsite are in contact with the bound substrate, and that the addition of SDPs significantly improves functional site prediction accuracy. We apply SDPsite to various families of proteins containing known three-dimensional structures, but lacking clear functional annotations, and discusse several illustrative examples. CONCLUSION: The results suggest a better means to predict functional details for the thousands of protein structures determined prior to a clear understanding of molecular function. Olga V. Kalinina, Mikhail S. Gelfand, Robert B. Russell |
BMC Bioinform. | 2 |
| 2008 | Identification of replication origins in prokaryotic genomesabstractThe availability of hundreds of complete bacterial genomes has created new challenges and simultaneously opportunities for bioinformatics. In the area of statistical analysis of genomic sequences, the studies of nucleotide compositional bias and gene bias between strands and replichores paved way to the development of tools for prediction of bacterial replication origins. Only a few (about 20) origin regions for eubacteria and archaea have been proven experimentally. One reason for that may be that this is now considered as an essentially bioinformatics problem, where predictions are sufficiently reliable not to run labor-intensive experiments, unless specifically needed. Here we describe the main existing approaches to the identification of replication origin (oriC) and termination (terC) loci in prokaryotic chromosomes and characterize a number of computational tools based on various skew types and other types of evidence. We also classify the eubacterial and archaeal chromosomes by predictability of their replication origins using skew plots. Finally, we discuss possible combined approaches to the identification of the oriC sites that may be used to improve the prediction tools, in particular, the analysis of DnaA binding sites using the comparative genomic methods. Natalia V. Sernova, Mikhail S. Gelfand |
Briefings Bioinform. | 2 |
| 2006 | Computational Reconstruction of Iron- and Manganese-Responsive Transcriptional Networks in α-ProteobacteriaabstractWe used comparative genomics to investigate the distribution of conserved DNA-binding motifs in the regulatory regions of genes involved in iron and manganese homeostasis in alpha-proteobacteria. Combined with other computational approaches, this allowed us to reconstruct the metal regulatory network in more than three dozen species with available genome sequences. We identified several classes of cis-acting regulatory DNA motifs (Irr-boxes or ICEs, RirA-boxes, Iron-Rhodo-boxes, Fur-alpha-boxes, Mur-box or MRS, MntR-box, and IscR-boxes) in regulatory regions of various genes involved in iron and manganese uptake, Fe-S and heme biosynthesis, iron storage, and usage. Despite the different nature of the iron regulons in selected lineages of alpha-proteobacteria, the overall regulatory network is consistent with, and confirmed by, many experimental observations. This study expands the range of genes involved in iron homeostasis and demonstrates considerable interconnection between iron-responsive regulatory systems. The detailed comparative and phylogenetic analyses of the regulatory systems allowed us to propose a theory about the possible evolution of Fe and Mn regulons in alpha-proteobacteria. The main evolutionary event likely occurred in the common ancestor of the Rhizobiales and Rhodobacterales, where the Fur protein switched to regulating manganese transporters (and hence Fur had become Mur). In these lineages, the role of global iron homeostasis was taken by RirA and Irr, two transcriptional regulators that act by sensing the physiological consequence of the metal availability rather than its concentration per se, and thus provide for more flexible regulation. Dmitry A. Rodionov, Mikhail S. Gelfand, Jonathan D. Todd, Andrew R. J. Curson, Andrew W. B. Johnston |
PLoS Comput. Biol. | 2 |
| 2005 | Mining sequence annotation databanks for association patternsabstractMOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm. Irena I. Artamonova, Goar Frishman, Mikhail S. Gelfand, Dmitrij Frishman |
Bioinform. | 3 |
| 2005 | A Gibbs sampler for identification of symmetrically structured, spaced DNA motifs with improved estimation of the signal lengthabstractMOTIVATION: Transcription regulatory protein factors often bind DNA as homo-dimers or hetero-dimers. Thus they recognize structured DNA motifs that are inverted or direct repeats or spaced motif pairs. However, these motifs are often difficult to identify owing to their high divergence. The motif structure included explicitly into the motif recognition algorithm improves recognition efficiency for highly divergent motifs as well as estimation of motif geometric parameters. RESULT: We present a modification of the Gibbs sampling motif extraction algorithm, SeSiMCMC (Sequence Similarities by Markov Chain Monte Carlo), which finds structured motifs of these types, as well as non-structured motifs, in a set of unaligned DNA sequences. It employs improved estimators of motif and spacer lengths. The probability that a sequence does not contain any motif is accounted for in a rigorous Bayesian manner. We have applied the algorithm to a set of upstream regions of genes from two Escherichia coli regulons involved in respiration. We have demonstrated that accounting for a symmetric motif structure allows the algorithm to identify weak motifs more accurately. In the examples studied, ArcA binding sites were demonstrated to have the structure of a direct spaced repeat, whereas NarP binding sites exhibited the palindromic structure. AVAILABILITY: The WWW interface of the program, its FreeBSD (4.0) and Windows 32 console executables are available at http://bioinform.genetika.ru/SeSiMCMC Alexander V. Favorov, Mikhail S. Gelfand, Anna V. Gerasimova, Dmitry A. Ravcheev, Andrey A. Mironov, Vsevolod J. Makeev |
Bioinform. | 2 |
| 2005 | Alternative splicing and protein functionabstractBACKGROUND: Alternative splicing is a major mechanism of generating protein diversity in higher eukaryotes. Although at least half, and probably more, of mammalian genes are alternatively spliced, it was not clear, whether the frequency of alternative splicing is the same in different functional categories. The problem is obscured by uneven coverage of genes by ESTs and a large number of artifacts in the EST data. RESULTS: We have developed a method that generates possible mRNA isoforms for human genes contained in the EDAS database, taking into account the effects of nonsense-mediated decay and translation initiation rules, and a procedure for offsetting the effects of uneven EST coverage. Then we computed the number of mRNA isoforms for genes from different functional categories. Genes encoding ribosomal proteins and genes in the category "Small GTPase-mediated signal transduction" tend to have fewer isoforms than the average, whereas the genes in the category "DNA replication and chromosome cycle" have more isoforms than the average. Genes encoding proteins involved in protein-protein interactions tend to be alternatively spliced more often than genes encoding non-interacting proteins, although there is no significant difference in the number of isoforms of alternatively spliced genes. CONCLUSION: Filtering for functional isoforms satisfying biological constraints and accounting for uneven EST coverage allowed us to describe differences in alternative splicing of genes from different functional categories. The observations seem to be consistent with expectations based on current biological knowledge: less isoforms for ribosomal and signal transduction proteins, and more alternative splicing of interacting and cell cycle proteins. A. D. Neverov, Irena I. Artamonova, Ramil N. Nurtdinov, Dmitrij Frishman, Mikhail S. Gelfand, Andrey A. Mironov |
BMC Bioinform. | 5 |
| 2005 | Dissimilatory Metabolism of Nitrogen Oxides in Bacteria: Comparative Reconstruction of Transcriptional NetworksabstractBacterial response to nitric oxide (NO) is of major importance since NO is an obligatory intermediate of the nitrogen cycle. Transcriptional regulation of the dissimilatory nitric oxides metabolism in bacteria is diverse and involves FNR-like transcription factors HcpR, DNR, and NnrR; two-component systems NarXL and NarQP; NO-responsive activator NorR; and nitrite-sensitive repressor NsrR. Using comparative genomics approaches, we predict DNA-binding motifs for these transcriptional factors and describe corresponding regulons in available bacterial genomes. Within the FNR family of regulators, we observed a correlation of two specificity-determining amino acids and contacting bases in corresponding DNA recognition motif. Highly conserved regulon HcpR for the hybrid cluster protein and some other redox enzymes is present in diverse anaerobic bacteria, including Clostridia, Thermotogales, and delta-proteobacteria. NnrR and DNR control denitrification in alpha- and beta-proteobacteria, respectively. Sigma-54-dependent NorR regulon found in some gamma- and beta-proteobacteria contains various enzymes involved in the NO detoxification. Repressor NsrR, which was previously known to control only nitrite reductase operon in Nitrosomonas spp., appears to be the master regulator of the nitric oxides' metabolism, not only in most gamma- and beta-proteobacteria (including well-studied species such as Escherichia coli), but also in gram-positive Bacillus and Streptomyces species. Positional analysis and comparison of regulatory regions of NO detoxification genes allows us to propose the candidate NsrR-binding motif. The most conserved member of the predicted NsrR regulon is the NO-detoxifying flavohemoglobin Hmp. In enterobacteria, the regulon also includes two nitrite-responsive loci, nipAB (hcp-hcr) and nipC (dnrN), thus confirming the identity of the effector, i.e. nitrite. The proposed NsrR regulons in Neisseria and some other species are extended to include denitrification genes. As the result, we demonstrate considerable interconnection between various nitrogen-oxides-responsive regulatory systems for the denitrification and NO detoxification genes and evolutionary plasticity of this transcriptional network. Dmitry A. Rodionov, Inna Dubchak, Adam P. Arkin, Eric J. Alm, Mikhail S. Gelfand |
PLoS Comput. Biol. | 5 |
| 2004 | No statistical support for correlation between the positions of protein interaction sites and alternatively spliced regionsabstractBACKGROUND: Alternative splicing is an efficient mechanism for increasing the variety of functions fulfilled by proteins in a living cell. It has been previously demonstrated that alternatively spliced regions often comprise functionally important and conserved sequence motifs. The objective of this work was to test the hypothesis that alternative splicing is correlated with contact regions of protein-protein interactions. RESULTS: Protein sequence spans involved in contacts with an interaction partner were delineated from atomic structures of transient interaction complexes and juxtaposed with the location of alternatively spliced regions detected by comparative genome analysis and spliced alignment. The total of 42 alternatively spliced isoforms were identified in 21 amino acid chains involved in biomolecular interactions. Using this limited dataset and a variety of sophisticated counting procedures we were not able to establish a statistically significant correlation between the positions of protein interaction sites and alternatively spliced regions. CONCLUSIONS: This finding contradicts a naïve hypothesis that alternatively spliced regions would correlate with points of contact. One possible explanation for that could be that all alternative splicing events change the spatial structure of the interacting domain to a sufficient degree to preclude interaction. This is indirectly supported by the observed lack of difference in the behaviour of relatively short regions affected by alternative splicing and cases when large portions of proteins are removed. More structural data on complexes of interacting proteins, including structures of alternative isoforms, are needed to test this conjecture. Marc N. Offman, Ramil N. Nurtdinov, Mikhail S. Gelfand, Dmitrij Frishman |
BMC Bioinform. | 3 |
| 2002 | Exact mapping of prokaryotic gene startsabstractIt is known that while the programs used to find genes in prokaryotic genomes reliably map protein-coding regions, they often fail in the exact determination of gene starts. This problem is further aggravated by sequencing errors, most notably insertions and deletions leading to frame-shifts. Therefore, the exact mapping of gene starts and identification of frame-shifts are important problems of the computer-assisted functional analysis of newly sequenced genomes. Here we review methods of gene recognition and describe a new algorithm for correction of gene starts and identification of frame-shifts in prokaryotic genomes. The algorithm is based on the comparison of nucleotide and protein sequences of homologous genes from related organisms, using the assumption that the rate of evolutionary changes in protein-coding regions is lower than that in non-coding regions. A dynamic programming algorithm is used to align protein sequences obtained by formal translation of genomic nucleotide sequences. The possibility of frame-shifts is taken into account. The algorithm was tested on several groups of related organisms: gamma-proteobacteria, the Bacillus/Clostridium group, and three Pyrococcus genomes. The testing demonstrated that, dependent or a genome, 1-10 per cent of genes have incorrect starts or contain frame-shifts. The algorithm is implemented in the program package Orthologator-GeneCorrector. M. V. Baytaluk, Mikhail S. Gelfand, Andrey A. Mironov |
Briefings Bioinform. | 2 |
| 2001 | Pro-Frame: similarity-based gene recognition in eukaryotic DNA sequences with errorsabstractAbstract Summary: Performance of existing algorithms for similarity-based gene recognition in eukaryotes drops when the genomic DNA has been sequenced with errors. A modification of the spliced alignment algorithm allows for gene recognition in sequences with errors, in particular frameshifts. It tolerates up to 5% of sequencing errors without considerable drop of prediction reliability when a sufficiently close homologous protein is available (normalized evolutionary distance similarity score 50% or higher). Availability: The program is free for academic users and available upon request at http://www.anchorgen.com Contact: [email protected] Andrey A. Mironov, Pavel S. Novichkov, Mikhail S. Gelfand |
Bioinform. | 3 |
| 2001 | Gene recognition in eukaryotic DNA by comparison of genomic sequencesabstractMOTIVATION: Sequencing of complete eukaryotic genomes and large syntenic fragments of genomes makes it possible to apply genomic comparison for gene recognition. RESULTS: This paper describes a spliced alignment algorithm that aligns candidate exon chains of two homologous genomic sequence fragments from different species. The algorithm is implemented in Pro-Gen software. Unlike other algorithms, Pro-Gen does not assume conservation of the exon-intron structure. Amino acid sequences obtained by the formal translation of candidate exons are aligned instead of nucleotide sequences, which allows for distant comparisons. The algorithm was tested on a sample of human-mammal (mouse), human-vertebrate (Xenopus ) and human-invertebrate (Drosophila ) gene pairs. Surprisingly, the best results, 97-98% correlation between the actual and predicted genes, were obtained for more distant comparisons, whereas the correlation on the human-mouse sample was only 93%. The latter value increases to 95% if conservation of the exon-intron structure is assumed. This is caused by a large amount of sequence conservation in non-coding regions of the human and mouse genes probably due to regulatory elements. AVAILABILITY: Pro-Gen v. 3.0 is available to academic researchers free of charge at http://www.anchorgen.com/pro_gen/pro_gen.html. Pavel S. Novichkov, Mikhail S. Gelfand, Andrey A. Mironov |
Bioinform. | 2 |
| 2000 | Comparative Analysis of Regulatory Patterns in Bacterial GenomesabstractRecognition of transcription regulatory sites in bacterial genomes is a notoriously difficult problem. There are no algorithms capable of making reliable predictions even for well-studied sites such as the CRP (cyclic AMP receptor protein) box. However, availability of complete bacterial genomes makes it possible to make reliable predictions with bad rules. This comparative approach is based on the assumption that sets of co-regulated genes are conserved in related bacteria. Thus true sites occur upstream of orthologous genes, whereas false candidates are scattered at random. This means not only that knowledge about regulation in well-studied genomes can be transferred to newly sequenced ones, but also that new members of regulons can be found. This paper reviews several recent studies. In particular, a detailed analysis of catabolite repression in gamma-purple bacteria is presented. Mikhail S. Gelfand, Pavel S. Novichkov, Elena S. Novichkova, Andrey A. Mironov |
Briefings Bioinform. | 1 |
| 1999 | Segmentation of yeast DNA using hidden Markov modelsabstractMOTIVATION: Compositionally homogeneous segments of genomic DNA often correspond to meaningful biological units. Simple sliding window analysis is usually insufficient for compositional segmentation of natural sequences. Hidden Markov models (HMM) with a small number of states are a natural language for description of compositional properties of chromosome-size DNA sequences. RESULTS: The algorithms were applied to yeast Saccharomyces cerevisiae chromosomes (YC) I, III, IV, VI and IX. The optimal number of HMM states is found to be four. The optimal four-state HMMs for all chromosomes are very similar, as well as the reconstructed segmentations. In most cases the models with k + 1 states are obtained by 'splitting' one of the states in the model with k states, and the corresponding increase of the level of detail in segmentation. The high AT states usually correspond to intergenic regions. We also explore the model's likelihood landscape and analyze the dynamics of the optimization process, thus addressing the problem of reliability of the obtained optima and efficiency of the algorithms. Leonid Peshkin, Mikhail S. Gelfand |
Bioinform. | 2 |
| 1998 | Algorithms and software for support of gene identification experimentsabstractMOTIVATION: Gene annotation is the final goal of gene prediction algorithms. However, these algorithms frequently make mistakes and therefore the use of gene predictions for sequence annotation is hardly possible. As a result, biologists are forced to conduct time-consuming gene identification experiments by designing appropriate PCR primers to test cDNA libraries or applying RT-PCR, exon trapping/amplification, or other techniques. This process frequently amounts to 'guessing' PCR primers on top of unreliable gene predictions and frequently leads to wasting of experimental efforts. RESULTS: The present paper proposes a simple and reliable algorithm for experimental gene identification which bypasses the unreliable gene prediction step. Studies of the performance of the algorithm on a sample of human genes indicate that an experimental protocol based on the algorithm's predictions achieves an accurate gene identification with relatively few PCR primers. Predictions of PCR primers may be used for exon amplification in preliminary mutation analysis during an attempt to identify a gene responsible for a disease. We propose a simple approach to find a short region from a genomic sequence that with high probability overlaps with some exon of the gene. The algorithm is enhanced to find one or more segments that are probably contained in the translated region of the gene and can be used as PCR primers to select appropriate clones in cDNA libraries by selective amplification. The algorithm is further extended to locate a set of PCR primers that uniformly cover all translated regions and can be used for RT-PCR and further sequencing of (unknown) mRNA. Sing-Hoi Sze, Mikhail A. Roytberg, Mikhail S. Gelfand, Andrey A. Mironov, Tatiana V. Astakhova, Pavel A. Pevzner |
Bioinform. | 3 |
| 1996 | Spliced Alignment: A New Approach to Gene Recognition
Mikhail S. Gelfand, Andrey A. Mironov, Pavel A. Pevzner |
CPM | 1 |
| 1995 | FANS-REF: a bibliography on statistics and functional analysis of nucleotide sequencesabstractM.S. Gelfand; FANS-REF: a bibliography on statistics and functional analysis of nucleotide sequences, Bioinformatics, Volume 11, Issue 5, 1 October 1995, Pages Mikhail S. Gelfand |
Comput. Appl. Biosci. | 1 |
| 1992 | Extendable words in nucleotide sequencesabstractPrevious statistical analyses revealed several peculiarities of nucleotide sequences that preclude their description by existing models and thus allow one to distinguish DNA and RNA sequences from random A,T,G,C-texts. This is a consequence of the unusual distribution of certain words in nucleotide sequences: while the distribution of (most) words is consistent with Markov models of small orders, the distribution of certain words cannot be described by any previous model (anomalies in distribution of homonucleotide/homopurine/homopyrimidine runs, complementary and mirror palindromes, and non-stationary words). In this work we introduce a probabilistic approach that is partly motivated by analogy with linguistics. We also describe another important feature of DNA/RNA sequences: anomalies in distribution of words of poor nucleotide composition. We show that some classes of these words are the major obstacle for the simple Markov description of nucleotide sequences. Mikhail S. Gelfand, C. G. Kozhukhin, Pavel A. Pevzner |
Comput. Appl. Biosci. | 1 |