Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Andrey A. Mironov

dblp:65/4184 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
10 papers
Bioinformatics and computational biology · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › gene expression analysis › gene expression quantification
allele-specific expression
0.712023
Foreign RNA spike-ins enable accurate allele-specific expression analysis at scale · Bioinform. 2023
Bioinformatics and computational biology
gene expression analysis
0.712023
Foreign RNA spike-ins enable accurate allele-specific expression analysis at scale · Bioinform. 2023
Bioinformatics and computational biology › gene regulation
regulatory genomics
0.312017
StereoGene: rapid estimation of genome-wide correlation of continuous or interval feature data · Bioinform. 2017
Bioinformatics and computational biology
RNA sequencing
0.212023
Foreign RNA spike-ins enable accurate allele-specific expression analysis at scale · Bioinform. 2023
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure
RNA secondary structure
0.212014
RNASurface: fast and accurate detection of locally optimal potentially structured RNA segments · Bioinform. 2014
Bioinformatics and computational biology › transcriptomics
alternative splicing analysis
0.112012
Evidence for Widespread Association of Mammalian Splicing and Conserved Long-Range RNA Structures · RECOMB 2012
Bioinformatics and computational biology › RNA biology
RNA processing
0.112012
Evidence for Widespread Association of Mammalian Splicing and Conserved Long-Range RNA Structures · RECOMB 2012
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA structure analysis
0.112012
Evidence for Widespread Association of Mammalian Splicing and Conserved Long-Range RNA Structures · RECOMB 2012
Bioinformatics and computational biology › genome annotation
gene prediction
0.132001
Gene recognition in eukaryotic DNA by comparison of genomic sequences · Bioinform. 2001
Pro-Frame: similarity-based gene recognition in eukaryotic DNA sequences with errors · Bioinform. 2001
Algorithms and software for support of gene identification experiments · Bioinform. 1998
Bioinformatics and computational biology › gene regulation › regulatory element discovery
regulatory element prediction
0.112005
A Gibbs sampler for identification of symmetrically structured, spaced DNA motifs with improved estimation of the signal length · Bioinform. 2005
Bioinformatics and computational biology
comparative genomics
0.012001
Gene recognition in eukaryotic DNA by comparison of genomic sequences · Bioinform. 2001
Bioinformatics and computational biology
genomics
0.012001
Pro-Frame: similarity-based gene recognition in eukaryotic DNA sequences with errors · Bioinform. 2001
Bioinformatics and computational biology › genome annotation › gene prediction
homology-based gene prediction
0.012001
Pro-Frame: similarity-based gene recognition in eukaryotic DNA sequences with errors · Bioinform. 2001
Bioinformatics and computational biology › sequence analysis › read mapping
spliced alignment
0.012001
Gene recognition in eukaryotic DNA by comparison of genomic sequences · Bioinform. 2001
Bioinformatics and computational biology
sequence analysis
0.011995
DNASUN: a package of computer programs for the biotechnology laboratory · Comput. Appl. Biosci. 1995
Bioinformatics and computational biology › genomics › genome analysis
genome mapping
0.011990
Mapping DNA by stochastic relaxation: a new approach to fragment sizes · Comput. Appl. Biosci. 1990
Bioinformatics and computational biology › sequence analysis › DNA sequence analysis
restriction mapping
0.011990
Mapping DNA by stochastic relaxation: a new approach to fragment sizes · Comput. Appl. Biosci. 1990
Bioinformatics and computational biology
genome annotation
0.011998
Algorithms and software for support of gene identification experiments · Bioinform. 1998

Methods — techniques the papers use, named apart from their topics

statistical estimation · 0.7spike-in normalization · 0.7partial correlation · 0.3kernel correlation · 0.3significance estimation · 0.2minimum free energy matrix · 0.2spliced alignment algorithm · 0.1markov chain monte carlo · 0.1gibbs sampling · 0.1bayesian estimation · 0.1
YearPublicationVenuePosition
2023 Foreign RNA spike-ins enable accurate allele-specific expression analysis at scale
abstract
MOTIVATION: Analysis of allele-specific expression is strongly affected by the technical noise present in RNA-seq experiments. Previously, we showed that technical replicates can be used for precise estimates of this noise, and we provided a tool for correction of technical noise in allele-specific expression analysis. This approach is very accurate but costly due to the need for two or more replicates of each library. Here, we develop a spike-in approach which is highly accurate at only a small fraction of the cost. RESULTS: We show that a distinct RNA added as a spike-in before library preparation reflects technical noise of the whole library and can be used in large batches of samples. We experimentally demonstrate the effectiveness of this approach using combinations of RNA from species distinguishable by alignment, namely, mouse, human, and Caenorhabditis elegans. Our new approach, controlFreq, enables highly accurate and computationally efficient analysis of allele-specific expression in (and between) arbitrarily large studies at an overall cost increase of ∼5%. AVAILABILITY AND IMPLEMENTATION: Analysis pipeline for this approach is available at GitHub as R package controlFreq (github.com/gimelbrantlab/controlFreq).
Asia Mendelevich, Saumya Gupta, Aleksei Pakharev, Athanasios Teodosiadis, Andrey A. Mironov, Alexander A. Gimelbrant
Bioinform.5
2021 An extended catalogue of tandem alternative splice sites in human tissue transcriptomes
abstract
Tandem alternative splice sites (TASS) is a special class of alternative splicing events that are characterized by a close tandem arrangement of splice sites. Most TASS lack functional characterization and are believed to arise from splicing noise. Based on the RNA-seq data from the Genotype Tissue Expression project, we present an extended catalogue of TASS in healthy human tissues and analyze their tissue-specific expression. The expression of TASS is usually dominated by one major splice site (maSS), while the expression of minor splice sites (miSS) is at least an order of magnitude lower. Among 46k miSS with sufficient read support, 9k (20%) are significantly expressed above the expected noise level, and among them 2.5k are expressed tissue-specifically. We found significant correlations between tissue-specific expression of RNA-binding proteins (RBP), tissue-specific expression of miSS, and miSS response to RBP inactivation by shRNA. In combination with RBP profiling by eCLIP, this allowed prediction of novel cases of tissue-specific splicing regulation including a miSS in QKI mRNA that is likely regulated by PTBP1. The analysis of human primary cell transcriptomes suggested that both tissue-specific and cell-type-specific factors contribute to the regulation of miSS expression. More than 20% of tissue-specific miSS affect structured protein regions and may adjust protein-protein interactions or modify the stability of the protein core. The significantly expressed miSS evolve under the same selection pressure as maSS, while other miSS lack signatures of evolutionary selection and conservation. Using mixture models, we estimated that not more than 15% of maSS and not more than 54% of tissue-specific miSS are noisy, while the proportion of noisy splice sites among non-significantly expressed miSS is above 63%.
Andrey A. Mironov, Stepan Denisov, Alexander Greß, Olga V. Kalinina, Dmitri D. Pervouchine
PLoS Comput. Biol.1
2017 StereoGene: rapid estimation of genome-wide correlation of continuous or interval feature data
abstract
MOTIVATION: Genomics features with similar genome-wide distributions are generally hypothesized to be functionally related, for example, colocalization of histones and transcription start sites indicate chromatin regulation of transcription factor activity. Therefore, statistical algorithms to perform spatial, genome-wide correlation among genomic features are required. RESULTS: Here, we propose a method, StereoGene, that rapidly estimates genome-wide correlation among pairs of genomic features. These features may represent high-throughput data mapped to reference genome or sets of genomic annotations in that reference genome. StereoGene enables correlation of continuous data directly, avoiding the data binarization and subsequent data loss. Correlations are computed among neighboring genomic positions using kernel correlation. Representing the correlation as a function of the genome position, StereoGene outputs the local correlation track as part of the analysis. StereoGene also accounts for confounders such as input DNA by partial correlation. We apply our method to numerous comparisons of ChIP-Seq datasets from the Human Epigenome Atlas and FANTOM CAGE to demonstrate its wide applicability. We observe the changes in the correlation between epigenomic features across developmental trajectories of several tissue types consistent with known biology and find a novel spatial correlation of CAGE clusters with donor splice sites and with poly(A) sites. These analyses provide examples for the broad applicability of StereoGene for regulatory genomics. AVAILABILITY AND IMPLEMENTATION: The StereoGene C ++ source code, program documentation, Galaxy integration scripts and examples are available from the project homepage http://stereogene.bioinf.fbb.msu.ru/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Elena D. Stavrovskaya, Tejasvi Niranjan, Elana J. Fertig, Sarah J. Wheelan, Alexander V. Favorov, Andrey A. Mironov
Bioinform.6
2014 RNASurface: fast and accurate detection of locally optimal potentially structured RNA segments
abstract
MOTIVATION: During the past decade, new classes of non-coding RNAs (ncRNAs) and their unexpected functions were discovered. Stable secondary structure is the key feature of many non-coding RNAs. Taking into account huge amounts of genomic data, development of computational methods to survey genomes for structured RNAs remains an actual problem, especially when homologous sequences are not available for comparative analysis. Existing programs scan genomes with a fixed window by efficiently constructing a matrix of RNA minimum free energies. A wide range of lengths of structured RNAs necessitates the use of many different window lengths that substantially increases the output size and computational efforts. RESULTS: In this article, we present an algorithm RNASurface to efficiently scan genomes by constructing a matrix of significance of RNA secondary structures and to identify all locally optimal structured RNA segments up to a predefined size. RNASurface significantly improves precision of identification of known ncRNA in Bacillus subtilis. AVAILABILITY AND IMPLEMENTATION: RNASurface C source code is available from http://bioinf.fbb.msu.ru/RNASurface/downloads.html.
Ruslan A. Soldatov, Svetlana V. Vinogradova, Andrey A. Mironov
Bioinform.3
2012 Evidence for Widespread Association of Mammalian Splicing and Conserved Long-Range RNA Structures
Dmitri D. Pervouchine, Ekaterina Khrameeva, Marina Pichugina, Olexii Nikolaienko, Mikhail S. Gelfand, Petr Rubtsov, Andrey A. Mironov
RECOMB7
2012 Exploring Massive, Genome Scale Datasets with the GenometriCorr Package
abstract
UNLABELLED: We have created a statistically grounded tool for determining the correlation of genomewide data with other datasets or known biological features, intended to guide biological exploration of high-dimensional datasets, rather than providing immediate answers. The software enables several biologically motivated approaches to these data and here we describe the rationale and implementation for each approach. Our models and statistics are implemented in an R package that efficiently calculates the spatial correlation between two sets of genomic intervals (data and/or annotated features), for use as a metric of functional interaction. The software handles any type of pointwise or interval data and instead of running analyses with predefined metrics, it computes the significance and direction of several types of spatial association; this is intended to suggest potentially relevant relationships between the datasets. AVAILABILITY AND IMPLEMENTATION: The package, GenometriCorr, can be freely downloaded at http://genometricorr.sourceforge.net/. Installation guidelines and examples are available from the sourceforge repository. The package is pending submission to Bioconductor.
Alexander V. Favorov, Loris Mularoni, Leslie Cope, Yulia A. Medvedeva, Andrey A. Mironov, Vsevolod J. Makeev, Sarah J. Wheelan
PLoS Comput. Biol.5
2009 Evolution of Regulatory Systems in Bacteria (Invited Keynote Talk)
Mikhail S. Gelfand, Alexei E. Kazakov, Yuri D. Korostelev, Olga N. Laikova, Andrey A. Mironov, Aleksandra B. Rakhmaninova, Dmitry A. Ravcheev, Dmitry A. Rodionov, Alexei G. Vitreschak
ISBRA5
2005 A Gibbs sampler for identification of symmetrically structured, spaced DNA motifs with improved estimation of the signal length
abstract
MOTIVATION: Transcription regulatory protein factors often bind DNA as homo-dimers or hetero-dimers. Thus they recognize structured DNA motifs that are inverted or direct repeats or spaced motif pairs. However, these motifs are often difficult to identify owing to their high divergence. The motif structure included explicitly into the motif recognition algorithm improves recognition efficiency for highly divergent motifs as well as estimation of motif geometric parameters. RESULT: We present a modification of the Gibbs sampling motif extraction algorithm, SeSiMCMC (Sequence Similarities by Markov Chain Monte Carlo), which finds structured motifs of these types, as well as non-structured motifs, in a set of unaligned DNA sequences. It employs improved estimators of motif and spacer lengths. The probability that a sequence does not contain any motif is accounted for in a rigorous Bayesian manner. We have applied the algorithm to a set of upstream regions of genes from two Escherichia coli regulons involved in respiration. We have demonstrated that accounting for a symmetric motif structure allows the algorithm to identify weak motifs more accurately. In the examples studied, ArcA binding sites were demonstrated to have the structure of a direct spaced repeat, whereas NarP binding sites exhibited the palindromic structure. AVAILABILITY: The WWW interface of the program, its FreeBSD (4.0) and Windows 32 console executables are available at http://bioinform.genetika.ru/SeSiMCMC
Alexander V. Favorov, Mikhail S. Gelfand, Anna V. Gerasimova, Dmitry A. Ravcheev, Andrey A. Mironov, Vsevolod J. Makeev
Bioinform.5
2005 Alternative splicing and protein function
abstract
BACKGROUND: Alternative splicing is a major mechanism of generating protein diversity in higher eukaryotes. Although at least half, and probably more, of mammalian genes are alternatively spliced, it was not clear, whether the frequency of alternative splicing is the same in different functional categories. The problem is obscured by uneven coverage of genes by ESTs and a large number of artifacts in the EST data. RESULTS: We have developed a method that generates possible mRNA isoforms for human genes contained in the EDAS database, taking into account the effects of nonsense-mediated decay and translation initiation rules, and a procedure for offsetting the effects of uneven EST coverage. Then we computed the number of mRNA isoforms for genes from different functional categories. Genes encoding ribosomal proteins and genes in the category "Small GTPase-mediated signal transduction" tend to have fewer isoforms than the average, whereas the genes in the category "DNA replication and chromosome cycle" have more isoforms than the average. Genes encoding proteins involved in protein-protein interactions tend to be alternatively spliced more often than genes encoding non-interacting proteins, although there is no significant difference in the number of isoforms of alternatively spliced genes. CONCLUSION: Filtering for functional isoforms satisfying biological constraints and accounting for uneven EST coverage allowed us to describe differences in alternative splicing of genes from different functional categories. The observations seem to be consistent with expectations based on current biological knowledge: less isoforms for ribosomal and signal transduction proteins, and more alternative splicing of interacting and cell cycle proteins.
A. D. Neverov, Irena I. Artamonova, Ramil N. Nurtdinov, Dmitrij Frishman, Mikhail S. Gelfand, Andrey A. Mironov
BMC Bioinform.6
2002 Exact mapping of prokaryotic gene starts
abstract
It is known that while the programs used to find genes in prokaryotic genomes reliably map protein-coding regions, they often fail in the exact determination of gene starts. This problem is further aggravated by sequencing errors, most notably insertions and deletions leading to frame-shifts. Therefore, the exact mapping of gene starts and identification of frame-shifts are important problems of the computer-assisted functional analysis of newly sequenced genomes. Here we review methods of gene recognition and describe a new algorithm for correction of gene starts and identification of frame-shifts in prokaryotic genomes. The algorithm is based on the comparison of nucleotide and protein sequences of homologous genes from related organisms, using the assumption that the rate of evolutionary changes in protein-coding regions is lower than that in non-coding regions. A dynamic programming algorithm is used to align protein sequences obtained by formal translation of genomic nucleotide sequences. The possibility of frame-shifts is taken into account. The algorithm was tested on several groups of related organisms: gamma-proteobacteria, the Bacillus/Clostridium group, and three Pyrococcus genomes. The testing demonstrated that, dependent or a genome, 1-10 per cent of genes have incorrect starts or contain frame-shifts. The algorithm is implemented in the program package Orthologator-GeneCorrector.
M. V. Baytaluk, Mikhail S. Gelfand, Andrey A. Mironov
Briefings Bioinform.3
2001 Pro-Frame: similarity-based gene recognition in eukaryotic DNA sequences with errors
abstract
Abstract Summary: Performance of existing algorithms for similarity-based gene recognition in eukaryotes drops when the genomic DNA has been sequenced with errors. A modification of the spliced alignment algorithm allows for gene recognition in sequences with errors, in particular frameshifts. It tolerates up to 5% of sequencing errors without considerable drop of prediction reliability when a sufficiently close homologous protein is available (normalized evolutionary distance similarity score 50% or higher). Availability: The program is free for academic users and available upon request at http://www.anchorgen.com Contact: [email protected]
Andrey A. Mironov, Pavel S. Novichkov, Mikhail S. Gelfand
Bioinform.1
2001 Gene recognition in eukaryotic DNA by comparison of genomic sequences
abstract
MOTIVATION: Sequencing of complete eukaryotic genomes and large syntenic fragments of genomes makes it possible to apply genomic comparison for gene recognition. RESULTS: This paper describes a spliced alignment algorithm that aligns candidate exon chains of two homologous genomic sequence fragments from different species. The algorithm is implemented in Pro-Gen software. Unlike other algorithms, Pro-Gen does not assume conservation of the exon-intron structure. Amino acid sequences obtained by the formal translation of candidate exons are aligned instead of nucleotide sequences, which allows for distant comparisons. The algorithm was tested on a sample of human-mammal (mouse), human-vertebrate (Xenopus ) and human-invertebrate (Drosophila ) gene pairs. Surprisingly, the best results, 97-98% correlation between the actual and predicted genes, were obtained for more distant comparisons, whereas the correlation on the human-mouse sample was only 93%. The latter value increases to 95% if conservation of the exon-intron structure is assumed. This is caused by a large amount of sequence conservation in non-coding regions of the human and mouse genes probably due to regulatory elements. AVAILABILITY: Pro-Gen v. 3.0 is available to academic researchers free of charge at http://www.anchorgen.com/pro_gen/pro_gen.html.
Pavel S. Novichkov, Mikhail S. Gelfand, Andrey A. Mironov
Bioinform.3
2000 Comparative Analysis of Regulatory Patterns in Bacterial Genomes
abstract
Recognition of transcription regulatory sites in bacterial genomes is a notoriously difficult problem. There are no algorithms capable of making reliable predictions even for well-studied sites such as the CRP (cyclic AMP receptor protein) box. However, availability of complete bacterial genomes makes it possible to make reliable predictions with bad rules. This comparative approach is based on the assumption that sets of co-regulated genes are conserved in related bacteria. Thus true sites occur upstream of orthologous genes, whereas false candidates are scattered at random. This means not only that knowledge about regulation in well-studied genomes can be transferred to newly sequenced ones, but also that new members of regulons can be found. This paper reviews several recent studies. In particular, a detailed analysis of catabolite repression in gamma-purple bacteria is presented.
Mikhail S. Gelfand, Pavel S. Novichkov, Elena S. Novichkova, Andrey A. Mironov
Briefings Bioinform.4
1998 SST versus EST in Gene Recognition (Invited Paper)
abstract
The EST data provide a powerful tool for identification of transcribed DNA sequences. However, since ESTs are relatively short, many exons are poorly covered by ESTs thus reducing the utility of EST data. Recently, SST (Signature Sequence Tags) fingerprints were proposed as an alternative to EST fingerprints. Given a fingerprint set of probes, SST of a clone is a subset of probes from the fingerprint set that hybridize with the clone. We demonstrate that besides being a powerful technique for screening cDNA libraries, SST technology provides for very accurate gene predictions. Even with a small fingerprint set (600-800 probes) SST-based gene recognition outperforms many conventional and EST-based methods. The increase in the size of fingerprint set to 1500 probes provides almost perfect gene recognition. Even more importantly, SST-based gene predictions miss very few exons and therefore provide an opportunity to bypass cDNA sequencing step on the way from finished genomic sequence to mutation detection in gene hunting projects. Since SST data can be obtained in a highly parallel and inexpensive way, SST technology has a potential of substituting EST technology for gene hunting.
Andrey A. Mironov, Pavel A. Pevzner
SPIRE1
1998 Algorithms and software for support of gene identification experiments
abstract
MOTIVATION: Gene annotation is the final goal of gene prediction algorithms. However, these algorithms frequently make mistakes and therefore the use of gene predictions for sequence annotation is hardly possible. As a result, biologists are forced to conduct time-consuming gene identification experiments by designing appropriate PCR primers to test cDNA libraries or applying RT-PCR, exon trapping/amplification, or other techniques. This process frequently amounts to 'guessing' PCR primers on top of unreliable gene predictions and frequently leads to wasting of experimental efforts. RESULTS: The present paper proposes a simple and reliable algorithm for experimental gene identification which bypasses the unreliable gene prediction step. Studies of the performance of the algorithm on a sample of human genes indicate that an experimental protocol based on the algorithm's predictions achieves an accurate gene identification with relatively few PCR primers. Predictions of PCR primers may be used for exon amplification in preliminary mutation analysis during an attempt to identify a gene responsible for a disease. We propose a simple approach to find a short region from a genomic sequence that with high probability overlaps with some exon of the gene. The algorithm is enhanced to find one or more segments that are probably contained in the translated region of the gene and can be used as PCR primers to select appropriate clones in cDNA libraries by selective amplification. The algorithm is further extended to locate a set of PCR primers that uniformly cover all translated regions and can be used for RT-PCR and further sequencing of (unknown) mRNA.
Sing-Hoi Sze, Mikhail A. Roytberg, Mikhail S. Gelfand, Andrey A. Mironov, Tatiana V. Astakhova, Pavel A. Pevzner
Bioinform.4
1996 Spliced Alignment: A New Approach to Gene Recognition
Mikhail S. Gelfand, Andrey A. Mironov, Pavel A. Pevzner
CPM2
1995 DNASUN: a package of computer programs for the biotechnology laboratory
abstract
The paper describes a new software package DNASUN developed for supporting gene engineering laboratories. The package provides a user-friendly interface for experimental researches and supports the traditional nucleotide/protein sequence analysis as well as physical mapping, sequencing, plasmid manipulations, optimal oligonucleotide probe selection and other common molecular biology procedures.
Andrey A. Mironov, N. N. Alexandrov, N. Yu. Bogodarova, A. Grigorjev, V. F. Lebedev, L. V. Lunovskaya, M. E. Truchan, Pavel A. Pevzner
Comput. Appl. Biosci.1
1990 Mapping DNA by stochastic relaxation: a new approach to fragment sizes
abstract
Instead of the traditional manipulations with given fixed fragment lengths in the restriction map construction a method of varying the lengths is proposed and realized under the simulated annealing algorithm scheme. The described approach has no upper limit on the number of fragments mapped with even ordinary hardware. A program has been derived from the algorithm combined with the last-squares refinement procedure for both linear and circular maps. The algorithm's ability to pick up missed maps is illustrated and the problem of reducing the number of solutions is discussed.
A. V. Grigorjev, Andrey A. Mironov
Comput. Appl. Biosci.2