EDBT 2026 Demo / reviewers in the wild / expert
J. Michael Cherry
dblp:61/6653
· DBLP profile ↗
8ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0001-9163-5180ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › omics data analysis
batch effect correction |
0.9 | 1 | 2025 | Gene spatial integration: enhancing spatial transcriptomics analysis via deep learning and batch effect mitigation · Bioinform. 2025 |
Bioinformatics and computational biology
data integration |
0.9 | 1 | 2025 | Gene spatial integration: enhancing spatial transcriptomics analysis via deep learning and batch effect mitigation · Bioinform. 2025 |
Bioinformatics and computational biology › transcriptomics › spatial transcriptomics
spatial transcriptomics analysis |
0.9 | 1 | 2025 | Gene spatial integration: enhancing spatial transcriptomics analysis via deep learning and batch effect mitigation · Bioinform. 2025 |
Bioinformatics and computational biology
biomedical text mining |
0.1 | 1 | 2007 | Mining experimental evidence of molecular function claims from the literature · Bioinform. 2007 |
Bioinformatics and computational biology
biological database |
0.0 | 1 | 2004 | Database model and specification of GermOnline Release 2.0, a cross-species community annotation knowledgebase on germ cell differentiation · Bioinform. 2004 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene ontology analysis |
0.0 | 1 | 2004 | GO: : TermFinder--open source software for accessing Gene Ontology information and finding significantly enriched Gene Ontology terms associated with a list of genes · Bioinform. 2004 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.0 | 1 | 2004 | GO: : TermFinder--open source software for accessing Gene Ontology information and finding significantly enriched Gene Ontology terms associated with a list of genes · Bioinform. 2004 |
Bioinformatics and computational biology
genome annotation |
0.0 | 1 | 2004 | Database model and specification of GermOnline Release 2.0, a cross-species community annotation knowledgebase on germ cell differentiation · Bioinform. 2004 |
Bioinformatics and computational biology › gene expression analysis
cluster visualization |
0.0 | 1 | 2001 | Visualization of expression clusters using Sammon's non-linear mapping · Bioinform. 2001 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2001 | Visualization of expression clusters using Sammon's non-linear mapping · Bioinform. 2001 |
Bioinformatics and computational biology › genome annotation
gene annotation |
0.0 | 1 | 2007 | Mining experimental evidence of molecular function claims from the literature · Bioinform. 2007 |
Bioinformatics and computational biology › genome annotation
functional annotation |
0.0 | 1 | 2004 | GO: : TermFinder--open source software for accessing Gene Ontology information and finding significantly enriched Gene Ontology terms associated with a list of genes · Bioinform. 2004 |
Methods — techniques the papers use, named apart from their topics
representation learning · 0.9deep learning · 0.9autoencoder · 0.9text mining · 0.1information extraction · 0.1statistical significance testing · 0.0relational database design · 0.0microarray data storage and visualization · 0.0sammon's non-linear mapping · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gene spatial integration: enhancing spatial transcriptomics analysis via deep learning and batch effect mitigationabstractMOTIVATION: Spatial transcriptomics (ST) is a groundbreaking technique for studying the correlation between cellular organization within a tissue and its physiological and pathological properties. Every facet of spatial information, including cell/spot proximity, distribution, and dimensionality, is significant. Most methods lean heavily on proximity for ST analysis, each resulting in useful insights but still leaving other aspects untapped. In addition, samples procured at different times, by different donors, and by different technologies introduce a batch effects problem that hinders the statistical approach employed by most analysis tools. Addressing these challenges, we have developed a deep learning method for analyzing integrated multiple ST data, focusing on the distribution aspect. Furthermore, our method aims to leverage single-cell analysis tools. RESULTS: Our study introduces Gene Spatial Integration (GSI), a data integration pipeline utilizing a representation learning approach to extract the spatial distribution of genes into the same feature space as gene expression features. We employ an autoencoder network to extract spatial embedding, facilitating the projection of spatial features into gene expression feature space. Our approach allows for seamless integration of multiple samples with minimum detriment, increasing the performance of the ST data analysis tool. We show the application of our method on the human dorsolateral prefrontal cortex dataset. Our method consistently improves the performance of the clustering of Seurat tools, with the most significant increase observed in sample 151673, almost doubling the ARI score from 0.225 to 0.405. We also combine our pipeline with the clustering of GraphST, achieving a significantly higher ARI score in sample 151672 from 0.614 to 0.795. This result reveals the potential of gene distribution spatial aspect, also emphasizes the impact of integration and batch effect removal in developing a refined analysis in understanding tissue characteristics. AVAILABILITY AND IMPLEMENTATION: Implementation of GSI is accessible at https://github.com/Riandanis/Spatial_Integration_GSI. Rian Pratama, Jason A. Hilton, J. Michael Cherry, Giltae Song |
Bioinform. | 3 |
| 2011 | Toward an interactive article: integrating journals and biological databasesabstractBACKGROUND: Journal articles and databases are two major modes of communication in the biological sciences, and thus integrating these critical resources is of urgent importance to increase the pace of discovery. Projects focused on bridging the gap between journals and databases have been on the rise over the last five years and have resulted in the development of automated tools that can recognize entities within a document and link those entities to a relevant database. Unfortunately, automated tools cannot resolve ambiguities that arise from one term being used to signify entities that are quite distinct from one another. Instead, resolving these ambiguities requires some manual oversight. Finding the right balance between the speed and portability of automation and the accuracy and flexibility of manual effort is a crucial goal to making text markup a successful venture. RESULTS: We have established a journal article mark-up pipeline that links GENETICS journal articles and the model organism database (MOD) WormBase. This pipeline uses a lexicon built with entities from the database as a first step. The entity markup pipeline results in links from over nine classes of objects including genes, proteins, alleles, phenotypes and anatomical terms. New entities and ambiguities are discovered and resolved by a database curator through a manual quality control (QC) step, along with help from authors via a web form that is provided to them by the journal. New entities discovered through this pipeline are immediately sent to an appropriate curator at the database. Ambiguous entities that do not automatically resolve to one link are resolved by hand ensuring an accurate link. This pipeline has been extended to other databases, namely Saccharomyces Genome Database (SGD) and FlyBase, and has been implemented in marking up a paper with links to multiple databases. CONCLUSIONS: Our semi-automated pipeline hyperlinks articles published in GENETICS to model organism databases such as WormBase. Our pipeline results in interactive articles that are data rich with high accuracy. The use of a manual quality control step sets this pipeline apart from other hyperlinking tools and results in benefits to authors, journals, readers and databases. Arun Rangarajan, Tim Schedl, Karen Yook, Juancarlos Chan, Stephen Haenel, Lolly Otis, Sharon Faelten, Tracey DePellegrin-Connelly, Ruth Isaacson, Marek S. Skrzypek, J. Michael Cherry, Paul W. Sternberg, Hans-Michael Müller |
BMC Bioinform. | 11 |
| 2007 | Mining experimental evidence of molecular function claims from the literatureabstractMOTIVATION: The rate at which gene-related findings appear in the scientific literature makes it difficult if not impossible for biomedical scientists to keep fully informed and up to date. The importance of these findings argues for the development of automated methods that can find, extract and summarize this information. This article reports on methods for determining the molecular function claims that are being made in a scientific article, specifically those that are backed by experimental evidence. RESULTS: The most significant result is that for molecular function claims based on direct assays, our methods achieved recall of 70.7% and precision of 65.7%. Furthermore, our methods correctly identified in the text 44.6% of the specific molecular function claims backed up by direct assays, but with a precision of only 0.92%, a disappointing outcome that led to an examination of the different kinds of errors. These results were based on an analysis of 1823 articles from the literature of Saccharomyces cerevisiae (budding yeast). AVAILABILITY: The annotation files for S.cerevisiae are available from ftp://genome-ftp.stanford.edu/pub/yeast/data_download/literature_curation/gene_association.sgd.gz. The draft protocol vocabulary is available by request from the first author. Colleen E. Crangle, J. Michael Cherry, Eurie L. Hong, Alex Zbyslaw |
Bioinform. | 2 |
| 2004 | Saccharomyces genome database: Underlying principles and organisationabstractA scientific database can be a powerful tool for biologists in an era where large-scale genomic analysis, combined with smaller-scale scientific results, provides new insights into the roles of genes and their products in the cell. However, the collection and assimilation of data is, in itself, not enough to make a database useful. The data must be incorporated into the database and presented to the user in an intuitive and biologically significant manner. Most importantly, this presentation must be driven by the user's point of view; that is, from a biological perspective. The success of a scientific database can therefore be measured by the response of its users - statistically, by usage numbers and, in a less quantifiable way, by its relationship with the community it serves and its ability to serve as a model for similar projects. Since its inception ten years ago, the Saccharomyces Genome Database (SGD) has seen a dramatic increase in its usage, has developed and maintained a positive working relationship with the yeast research community, and has served as a template for at least one other database. The success of SGD, as measured by these criteria, is due in large part to philosophies that have guided its mission and organisation since it was established in 1993. This paper aims to detail these philosophies and how they shape the organisation and presentation of the database. Selina S. Dwight, Rama Balakrishnan, Karen R. Christie, Maria C. Costanzo, Kara Dolinski, Stacia R. Engel, Becket Feierbach, Dianna G. Fisk, Jodi E. Hirschman, Eurie L. Hong, Laurie Issel-Tarver, Robert S. Nash, Anand Sethuraman, Barry Starr, Chandra L. Theesfeld, Rey Andrada, Gail Binkley, Qing Dong 0003, Christopher Lane, Mark Schroeder, Shuai Weng, David Botstein, J. Michael Cherry |
Briefings Bioinform. | 23 |
| 2004 | GO: : TermFinder--open source software for accessing Gene Ontology information and finding significantly enriched Gene Ontology terms associated with a list of genesabstractSUMMARY: GO::TermFinder comprises a set of object-oriented Perl modules for accessing Gene Ontology (GO) information and evaluating and visualizing the collective annotation of a list of genes to GO terms. It can be used to draw conclusions from microarray and other biological data, calculating the statistical significance of each annotation. GO::TermFinder can be used on any system on which Perl can be run, either as a command line application, in single or batch mode, or as a web-based CGI script. AVAILABILITY: The full source code and documentation for GO::TermFinder are freely available from http://search.cpan.org/dist/GO-TermFinder/. Elizabeth I. Boyle, Shuai Weng, Jeremy Gollub, Heng Jin, David Botstein, J. Michael Cherry, Gavin Sherlock |
Bioinform. | 6 |
| 2004 | Database model and specification of GermOnline Release 2.0, a cross-species community annotation knowledgebase on germ cell differentiationabstractUNLABELLED: GermOnline is a web-accessible relational database that enables life scientists to make a significant and sustained contribution to the annotation of genes relevant for the fields of mitosis, meiosis, germ line development and gametogenesis across species. This novel approach to genome annotation includes a platform for knowledge submission and curation as well as microarray data storage and visualization hosted by a global network of servers. AVAILABILITY: The database is accessible at http://www.germonline.org/. For convenient world-wide access we have set up a network of servers in Europe (http://germonline.unibas.ch/; http://germonline.igh.cnrs.fr/), Japan (http://germonline.biochem.s.u-tokyo.ac.jp/) and USA (http://germonline.yeastgenome.org/). SUPPLEMENTARY INFORMATION: Extended documentation of the database is available through the link 'About GermOnline' at the websites. Christa Niederhauser-Wiederkehr, R. Basavaraj, Cyril Sarrauste de Menthière, R. Koch, U. Schlecht, Leandro Hermida, Benjamin Masdoua, Ryohei Ishii, V. Cassen, Christopher Lane, J. Michael Cherry, N. Lamb, Michael Primig |
Bioinform. | 12 |
| 2004 | A short study on the success of the Gene Ontology
Michael Bada, Robert Stevens 0001, Carole A. Goble, Yolanda Gil, Michael Ashburner, Judith A. Blake, J. Michael Cherry, Midori A. Harris, Suzanna Lewis |
J. Web Semant. | 7 |
| 2001 | Visualization of expression clusters using Sammon's non-linear mappingabstractAbstract Summary: A method of exploratory analysis and visualization of multi-dimensional gene expression data using Sammon’s Non-Linear Mapping (NLM) is presented. Availability: Scripts are available from the authors. Contact: [email protected] * To whom correspondence should be addressed. Rob M. Ewing, J. Michael Cherry |
Bioinform. | 2 |