EDBT 2026 Demo / reviewers in the wild / expert
Robert J. Genco
dblp:235/6479
· DBLP profile ↗
3ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0001-7604-506XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › sequence analysis › sequence comparison
alignment-free sequence comparison |
0.4 | 1 | 2019 | SENSE: Siamese neural network for sequence embedding and alignment-free comparison · Bioinform. 2019 |
Bioinformatics and computational biology
sequence analysis |
0.4 | 1 | 2019 | SENSE: Siamese neural network for sequence embedding and alignment-free comparison · Bioinform. 2019 |
Bioinformatics and computational biology › sequence analysis
sequence clustering |
0.4 | 1 | 2019 | A parallel computational framework for ultra-large-scale sequence clustering analysis · Bioinform. 2019 |
Bioinformatics and computational biology › sequence analysis › sequence feature extraction
sequence embedding |
0.4 | 1 | 2019 | SENSE: Siamese neural network for sequence embedding and alignment-free comparison · Bioinform. 2019 |
Parallel and multicore computing
parallel computing |
0.1 | 1 | 2019 | A parallel computational framework for ultra-large-scale sequence clustering analysis · Bioinform. 2019 |
Methods — techniques the papers use, named apart from their topics
landmark-based divisive clustering · 0.8apache spark · 0.8siamese neural network · 0.4deep metric learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Computational approach to modeling microbiome landscapes associated with chronic human disease progressionabstractA microbial community is a dynamic system undergoing constant change in response to internal and external stimuli. These changes can have significant implications for human health. However, due to the difficulty in obtaining longitudinal samples, the study of the dynamic relationship between the microbiome and human health remains a challenge. Here, we introduce a novel computational strategy that uses massive cross-sectional sample data to model microbiome landscapes associated with chronic disease development. The strategy is based on the rationale that each static sample provides a snapshot of the disease process, and if the number of samples is sufficiently large, the footprints of individual samples populate progression trajectories, which enables us to recover disease progression paths along a microbiome landscape by using computational approaches. To demonstrate the validity of the proposed strategy, we developed a bioinformatics pipeline and applied it to a gut microbiome dataset available from a Crohn's disease study. Our analysis resulted in one of the first working models of microbial progression for Crohn's disease. We performed a series of interrogations to validate the constructed model. Our analysis suggested that the model recapitulated the longitudinal progression of microbial dysbiosis during the known clinical trajectory of Crohn's disease. By overcoming restrictions associated with complex longitudinal sampling, the proposed strategy can provide valuable insights into the role of the microbiome in the pathogenesis of chronic disease and facilitate the shift of the field from descriptive research to mechanistic studies. Lu Li 0009, Jiho Sohn, Robert J. Genco, Jean Wactawski-Wende, Steve Goodison, Patricia I. Diaz, Yijun Sun |
PLoS Comput. Biol. | 3 |
| 2019 | A parallel computational framework for ultra-large-scale sequence clustering analysisabstractMotivation: The rapid development of sequencing technology has led to an explosive accumulation of genomic data. Clustering is often the first step to be performed in sequence analysis. However, existing methods scale poorly with respect to the unprecedented growth of input data size. As high-performance computing systems are becoming widely accessible, it is highly desired that a clustering method can easily scale to handle large-scale sequence datasets by leveraging the power of parallel computing. Results: In this paper, we introduce SLAD (Separation via Landmark-based Active Divisive clustering), a generic computational framework that can be used to parallelize various de novo operational taxonomic unit (OTU) picking methods and comes with theoretical guarantees on both accuracy and efficiency. The proposed framework was implemented on Apache Spark, which allows for easy and efficient utilization of parallel computing resources. Experiments performed on various datasets demonstrated that SLAD can significantly speed up a number of popular de novo OTU picking methods and meanwhile maintains the same level of accuracy. In particular, the experiment on the Earth Microbiome Project dataset (∼2.2B reads, 437 GB) demonstrated the excellent scalability of the proposed method. Availability and implementation: Open-source software for the proposed method is freely available at https://www.acsu.buffalo.edu/~yijunsun/lab/SLAD.html. Supplementary information: Supplementary data are available at Bioinformatics online. Wei Zheng 0010, Qi Mao 0001, Robert J. Genco, Jean Wactawski-Wende, Michael J. Buck, Yunpeng Cai, Yijun Sun |
Bioinform. | 3 |
| 2019 | SENSE: Siamese neural network for sequence embedding and alignment-free comparisonabstractMOTIVATION: Sequence analysis is arguably a foundation of modern biology. Classic approaches to sequence analysis are based on sequence alignment, which is limited when dealing with large-scale sequence data. A dozen of alignment-free approaches have been developed to provide computationally efficient alternatives to alignment-based approaches. However, existing methods define sequence similarity based on various heuristics and can only provide rough approximations to alignment distances. RESULTS: In this article, we developed a new approach, referred to as SENSE (SiamEse Neural network for Sequence Embedding), for efficient and accurate alignment-free sequence comparison. The basic idea is to use a deep neural network to learn an explicit embedding function based on a small training dataset to project sequences into an embedding space so that the mean square error between alignment distances and pairwise distances defined in the embedding space is minimized. To the best of our knowledge, this is the first attempt to use deep learning for alignment-free sequence analysis. A large-scale experiment was performed that demonstrated that our method significantly outperformed the state-of-the-art alignment-free methods in terms of both efficiency and accuracy. AVAILABILITY AND IMPLEMENTATION: Open-source software for the proposed method is developed and freely available at https://www.acsu.buffalo.edu/∼yijunsun/lab/SENSE.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Zheng 0010, Robert J. Genco, Jean Wactawski-Wende, Michael J. Buck, Yijun Sun |
Bioinform. | 3 |