EDBT 2026 Demo / reviewers in the wild / expert
Oluwatosin Oluwadare
dblp:170/4568
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-5264-2342ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spliceread: improving canonical and non-canonical splice site prediction with residual blocks and synthetic data augmentationabstractAccurate splice site prediction is fundamental to understanding gene expression and its associated disorders. However, most existing models are biased toward frequent canonical sites, limiting their ability to detect rare but biologically important non-canonical variants. These models often rely heavily on large, imbalanced datasets that fail to capture the sequence diversity of non-canonical sites, leading to high false-negative rates. Here, we present SpliceRead, a novel deep learning model designed to improve the classification of both canonical and non-canonical splice sites using a combination of residual convolutional blocks and synthetic data augmentation. SpliceRead employs a data augmentation method to generate diverse non-canonical sequences and uses residual connections to enhance gradient flow and capture subtle genomic features. Trained and tested on a multi-species dataset of 400- and 600-nucleotide sequences, SpliceRead consistently outperforms state-of-the-art models across all key metrics, including F1-score, accuracy, precision, and recall. Notably, it achieves a substantially lower non-canonical misclassification rate than baseline methods. Extensive evaluations, including cross-validation, cross-species testing, and input-length generalization, confirm its robustness and adaptability. We further evaluated the adaptability of this augmentation technique by applying it to other state-of-the-art models, demonstrating consistent improvements and effective generalization. SpliceRead offers a powerful, generalizable framework for splice site prediction, particularly in challenging, low-frequency sequence scenarios, and paves the way for more accurate gene annotation in both model and non-model organisms. The open-sourced code of SpliceRead and a detailed documentation is available at https://github.com/OluwadareLab/SpliceRead . Sahil Thapa, Khushali Samderiya, Rohit Menon, Oluwatosin Oluwadare |
BMC Bioinform. | 4 |
| 2025 | Unicorn: enhancing single-cell Hi-C data with blind super-resolution for 3D genome structure reconstructionabstractMOTIVATION: Single-cell Hi-C (scHi-C) data provide critical insights into chromatin interactions at individual cell levels, uncovering unique genomic 3D structures. However, scHi-C datasets are characterized by sparsity and noise, complicating efforts to accurately reconstruct high-resolution chromosomal structures. In this study, we present ScUnicorn, a novel blind super-resolution framework for scHi-C data enhancement. ScUnicorn uses an iterative degradation kernel optimization process, unlike traditional super-resolution approaches, which rely on downsampling, predefined degradation ratios, or constant assumptions about the input data to reconstruct high-resolution interaction matrices. Hence, our approach more reliably preserves critical biological patterns and minimizes noise. Additionally, we propose 3DUnicorn, a maximum likelihood algorithm that leverages the enhanced scHi-C data to infer precise 3D chromosomal structures. RESULTS: Our evaluation demonstrates that ScUnicorn achieves superior performance over the state-of-the-art methods in terms of Peak Signal-to-Noise Ratio, Structural Similarity Index Measure, and GenomeDisco scores. Moreover, 3DUnicorn's reconstructed structures align closely with experimental 3D-FISH data, underscoring its biological relevance. Together, ScUnicorn and 3DUnicorn provide a robust framework for advancing genomic research by enhancing scHi-C data fidelity and enabling accurate 3D genome structure reconstruction. AVAILABILITY AND IMPLEMENTATION: Unicorn implementation is publicly accessible at https://github.com/OluwadareLab/Unicorn. Mohan Kumar B. Chandrashekar, Rohit Menon, Samuel Olowofila, Oluwatosin Oluwadare |
Bioinform. | 4 |
| 2025 | DiCARN-DNase: enhancing cell-to-cell Hi-C resolution using dilated cascading ResNet with self-attention and DNase-seq chromatin accessibility dataabstractMOTIVATION: The spatial organization of chromatin is fundamental to gene regulation and essential for proper cellular function. The Hi-C technique remains the leading method for unraveling 3D genome structures, but the limited availability of high-resolution (HR) Hi-C data poses significant challenges for comprehensive analysis. Deep learning models have been developed to predict HR Hi-C data from low-resolution counterparts. Early Convolutional Neural Network (CNN)-based models improved resolution but struggled with issues like blurring and capturing fine details. In contrast, Generative Adversarial Network (GAN)-based methods encountered difficulties in maintaining diversity and generalization. Additionally, most existing algorithms perform poorly in cross-cell line generalization, where a model trained on one cell type is used to enhance HR data in another cell type. RESULTS: In this work, we propose Dilated Cascading Residual Network (DiCARN) to overcome these challenges and improve Hi-C data resolution. DiCARN leverages dilated convolutions and cascading residuals to capture a broader context while preserving fine-grained genomic interactions. Additionally, we incorporate DNase-seq data into our model, providing a robust framework that demonstrates superior generalizability across cell lines in HR Hi-C data reconstruction. AVAILABILITY AND IMPLEMENTATION: DiCARN is publicly available at https://github.com/OluwadareLab/DiCARN. Samuel Olowofila, Oluwatosin Oluwadare |
Bioinform. | 2 |
| 2025 | HiCForecast: dynamic network optical flow estimation algorithm for spatiotemporal Hi-C data forecastingabstractMOTIVATION: The exploration of the 3D organization of DNA within the nucleus in relation to various stages of cellular development has led to experiments generating spatiotemporal Hi-C data. However, there is limited spatiotemporal Hi-C data for many organisms, impeding the study of 3D genome dynamics. To overcome this limitation and advance our understanding of genome organization, it is crucial to develop methods for forecasting Hi-C data at future time points from existing timeseries Hi-C data. RESULT: In this work, we designed a novel framework named HiCForecast, adopting a dynamic voxel flow algorithm to forecast future spatiotemporal Hi-C data. We evaluated how well our method generalizes forecasting data across different species and systems, ensuring performance in homogeneous, heterogeneous, and general contexts. Using both computational and biological evaluation metrics, our results show that HiCForecast outperforms the current state-of-the-art algorithm, emerging as an efficient and powerful tool for forecasting future spatiotemporal Hi-C datasets. AVAILABILITY AND IMPLEMENTATION: HiCForecast is publicly available at https://github.com/OluwadareLab/HiCForecast. Dmitry Pinchuk, H. M. A. Mohit Chowdhury, Abhishek Pandeya, Oluwatosin Oluwadare |
Bioinform. | 4 |
| 2024 | Comparative study on chromatin loop callers using Hi-C data reveals their effectivenessabstractAbstract Background Chromosome is one of the most fundamental part of cell biology where DNA holds the hierarchical information. DNA compacts its size by forming loops, and these regions house various protein particles, including CTCF, SMC3, H3 histone. Numerous sequencing methods, such as Hi-C, ChIP-seq, and Micro-C, have been developed to investigate these properties. Utilizing these data, scientists have developed a variety of loop prediction techniques that have greatly improved their methods for characterizing loop prediction and related aspects. Results In this study, we categorized 22 loop calling methods and conducted a comprehensive study of 11 of them. Additionally, we have provided detailed insights into the methodologies underlying these algorithms for loop detection, categorizing them into five distinct groups based on their fundamental approaches. Furthermore, we have included critical information such as resolution, input and output formats, and parameters. For this analysis, we utilized the GM12878 Hi-C datasets at 5 KB, 10 KB, 100 KB and 250 KB resolutions. Our evaluation criteria encompassed various factors, including memory usages, running time, sequencing depth, and recovery of protein-specific sites such as CTCF, H3K27ac, and RNAPII. Conclusion This analysis offers insights into the loop detection processes of each method, along with the strengths and weaknesses of each, enabling readers to effectively choose suitable methods for their datasets. We evaluate the capabilities of these tools and introduce a novel Biological, Consistency, and Computational robustness score ( $$BCC_{score}$$ B C C score ) to measure their overall robustness ensuring a comprehensive evaluation of their performance. H. M. A. Mohit Chowdhury, Terrance E. Boult, Oluwatosin Oluwadare |
BMC Bioinform. | 3 |
| 2022 | HiCARN: resolution enhancement of Hi-C data using cascading residual networksabstractMOTIVATION: High throughput chromosome conformation capture (Hi-C) contact matrices are used to predict 3D chromatin structures in eukaryotic cells. High-resolution Hi-C data are less available than low-resolution Hi-C data due to sequencing costs but provide greater insight into the intricate details of 3D chromatin structures such as enhancer-promoter interactions and sub-domains. To provide a cost-effective solution to high-resolution Hi-C data collection, deep learning models are used to predict high-resolution Hi-C matrices from existing low-resolution matrices across multiple cell types. RESULTS: Here, we present two Cascading Residual Networks called HiCARN-1 and HiCARN-2, a convolutional neural network and a generative adversarial network, that use a novel framework of cascading connections throughout the network for Hi-C contact matrix prediction from low-resolution data. Shown by image evaluation and Hi-C reproducibility metrics, both HiCARN models, overall, outperform state-of-the-art Hi-C resolution enhancement algorithms in predictive accuracy for both human and mouse 1/16, 1/32, 1/64 and 1/100 downsampled high-resolution Hi-C data. Also, validation by extracting topologically associating domains, chromosome 3D structure and chromatin loop predictions from the enhanced data shows that HiCARN can proficiently reconstruct biologically significant regions. AVAILABILITY AND IMPLEMENTATION: HiCARN can be accessed and utilized as an open-sourced software at: https://github.com/OluwadareLab/HiCARN and is also available as a containerized application that can be run on any platform. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Parker Hicks, Oluwatosin Oluwadare |
Bioinform. | 2 |
| 2022 | EnsembleSplice: ensemble deep learning model for splice site predictionabstractBACKGROUND: Identifying splice site regions is an important step in the genomic DNA sequencing pipelines of biomedical and pharmaceutical research. Within this research purview, efficient and accurate splice site detection is highly desirable, and a variety of computational models have been developed toward this end. Neural network architectures have recently been shown to outperform classical machine learning approaches for the task of splice site prediction. Despite these advances, there is still considerable potential for improvement, especially regarding model prediction accuracy, and error rate. RESULTS: Given these deficits, we propose EnsembleSplice, an ensemble learning architecture made up of four (4) distinct convolutional neural networks (CNN) model architecture combination that outperform existing splice site detection methods in the experimental evaluation metrics considered including the accuracies and error rates. We trained and tested a variety of ensembles made up of CNNs and DNNs using the five-fold cross-validation method to identify the model that performed the best across the evaluation and diversity metrics. As a result, we developed our diverse and highly effective splice site (SS) detection model, which we evaluated using two (2) genomic Homo sapiens datasets and the Arabidopsis thaliana dataset. The results showed that for of the Homo sapiens EnsembleSplice achieved accuracies of 94.16% for one of the acceptor splice sites and 95.97% for donor splice sites, with an error rate for the same Homo sapiens dataset, 4.03% for the donor splice sites and 5.84% for the acceptor splice sites datasets. CONCLUSIONS: Our five-fold cross validation ensured the prediction accuracy of our models are consistent. For reproducibility, all the datasets used, models generated, and results in our work are publicly available in our GitHub repository here: https://github.com/OluwadareLab/EnsembleSplice. Victor Akpokiro, Trevor P. Martin, Oluwatosin Oluwadare |
BMC Bioinform. | 3 |
| 2022 | TADMaster: a comprehensive web-based tool for the analysis of topologically associated domainsabstractBACKGROUND: Chromosome conformation capture and its derivatives have provided substantial genetic data for understanding how chromatin self-organizes. These techniques have identified regions of high intrasequence interactions called topologically associated domains (TADs). TADs are structural and functional units that shape chromosomes and influence genomic expression. Many of these domains differ across cell development and can be impacted by diseases. Thus, analysis of the identified domains can provide insight into genome regulation. Hence, there are many approaches to identifying such domains across many cell lines. Despite the availability of multiple tools for TAD detection, TAD callers' speed, flexibility, result inconsistency, and reproducibility remain challenges in this research area. RESULTS: In this work, we developed a computational webserver called TADMaster that provides an analysis suite to directly evaluate the concordance level and robustness of two or more TAD data on any given genome region. The suite provides multiple visual and quantitative metrics to compare the identified domains' number, size, and various comparisons of shared domains, domain boundaries, and domain overlap. CONCLUSIONS: TADMaster is an efficient and easy-to-use web application that provides a set of consensus and unique TADs to inform the choice of TADs. It can be accessed at http://tadmaster.io and is also available as a containerized application that can be deployed and run locally on any platform or operating system. Sean Higgins, Victor Akpokiro, Allen Westcott, Oluwatosin Oluwadare |
BMC Bioinform. | 4 |
| 2021 | DeepSplicer: An Improved Method of Splice Sites Prediction using Deep LearningabstractPost-transcriptional splicing of ribonucleic acid (mRNA) entails removing regions of RNA sequences (Introns) that do not include information for protein synthesis. Thus, accurate splicing site detection is integral for understanding gene structure and, as a result, protein synthesis for biological and medicinal applications. However, the necessity to develop an advanced computational algorithm arises because existing splice site (SS) prediction methods are either computationally inefficient or expensive. Considering this, we present DeepSplicer-a deep learning-based Convolutional Neural Network (CNN) model for locating splice sites. In this work, we compared the ability of the existing SS prediction algorithms model to identify SS in organisms-Homo sapiens, Oryza sativa japonica, Arabidopsis thaliana, DrosophUa melanogaster, and Caenorhabditis elegans-to ours. Using a 5-fold cross-validation test, DeepSplicer achieves an accuracy of 96.65% for acceptor homo sapiens dataset and 94.75% for donor homo sapiens dataset. The datasets used and models generated are available at our GitHub repository here: https://github.com/OluwadareLab/DeeoSolicer. Victor Akpokiro, Oluwatosin Oluwadare, Jugal K. Kalita |
ICMLA | 2 |
| 2019 | GenomeFlow: a comprehensive graphical tool for modeling and analyzing 3D genome structureabstractMOTIVATION: Three-dimensional (3D) genome organization plays important functional roles in cells. User-friendly tools for reconstructing 3D genome models from chromosomal conformation capturing data and analyzing them are needed for the study of 3D genome organization. RESULTS: We built a comprehensive graphical tool (GenomeFlow) to facilitate the entire process of modeling and analysis of 3D genome organization. This process includes the mapping of Hi-C data to one-dimensional (1D) reference genomes, the generation, normalization and visualization of two-dimensional (2D) chromosomal contact maps, the reconstruction and the visualization of the 3D models of chromosome and genome, the analysis of 3D models and the integration of these models with functional genomics data. This graphical tool is the first of its kind in reconstructing, storing, analyzing and annotating 3D genome models. It can reconstruct 3D genome models from Hi-C data and visualize them in real-time. This tool also allows users to overlay gene annotation, gene expression data and genome methylation data on top of 3D genome models. AVAILABILITY AND IMPLEMENTATION: The source code and user manual: https://github.com/jianlin-cheng/GenomeFlow. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tuan Trieu, Oluwatosin Oluwadare, Julia Wopata, Jianlin Cheng |
Bioinform. | 2 |
| 2017 | ClusterTAD: an unsupervised machine learning approach to detecting topologically associated domains of chromosomes from Hi-C dataabstractBACKGROUND: With the development of chromosomal conformation capturing techniques, particularly, the Hi-C technique, the study of the spatial conformation of a genome is becoming an important topic in bioinformatics and computational biology. The Hi-C technique can generate genome-wide chromosomal interaction (contact) data, which can be used to investigate the higher-level organization of chromosomes, such as Topologically Associated Domains (TAD), i.e., locally packed chromosome regions bounded together by intra chromosomal contacts. The identification of the TADs for a genome is useful for studying gene regulation, genomic interaction, and genome function. RESULTS: Here, we formulate the TAD identification problem as an unsupervised machine learning (clustering) problem, and develop a new TAD identification method called ClusterTAD. We introduce a novel method to represent chromosomal contacts as features to be used by the clustering algorithm. Our results show that ClusterTAD can accurately predict the TADs on a simulated Hi-C data. Our method is also largely complementary and consistent with existing methods on the real Hi-C datasets of two mouse cells. The validation with the chromatin immunoprecipitation (ChIP) sequencing (ChIP-Seq) data shows that the domain boundaries identified by ClusterTAD have a high enrichment of CTCF binding sites, promoter-related marks, and enhancer-related histone modifications. CONCLUSIONS: As ClusterTAD is based on a proven clustering approach, it opens a new avenue to apply a large array of clustering methods developed in the machine learning field to the TAD identification problem. The source code, the results, and the TADs generated for the simulated and real Hi-C datasets are available here: https://github.com/BDM-Lab/ClusterTAD . Oluwatosin Oluwadare, Jianlin Cheng |
BMC Bioinform. | 1 |
| 2015 | Iterative reconstruction of three-dimensional models of human chromosomes from chromosomal contact dataabstractBACKGROUND: The entire collection of genetic information resides within the chromosomes, which themselves reside within almost every cell nucleus of eukaryotic organisms. Each individual chromosome is found to have its own preferred three-dimensional (3D) structure independent of the other chromosomes. The structure of each chromosome plays vital roles in controlling certain genome operations, including gene interaction and gene regulation. As a result, knowing the structure of chromosomes assists in the understanding of how the genome functions. Fortunately, the 3D structure of chromosomes proves possible to construct through computational methods via contact data recorded from the chromosome. We developed a unique computational approach based on optimization procedures known as adaptation, simulated annealing, and genetic algorithm to construct 3D models of human chromosomes, using chromosomal contact data. RESULTS: Our models were evaluated using a percentage-based scoring function. Analysis of the scores of the final 3D models demonstrated their effective construction from our computational approach. Specifically, the models resulting from our approach yielded an average score of 80.41%, with a high of 91%, across models for all chromosomes of a normal human B-cell. Comparisons made with other methods affirmed the effectiveness of our strategy. Particularly, juxtaposition with models generated through the publicly available method Markov chain Monte Carlo 5C (MCMC5C) illustrated the outperformance of our approach, as seen through a higher average score for all chromosomes. Our methodology was further validated using two consistency checking techniques known as convergence testing and robustness checking, which both proved successful. CONCLUSIONS: The pursuit of constructing accurate 3D chromosomal structures is fueled by the benefits revealed by the findings as well as any possible future areas of study that arise. This motivation has led to the development of our computational methodology. The implementation of our approach proved effective in constructing 3D chromosome models and proved consistent with, and more effective than, some other methods thereby achieving our goal of creating a tool to help advance certain research efforts. The source code, test data, test results, and documentation of our method, Gen3D, are available at our sourceforge site at: http://sourceforge.net/projects/gen3d/. Jackson Nowotny, Sharif Ahmed, Lingfei Xu, Oluwatosin Oluwadare, Hannah Chen 0002, Noelan Hensley, Tuan Trieu, Renzhi Cao, Jianlin Cheng |
BMC Bioinform. | 4 |