EDBT 2026 Demo / reviewers in the wild / expert
Rajesh Kumar Sani
dblp:311/0007
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Machine Learning Approach to Estimate Volumetric Quantification of 2D SEM Images of BiofilmsabstractScanning Electron Microscopy (SEM) is one of the most important imaging techniques to understand the dynamics of microscale objects. However, SEM images could mostly produce two-dimensional images which limit the exploration of volumetric information. To address this limitation, we propose the integration of a state-of-the-art machine learning approach that leverages Confocal Laser Scanning Microscopy (CLSM) image data for training a deep learning model. The objective of this model is to approximate volumetric information by taking a 2D SEM image as input and generating corresponding depth information. The depth of information then be utilized to estimate volumetric quantification such as biofilm density and bacteria cell counts. Dilanga Abeyrathna, Ram Singh, Samira Badrloo, Rajesh Kumar Sani, Mahadevan Subramaniam |
BIBM | 4 |
| 2024 | NFB-Checker: An AI/ML-Powered Microorganism Nitrogen Fixation Susceptibility Prediction from Gene CollectionabstractNitrogen fixation is a crucial process involved in many aspects of the ecological life cycle. However, few organismsparticularly kingdom bacteria-are known to fix N2in a form that could be used in agribusiness. The ability to identify an organism as an N2fixer is a novel case for more identification and application. This study employs functional genes and an XGBoost machine learning model to predict an organism's ability to fix nitrogen, focusing on key orthologs like nifH, nifD, and nifK, essential in the nitrogenase enzyme complex. The model, trained on a dataset of 923 N2fixers and 981 non-fixers, achieved an accuracy of 92.6% and an F1 score of 0.92. This research highlights the potential of machine learning in identifying genetic markers for N2fixation, providing an efficient alternative to traditional methods, and offers new insights for ecological and agricultural research. The inclusion of specific orthologs enhances the model's predictive accuracy, demonstrating the importance of targeted genetic markers in computational biology. Tuyen Do, Shiva Aryal, Bichar Dip Shrestha Gurung, Diing D. M. Agany, Nick Klein, Ruanbao Zhou, Rajesh Kumar Sani, Etienne Z. Gnimpieba |
BIBM | 7 |
| 2024 | Biofilm marker discovery with cloud-based dockerized metagenomics analysis of microbial communitiesabstractIn an environment, microbes often work in communities to achieve most of their essential functions, including the production of essential nutrients. Microbial biofilms are communities of microbes that attach to a nonliving or living surface by embedding themselves into a self-secreted matrix of extracellular polymeric substances. These communities work together to enhance their colonization of surfaces, produce essential nutrients, and achieve their essential functions for growth and survival. They often consist of diverse microbes including bacteria, viruses, and fungi. Biofilms play a critical role in influencing plant phenotypes and human microbial infections. Understanding how these biofilms impact plant health, human health, and the environment is important for analyzing genotype-phenotype-driven rule-of-life functions. Such fundamental knowledge can be used to precisely control the growth of biofilms on a given surface. Metagenomics is a powerful tool for analyzing biofilm genomes through function-based gene and protein sequence identification (functional metagenomics) and sequence-based function identification (sequence metagenomics). Metagenomic sequencing enables a comprehensive sampling of all genes in all organisms present within a biofilm sample. However, the complexity of biofilm metagenomic study warrants the increasing need to follow the Findability, Accessibility, Interoperability, and Reusable (FAIR) Guiding Principles for scientific data management. This will ensure that scientific findings can be more easily validated by the research community. This study proposes a dockerized, self-learning bioinformatics workflow to increase the community adoption of metagenomics toolkits in a metagenomics and meta-transcriptomics investigation. Our biofilm metagenomics workflow self-learning module includes integrated learning resources with an interactive dockerized workflow. This module will allow learners to analyze resources that are beneficial for aggregating knowledge about biofilm marker genes, proteins, and metabolic pathways as they define the composition of specific microbial communities. Cloud and dockerized technology can allow novice learners-even those with minimal knowledge in computer science-to use complicated bioinformatics tools. Our cloud-based, dockerized workflow splits biofilm microbiome metagenomics analyses into four easy-to-follow submodules. A variety of tools are built into each submodule. As students navigate these submodules, they learn about each tool used to accomplish the task. The downstream analysis is conducted using processed data obtained from online resources or raw data processed via Nextflow pipelines. This analysis takes place within Vertex AI's Jupyter notebook instance with R and Python kernels. Subsequently, results are stored and visualized in Google Cloud storage buckets, alleviating the computational burden on local resources. The result is a comprehensive tutorial that guides bioinformaticians of any skill level through the entire workflow. It enables them to comprehend and implement the necessary processes involved in this integrated workflow from start to finish. This manuscript describes the development of a resource module that is part of a learning platform named "NIGMS Sandbox for Cloud-based Learning" https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses. Etienne Z. Gnimpieba, Timothy W. Hartman, Tuyen Do, Jessica Zylla, Shiva Aryal, Samuel J. Haas, Diing D. M. Agany, Bichar Dip Shrestha Gurung, Valena Doe, Zelaikha B. Yosufzai, Daniel Pan, Ross Campbell, Victor C. Huber, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Carol Lushbough |
Briefings Bioinform. | 14 |
| 2023 | Genome wide computational prediction and analysis of noncoding genome of corrosive biofilm forming Oleidesulfovibrio alaskensis G20abstractBiofilm, is a special-complex organization of bacterial cells with multiple layers, formed in certain environmental conditions. In the complex biofilm different bacterial cells perform different functions and contribute to overall biofilm activity, in different physiological states of growth, stress and signaling. Sulphate reducing bacteria (SRB) contributes to huge economic loss ($4B) causing microbial induced corrosion. To effectively combat the challenges posed by SRB, it is essential to understand their molecular mechanisms of biofilm formation and biocorrosion. Regulation of all key pathways and mechanisms involved in growth, stress management and biofilm formation performed by non-coding genomes. Here, in this study we have identified genome wide distribution of non-coding genome (ncRNAs) in a biofilm forming and corrosive SRB Oleidesulfovibrio alaskensis (earlier Desulfovibrio alaskensis) strain G20 (OA-G20). It is understood that regulatory small RNAs play key role in therefore objective of this study was to identify the noncoding RNAs in OA-G20 genome and unveil their role in biofilm formation and corrosive nature. This is the first study to identify ncRNAs in SRB. Here, we are presenting the results using computational methods to identify noncoding genome of OA-G20 and potential role in biofilm formation. Ram Nageena Singh, Etienne Z. Gnimpieba, Rajesh Kumar Sani |
BIBM | 3 |
| 2022 | Attention model-based and multi-organism driven gene recognition from text: application to a microbial biofilm organism setabstractNowadays, online databases such as PUBMED and PMC are experiencing an explosion of publications in the field of biomedical sciences. With so much information available online, one of the biggest challenges is managing all that raw, unstructured data and making it machine-readable. Name entity recognition is nowadays a prerequisite for data identification and extraction in biosciences. One of the areas that allows automatic extraction of information from biomedical literature today is Name Entity Recognition. Indeed, it makes it possible to simplify the workflow analysis and automatic extraction of name entities, thus improving the various existing models. There is in the literature a lot of tools for this purpose, but they are unable to extract microbial genes accurately. Moreover, current goal standard corpora such as BIOCREATIVE I to IV have limited representation of microbial knowledge. In this paper, we proposed a new method to recognize biofilm gene mentions from free text. This method relies on a context-specific dictionary to annotate a consistent corpus necessary to train an efficient recognition model. Indeed, this method provides a new workflow for dataset collection generation for microbial biofilm gene. Trained on a set of biofilm organisms our method achieves a score of up to 94%, outperforming state-of-the-art frameworks. Alain Bertrand Bomgni, Ernest Basile Fotseu Fotseu, Daril Raoul Kengne Wambo, Rajesh Kumar Sani, Carol Lushbough, Etienne Z. Gnimpieba |
BIBM | 4 |
| 2022 | Using BASIN-ML for Machine Learning-Based Statistical Analysis and Reporting for Biofilm DatasetsabstractBiological and biomedical microscope image (bioimage) comparison remains useful to approach many research challenges—from biofilms to human diseases. This powerful technology allows researchers to provide the community with a quick visual snapshot of varying experimental conditions. But a two-condition comparison still relies on a researcher’s eyes to draw conclusions despite the availability of multiple— often complex—digital image analysis tools. Our Bioimage Analysis, Statistic, and Comparison (BASIN) software provides an easy, objective, reproducible comparison leveraging inferential statistics to bridge image data analysis with other biomedical data modalities such as gene expression. Users have access to a machine learning module to assist with image segmentation using modern, trainable algorithms. BASIN also provides several key data points including images’ object counts, net and mean pixel intensities, net and mean object surface areas, plus a variety of other potentially useful data. Hypothesis testing is performed on mean object intensities and surface areas using the statistical power of the R programming language. These features allow BASIN to extend the current scope of image comparison. It gives researchers a multi-model knowledge about matters such as drug protein marker response, the significance of cell population changes, and changes in cell morphology. To improve BASIN’s accessibility and transparency we implemented it in R using Shiny framework and provided both an online trial version and a customizable offline version. We also have a batch version to run on datasets with hundreds of biomedical images. BASIN workflows consist of five core modules including image upload, feature extraction, statistical analysis, visualization, and report generation. Sam Haas, Timothy W. Hartman, Bichar Dip Shrestha Gurung, Tuyen Do, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 5 |
| 2022 | Title: Challenges in single cells sequencing Microbial community and biofilm: A case of Oleidesulfovibrio alaskensis G20 NGS protocolabstractBiofilm, is a special-complex organization of bacterial cells with multiple layers, formed in certain environmental conditions. In the complex biofilm different bacterial (single) cells perform different functions and contribute to overall biofilm activity, where single bacterial cells are in different physiological states of growth, stress and signaling. Sulphate reducing bacteria (SRB) contributes to huge economic loss (${\$}$4B) causing microbial induced corrosion. To effectively combat the challenges posed by SRB, it is essential to understand their molecular mechanisms of biofilm formation and biocorrosion. Single Cell Genomics (single cell transcriptomics) is a technique which investigates the gene expression at the level of a single biological cell and could not be attainable in bulk analysis (RNASeq) could decipher the uniqueness of each cell in complex group. We are developing a NGS (oxford Nanopore)-workflow for single cell transcriptomics of biofilms using Oleidesulfovibrio alaskensis (earlier Desulfovibrio alaskensis) strain G20 (OA-G20) as a model. We encountered several challenges to design and develop the workflow such 1) absence of ploy(A) tail in mRNAs, 2) amount of total RNA, 3) half-lives of mRNAs, 4) perforation of cell membrane, 5) low copy number of mRNAs, 6) single cell suspension and cell sorting, 7) quantification of bacterial cells, 8) fixation of cells, 9) designing of probes, 10) hybridization of probes and 11) amplification of mRNAs. Here, we are presenting how we resolved some of these challenges to design and develop a single cell genomics (transcriptomics) workflow for SRB biofilms. Ram Nageena Singh, Etienne Z. Gnimpieba, Rajesh Kumar Sani |
BIBM | 3 |
| 2021 | Segmentation of Bacterial Cells in Biofilms Using an Overlapped Ellipse Fitting TechniqueabstractAccurate detection and segmentation of bacterial cells in the microscopy images of a biofilm is essential to develop technologies to resist microbial corrosion. The traditional approach of manually identifying cell regions in microscopy images is a time-consuming and error-prone task. Nonetheless, many of the existing approaches, including automated systems adopting advanced machine learning models, find it challenging to detect and segment cell instances in clustered biofilms where cells are overlapping and touching each other. In this paper, we develop a method to segment and extract the size properties of all cells. The proposed method consists of two stages, a semantic segmentation stage based on a U-Net architecture followed by a region-based ellipse fitting technique for instance segmentation and size property extraction. We compared the performance of our approach against a widely used object segmentation approach namely Mask R-CNN and found that our algorithm outperformed Mask R-CNN in terms of the segmentation efficiency and cell size estimation for images of Bacillus subtilis biofilms. Dilanga Abeyrathna, Terrance Life, Shailabh Rauniyar, Shankarachary Ragi, Rajesh Kumar Sani, Parvathi Chundi |
BIBM | 5 |
| 2021 | GenNER - A highly scalable and optimal NER method for text-based gene and protein recognitionabstractNowadays, there are a large number of models in the scientific literature capable of recognizing and extracting gene mentions from a given text. Several data sets have been developed to facilitate the learning process of these models. However, very few models are able to increase their knowledge and performance progressively from new annotated text but also to take into account the granularity of the input text of the model. Our proposed solution, GenNER, is a method for recognizing gene/protein mentions from free text. GenNER relies on continuous learning and a text granularization algorithm as input to the model, which allows it to achieve better performance. Its evaluation process was done around BioCreative II annotated datasets; we obtained an average F1-score of 0.9704, which outperforms current methods. Ernest Basile Fotseu Fotseu, Thierry Kongne Nembot, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba, Alain Bertrand Bomgni |
BIBM | 3 |
| 2021 | Prediction of essential genes in G20 using machine learning modelabstractDespite the exponential growth in bioscience data, one of the key challenges for machine learning engineers remains the incompleteness of bioscience dataset (biodata). For a specific bioscience problem such as (e.g. biofilm formation, drug response, organism survival), it is very difficult to find a good consistent dataset capturing the numerous variables involved in each of these processes. Each systems biology data point is measured with different protocols in different settings, making their integration hard and not reliable. This paper focuses on using machine learning (ML) models and data mining (DM) workflow to perform gene essential prediction in G20. Actually, developing next-generation and nano-scale coatings to control biofilm formation on technologically relevant materials is a great challenge today. This can help to control microbial corrosion on material or engineer better relevant material. To tackle this relevant problem, a detailed understanding of the bacterial survival mechanisms is crucial. Computational methods for predicting essential genes can make it easier and faster to obtain reliable results. Method: The main hypothesis of our work is that a minimal information-driven specific Machine Learning model can outperform an interesting prediction score. To reach our goal, we set up first a completed data mining workflow to extract gene features from G20. We then derive 10192 features from gene sequence and protein sequence divided into 25 relevant subgroups. From each subgroup, we build a couple of interesting machine learning models. Result: We identified 69 relevant subgroups of features using our features selection algorithm. We tested the model performance on each of these subgroups and our predictive result achieved up to 98% accuracy score. These subgroups of features can be used to assist researchers to select good variables for their respective experiments. Thierry Kongne Nembot, Ernest Basile Fotseu Fotseu, Rajesh Kumar Sani, Etienne Z. Gnimpieba, Carol Lushbough, Alain Bertrand Bomgni |
BIBM | 3 |
| 2021 | Integration of text mining and biological network analysis to access essential genes in Desulfovibrio alaskensis G20abstractEssential genes are crucial for the survival and growth of any organism, and therefore alteration of such genes could result in unexpected behavioral change. Identification of essential genes and their role in functioning of organisms is a basic knowledge requirement for any research, which could be manipulated to understand the mechanisms of survival and growth [1]. Several decades have witnessed the virtues and iniquities of the gram-negative facultative anaerobes, sulfate reducing bacteria (SRB) in both ecological and commercial arena. Despite of relentless increase in the number of published articles that belong to diverse research areas-from industrial biotechnology (removal of heavy metals and waste valorization) to molecular biology (genetic architecture of the genes in biocorrosion and biofilm formation on metals), not much information about the essential genes of SRB community is known yet [2]. The Desulfovibrio alaskensis G20 (DA-G20) is a well-known SRB; its genes have been annotated but have large numbers that encode for hypothetical proteins. Till date no categorization is available for the genes of DA-G20 with reference to essentiality [3]. The in-vitro prediction of essential genes relies highly on the exhaustive multi-omics strategies. In the era of big-data and artificial intelligence research, demand of abstraction and interpretation of complex relationships of biological importance using text mining has increased. Therefore, to propose an alternative and economic method, text mining is a comparable method for the prediction of the essential genes. In this study, we reported the essential genes of DA-G20 using text mining and biological network analysis. Moreover, the present work provides a foundation for the expansion of genome wide investigation and identification of essential genes in prokaryotes using machine learning and data science approaches. Priya Saxena, Abhilash Kumar Tripathi, Payal Thakur, Shailabh Rauniyar, Vinoj Gopalakrishnan, Ram Nageena Singh, Mathew A. Olakunle, Etienne Z. Gnimpieba, Rajesh Kumar Sani |
BIBM | 9 |
| 2021 | Identifying genes involved in biocorrosion from the literature using text-miningabstractBackground: The central repository of scientific models from the literature; however, the manual knowledge is research publications which also plays a selection of pertinent information from databases like crucial role in communication within the scientific PubMed, PMC, Dimension, Google scholar, and community. The public repositories like PubMed, Semantic scholar can be tedious; therefore, a robust PMC, Dimension, Google scholar, and Semantic approach like text mining can be used for this process. scholar act as a storehouse of biological systems data. Text mining can be defined as a practical approach to A substantial amount of information can be recovered extracting biologically relevant information from the in a semi-structured form in the literature. The main growing amount of published literature. It comprises obstacle to large-scale analysis of this kind of data is three main tasks: information retrieval from relevant their highly unstructured and heterogenous format, documents, extraction of information of interest, and making it even harder to extract information contained data mining, which allows identifying new within the literature. Nonetheless, this information is associations among the extracted set of information. inherently helpful in a variety of genomics and Here, we show that a text mining approach can exploit systems biology contexts. For example, it is a standard large literature databases like PubMed and PMC to practice in the genomics community to manually extract genes/proteins related to biocorrosion by curate and extract literature-derived protein-protein Sulfate-reducing bacteria(SRB). The corrosion of metal due to microbial activity is known as biocorrosion or MIC(Microbial Induced corrosion). The primary class of bacteria associated with corrosion of metals in aquatic and terrestrial habitats is Sulfur Reducing Bacteria(SRB). Biocorrosion results from collaborative interactions between the metal surface, corrosion products, and bacterial cells and their metabolites. SRB are nonpathogenic and anaerobic bacteria, but SRB can act as a catalyst in the reduction reaction of sulfate to sulfide. It means they can make severe corrosion of metals in a water system by producing enzymes, which can accelerate the reduction of sulphate compounds to hydrogen sulfide. MIC is also known as metabolite corrosion or chemical microbially influenced corrosion (CMIC) owing to the generation of corrosive metabolite (hydrogen sulfide).In contrast, corrosion through direct withdrawal of electrons is called electrical microbial influenced corrosion. The presence of biofilm affects microbial corrosion; It is recognized that under the biofilm at the metal/biofilm interface, the concentrations of acidic metabolites are much greater, and their impact is amplified, leading to higher metal corrosion. It is also becoming apparent that one predominant mechanism of biocorrosion does not exist, and experimentally validating each of these theories can be laborious. Therefore, there is a need for an advanced technique for identifying genes and proteins of SRB involved in biocorrosion; this can help construct other biological processes, related pathways, and other processes associated with these genes. Payal Thakur, Shailabh Rauniyar, Abhilash Kumar Tripathi, Priya Saxena, Vinoj Gopalakrishnan, Ram Nageena Singh, Mathew A. Olakunle, Etienne Z. Gnimpieba, Rajesh Kumar Sani |
BIBM | 9 |
| 2021 | Discovery of genes associated with sulfate-reducing bacteria biofilm using text mining and biological network analysisabstractBacterial biofilms are complex surface attached communities of bacteria glued together by extracellular polymeric substance (EPS) matrix, secreted proteins, and extracellular DNAs [1]. Biofilm show reduced growth rates and metabolism. Biofilm formation is a survival mechanism that provides with better options compared to their planktonic counterparts. It impart bacterial communities stronger ability to grow in oligotrophic environments, greater access to nutritional resources, and enhanced syntropic interactions as well as greater tolerance towards environmental stress [2]. Biofilm play a detrimental role in many areas such as healthcare, food industry, water distribution systems, oil and gas industry etc. The composition of biofilm microbial community is varies depending on the environment in which it is formed. Biofilms are stratified formations where deeper layers maintain anoxic conditions. These anoxic niches promote the growth of certain specific groups, including sulfate reducing bacteria (SRB), that use the surface (usually metal) as resources for their survival. Abhilash Kumar Tripathi, Priya Saxena, Payal Thakur, Shailabh Rauniyar, Vinoj Gopalakrishnan, Ram Nageena Singh, Mathew A. Olakunle, Etienne Z. Gnimpieba, Rajesh Kumar Sani |
BIBM | 9 |