EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Ahuja
dblp:141/0641
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-2837-9361ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Network based simultaneous embedding of cells and marker genes from scRNA-seq studiesabstractThe complexity of scRNA-sequencing datasets highlights the urgent need for enhanced clustering and visualization methods. Here, we propose Stardust, an iterative, force-directed graph layout algorithm that enables the simultaneous embedding of cells and marker genes. Stardust, for the first time, allows a single-stop visualization of cells and marker genes on a single 2D map. While Stardust provides its own visualization pipeline, it can be plugged in with state-of-the-art methods such as Uniform Manifold Approximation and Projection (UMAP) and t-Distributed Stochastic Neighbor Embedding (t-SNE). We benchmarked Stardust against popular visualization and clustering tools on both scRNA-seq and spatial transcriptomics datasets. In all cases, Stardust performs competitively in identifying and visualizing cell types in an accurate and spatially coherent manner. Namrata Bhattacharya, Swagatam Chakraborti, Stuti Kumari, Bernadette Mathew, Abhishek Halder, Sakshi Gujral, Krishan Gupta, Aayushi Mittal, Debajyoti Sinha, Colleen C. Nelson, Tanmoy Chakraborty 0002, Gaurav Ahuja, Debarka Sengupta |
Briefings Bioinform. | 12 |
| 2024 | Literature mining discerns latent disease-gene relationshipsabstractMOTIVATION: Dysregulation of a gene's function, either due to mutations or impairments in regulatory networks, often triggers pathological states in the affected tissue. Comprehensive mapping of these apparent gene-pathology relationships is an ever-daunting task, primarily due to genetic pleiotropy and lack of suitable computational approaches. With the advent of high throughput genomics platforms and community scale initiatives such as the Human Cell Landscape project, researchers have been able to create gene expression portraits of healthy tissues resolved at the level of single cells. However, a similar wealth of knowledge is currently not at our finger-tip when it comes to diseases. This is because the genetic manifestation of a disease is often quite diverse and is confounded by several clinical and demographic covariates. RESULTS: To circumvent this, we mined ∼18 million PubMed abstracts published till May 2019 and automatically selected ∼4.5 million of them that describe roles of particular genes in disease pathogenesis. Further, we fine-tuned the pretrained bidirectional encoder representations from transformers (BERT) for language modeling from the domain of natural language processing to learn vector representation of entities such as genes, diseases, tissues, cell-types, etc., in a way such that their relationship is preserved in a vector space. The repurposed BERT predicted disease-gene associations that are not cited in the training data, thereby highlighting the feasibility of in silico synthesis of hypotheses linking different biological entities such as genes and conditions. AVAILABILITY AND IMPLEMENTATION: PathoBERT pretrained model: https://github.com/Priyadarshini-Rai/Pathomap-Model. BioSentVec-based abstract classification model: https://github.com/Priyadarshini-Rai/Pathomap-Model. Pathomap R package: https://github.com/Priyadarshini-Rai/Pathomap. Priyadarshini Rai, Atishay Jain, Shivani Kumar, Divya Sharma 0002, Neha Jha, Smriti Chawla, Abhijit Raj, Apoorva Gupta, Sarita Poonia, Angshul Majumdar, Tanmoy Chakraborty 0002, Gaurav Ahuja, Debarka Sengupta |
Bioinform. | 12 |
| 2022 | deepGraphh: AI-driven web service for graph-based quantitative structure-activity relationship analysisabstractArtificial intelligence (AI)-based computational techniques allow rapid exploration of the chemical space. However, representation of the compounds into computational-compatible and detailed features is one of the crucial steps for quantitative structure-activity relationship (QSAR) analysis. Recently, graph-based methods are emerging as a powerful alternative to chemistry-restricted fingerprints or descriptors for modeling. Although graph-based modeling offers multiple advantages, its implementation demands in-depth domain knowledge and programming skills. Here we introduce deepGraphh, an end-to-end web service featuring a conglomerate of established graph-based methods for model generation for classification or regression tasks. The graphical user interface of deepGraphh supports highly configurable parameter support for model parameter tuning, model generation, cross-validation and testing of the user-supplied query molecules. deepGraphh supports four widely adopted methods for QSAR analysis, namely, graph convolution network, graph attention network, directed acyclic graph and Attentive FP. Comparative analysis revealed that deepGraphh supported methods are comparable to the descriptors-based machine learning techniques. Finally, we used deepGraphh models to predict the blood-brain barrier permeability of human and microbiome-generated metabolites. In summary, deepGraphh offers a one-stop web service for graph-based methods for chemoinformatics. Vishakha Gautam, Deepti Gupta, Anubhav Ruhela, Aayushi Mittal, Sanjay Kumar Mohanty, Ria Gupta, Chandan Saini, Debarka Sengupta, Natarajan Arul Murugan, Gaurav Ahuja |
Briefings Bioinform. | 12 |
| 2021 | EcTracker: Tracking and elucidating ectopic expression leveraging large-scale scRNA-seq studiesabstractDramatic genomic alterations, either inducible or in a pathological state, dismantle the core regulatory networks, leading to the activation of normally silent genes. Despite possessing immense therapeutic potential, accurate detection of these transcripts is an ever-challenging task, as it requires prior knowledge of the physiological gene expression levels. Here, we introduce EcTracker, an R-/Shiny-based single-cell data analysis web server that bestows a plethora of functionalities that collectively enable the quantitative and qualitative assessments of bona fide cell types or tissue-specific transcripts and, conversely, the ectopically expressed genes in the single-cell ribonucleic acid sequencing datasets. Moreover, it also allows regulon analysis to identify the key transcriptional factors regulating the user-selected gene signatures. To demonstrate the EcTracker functionality, we reanalyzed the CRISPR interference (CRISPRi) dataset of the human embryonic stem cells differentiated into endoderm lineage and identified the prominent enrichment of a specific gene signature in the SMAD2 knockout cells whose identity was ambiguous in the original study. The key distinguishing features of EcTracker lie within its processing speed, availability of multiple add-on modules, interactive graphical user interface and comprehensiveness. In summary, EcTracker provides an easy-to-perform, integrative and end-to-end single-cell data analysis platform that allows decoding of cellular identities, identification of ectopically expressed genes and their regulatory networks, and therefore, collectively imparts a novel dimension for analyzing single-cell datasets. Vishakha Gautam, Aayushi Mittal, Siddhant Kalra, Sanjay Kumar Mohanty, Krishan Gupta, Komal Rani, Srivatsava Naidu, Tripti Mishra, Debarka Sengupta, Gaurav Ahuja |
Briefings Bioinform. | 10 |
| 2021 | The Cellular basis of loss of smell in 2019-nCoV-infected individualsabstractA prominent clinical symptom of 2019-novel coronavirus (nCoV) infection is hyposmia/anosmia (decrease or loss of sense of smell), along with general symptoms such as fatigue, shortness of breath, fever and cough. The identity of the cell lineages that underpin the infection-associated loss of olfaction could be critical for the clinical management of 2019-nCoV-infected individuals. Recent research has confirmed the role of angiotensin-converting enzyme 2 (ACE2) and transmembrane protease serine 2 (TMPRSS2) as key host-specific cellular moieties responsible for the cellular entry of the virus. Accordingly, the ongoing medical examinations and the autopsy reports of the deceased individuals indicate that organs/tissues with high expression levels of ACE2, TMPRSS2 and other putative viral entry-associated genes are most vulnerable to the infection. We studied if anosmia in 2019-nCoV-infected individuals can be explained by the expression patterns associated with these host-specific moieties across the known olfactory epithelial cell types, identified from a recently published single-cell expression study. Our findings underscore selective expression of these viral entry-associated genes in a subset of sustentacular cells (SUSs), Bowman's gland cells (BGCs) and stem cells of the olfactory epithelium. Co-expression analysis of ACE2 and TMPRSS2 and protein-protein interaction among the host and viral proteins elected regulatory cytoskeleton protein-enriched SUSs as the most vulnerable cell type of the olfactory epithelium. Furthermore, expression, structural and docking analyses of ACE2 revealed the potential risk of olfactory dysfunction in four additional mammalian species, revealing an evolutionarily conserved infection susceptibility. In summary, our findings provide a plausible cellular basis for the loss of smell in 2019-nCoV-infected patients. Krishan Gupta, Sanjay Kumar Mohanty, Aayushi Mittal, Siddhant Kalra, Suvendu Kumar, Tripti Mishra, Jatin Ahuja, Debarka Sengupta, Gaurav Ahuja |
Briefings Bioinform. | 9 |
| 2021 | Machine-OlF-Action: a unified framework for developing and interpreting machine-learning models for chemosensory researchabstractSUMMARY: Machine Learning-based techniques are emerging as state-of-the-art methods in chemoinformatics to selectively, effectively and speedily identify biologically relevant molecules from large databases. So far, a multitude of such techniques have been proposed, but unfortunately due to their sparse availability, and the dependency on high-end computational literacy, their wider adaptation faces challenges, at least in the context of G-Protein Coupled Receptors (GPCRs)-associated chemosensory research. Here, we report Machine-OlF-Action (MOA), a user-friendly, open-source computational framework, that utilizes user-supplied SMILES (simplified molecular input line entry system) of the chemicals, along with their activation status, to synthesize classification models. MOA integrates a number of popular chemical databases collectively harboring approximately 103 million chemical moieties. MOA also facilitates customized screening of user-supplied chemical datasets. A key feature of MOA is its ability to embed molecules based on the similarity of their local neighborhood, by utilizing a state-of-the-art model interpretability framework LIME. We demonstrate the utility of MOA in identifying previously unreported agonists for human and mouse olfactory receptors OR1A1 and MOR174-9 by leveraging the chemical features of their known agonists and non-agonists. In summary, here we develop an ML-powered software playground for performing supervisory learning tasks involving chemical compounds. AVAILABILITY AND IMPLEMENTATION: MOA is available for Windows, Mac and Linux operating systems. It's accessible at (https://ahuja-lab.in/). Source code, user manual, step-by-step guide and support is available at GitHub (https://github.com/the-ahuja-lab/Machine-Olf-Action). For results, reproducibility and hyperparameters, refer to Supplementary Notes. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anku Gupta, Mohit Choudhary, Sanjay Kumar Mohanty, Aayushi Mittal, Krishan Gupta, Aditya Arya, Suvendu Kumar, Nikhil Katyayan, Nilesh Kumar Dixit, Siddhant Kalra, Manshi Goel, Megha Sahni, Vrinda Singhal, Tripti Mishra, Debarka Sengupta, Gaurav Ahuja |
Bioinform. | 16 |