EDBT 2026 Demo / reviewers in the wild / expert
Debarka Sengupta
dblp:14/11132
· DBLP profile ↗
19ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Network based simultaneous embedding of cells and marker genes from scRNA-seq studiesabstractThe complexity of scRNA-sequencing datasets highlights the urgent need for enhanced clustering and visualization methods. Here, we propose Stardust, an iterative, force-directed graph layout algorithm that enables the simultaneous embedding of cells and marker genes. Stardust, for the first time, allows a single-stop visualization of cells and marker genes on a single 2D map. While Stardust provides its own visualization pipeline, it can be plugged in with state-of-the-art methods such as Uniform Manifold Approximation and Projection (UMAP) and t-Distributed Stochastic Neighbor Embedding (t-SNE). We benchmarked Stardust against popular visualization and clustering tools on both scRNA-seq and spatial transcriptomics datasets. In all cases, Stardust performs competitively in identifying and visualizing cell types in an accurate and spatially coherent manner. Namrata Bhattacharya, Swagatam Chakraborti, Stuti Kumari, Bernadette Mathew, Abhishek Halder, Sakshi Gujral, Krishan Gupta, Aayushi Mittal, Debajyoti Sinha, Colleen C. Nelson, Tanmoy Chakraborty 0002, Gaurav Ahuja, Debarka Sengupta |
Briefings Bioinform. | 13 |
| 2024 | Literature mining discerns latent disease-gene relationshipsabstractMOTIVATION: Dysregulation of a gene's function, either due to mutations or impairments in regulatory networks, often triggers pathological states in the affected tissue. Comprehensive mapping of these apparent gene-pathology relationships is an ever-daunting task, primarily due to genetic pleiotropy and lack of suitable computational approaches. With the advent of high throughput genomics platforms and community scale initiatives such as the Human Cell Landscape project, researchers have been able to create gene expression portraits of healthy tissues resolved at the level of single cells. However, a similar wealth of knowledge is currently not at our finger-tip when it comes to diseases. This is because the genetic manifestation of a disease is often quite diverse and is confounded by several clinical and demographic covariates. RESULTS: To circumvent this, we mined ∼18 million PubMed abstracts published till May 2019 and automatically selected ∼4.5 million of them that describe roles of particular genes in disease pathogenesis. Further, we fine-tuned the pretrained bidirectional encoder representations from transformers (BERT) for language modeling from the domain of natural language processing to learn vector representation of entities such as genes, diseases, tissues, cell-types, etc., in a way such that their relationship is preserved in a vector space. The repurposed BERT predicted disease-gene associations that are not cited in the training data, thereby highlighting the feasibility of in silico synthesis of hypotheses linking different biological entities such as genes and conditions. AVAILABILITY AND IMPLEMENTATION: PathoBERT pretrained model: https://github.com/Priyadarshini-Rai/Pathomap-Model. BioSentVec-based abstract classification model: https://github.com/Priyadarshini-Rai/Pathomap-Model. Pathomap R package: https://github.com/Priyadarshini-Rai/Pathomap. Priyadarshini Rai, Atishay Jain, Shivani Kumar, Divya Sharma 0002, Neha Jha, Smriti Chawla, Abhijit Raj, Apoorva Gupta, Sarita Poonia, Angshul Majumdar, Tanmoy Chakraborty 0002, Gaurav Ahuja, Debarka Sengupta |
Bioinform. | 13 |
| 2023 | A combined experimental-computational approach uncovers a role for the Golgi matrix protein Giantin in breast cancer progressionabstractOur understanding of how speed and persistence of cell migration affects the growth rate and size of tumors remains incomplete. To address this, we developed a mathematical model wherein cells migrate in two-dimensional space, divide, die or intravasate into the vasculature. Exploring a wide range of speed and persistence combinations, we find that tumor growth positively correlates with increasing speed and higher persistence. As a biologically relevant example, we focused on Golgi fragmentation, a phenomenon often linked to alterations of cell migration. Golgi fragmentation was induced by depletion of Giantin, a Golgi matrix protein, the downregulation of which correlates with poor patient survival. Applying the experimentally obtained migration and invasion traits of Giantin depleted breast cancer cells to our mathematical model, we predict that loss of Giantin increases the number of intravasating cells. This prediction was validated, by showing that circulating tumor cells express significantly less Giantin than primary tumor cells. Altogether, our computational model identifies cell migration traits that regulate tumor progression and uncovers a role of Giantin in breast cancer progression. Salim Ghannoum, Damiano Fantini, Muhammad Zahoor, Veronika Reiterer, Santosh Phuyal, Waldir Leoncio Netto, Øystein Sørensen, Arvind Iyer, Debarka Sengupta, Lina Prasmickaite, Gunhild Mari Mælandsmo, Alvaro Köhn-Luque, Hesso Farhan |
PLoS Comput. Biol. | 9 |
| 2022 | deepGraphh: AI-driven web service for graph-based quantitative structure-activity relationship analysisabstractArtificial intelligence (AI)-based computational techniques allow rapid exploration of the chemical space. However, representation of the compounds into computational-compatible and detailed features is one of the crucial steps for quantitative structure-activity relationship (QSAR) analysis. Recently, graph-based methods are emerging as a powerful alternative to chemistry-restricted fingerprints or descriptors for modeling. Although graph-based modeling offers multiple advantages, its implementation demands in-depth domain knowledge and programming skills. Here we introduce deepGraphh, an end-to-end web service featuring a conglomerate of established graph-based methods for model generation for classification or regression tasks. The graphical user interface of deepGraphh supports highly configurable parameter support for model parameter tuning, model generation, cross-validation and testing of the user-supplied query molecules. deepGraphh supports four widely adopted methods for QSAR analysis, namely, graph convolution network, graph attention network, directed acyclic graph and Attentive FP. Comparative analysis revealed that deepGraphh supported methods are comparable to the descriptors-based machine learning techniques. Finally, we used deepGraphh models to predict the blood-brain barrier permeability of human and microbiome-generated metabolites. In summary, deepGraphh offers a one-stop web service for graph-based methods for chemoinformatics. Vishakha Gautam, Deepti Gupta, Anubhav Ruhela, Aayushi Mittal, Sanjay Kumar Mohanty, Ria Gupta, Chandan Saini, Debarka Sengupta, Natarajan Arul Murugan, Gaurav Ahuja |
Briefings Bioinform. | 10 |
| 2022 | SelfE: Gene Selection via Self-Expression for Single-Cell DataabstractSingle-cell RNA sequencing has been proved to be advantageous in discerning molecular heterogeneity in seemingly similar cells in a tissue. Due to the paucity of starting RNA, a large fraction of transcripts fail to amplify during the polymerase chain reaction cycle. This gets compounded by trivial biological noise such as variability in the cell cycle specific genes. As a result expression matrix obtained from a single-cell study is highly sparse with a large number of missing values. This hinders downstream analysis of single-cell expression data. It has been observed that feature engineering significantly improves the analysis outcomes. Feature extraction methods such as principal component analysis and zero-inflated factor analysis have been shown to be useful for subsequent steps of data analysis including clustering. However, too little or no visible efforts have been observed for developing feature selection techniques, which offer transparency for the analyst's consumption. We propose SelfE, a novel$l_{2,0}$-minimization algorithm that determines an optimal subset of feature vectors that preserves sub-space structures as observed in the data. We compared SelfE with the commonly used feature selection methods for single-cell expression data analysis. Priyadarshini Rai, Debarka Sengupta, Angshul Majumdar |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Enhash: A Fast Streaming Algorithm For Concept Drift DetectionabstractWe propose Enhash, a fast ensemble learner that detects concept drift in a data stream.A stream may consist of abrupt, gradual, virtual, or recurring events, or a mixture of various types of drift.Enhash employs projection hash to insert an incoming sample.Benchmark tests on 6 artificial and 4 real data sets consisting of various types of drift show that Enhash is competitive with stateof-the-art ensemble learners while being significantly faster.It also has moderate resource requirements. 59 Aashi Jindal, Debarka Sengupta, Jayadeva |
ESANN | 3 |
| 2021 | EcTracker: Tracking and elucidating ectopic expression leveraging large-scale scRNA-seq studiesabstractDramatic genomic alterations, either inducible or in a pathological state, dismantle the core regulatory networks, leading to the activation of normally silent genes. Despite possessing immense therapeutic potential, accurate detection of these transcripts is an ever-challenging task, as it requires prior knowledge of the physiological gene expression levels. Here, we introduce EcTracker, an R-/Shiny-based single-cell data analysis web server that bestows a plethora of functionalities that collectively enable the quantitative and qualitative assessments of bona fide cell types or tissue-specific transcripts and, conversely, the ectopically expressed genes in the single-cell ribonucleic acid sequencing datasets. Moreover, it also allows regulon analysis to identify the key transcriptional factors regulating the user-selected gene signatures. To demonstrate the EcTracker functionality, we reanalyzed the CRISPR interference (CRISPRi) dataset of the human embryonic stem cells differentiated into endoderm lineage and identified the prominent enrichment of a specific gene signature in the SMAD2 knockout cells whose identity was ambiguous in the original study. The key distinguishing features of EcTracker lie within its processing speed, availability of multiple add-on modules, interactive graphical user interface and comprehensiveness. In summary, EcTracker provides an easy-to-perform, integrative and end-to-end single-cell data analysis platform that allows decoding of cellular identities, identification of ectopically expressed genes and their regulatory networks, and therefore, collectively imparts a novel dimension for analyzing single-cell datasets. Vishakha Gautam, Aayushi Mittal, Siddhant Kalra, Sanjay Kumar Mohanty, Krishan Gupta, Komal Rani, Srivatsava Naidu, Tripti Mishra, Debarka Sengupta, Gaurav Ahuja |
Briefings Bioinform. | 9 |
| 2021 | The Cellular basis of loss of smell in 2019-nCoV-infected individualsabstractA prominent clinical symptom of 2019-novel coronavirus (nCoV) infection is hyposmia/anosmia (decrease or loss of sense of smell), along with general symptoms such as fatigue, shortness of breath, fever and cough. The identity of the cell lineages that underpin the infection-associated loss of olfaction could be critical for the clinical management of 2019-nCoV-infected individuals. Recent research has confirmed the role of angiotensin-converting enzyme 2 (ACE2) and transmembrane protease serine 2 (TMPRSS2) as key host-specific cellular moieties responsible for the cellular entry of the virus. Accordingly, the ongoing medical examinations and the autopsy reports of the deceased individuals indicate that organs/tissues with high expression levels of ACE2, TMPRSS2 and other putative viral entry-associated genes are most vulnerable to the infection. We studied if anosmia in 2019-nCoV-infected individuals can be explained by the expression patterns associated with these host-specific moieties across the known olfactory epithelial cell types, identified from a recently published single-cell expression study. Our findings underscore selective expression of these viral entry-associated genes in a subset of sustentacular cells (SUSs), Bowman's gland cells (BGCs) and stem cells of the olfactory epithelium. Co-expression analysis of ACE2 and TMPRSS2 and protein-protein interaction among the host and viral proteins elected regulatory cytoskeleton protein-enriched SUSs as the most vulnerable cell type of the olfactory epithelium. Furthermore, expression, structural and docking analyses of ACE2 revealed the potential risk of olfactory dysfunction in four additional mammalian species, revealing an evolutionarily conserved infection susceptibility. In summary, our findings provide a plausible cellular basis for the loss of smell in 2019-nCoV-infected patients. Krishan Gupta, Sanjay Kumar Mohanty, Aayushi Mittal, Siddhant Kalra, Suvendu Kumar, Tripti Mishra, Jatin Ahuja, Debarka Sengupta, Gaurav Ahuja |
Briefings Bioinform. | 8 |
| 2021 | Machine-OlF-Action: a unified framework for developing and interpreting machine-learning models for chemosensory researchabstractSUMMARY: Machine Learning-based techniques are emerging as state-of-the-art methods in chemoinformatics to selectively, effectively and speedily identify biologically relevant molecules from large databases. So far, a multitude of such techniques have been proposed, but unfortunately due to their sparse availability, and the dependency on high-end computational literacy, their wider adaptation faces challenges, at least in the context of G-Protein Coupled Receptors (GPCRs)-associated chemosensory research. Here, we report Machine-OlF-Action (MOA), a user-friendly, open-source computational framework, that utilizes user-supplied SMILES (simplified molecular input line entry system) of the chemicals, along with their activation status, to synthesize classification models. MOA integrates a number of popular chemical databases collectively harboring approximately 103 million chemical moieties. MOA also facilitates customized screening of user-supplied chemical datasets. A key feature of MOA is its ability to embed molecules based on the similarity of their local neighborhood, by utilizing a state-of-the-art model interpretability framework LIME. We demonstrate the utility of MOA in identifying previously unreported agonists for human and mouse olfactory receptors OR1A1 and MOR174-9 by leveraging the chemical features of their known agonists and non-agonists. In summary, here we develop an ML-powered software playground for performing supervisory learning tasks involving chemical compounds. AVAILABILITY AND IMPLEMENTATION: MOA is available for Windows, Mac and Linux operating systems. It's accessible at (https://ahuja-lab.in/). Source code, user manual, step-by-step guide and support is available at GitHub (https://github.com/the-ahuja-lab/Machine-Olf-Action). For results, reproducibility and hyperparameters, refer to Supplementary Notes. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anku Gupta, Mohit Choudhary, Sanjay Kumar Mohanty, Aayushi Mittal, Krishan Gupta, Aditya Arya, Suvendu Kumar, Nikhil Katyayan, Nilesh Kumar Dixit, Siddhant Kalra, Manshi Goel, Megha Sahni, Vrinda Singhal, Tripti Mishra, Debarka Sengupta, Gaurav Ahuja |
Bioinform. | 15 |
| 2021 | Linear time identification of local and global outliers
Aashi Jindal, Jayadeva, Debarka Sengupta |
Neurocomputing | 4 |
| 2021 | Stable feature selection using copula based mutual information
Snehalika Lall, Debajyoti Sinha, Abhik Ghosh, Debarka Sengupta, Sanghamitra Bandyopadhyay |
Pattern Recognit. | 4 |
| 2021 | Hide and Seek: Outwitting Community Detection AlgorithmsabstractCommunity affiliation of a node plays an important role in determining its contextual position in the network, which may raise privacy concerns when a sensitive node wants to hide its identity in a network. Oftentimes, a target community seeks to protect itself from adversaries so that its constituent members remain hidden inside the network. The current study focuses on hiding such sensitive communities so that the community affiliation of the targeted nodes can be concealed. This leads to the problem of community deception, which investigates the avenues of minimally rewiring nodes in a network so that a given target community maximally hides from a community detection algorithm (CDA). We formalize the problem of community deception and introduce Network deception using permanence loss (NEURAL), a novel method that greedily optimizes a node-centric objective function to determine the rewiring strategy. Theoretical settings pose a restriction on the number of strategies that can be employed to optimize the objective function, which in turn reduces the overhead of choosing the best strategy from multiple options. We also show that our objective function is submodular and monotone. When tested on both synthetic and seven real-world networks, NEURAL is able to deceive six widely used CDAs. We benchmark its performance with respect to four state-of-the-art methods on four evaluation metrics. In addition, our qualitative analysis on three other attributed real-world networks reveals that NEURAL, quite strikingly, captures important metainformation about edges that otherwise could not be inferred by observing only their topological structures. Shravika Mittal, Debarka Sengupta, Tanmoy Chakraborty 0002 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2020 | Cluster Aware Deep Dictionary Learning for Single Cell Analysis
Priyadarshini Rai, Angshul Majumdar, Debarka Sengupta |
ICONIP (5) | 3 |
| 2020 | Improved dropClust R package with integrative analysis support for scRNA-seq dataabstractSUMMARY: DropClust leverages Locality Sensitive Hashing (LSH) to speed up clustering of large scale single cell expression data. Here we present the improved dropClust, a complete R package that is, fast, interoperable and minimally resource intensive. The new dropClust features a novel batch effect removal algorithm that allows integrative analysis of single cell RNA-seq (scRNA-seq) datasets. AVAILABILITY AND IMPLEMENTATION: dropClust is freely available at https://github.com/debsin/dropClust as an R package. A lightweight online version of the dropClust is available at https://debsinha.shinyapps.io/dropClust/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Debajyoti Sinha, Pradyumn Sinha, Ritwik Saha, Sanghamitra Bandyopadhyay, Debarka Sengupta |
Bioinform. | 5 |
| 2017 | A Scoring Scheme for Online Feature Selection: Simulating Model Performance Without RetrainingabstractIncreasing the number of features increases the complexity of a model even if the additional feature does not improve its decision-making capacity. Irrelevant features may also cause overfitting and reduce interpretability of the concerned model. It is, therefore, important that the features are optimally selected before a model is built. In the case of online learning, new instances are periodically discovered, and the respective model is tactically retrained as required. Similarly, there are many real-life situations where hundreds of new features are discovered periodically, and the existing model needs to be retrained or tested for its performance improvement. Supervised selection of feature subset usually requires creation of multiple suboptimal models, thus incurring time-intensive computations. Unsupervised selections, although faster, largely rely on some subjective definition of feature relevance. In this paper, we introduce a score that accurately determines the importance of the features. The proposed score is appropriate for online feature selection scenarios for its low time complexity and ability to interpret performance improvement of the current model after the addition of a new feature, without invoking a retraining. Debarka Sengupta, Sanghamitra Bandyopadhyay, Debajyoti Sinha |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | FOCS: Fast Overlapped Community SearchabstractDiscovery of natural groups of similarly functioning individuals is a key task in analysis of real world networks. Also, overlap between community pairs is commonplace in large social and biological graphs, in particular. In fact, overlaps between communities are known to be denser than the non-overlapped regions of the communities. However, most of the existing algorithms that detect overlapping communities assume that the communities are denser than their surrounding regions, and falsely identify overlaps as communities. Further, many of these algorithms are computationally demanding and thus, do not scale reasonably with varying network sizes. In this article, we propose Fast Overlapped Community Search (FOCS), an algorithm that accounts for local connectedness in order to identify overlapped communities. FOCS is shown to be linear in number of edges and nodes. It additionally gains in speed via simultaneous selection of multiple near-best communities rather than merely the best, at each iteration. FOCS outperforms some popular overlapped community finding algorithms in terms of computational time while not compromising with quality. Sanghamitra Bandyopadhyay, Garisha Chowdhary, Debarka Sengupta |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Reformulated Kemeny Optimal Aggregation with Application in Consensus Ranking of microRNA TargetsabstractMicroRNAs are very recently discovered small noncoding RNAs, responsible for negative regulation of gene expression. Members of this endogenous family of small RNA molecules have been found implicated in many genetic disorders. Each microRNA targets tens to hundreds of genes. Experimental validation of target genes is a time- and cost-intensive procedure. Therefore, prediction of microRNA targets is a very important problem in computational biology. Though, dozens of target prediction algorithms have been reported in the past decade, they disagree significantly in terms of target gene ranking (based on predicted scores). Rank aggregation is often used to combine multiple target orderings suggested by different algorithms. This technique has been used in diverse fields including social choice theory, meta search in web, and most recently, in bioinformatics. Kemeny optimal aggregation (KOA) is considered the more profound objective for rank aggregation. The consensus ordering obtained through Kemeny optimal aggregation incurs minimum pairwise disagreement with the input orderings. Because of its computational intractability, heuristics are often formulated to obtain a near optimal consensus ranking. Unlike its real time use in meta search, there are a number of scenarios in bioinformatics (e.g., combining microRNA target rankings, combining disease-related gene rankings obtained from microarray experiments) where evolutionary approaches can be afforded with the ambition of better optimization. We conjecture that an ideal consensus ordering should have its total disagreement shared, as equally as possible, with the input orderings. This is also important to refrain the evolutionary processes from getting stuck to local extremes. In the current work, we reformulate Kemeny optimal aggregation while introducing a trade-off between the total pairwise disagreement and its distribution. A simulated annealing-based implementation of the proposed objective has been found effective in context of microRNA target ranking. Supplementary data and source code link are available at: >http://www.isical.ac.in/bioinfo_miu/ieee_tcbb_kemeny.rar. Debarka Sengupta, Aroonalok Pyne, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2012 | Score Based Aggregation of microRNA Target Orderings
Debarka Sengupta, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
ISBRA | 1 |
| 2012 | Weighted Markov Chain Based Aggregation of Biomolecule OrderingsabstractThe scope and effectiveness of Rank Aggregation (RA) have already been established in contemporary bioinformatics research. Rank aggregation helps in meta-analysis of putative results collected from different analytic or experimental sources. For example, we often receive considerably differing ranked lists of genes or microRNAs from various target prediction algorithms or microarray studies. Sometimes combining them all, in some sense, yields more effective ordering of the set of objects. Also, assigning a certain level of confidence to each source of ranking is a natural demand of aggregation. Assignment of weights to the sources of orderings can be performed by experts. Several rank aggregation approaches like those based on Markov Chains (MCs), evolutionary algorithms, etc., exist in the literature. Markov chains, in general, are faster than the evolutionary approaches. Unlike the evolutionary computing approaches Markov chains have not been used for weighted aggregation scenarios. This is because of the absence of a formal framework of Weighted Markov Chain (WMC). In this paper, we propose the use of a modified version of MC4 (one of the Markov chains proposed by Dwork et al., 2001), followed by the weighted analog of local Kemenization for performing rank aggregation, where the sources of rankings can be prioritized by an expert. Effectiveness of the weighted Markov chain approach over the very recently proposed Genetic Algorithm (GA) and Cross-Entropy Monte Carlo (MC) algorithm-based techniques, has been established for gene orderings from microarray analysis and orderings of predicted microRNA targets. Debarka Sengupta, Ujjwal Maulik, Sanghamitra Bandyopadhyay |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |