VLDB 2026 Research / reviewers in the wild / expert
Debashis Ghosh
dblp:85/5810
· DBLP profile ↗
57ranked-venue papers
7as first author
21since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Transfer Learning Based Decision Level Multimodal Framework for Continuous Sign Language RecognitionabstractMultimodal frameworks have appeared as a potential solution to achieve breakthrough results in the field of sign language and hand gesture recognition. In this paper, we propose a classifier combination based multimodal framework for continuous sign language recognition. In this work, we propose to use transfer learning to perform classification task on individual modalities. Further, we apply majority voting scheme to combine all the individual classification performances to obtain the final classification accuracy. For transfer learning, pretrained deep neural networks like GoogleNet, MobileNet-v2 and EfficientNet-b0 are used independently from one another. We applied these networks independently on individual modalities of IPN Hand dataset and combined the classification results obtained on these modalities to obtain the final classification result. IPN Hand dataset is one of the most challenging and dynamic continuous hand gesture datasets. Among the used pre-trained networks, EfficientNet-b0 appeared as the best network in terms of accuracy whereas GoogleNet took the least computational time during classification. On this dataset, our proposed approach performs exceptionally well with individual modalities. After combining the classification results of individual modalities according to our employed algorithm, it is observed that our proposed approach performs better than the earlier reported results. Our method surpasses the reported benchmark performance as well as performs superior to many state-of-the-art results. Navneet Nayan, Debashis Ghosh, Pyari Mohan Pradhan |
TENCON | 2 |
| 2025 | Multimodal Approach Based Sentence Level Sign Language SynthesisabstractIn this paper, we present a multimodal sentence level sign language synthesis system. Sentence level sign language synthesis requires synthesis of signs as well as synthesis of transition segments between the signs. In this paper, our focus is to develop. an efficient transition segment between two signs. For this, we propose to use edge features and trajectory features obtained from the transition segment of the query sign sentences. Edge features are extracted using the morphological operations on the hand images, whereas trajectory features are obtained from centroid detection of consecutive hand images. From trajectory, we obtain the direction of motion and shape and co-ordinates of the trajectory. Further, with the help of interpolation techniques we synthesize the hand gestures between the two signs. The correctness of the synthesis is analyzed and modified with the help of edge features and trajectory features. Our proposed approach is tested on some Indian Sign Language phrases and sentences. The proposed method is evaluated based on the mean opinion scores obtained from several users. Evaluation was based on three criteria, namely clarity in understanding, smoothness of the generated videos and similarity to the original sign videos. We obtained decent mean scores and encouraging feedbacks from the users. Navneet Nayan, Debashis Ghosh, Pyari Mohan Pradhan |
TENCON | 2 |
| 2025 | ConvMixFormer- A Resource-Efficient Convolution Mixer for Transformer-Based Dynamic Hand Gesture RecognitionabstractTransformer models have demonstrated remarkable success in many domains such as natural language processing (NLP) and computer vision. With the growing interest in transformer-based architectures, they are now utilized for gesture recognition. So, we also explore and devise a novel ConvMixFormer architecture for dynamic hand gestures. The transformers use quadratic scaling of the attention features with the sequential data, due to which these models are computationally complex and heavy. We have considered this drawback of the transformer and designed a resource-efficient model that replaces the self-attention in the transformer with the simple convolutional layer-based token mixer. The computational cost and the parameters used for the convolution-based mixer are comparatively less than the quadratic self-attention. Convolution-mixer helps the model capture the local spatial features that self-attention struggles to capture due to their sequential processing nature. Further, an efficient gate mechanism is employed instead of a conventional feed-forward network in the transformer to help the model control the flow of features within different stages of the proposed model. This design uses fewer learnable parameters which is nearly half the vanilla transformer that helps in fast and efficient training. The proposed method is evaluated on NVidia Dynamic Hand Gesture and Briareo datasets and our model has achieved state-of-the-art results on single and multimodal inputs. We have also shown the parameter efficiency of the proposed ConvMixFormer model compared to other methods. The source code is available at https://github.com/mallikagarg/ConvMixFormer. Mallika, Debashis Ghosh, Pyari Mohan Pradhan |
WACV | 2 |
| 2025 | cytoKernel: robust kernel embeddings for assessing differential expression of single-cell dataabstractMOTIVATION: High-throughput sequencing of single-cell data can be used to rigorously evaluate cell specification and enable intricate variations between groups or conditions to be identified. Many popular existing methods for differential expression target differences in aggregate measurement (mean, median, sum) and limit their approaches to detect only global differential changes. RESULTS: We present a robust method for differential expression of single-cell data using a kernel-based score test, cytoKernel. CytoKernel is specifically designed to assess the differential expression of single-cell RNA sequencing and high-dimensional flow or mass cytometry data using the full probability distribution pattern. cytoKernel is based on kernel embeddings which employs the probability distributions of the single-cell data, by calculating the pairwise divergence/distance between distributions of subjects. It can detect both patterns involving changes in the aggregate, as well as more elusive variations that are often overlooked due to the multimodal characteristics of single-cell data. We performed extensive benchmarks across both simulated and real data sets from mass cytometry data and single-cell RNA sequencing. The cytoKernel procedure effectively controls the false discovery rate and shows favorable performance compared to existing methods. The method is able to identify more differential patterns than existing approaches. We apply cytoKernel to assess gene expression and protein marker expression differences from cell subpopulations in various publicly available single-cell RNAseq and mass cytometry datasets. AVAILABILITY AND IMPLEMENTATION: The methods described in this paper are implemented in the open-source R package cytoKernel, which is freely available from Bioconductor at http://bioconductor.org/packages/cytoKernel. Tusharkanti Ghosh, Ryan M. Baxter, Souvik Seal, Victor G. Lui, Pratyaydipta Rudra, Thao Vu, Elena W. Y. Hsieh, Debashis Ghosh |
Bioinform. | 8 |
| 2025 | Dynamic group decision-making for enterprise resource planning selection using two-tuples Pythagorean fuzzy MOORA approach
Biplab Sinha Mahapatra, Debashis Ghosh, Dragan Pamucar, G. S. Mahapatra 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Impact of Rician Phase Shifts on Multi-Pair Two-Way Full-Duplex Massive MIMO RelayingabstractMost of the prior research in multipair two-way full-duplex (FD) massive multiple input multiple output (mMIMO) relaying systems considers a Rician channel with a deterministic line of sight component assuming fixed locations of the transceivers. We consider random phase shifts due to phase noise and user mobility in the LoS component to capture a more realistic scenario. We estimate the phase aware and phase unaware Rician channel via the minimum mean square error and linear minimum mean square error-based estimators, respectively, and use a maximal ratio combiner/maximal ratio transmitter to derive the closed-form spectral efficiency (SE) for an arbitrary number of relay antennas. The numerical outcomes validate the correctness of the derived closed-form SE expression when compared with its ergodic equivalent. We numerically investigate the impact of phase shift and loop interference on the SE. Malay Chakraborty, Ekant Sharma, Debashis Ghosh |
WCNC | 4 |
| 2024 | A multi-modal framework for continuous and isolated hand gesture recognition utilizing movement epenthesis detection
Navneet Nayan, Debashis Ghosh, Pyari Mohan Pradhan |
Mach. Vis. Appl. | 2 |
| 2024 | Kernel machine tests of association using extrinsic and intrinsic cluster evaluation metricsabstractModeling the network topology of the human brain within the mesoscale has become an increasing focus within the neuroscientific community due to its variation across diverse cognitive processes, in the presence of neuropsychiatric disease or injury, and over the lifespan. Much research has been done on the creation of algorithms to detect these mesoscopic structures, called communities or modules, but less has been done to conduct inference on these structures. The literature on analysis of these community detection algorithms has focused on comparing them within the same subject. These approaches, however, either do not accomodate a more general association between community structure and an outcome or cannot accommodate additional covariates that may confound the association of interest. We propose a semiparametric kernel machine regression model for either a continuous or binary outcome, where covariate effects are modeled parametrically and brain connectivity measures are measured nonparametrically. By incorporating notions of similarity between network community structures into a kernel distance function, the high-dimensional feature space of brain networks, defined on input pairs, can be generalized to non-linear spaces, allowing for a wider class of distance-based algorithms. We evaluate our proposed methodology on both simulated and real datasets. Alexandria M. Jensen, Peter Dewitt, Brianne M. Bettcher, Julia Wrobel, Katerina J. Kechris, Debashis Ghosh |
PLoS Comput. Biol. | 6 |
| 2023 | Learning with Synthesized Data for Generalizable Lesion Detection in Real PET Images
Bennett B. Chin, Michael Silosky, Daniel Litwiller, Debashis Ghosh, Fuyong Xing |
MICCAI (5) | 5 |
| 2023 | TreeKernel: interpretable kernel machine tests for interactions between -omics and clinical predictors with applications to metabolomics and COPD phenotypesabstractBACKGROUND: In this paper, we are interested in interactions between a high-dimensional -omics dataset and clinical covariates. The goal is to evaluate the relationship between a phenotype of interest and a high-dimensional omics pathway, where the effect of the omics data depends on subjects' clinical covariates (age, sex, smoking status, etc.). For instance, metabolic pathways can vary greatly between sexes which may also change the relationship between certain metabolic pathways and a clinical phenotype of interest. We propose partitioning the clinical covariate space and performing a kernel association test within those partitions. To illustrate this idea, we focus on hierarchical partitions of the clinical covariate space and kernel tests on metabolic pathways. RESULTS: We see that our proposed method outperforms competing methods in most simulation scenarios. It can identify different relationships among clinical groups with higher power in most scenarios while maintaining a proper Type I error rate. The simulation studies also show a robustness to the grouping structure within the clinical space. We also apply the method to the COPDGene study and find several clinically meaningful interactions between metabolic pathways, the clinical space, and lung function. CONCLUSION: TreeKernel provides a simple and interpretable process for testing for relationships between high-dimensional omics data and clinical outcomes in the presence of interactions within clinical cohorts. The method is broadly applicable to many studies. Charlie M. Carpenter, Lucas A. Gillenwater, Russell Bowler, Katerina J. Kechris, Debashis Ghosh |
BMC Bioinform. | 5 |
| 2023 | Learning with limited target data to detect cells in cross-modality images
Fuyong Xing, Toby C. Cornish, Debashis Ghosh |
Medical Image Anal. | 4 |
| 2023 | A platform-independent framework for phenotyping of multiplex tissue imaging dataabstractMultiplex imaging is a powerful tool to analyze the structural and functional states of cells in their morphological and pathological contexts. However, hypothesis testing with multiplex imaging data is a challenging task due to the extent and complexity of the information obtained. Various computational pipelines have been developed and validated to extract knowledge from specific imaging platforms. A common problem with customized pipelines is their reduced applicability across different imaging platforms: Every multiplex imaging technique exhibits platform-specific characteristics in terms of signal-to-noise ratio and acquisition artifacts that need to be accounted for to yield reliable and reproducible results. We propose a pixel classifier-based image preprocessing step that aims to minimize platform-dependency for all multiplex image analysis pipelines. Signal detection and noise reduction as well as artifact removal can be posed as a pixel classification problem in which all pixels in multiplex images can be assigned to two general classes of either I) signal of interest or II) artifacts and noise. The resulting feature representation maps contain pixel-scale representations of the input data, but exhibit significantly increased signal-to-noise ratios with normalized pixel values as output data. We demonstrate the validity of our proposed image preprocessing approach by comparing the results of two well-accepted and widely-used image analysis pipelines. Mansooreh Ahmadian, Christian Rickert, Angela Minic, Julia Wrobel, Benjamin G. Bitler, Fuyong Xing, Michael Angelo, Elena W. Y. Hsieh, Debashis Ghosh, Kimberly R. Jordan |
PLoS Comput. Biol. | 9 |
| 2023 | FunSpace: A functional and spatial analytic approach to cell imaging data using entropy measuresabstractSpatial heterogeneity in the tumor microenvironment (TME) plays a critical role in gaining insights into tumor development and progression. Conventional metrics typically capture the spatial differential between TME cellular patterns by either exploring the cell distributions in a pairwise fashion or aggregating the heterogeneity across multiple cell distributions without considering the spatial contribution. As such, none of the existing approaches has fully accounted for the simultaneous heterogeneity caused by both cellular diversity and spatial configurations of multiple cell categories. In this article, we propose an approach to leverage spatial entropy measures at multiple distance ranges to account for the spatial heterogeneity across different cellular organizations. Functional principal component analysis (FPCA) is applied to estimate FPC scores which are then served as predictors in a Cox regression model to investigate the impact of spatial heterogeneity in the TME on survival outcome, potentially adjusting for other confounders. Using a non-small cell lung cancer dataset (n = 153) as a case study, we found that the spatial heterogeneity in the TME cellular composition of CD14+ cells, CD19+ B cells, CD4+ and CD8+ T cells, and CK+ tumor cells, had a significant non-zero effect on the overall survival (p = 0.027). Furthermore, using a publicly available multiplexed ion beam imaging (MIBI) triple-negative breast cancer dataset (n = 33), our proposed method identified a significant impact of cellular interactions between tumor and immune cells on the overall survival (p = 0.046). In simulation studies under different spatial configurations, the proposed method demonstrated a high predictive power by accounting for both clinical effect and the impact of spatial heterogeneity. Thao Vu, Souvik Seal, Tusharkanti Ghosh, Mansooreh Ahmadian, Julia Wrobel, Debashis Ghosh |
PLoS Comput. Biol. | 6 |
| 2023 | Multiscaled Multi-Head Attention-Based Video Transformer Network for Hand Gesture RecognitionabstractDynamic gesture recognition is one of the challenging research areas due to variations in pose, size, and shape of the signer's hand. In this letter, Multiscaled Multi-Head Attention Video Transformer Network (MsMHA-VTN) for dynamic hand gesture recognition is proposed. A pyramidal hierarchy of multiscale features is extracted using the transformer multiscaled head attention model. The proposed model employs different attention dimensions for each head of the transformer which enables it to provide attention at the multiscale level. Further, in addition to single modality, recognition performance using multiple modalities is examined. Extensive experiments demonstrate the superior performance of the proposed MsMHA-VTN with an overall accuracy of 88.22% and 99.10% on NVGesture and Briareo datasets, respectively. Mallika, Debashis Ghosh, Pyari Mohan Pradhan |
IEEE Signal Process. Lett. | 2 |
| 2022 | An Unsupervised Learning Approach to Handle Movement Epenthesis in Continuous Sign Language RecognitionabstractIn this paper, the problem of movement epenthesis in continuous sign language sentences is considered. Movement epenthesis caused due to unwanted but unavoidable hand movement in between two sign gestures in continuous signing has emerged as one of the most challenging problems in automatic sign language recognition. To handle this problem, a novel method based on unsupervised learning approach has been proposed in this paper to separate out video frames corresponding to meaningful sign gestures from meaningless movement epenthesis segments in a continuous signing gesture video clip. Our proposed method is based on K-means clustering of the norm values of the absolute difference between current frames and the reference frame to detect the movement epenthesis frame and then classify the frames of the video as movement epenthesis frames and sign frames. Exhaustive experimentation on the publicly available standard ChaLearn LAP ConGD gesture video dataset was carried out to test our algorithm. Experimental results demonstrate that the proposed method is good enough to detect and separate out the movement epenthesis frames in sentence level sign language videos. Our proposed approach performs the movement epenthesis detection task accurately in 91% of the videos of the ChaLearn LAP ConGD dataset. Navneet Nayan, Debashis Ghosh, Pyari Mohan Pradhan |
ICARCV | 2 |
| 2022 | MIAMI: mutual information-based analysis of multiplex imaging dataabstractMOTIVATION: Studying the interaction or co-expression of the proteins or markers in the tumor microenvironment of cancer subjects can be crucial in the assessment of risks, such as death or recurrence. In the conventional approach, the cells need to be declared positive or negative for a marker based on its intensity. For multiple markers, manual thresholds are required for all the markers, which can become cumbersome. The performance of the subsequent analysis relies heavily on this step and thus suffers from subjectivity and lacks robustness. RESULTS: We present a new method where different marker intensities are viewed as dependent random variables, and the mutual information (MI) between them is considered to be a metric of co-expression. Estimation of the joint density, as required in the traditional form of MI, becomes increasingly challenging as the number of markers increases. We consider an alternative formulation of MI which is conceptually similar but has an efficient estimation technique for which we develop a new generalization. With the proposed method, we analyzed a lung cancer dataset finding the co-expression of the markers, HLA-DR and CK to be associated with survival. We also analyzed a triple negative breast cancer dataset finding the co-expression of the immuno-regulatory proteins, PD1, PD-L1, Lag3 and IDO, to be associated with disease recurrence. We demonstrated the robustness of our method through different simulation studies. AVAILABILITY AND IMPLEMENTATION: The associated R package can be found here, https://github.com/sealx017/MIAMI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Souvik Seal, Debashis Ghosh |
Bioinform. | 2 |
| 2022 | Mechanism-aware imputation: a two-step approach in handling missing values in metabolomicsabstractWhen analyzing large datasets from high-throughput technologies, researchers often encounter missing quantitative measurements, which are particularly frequent in metabolomics datasets. Metabolomics, the comprehensive profiling of metabolite abundances, are typically measured using mass spectrometry technologies that often introduce missingness via multiple mechanisms: (1) the metabolite signal may be smaller than the instrument limit of detection; (2) the conditions under which the data are collected and processed may lead to missing values; (3) missing values can be introduced randomly. Missingness resulting from mechanism (1) would be classified as Missing Not At Random (MNAR), that from mechanism (2) would be Missing At Random (MAR), and that from mechanism (3) would be classified as Missing Completely At Random (MCAR). Two common approaches for handling missing data are the following: (1) omit missing data from the analysis; (2) impute the missing values. Both approaches may introduce bias and reduce statistical power in downstream analyses such as testing metabolite associations with clinical variables. Further, standard imputation methods in metabolomics often ignore the mechanisms causing missingness and inaccurately estimate missing values within a data set. We propose a mechanism-aware imputation algorithm that leverages a two-step approach in imputing missing values. First, we use a random forest classifier to classify the missing mechanism for each missing value in the data set. Second, we impute each missing value using imputation algorithms that are specific to the predicted missingness mechanism (i.e., MAR/MCAR or MNAR). Using complete data, we conducted simulations, where we imposed different missingness patterns within the data and tested the performance of combinations of imputation algorithms. Our proposed algorithm provided imputations closer to the original data than those using only one imputation algorithm for all the missing values. Consequently, our two-step approach was able to reduce bias for improved downstream analyses. Jonathan P. Dekermanjian, Elin Shaddox, Debmalya Nandy, Debashis Ghosh, Katerina J. Kechris |
BMC Bioinform. | 4 |
| 2022 | SPF: A spatial and functional data analytic approach to cell imaging dataabstractThe tumor microenvironment (TME), which characterizes the tumor and its surroundings, plays a critical role in understanding cancer development and progression. Recent advances in imaging techniques enable researchers to study spatial structure of the TME at a single-cell level. Investigating spatial patterns and interactions of cell subtypes within the TME provides useful insights into how cells with different biological purposes behave, which may consequentially impact a subject's clinical outcomes. We utilize a class of well-known spatial summary statistics, the K-function and its variants, to explore inter-cell dependence as a function of distances between cells. Using techniques from functional data analysis, we introduce an approach to model the association between these summary spatial functions and subject-level outcomes, while controlling for other clinical scalar predictors such as age and disease stage. In particular, we leverage the additive functional Cox regression model (AFCM) to study the nonlinear impact of spatial interaction between tumor and stromal cells on overall survival in patients with non-small cell lung cancer, using multiplex immunohistochemistry (mIHC) data. The applicability of our approach is further validated using a publicly available multiplexed ion beam imaging (MIBI) triple-negative breast cancer dataset. Thao Vu, Julia Wrobel, Benjamin G. Bitler, Erin L. Schenk, Kimberly R. Jordan, Debashis Ghosh |
PLoS Comput. Biol. | 6 |
| 2021 | Reproducibility of mass spectrometry based metabolomics dataabstractBACKGROUND: Assessing the reproducibility of measurements is an important first step for improving the reliability of downstream analyses of high-throughput metabolomics experiments. We define a metabolite to be reproducible when it demonstrates consistency across replicate experiments. Similarly, metabolites which are not consistent across replicates can be labeled as irreproducible. In this work, we introduce and evaluate the use (Ma)ximum (R)ank (R)eproducibility (MaRR) to examine reproducibility in mass spectrometry-based metabolomics experiments. We examine reproducibility across technical or biological samples in three different mass spectrometry metabolomics (MS-Metabolomics) data sets. RESULTS: We apply MaRR, a nonparametric approach that detects the change from reproducible to irreproducible signals using a maximal rank statistic. The advantage of using MaRR over model-based methods that it does not make parametric assumptions on the underlying distributions or dependence structures of reproducible metabolites. Using three MS Metabolomics data sets generated in the multi-center Genetic Epidemiology of Chronic Obstructive Pulmonary Disease (COPD) study, we applied the MaRR procedure after data processing to explore reproducibility across technical or biological samples. Under realistic settings of MS-Metabolomics data, the MaRR procedure effectively controls the False Discovery Rate (FDR) when there was a gradual reduction in correlation between replicate pairs for less highly ranked signals. Simulation studies also show that the MaRR procedure tends to have high power for detecting reproducible metabolites in most situations except for smaller values of proportion of reproducible metabolites. Bias (i.e., the difference between the estimated and the true value of reproducible signal proportions) values for simulations are also close to zero. The results reported from the real data show a higher level of reproducibility for technical replicates compared to biological replicates across all the three different datasets. In summary, we demonstrate that the MaRR procedure application can be adapted to various experimental designs, and that the nonparametric approach performs consistently well. CONCLUSIONS: This research was motivated by reproducibility, which has proven to be a major obstacle in the use of genomic findings to advance clinical practice. In this paper, we developed a data-driven approach to assess the reproducibility of MS-Metabolomics data sets. The methods described in this paper are implemented in the open-source R package marr, which is freely available from Bioconductor at http://bioconductor.org/packages/marr . Tusharkanti Ghosh, Daisy Philtron, Katerina J. Kechris, Debashis Ghosh |
BMC Bioinform. | 5 |
| 2021 | PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypesabstractHigh-throughput data such as metabolomics, genomics, transcriptomics, and proteomics have become familiar data types within the "-omics" family. For this work, we focus on subsets that interact with one another and represent these "pathways" as graphs. Observed pathways often have disjoint components, i.e., nodes or sets of nodes (metabolites, etc.) not connected to any other within the pathway, which notably lessens testing power. In this paper we propose the Pathway Integrated Regression-based Kernel Association Test (PaIRKAT), a new kernel machine regression method for incorporating known pathway information into the semi-parametric kernel regression framework. This work extends previous kernel machine approaches. This paper also contributes an application of a graph kernel regularization method for overcoming disconnected pathways. By incorporating a regularized or "smoothed" graph into a score test, PaIRKAT can provide more powerful tests for associations between biological pathways and phenotypes of interest and will be helpful in identifying novel pathways for targeted clinical research. We evaluate this method through several simulation studies and an application to real metabolomics data from the COPDGene study. Our simulation studies illustrate the robustness of this method to incorrect and incomplete pathway knowledge, and the real data analysis shows meaningful improvements of testing power in pathways. PaIRKAT was developed for application to metabolomic pathway data, but the techniques are easily generalizable to other data sources with a graph-like structure. Charlie M. Carpenter, Lucas A. Gillenwater, Cameron Severn, Tusharkanti Ghosh, Russell Bowler, Katerina J. Kechris, Debashis Ghosh |
PLoS Comput. Biol. | 8 |
| 2021 | Bidirectional Mapping-Based Domain Adaptation for Nucleus Detection in Cross-Modality Microscopy ImagesabstractCell or nucleus detection is a fundamental task in microscopy image analysis and has recently achieved state-of-the-art performance by using deep neural networks. However, training supervised deep models such as convolutional neural networks (CNNs) usually requires sufficient annotated image data, which is prohibitively expensive or unavailable in some applications. Additionally, when applying a CNN to new datasets, it is common to annotate individual cells/nuclei in those target datasets for model re-learning, leading to inefficient and low-throughput image analysis. To tackle these problems, we present a bidirectional, adversarial domain adaptation method for nucleus detection on cross-modality microscopy image data. Specifically, the method learns a deep regression model for individual nucleus detection with both source-to-target and target-to-source image translation. In addition, we explicitly extend this unsupervised domain adaptation method to a semi-supervised learning situation and further boost the nucleus detection performance. We evaluate the proposed method on three cross-modality microscopy image datasets, which cover a wide variety of microscopy imaging protocols or modalities, and obtain a significant improvement in nucleus detection compared to reference baseline approaches. In addition, our semi-supervised method is very competitive with recent fully supervised learning models trained with all real target training labels. Fuyong Xing, Toby C. Cornish, Tellen D. Bennett, Debashis Ghosh |
IEEE Trans. Medical Imaging | 4 |
| 2020 | An Invitation to System-wide Algorithmic FairnessabstractWe propose a framework for analyzing and evaluating system-wide algorithmic fairness. The core idea is to use simulation techniques in order to extend the scope of current fairness assessments by incorporating context and feedback to a phenomenon of interest. By doing so, we expect to better understand the interaction among the social behavior giving rise to discrimination, automated decision making tools, and fairness-inspired statistical constraints. In particular, we invite the community to use agent based models as an explanatory tool for causal mechanisms of population level properties. We also propose embedding these into a reinforcement learning algorithm to find optimal actions for meaningful change. As an incentive for taking a system-wide approach , we show through a simple model of predictive policing and trials that if we limit our attention to one portion of the system, we may determine some blatantly unfair practices as fair, and be blind to overall unfairness. Efren Cruz Cortes, Debashis Ghosh |
AIES | 2 |
| 2019 | Adversarial Domain Adaptation and Pseudo-Labeling for Cross-Modality Microscopy Image Quantification
Fuyong Xing, Tellen D. Bennett, Debashis Ghosh |
MICCAI (1) | 3 |
| 2018 | Accuracy of the Epic Sepsis Prediction Model in a Regional Health System
Tellen D. Bennett, Seth Russell, Lisa M. Schilling, Chan Voong, Nancy Rogers, Bonnie Adrian, Nicholas Bruce, Debashis Ghosh |
AMIA | 9 |
| 2017 | Examining the role of unmeasured confounding in mediation analysis with genetic and genomic applicationsabstractBACKGROUND: In mediation analysis if unmeasured confounding is present, the estimates for the direct and mediated effects may be over or under estimated. Most methods for the sensitivity analysis of unmeasured confounding in mediation have focused on the mediator-outcome relationship. RESULTS: The Umediation R package enables the user to simulate unmeasured confounding of the exposure-mediator, exposure-outcome, and mediator-outcome relationships in order to see how the results of the mediation analysis would change in the presence of unmeasured confounding. We apply the Umediation package to the Genetic Epidemiology of Chronic Obstructive Pulmonary Disease (COPDGene) study to examine the role of unmeasured confounding due to population stratification on the effect of a single nucleotide polymorphism (SNP) in the CHRNA5/3/B4 locus on pulmonary function decline as mediated by cigarette smoking. CONCLUSIONS: Umediation is a flexible R package that examines the role of unmeasured confounding in mediation analysis allowing for normally distributed or Bernoulli distributed exposures, outcomes, mediators, measured confounders, and unmeasured confounders. Umediation also accommodates multiple measured confounders, multiple unmeasured confounders, and allows for a mediator-exposure interaction on the outcome. Umediation is available as an R package at https://github.com/SharonLutz/Umediation A tutorial on how to install and use the Umediation package is available in the Additional file 1. Sharon Marie Lutz, Annie Thwing, Sarah Schmiege, Miranda Kroehl, Christopher D. Baker, Anne P. Starling, John E. Hokanson, Debashis Ghosh |
BMC Bioinform. | 8 |
| 2017 | An efficient low vision plant leaf shape identification system for smart phones
Shitala Prasad, Sateesh Kumar Peddoju, Debashis Ghosh |
Multim. Tools Appl. | 3 |
| 2017 | An adaptive plant leaf mobile informatics using RSSC
Shitala Prasad, Sateesh Kumar Peddoju, Debashis Ghosh |
Multim. Tools Appl. | 3 |
| 2016 | A novel copy number variants kernel association test with application to autism spectrum disorders studiesabstractMOTIVATION: Copy number variants (CNVs) have been implicated in a variety of neurodevelopmental disorders, including autism spectrum disorders, intellectual disability and schizophrenia. Recent advances in high-throughput genomic technologies have enabled rapid discovery of many genetic variants including CNVs. As a result, there is increasing interest in studying the role of CNVs in the etiology of many complex diseases. Despite the availability of an unprecedented wealth of CNV data, methods for testing association between CNVs and disease-related traits are still under-developed due to the low prevalence and complicated multi-scale features of CNVs. RESULTS: We propose a novel CNV kernel association test (CKAT) in this paper. To address the low prevalence, CNVs are first grouped into CNV regions (CNVR). Then, taking into account the multi-scale features of CNVs, we first design a single-CNV kernel which summarizes the similarity between two CNVs, and next aggregate the single-CNV kernel to a CNVR kernel which summarizes the similarity between two CNVRs. Finally, association between CNVR and disease-related traits is assessed by comparing the kernel-based similarity with the similarity in the trait using a score test for variance components in a random effect model. We illustrate the proposed CKAT using simulations and show that CKAT is more powerful than existing methods, while always being able to control the type I error. We also apply CKAT to a real dataset examining the association between CNV and autism spectrum disorders, which demonstrates the potential usefulness of the proposed method. AVAILABILITY AND IMPLEMENTATION: A R package to implement the proposed CKAT method is available at http://works.bepress.com/debashis_ghosh/ CONTACTS: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Xiang Zhan, Santhosh Girirajan, Ni Zhao, Michael C. Wu, Debashis Ghosh |
Bioinform. | 5 |
| 2015 | Kernel approaches for differential expression analysis of mass spectrometry-based metabolomics dataabstractBACKGROUND: Data generated from metabolomics experiments are different from other types of "-omics" data. For example, a common phenomenon in mass spectrometry (MS)-based metabolomics data is that the data matrix frequently contains missing values, which complicates some quantitative analyses. One way to tackle this problem is to treat them as absent. Hence there are two types of information that are available in metabolomics data: presence/absence of a metabolite and a quantitative value of the abundance level of a metabolite if it is present. Combining these two layers of information poses challenges to the application of traditional statistical approaches in differential expression analysis. RESULTS: In this article, we propose a novel kernel-based score test for the metabolomics differential expression analysis. In order to simultaneously capture both the continuous pattern and discrete pattern in metabolomics data, two new kinds of kernels are designed. One is the distance-based kernel and the other is the stratified kernel. While we initially describe the procedures in the case of single-metabolite analysis, we extend the methods to handle metabolite sets as well. CONCLUSIONS: Evaluation based on both simulated data and real data from a liver cancer metabolomics study indicates that our kernel method has a better performance than some existing alternatives. An implementation of the proposed kernel method in the R statistical computing environment is available at http://works.bepress.com/debashis_ghosh/60/ . Xiang Zhan, Andrew D. Patterson, Debashis Ghosh |
BMC Bioinform. | 3 |
| 2014 | Energy efficient mobile vision system for plant leaf disease identificationabstractClose monitoring, proper control and management of plant diseases are essential in the efficient cultivation of crops. This paper presents a scheme that uses mobile phones for real-time on-field imaging of diseased plants followed by disease diagnosis via analysis of visual phenotypes. A threshold based offloading scheme is employed for judicious sharing of the computational load between the mobile device and a central server at the plant pathology laboratory, thereby offering a trade-off between the power consumption in the mobile device and the transmission cost. The part of the processing carried out in the mobile device includes leaf image segmentation and spotting of disease patch using improved k-means clustering. The algorithm is simple and hence suitable for Android based mobile devices. The segmented image is subsequently communicated to the central server. This ensures reduced transmission cost compared to that in transmitting full leaf image. Shitala Prasad, Sateesh Kumar Peddoju, Debashis Ghosh |
WCNC | 3 |
| 2014 | MRHMMs: Multivariate Regression Hidden Markov Models and the variantSabstractSUMMARY: Hidden Markov models (HMMs) are flexible and widely used in scientific studies. Particularly in genomics and genetics, there are multiple distinct regimes in the genome within each of which the relationships among multivariate features are distinct. Examples include differential gene regulation depending on gene functions and experimental conditions, and varying combinatorial patterns of multiple transcription factors. We developed a software package called MRHMMs (Multivariate Regression Hidden Markov Models and the variantS) that accommodates a variety of HMMs that can be flexibly applied to many biological studies and beyond. MRHMMs supplements existing HMM software packages in two aspects. First, MRHMMs provides a diverse set of emission probability structures, including mixture of multivariate normal distributions and (logistic) regression models. Second, MRHMMs is computationally efficient for analyzing large data-sets generated in current genome-wide studies. Especially, the software is written in C for the speed advantage and further amenable to implement alternative models to meet users' own purposes. AVAILABILITY AND IMPLEMENTATION: http://sourceforge.net/projects/mrhmms/ Yeonok Lee, Debashis Ghosh, Ross C. Hardison, Yu Zhang 0002 |
Bioinform. | 2 |
| 2014 | A Two-Step Hierarchical Hypothesis Set Testing Framework, with Applications to Gene Expression Data on Ordered CategoriesabstractBACKGROUND: In complex large-scale experiments, in addition to simultaneously considering a large number of features, multiple hypotheses are often being tested for each feature. This leads to a problem of multi-dimensional multiple testing. For example, in gene expression studies over ordered categories (such as time-course or dose-response experiments), interest is often in testing differential expression across several categories for each gene. In this paper, we consider a framework for testing multiple sets of hypothesis, which can be applied to a wide range of problems. RESULTS: We adopt the concept of the overall false discovery rate (OFDR) for controlling false discoveries on the hypothesis set level. Based on an existing procedure for identifying differentially expressed gene sets, we discuss a general two-step hierarchical hypothesis set testing procedure, which controls the overall false discovery rate under independence across hypothesis sets. In addition, we discuss the concept of the mixed-directional false discovery rate (mdFDR), and extend the general procedure to enable directional decisions for two-sided alternatives. We applied the framework to the case of microarray time-course/dose-response experiments, and proposed three procedures for testing differential expression and making multiple directional decisions for each gene. Simulation studies confirm the control of the OFDR and mdFDR by the proposed procedures under independence and positive correlations across genes. Simulation results also show that two of our new procedures achieve higher power than previous methods. Finally, the proposed methodology is applied to a microarray dose-response study, to identify 17 β-estradiol sensitive genes in breast cancer cells that are induced at low concentrations. CONCLUSIONS: The framework we discuss provides a platform for multiple testing procedures covering situations involving two (or potentially more) sources of multiplicity. The framework is easy to use and adaptable to various practical settings that frequently occur in large-scale experiments. Procedures generated from the framework are shown to maintain control of the OFDR and mdFDR, quantities that are especially relevant in the case of multiple hypothesis set testing. The procedures work well in both simulations and real datasets, and are shown to have better power than existing methods. Debashis Ghosh |
BMC Bioinform. | 2 |
| 2014 | Meta-Analysis Based on Weighted Ordered P-values for Genomic Data with HeterogeneityabstractBACKGROUND: Meta-analysis has become increasingly popular in recent years, especially in genomic data analysis, due to the fast growth of available data and studies that target the same questions. Many methods have been developed, including classical ones such as Fisher's combined probability test and Stouffer's Z-test. However, not all meta-analyses have the same goal in mind. Some aim at combining information to find signals in at least one of the studies, while others hope to find more consistent signals across the studies. While many classical meta-analysis methods are developed with the former goal in mind, the latter goal has much more practicality for genomic data analysis. RESULTS: In this paper, we propose a class of meta-analysis methods based on summaries of weighted ordered p-values (WOP) that aim at detecting significance in a majority of studies. We consider weighted versions of classical procedures such as Fisher's method and Stouffer's method where the weight for each p-value is based on its order among the studies. In particular, we consider weights based on the binomial distribution, where the median of the p-values are weighted highest and the outlying p-values are down-weighted. We investigate the properties of our methods and demonstrate their strengths through simulations studies, comparing to existing procedures. In addition, we illustrate application of the proposed methodology by several meta-analysis of gene expression data. CONCLUSIONS: Our proposed weighted ordered p-value (WOP) methods displayed better performance compared to existing methods for testing the hypothesis that there is signal in the majority of studies. They also appeared to be much more robust in applications compared to the rth ordered p-value (rOP) method (Song and Tseng, Ann. Appl. Stat. 2014, 8(2):777-800). With the flexibility of incorporating different p-value combination methods and different weighting schemes, the weighted ordered p-values (WOP) methods have great potential in detecting consistent signal in meta-analysis with heterogeneity. Debashis Ghosh |
BMC Bioinform. | 2 |
| 2014 | Adaptive binarization of severely degraded and non-uniformly illuminated documents
Brij Mohan Singh, Rahul Sharma 0002, Debashis Ghosh, Ankush Mittal |
Int. J. Document Anal. Recognit. | 3 |
| 2014 | Spectrum Sensing for Cognitive Radios Based on Space-Time FRESH FilteringabstractIn this paper, we consider the problem of spectrum sensing of cyclostationary signals for cognitive radios. It is shown that the detection performance in this case may be improved by enhancing the cyclostationary features of the signal of interest. The optimal filter for cyclostationary features is revisited, following which an adaptive space-time structure exploiting the spatial, temporal and spectral coherence of a cyclostationary signal incident upon an antenna array is proposed for enhancing the signal. A low complexity adaptation algorithm for this structure is also proposed. The performance of this detector is evaluated and compared with standard spectrum sensing methods. Simulation results show that a suitable space-time FRESH filtering configuration may be used to yield gains of up to 10 dB over both the standard energy detector and the traditional cyclostationary detector. Ribhu Chopra, Debashis Ghosh, D. K. Mehra |
IEEE Trans. Wirel. Commun. | 2 |
| 2013 | Mobile augmented reality based interactive teaching & learning system with low computation approachabstractThis paper presents a fast and efficient hand gesture based mobile augmented reality (MAR) system for interactive classroom. It provides a complex visual augmented layer over static slides to understand the concepts more clearly without touching the computer devices or using whiteboard. Simple hand gestures are used to interact with slides while presenting in the classroom or in any conference room with high accuracy and efficiency without any expensive hardware. The gesture path is tracked continuously using a color tracking algorithm proposed. A decision tree is used to make the decisions based on the gestures. The preliminary result indicates that the gesture recognition rate is near about approximately 94% and it is mostly acceptable. This enhances the user's interaction level with immersive feeling in immersive environment. Shitala Prasad, Sateesh Kumar Peddoju, Debashis Ghosh |
CICA | 3 |
| 2013 | Parameter tuning for multi-prototype possibilistic classifier with reject optionsabstractFuzzy classifiers are suited for pattern classification when there exist a large amount of imprecision, uncertainty and ambiguity in the patterns. One such fuzzy classifier is based on the possibilistic fuzzy membership function used for measuring the degree of class belongingness. However, the performance of possibilistic classifier depends heavily on the cluster parameters such as the 3-dB point and the parameter that controls the degree of fuzziness in the cluster. In this paper, we develop an iterative method for tuning these parameters so that the performance of the classifier is improved. The classifier considered in our work is a multi-prototype classifier and includes options for rejecting patterns that are ambiguous and/or do not belong to any class. In our proposed scheme, the slopes of the membership function are suitably varied via parameter tuning so that the membership of a pattern to a cluster in which it actually belongs is maximized while that to other classes are forced to be as small as possible. We evaluate our method using the Wisconsin Breast Cancer Dataset (WBCD). The results show that the recognition rate is improved by as much as 8% when the cluster parameters are tuned. Debashis Ghosh, Ribhu Chopra, A. P. Shivaprasad |
FUZZ-IEEE | 1 |
| 2013 | Sparsely correlated hidden Markov models with application to genome-wide location studiesabstractMOTIVATION: Multiply correlated datasets have become increasingly common in genome-wide location analysis of regulatory proteins and epigenetic modifications. Their correlation can be directly incorporated into a statistical model to capture underlying biological interactions, but such modeling quickly becomes computationally intractable. RESULTS: We present sparsely correlated hidden Markov models (scHMM), a novel method for performing simultaneous hidden Markov model (HMM) inference for multiple genomic datasets. In scHMM, a single HMM is assumed for each series, but the transition probability in each series depends on not only its own hidden states but also the hidden states of other related series. For each series, scHMM uses penalized regression to select a subset of the other data series and estimate their effects on the odds of each transition in the given series. Following this, hidden states are inferred using a standard forward-backward algorithm, with the transition probabilities adjusted by the model at each position, which helps retain the order of computation close to fitting independent HMMs (iHMM). Hence, scHMM is a collection of inter-dependent non-homogeneous HMMs, capable of giving a close approximation to a fully multivariate HMM fit. A simulation study shows that scHMM achieves comparable sensitivity to the multivariate HMM fit at a much lower computational cost. The method was demonstrated in the joint analysis of 39 histone modifications, CTCF and RNA polymerase II in human CD4+ T cells. scHMM reported fewer high-confidence regions than iHMM in this dataset, but scHMM could recover previously characterized histone modifications in relevant genomic regions better than iHMM. In addition, the resulting combinatorial patterns from scHMM could be better mapped to the 51 states reported by the multivariate HMM method of Ernst and Kellis. AVAILABILITY: The scHMM package can be freely downloaded from http://sourceforge.net/p/schmm/ and is recommended for use in a linux environment. Hyungwon Choi, Damian Fermin, Alexey I. Nesvizhskii, Debashis Ghosh, Zhaohui S. Qin |
Bioinform. | 4 |
| 2013 | A new upper bound on the parameters of quasi-symmetric designs
Debashis Ghosh, Lakshmi Kanta Dey |
Inf. Process. Lett. | 1 |
| 2012 | Assumption weighting for incorporating heterogeneity into meta-analysis of genomic dataabstractMOTIVATION: There is now a large literature on statistical methods for the meta-analysis of genomic data from multiple studies. However, a crucial assumption for performing many of these analyses is that the data exhibit small between-study variation or that this heterogeneity can be sufficiently modelled probabilistically. RESULTS: In this article, we propose 'assumption weighting', which exploits a weighted hypothesis testing framework proposed by Genovese et al. to incorporate tests of between-study variation into the meta-analysis context. This methodology is fast and computationally simple to implement. Several weighting schemes are considered and compared using simulation studies. In addition, we illustrate application of the proposed methodology using data from several high-profile stem cell gene expression datasets. Debashis Ghosh |
Bioinform. | 2 |
| 2011 | Integrative set enrichment testing for multiple omics platformsabstractBACKGROUND: Enrichment testing assesses the overall evidence of differential expression behavior of the elements within a defined set. When we have measured many molecular aspects, e.g. gene expression, metabolites, proteins, it is desirable to assess their differential tendencies jointly across platforms using an integrated set enrichment test. In this work we explore the properties of several methods for performing a combined enrichment test using gene expression and metabolomics as the motivating platforms. RESULTS: Using two simulation models we explored the properties of several enrichment methods including two novel methods: the logistic regression 2-degree of freedom Wald test and the 2-dimensional permutation p-value for the sum-of-squared statistics test. In relation to their univariate counterparts we find that the joint tests can improve our ability to detect results that are marginal univariately. We also find that joint tests improve the ranking of associated pathways compared to their univariate counterparts. However, there is a risk of Type I error inflation with some methods and self-contained methods lose specificity when the sets are not representative of underlying association. CONCLUSIONS: In this work we show that consideration of data from multiple platforms, in conjunction with summarization via a priori pathway information, leads to increased power in detection of genomic associations with phenotypes. Laila M. Poisson, Jeremy M. G. Taylor, Debashis Ghosh |
BMC Bioinform. | 3 |
| 2010 | Script Recognition - A ReviewabstractA variety of different scripts are used in writing languages throughout the world. In a multiscript, multilingual environment, it is essential to know the script used in writing a document before an appropriate character recognition and document analysis algorithm can be chosen. In view of this, several methods for automatic script identification have been developed so far. They mainly belong to two broad categories-structure-based and visual-appearance-based techniques. This survey report gives an overview of the different script identification methodologies under each of these categories. Methods for script identification in online data and video-texts are also presented. It is noted that the research in this field is relatively thin and still more research is to be done, particularly in the case of handwritten documents. Debashis Ghosh, Tulika Dube, Adamane P. Shivaprasad |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Hierarchical hidden Markov model with application to joint analysis of ChIP-chip and ChIP-seq dataabstractMOTIVATION: Chromatin immunoprecipitation (ChIP) experiments followed by array hybridization, or ChIP-chip, is a powerful approach for identifying transcription factor binding sites (TFBS) and has been widely used. Recently, massively parallel sequencing coupled with ChIP experiments (ChIP-seq) has been increasingly used as an alternative to ChIP-chip, offering cost-effective genome-wide coverage and resolution up to a single base pair. For many well-studied TFs, both ChIP-seq and ChIP-chip experiments have been applied and their data are publicly available. Previous analyses have revealed substantial technology-specific binding signals despite strong correlation between the two sets of results. Therefore, it is of interest to see whether the two data sources can be combined to enhance the detection of TFBS. RESULTS: In this work, hierarchical hidden Markov model (HHMM) is proposed for combining data from ChIP-seq and ChIP-chip. In HHMM, inference results from individual HMMs in ChIP-seq and ChIP-chip experiments are summarized by a higher level HMM. Simulation studies show the advantage of HHMM when data from both technologies co-exist. Analysis of two well-studied TFs, NRSF and CCCTC-binding factor (CTCF), also suggests that HHMM yields improved TFBS identification in comparison to analyses using individual data sources or a simple merger of the two. AVAILABILITY: Source code for the software ChIPmeta is freely available for download at http://www.umich.edu/~hwchoi/HHMMsoftware.zip, implemented in C and supported on linux. Hyungwon Choi, Alexey I. Nesvizhskii, Debashis Ghosh, Zhaohui S. Qin |
Bioinform. | 3 |
| 2008 | Reconstructing tumor-wise protein expression in tissue microarray studies using a Bayesian cell mixture modelabstractMOTIVATION: Tissue microarrays (TMAs) quantify tissue-specific protein expression of cancer biomarkers via high-density immuno-histochemical staining assays. Standard analysis approach estimates a sample mean expression in the tumor, ignoring the complex tissue-specific staining patterns observed on tissue arrays. METHODS: In this article, a cell mixture model (CMM) is proposed to reconstruct tumor expression patterns in TMA experiments. The concept is to assemble the whole-tumor expression pattern by aggregating over the subpopulation of tissue specimens sampled by needle biopsies. The expression pattern in each individual tissue element is assumed to be a zero-augmented Gamma distribution to assimilate the non-staining areas and the staining areas. A hierarchical Bayes model is imposed to borrow strength across tissue specimens and across tumors. A joint model is presented to link the CMM expression model with a survival model for censored failure time observations. The implementation involves imputation steps within each Markov chain Monte Carlo iteration and Monte Carlo integration technique. RESULTS: The model-based approach provides estimates for various tumor expression characteristics including the percentage of staining, mean intensity of staining and a composite meanstaining to associate with patient survival outcome. AVAILABILITY: R package to fit CMM model is available at http://www.mskcc.org/mskcc/html/85130.cfm Ronglai Shen, Jeremy M. G. Taylor, Debashis Ghosh |
Bioinform. | 3 |
| 2008 | Estimation and testing for the effect of a genetic pathway on a disease outcome using logistic kernel machine regression via logistic mixed modelsabstractBACKGROUND: Growing interest on biological pathways has called for new statistical methods for modeling and testing a genetic pathway effect on a health outcome. The fact that genes within a pathway tend to interact with each other and relate to the outcome in a complicated way makes nonparametric methods more desirable. The kernel machine method provides a convenient, powerful and unified method for multi-dimensional parametric and nonparametric modeling of the pathway effect. RESULTS: In this paper we propose a logistic kernel machine regression model for binary outcomes. This model relates the disease risk to covariates parametrically, and to genes within a genetic pathway parametrically or nonparametrically using kernel machines. The nonparametric genetic pathway effect allows for possible interactions among the genes within the same pathway and a complicated relationship of the genetic pathway and the outcome. We show that kernel machine estimation of the model components can be formulated using a logistic mixed model. Estimation hence can proceed within a mixed model framework using standard statistical software. A score test based on a Gaussian process approximation is developed to test for the genetic pathway effect. The methods are illustrated using a prostate cancer data set and evaluated using simulations. An extension to continuous and discrete outcomes using generalized kernel machine models and its connection with generalized linear mixed models is discussed. CONCLUSION: Logistic kernel machine regression and its extension generalized kernel machine regression provide a novel and flexible statistical tool for modeling pathway effects on discrete and continuous outcomes. Their close connection to mixed models and attractive performance make them have promising wide applications in bioinformatics and other biomedical areas. Debashis Ghosh, Xihong Lin |
BMC Bioinform. | 2 |
| 2007 | A Latent Variable Approach for Meta-Analysis of Gene Expression Data from Multiple Microarray ExperimentsabstractBACKGROUND: With the explosion in data generated using microarray technology by different investigators working on similar experiments, it is of interest to combine results across multiple studies. RESULTS: In this article, we describe a general probabilistic framework for combining high-throughput genomic data from several related microarray experiments using mixture models. A key feature of the model is the use of latent variables that represent quantities that can be combined across diverse platforms. We consider two methods for estimation of an index termed the probability of expression (POE). The first, reported in previous work by the authors, involves Markov Chain Monte Carlo (MCMC) techniques. The second method is a faster algorithm based on the expectation-maximization (EM) algorithm. The methods are illustrated with application to a meta-analysis of datasets for metastatic cancer. CONCLUSION: The statistical methods described in the paper are available as an R package, metaArray 1.8.1, which is at Bioconductor, whose URL is http://www.bioconductor.org/. Hyungwon Choi, Ronglai Shen, Arul M. Chinnaiyan, Debashis Ghosh |
BMC Bioinform. | 4 |
| 2006 | Watershed Segmentation for Medical Ultrasound ImagesabstractThis paper presents an efficient method for biomedical ultrasound image segmentation based on watershed transformation. It consists of four major stages. These stages are pre-processing, multiscale morphological gradient, watershed segmentation and finally region merging. The proposed scheme is tested using a set of medical ultrasound images. Experimental results show that our proposed method can produce accurate contours in medical ultrasound images. Bhabesh Deka, Debashis Ghosh |
SMC | 2 |
| 2006 | Peptide length-based prediction of peptide-MHC class II bindingabstractMOTIVATION: Algorithms for predicting peptide-MHC class II binding are typically similar, if not identical, to methods for predicting peptide-MHC class I binding despite known differences between the two scenarios. We investigate whether representing one of these differences, the greater range of peptide lengths binding MHC class II, improves the performance of these algorithms. RESULTS: A non-linear relationship between peptide length and peptide-MHC class II binding affinity was identified in the data available for several MHC class II alleles. Peptide length was incorporated into existing prediction algorithms using one of several modifications: using regression to pre-process the data, using peptide length as an additional variable within the algorithm, or representing register shifting in longer peptides. For several datasets and at least two algorithms these modifications consistently improved prediction accuracy. AVAILABILITY: http://malthus.micro.med.umich.edu/Bioinformatics Stewart T. Chang, Debashis Ghosh, Denise E. Kirschner, Jennifer J. Linderman |
Bioinform. | 2 |
| 2006 | COPA - cancer outlier profile analysisabstractUNLABELLED: Chromosomal translocations are common in cancer, and in some cases may be causal in the progression of the disease. Using microarrays, in which the expression of thousands of genes are simultaneously measured, could potentially allow one to detect recurrent translocations for a particular cancer type. Standard statistical tests, such as the t-test are not suited for detecting these translocations, but a simple test based on robust centering and scaling of the data to help detect outlier samples, followed by a search for pairs of samples with mutually exclusive outliers, may be used to find genes involved in recurrent translocations. We have implemented this method, termed Cancer Outlier Profile Analysis (COPA) in an R package (that we call the copa package), and show its applicability on a publicly available dataset. AVAILABILITY: http://www.bioconductor.org James W. MacDonald, Debashis Ghosh |
Bioinform. | 2 |
| 2006 | Eigengene-based linear discriminant model for tumor classification using gene expression microarray dataabstractMOTIVATION: The nearest shrunken centroids classifier has become a popular algorithm in tumor classification problems using gene expression microarray data. Feature selection is an embedded part of the method to select top-ranking genes based on a univariate distance statistic calculated for each gene individually. The univariate statistics summarize gene expression profiles outside of the gene co-regulation network context, leading to redundant information being included in the selection procedure. RESULTS: We propose an Eigengene-based Linear Discriminant Analysis (ELDA) to address gene selection in a multivariate framework. The algorithm uses a modified rotated Spectral Decomposition (SpD) technique to select 'hub' genes that associate with the most important eigenvectors. Using three benchmark cancer microarray datasets, we show that ELDA selects the most characteristic genes, leading to substantially smaller classifiers than the univariate feature selection based analogues. The resulting de-correlated expression profiles make the gene-wise independence assumption more realistic and applicable for the shrunken centroids classifier and other diagonal linear discriminant type of models. Our algorithm further incorporates a misclassification cost matrix, allowing differential penalization of one type of error over another. In the breast cancer data, we show false negative prognosis can be controlled via a cost-adjusted discriminant function. AVAILABILITY: R code for the ELDA algorithm is available from author upon request. Ronglai Shen, Debashis Ghosh, Arul M. Chinnaiyan, Zhaoling Meng |
Bioinform. | 2 |
| 2006 | Finding cancer subtypes in microarray data using random projections
Debashis Ghosh |
Neurocomputing | 1 |
| 2005 | A model-based scan statistic for identifying extreme chromosomal regions of gene expression in human tumorsabstractMOTIVATION: The analysis of gene expression data in its chromosomal context has been a recent development in cancer research. However, currently available methods fail to account for variation in the distance between genes, gene density and genomic features (e.g. GC content) in identifying increased or decreased chromosomal regions of gene expression. RESULTS: We have developed a model-based scan statistic that accounts for these aspects of the complex landscape of the human genome in the identification of extreme chromosomal regions of gene expression. This method may be applied to gene expression data regardless of the microarray platform used to generate it. To demonstrate the accuracy and utility of this method, we applied it to a breast cancer gene expression dataset and tested its ability to predict regions containing medium-to-high level DNA amplification (DNA ratio values >2). A classifier was developed from the scan statistic results that had a 10-fold cross-validated classification rate of 93% and a positive predictive value of 88%. This result strongly suggests that the model-based scan statistic and the expression characteristics of an increased chromosomal region of gene expression can be used to accurately predict chromosomal regions containing amplified genes. AVAILABILITY: Functions in the R-language are available from the author upon request. CONTACT: [email protected]. Albert M. Levin, Debashis Ghosh, Kathleen R. Cho, Sharon L. R. Kardia |
Bioinform. | 2 |
| 2005 | Comparison of seven methods for producing Affymetrix expression scores based on False Discovery Rates in disease profiling dataabstractBACKGROUND: A critical step in processing oligonucleotide microarray data is combining the information in multiple probes to produce a single number that best captures the expression level of a RNA transcript. Several systematic studies comparing multiple methods for array processing have used tightly controlled calibration data sets as the basis for comparison. Here we compare performances for seven processing methods using two data sets originally collected for disease profiling studies. An emphasis is placed on understanding sensitivity for detecting differentially expressed genes in terms of two key statistical determinants: test statistic variability for non-differentially expressed genes, and test statistic size for truly differentially expressed genes. RESULTS: In the two data sets considered here, up to seven-fold variation across the processing methods was found in the number of genes detected at a given false discovery rate (FDR). The best performing methods called up to 90% of the same genes differentially expressed, had less variable test statistics under randomization, and had a greater number of large test statistics in the experimental data. Poor performance of one method was directly tied to a tendency to produce highly variable test statistic values under randomization. Based on an overall measure of performance, two of the seven methods (Dchip and a trimmed mean approach) are superior in the two data sets considered here. Two other methods (MAS5 and GCRMA-EB) are inferior, while results for the other three methods are mixed. CONCLUSIONS: Choice of processing method has a major impact on differential expression analysis of microarray data. Previously reported performance analyses using tightly controlled calibration data sets are not highly consistent with results reported here using data from human tissue samples. Performance of array processing methods in disease profiling and other realistic biological studies should be given greater consideration when comparing Affymetrix processing methods. Kerby Shedden, Wei Chen 0128, Rork Kuick, Debashis Ghosh, James W. MacDonald, Kathleen R. Cho, Thomas J. Giordano, Stephen B. Gruber, Eric R. Fearon, Jeremy M. G. Taylor, Samir Hanash |
BMC Bioinform. | 4 |
| 2004 | Mixture models for assessing differential expression in complex tissues using microarray dataabstractMOTIVATION: The use of DNA microarrays has become quite popular in many scientific and medical disciplines, such as in cancer research. One common goal of these studies is to determine which genes are differentially expressed between cancer and healthy tissue, or more generally, between two experimental conditions. A major complication in the molecular profiling of tumors using gene expression data is that the data represent a combination of tumor and normal cells. Much of the methodology developed for assessing differential expression with microarray data has assumed that tissue samples are homogeneous. RESULTS: In this paper, we outline a general framework for determining differential expression in the presence of mixed cell populations. We consider study designs in which paired tissues and unpaired tissues are available. A hierarchical mixture model is used for modeling the data; a combination of methods of moments procedures and the expectation-maximization algorithm are used to estimate the model parameters. The finite-sample properties of the methods are assessed in simulation studies; they are applied to two microarray datasets from cancer studies. Commands in the R language can be downloaded from the URL http://www.sph.umich.edu/~ghoshd/COMPBIO/COMPMIX/. Debashis Ghosh |
Bioinform. | 1 |
| 2003 | Cluster stability scores for microarray data in cancer studiesabstractBACKGROUND: A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will define subtypes of disease. Hierarchical clustering has been the primary analytical tool used to define disease subtypes from microarray experiments in cancer settings. Assessing cluster reliability poses a major complication in analyzing output from clustering procedures. While most work has focused on estimating the number of clusters in a dataset, the question of stability of individual-level clusters has not been addressed. RESULTS: We address this problem by developing cluster stability scores using subsampling techniques. These scores exploit the redundancy in biologically discriminatory information on the chip. Our approach is generic and can be used with any clustering method. We propose procedures for calculating cluster stability scores for situations involving both known and unknown numbers of clusters. We also develop cluster-size adjusted stability scores. The method is illustrated by application to data three cancer studies; one involving childhood cancers, the second involving B-cell lymphoma, and the final is from a malignant melanoma study. AVAILABILITY: Code implementing the proposed analytic method can be obtained at the second author's website. Mark Smolkin, Debashis Ghosh |
BMC Bioinform. | 2 |
| 2002 | Mixture modelling of gene expression data from microarray experimentsabstractAbstract Motivation: Hierarchical clustering is one of the major analytical tools for gene expression data from microarray experiments. A major problem in the interpretation of the output from these procedures is assessing the reliability of the clustering results. We address this issue by developing a mixture model-based approach for the analysis of microarray data. Within this framework, we present novel algorithms for clustering genes and samples. One of the byproducts of our method is a probabilistic measure for the number of true clusters in the data. Results: The proposed methods are illustrated by application to microarray datasets from two cancer studies; one in which malignant melanoma is profiled (Bittner et al. , Nature , 406, 536–540, 2000), and the other in which prostate cancer is profiled (Dhanasekaran et al. , 2001, submitted). Availability: Macros written in the R language implementing the methods in this report can be obtained at the first author’s website: http://www.sph.umich.edu/~ghoshd/COMPBIO/mixture1/index.html. Contact: [email protected] Debashis Ghosh, Arul M. Chinnaiyan |
Bioinform. | 1 |
| 1999 | An analytic approach for generation of artificial hand-printed character database from given generative models
Debashis Ghosh, A. P. Shivaprasad |
Pattern Recognit. | 1 |