EDBT 2026 Demo / reviewers in the wild / expert
Kun Huang 0001
dblp:10/4151-1
· DBLP profile ↗
67ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-8530-370XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identification of high-risk cells in single-cell spatially resolved transcriptomics data using Diagnostic Evidence GAuge of Single-cells with spatial smoothingabstractSUMMARY: The examination of high-risk cells and regions in tissue samples from spatially resolved transcriptomics platforms offers meaningful insights into specific disease processes. For existing methods, while cell types or clusters can be identified and associated with disease attributes, individual cells are unable to be associated in the same manner. METHOD: Diagnostic Evidence Gauge of Single-Cells and Spatial Transcriptomics (DEGAS) solves the above problem by employing latent representations of gene expression data and domain adaptation to transfer disease attributes from patients to individual cells from single-cell RNA sequencing datasets. In this research, we present and evaluate DEGAS's versatility in adapting to data arising from various single-cell spatially resolved transcriptomics (scSRT) platforms. DEGAS successfully identified high-risk cells and regions in liver hepatocellular carcinoma and skin cutaneous melanoma, which were validated through known markers. Additionally, DEGAS was applied to our newly generated Type II Diabetes Xenium dataset, revealing high-risk cells within the tissue samples. AVAILABILITY AND IMPLEMENTATION: The DEGAS software can be accessed at https://github.com/tsteelejohnson91/DEGAS. For the updated smoothing functions and associated codes, visit https://github.com/dchatter04/DEGAS-Spatial-Smoothing, which is archived at https://doi.org/10.5281/zenodo.18510221. Sources for the datasets reviewed are detailed in their respective sections. A description of some datasets, along with extra tables and figures, is provided in the Supplementary Materials file. Our newly generated Xenium data for Type II Diabetes can be found at https://doi.org/10.7303/syn68699752. Debolina Chatterjee, Justin L. Couetil, Kun Huang 0001, Chao Chen 0012, Jie Zhang 0010, Michael A Kalwat, Travis S. Johnson |
Bioinform. | 4 |
| 2025 | Leveraging transcription factor physical proximity for enhancing gene regulation inferenceabstractMOTIVATION: Gene regulation inference, a key challenge in systems biology, is crucial for understanding cell function, as it governs processes such as differentiation, cell state maintenance, signal transduction, and stress response. Leading methods utilize gene expression, chromatin accessibility, transcription factor (TF) DNA binding motifs, and prior knowledge. However, they overlook the fact that TFs must be in physical proximity to facilitate transcriptional gene regulation. RESULTS: To fill the gap, we develop GRIP-Gene Regulation Inference by considering TF Proximity-a gene regulation inference method that directly considers the physical proximity between regulating TFs. Specifically, we use the distance in a protein-protein interaction (PPI) network to estimate the physical proximity between TFs. We design a novel Boolean convex program, which can identify TFs that not only can explain the gene expression of target genes (TGs) but also stay close in the PPI network. We propose an efficient algorithm to solve the Boolean relaxation of the proposed model with a theoretical tightness guarantee. We compare our GRIP with state-of-the-art methods (SCENIC+, DirectNet, Pando, and CellOracle) on inferring cell-type-specific (CD4, CD8, and CD 14) gene regulation using the PBMC 3k scMultiome-seq data and demonstrate its out-performance in terms of the predictive power of the inferred TFs, the physical distance between the inferred TFs, and the agreement between the inferred gene regulation and PCHiC data. AVAILABILITY AND IMPLEMENTATION: https://github.com/EJIUB/GRIP. Xiaoqing Huang, Aamir R. Hullur, Elham Jafari, Kaushik Shridhar, Mu Zhou, Kenneth MacKie, Kun Huang 0001 |
Bioinform. | 7 |
| 2022 | Identify Consistent Imaging Genomic Biomarkers for Characterizing the Survival-Associated Interactions Between Tumor-Infiltrating Lymphocytes and Tumors
Yingli Zuo, Yawen Wu, Zixiao Lu, Qi Zhu 0001, Kun Huang 0001, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (2) | 5 |
| 2022 | SPCS: a spatial and pattern combined smoothing method for spatial transcriptomic expressionabstractHigh-dimensional, localized ribonucleic acid (RNA) sequencing is now possible owing to recent developments in spatial transcriptomics (ST). ST is based on highly multiplexed sequence analysis and uses barcodes to match the sequenced reads to their respective tissue locations. ST expression data suffer from high noise and dropout events; however, smoothing techniques have the promise to improve the data interpretability prior to performing downstream analyses. Single-cell RNA sequencing (scRNA-seq) data similarly suffer from these limitations, and smoothing methods developed for scRNA-seq can only utilize associations in transcriptome space (also known as one-factor smoothing methods). Since they do not account for spatial relationships, these one-factor smoothing methods cannot take full advantage of ST data. In this study, we present a novel two-factor smoothing technique, spatial and pattern combined smoothing (SPCS), that employs the k-nearest neighbor (kNN) technique to utilize information from transcriptome and spatial relationships. By performing SPCS on multiple ST slides from pancreatic ductal adenocarcinoma (PDAC), dorsolateral prefrontal cortex (DLPFC) and simulated high-grade serous ovarian cancer (HGSOC) datasets, smoothed ST slides have better separability, partition accuracy and biological interpretability than the ones smoothed by preexisting one-factor methods. Source code of SPCS is provided in Github (https://github.com/Usos/SPCS). Yusong Liu, Tongxin Wang, Ben Duggan, Michael F. Sharpnack, Kun Huang 0001, Jie Zhang 0010, Xiufen Ye, Travis S. Johnson |
Briefings Bioinform. | 5 |
| 2022 | TSAFinder: exhaustive tumor-specific antigen detection with RNAseqabstractMOTIVATION: Tumor-specific antigen (TSA) identification in human cancer predicts response to immunotherapy and provides targets for cancer vaccine and adoptive T-cell therapies with curative potential, and TSAs that are highly expressed at the RNA level are more likely to be presented on major histocompatibility complex (MHC)-I. Direct measurements of the RNA expression of peptides would allow for generalized prediction of TSAs. Human leukocyte antigen (HLA)-I genotypes were predicted with seq2HLA. RNA sequencing (RNAseq) fastq files were translated into all possible peptides of length 8-11, and peptides with high and low expressions in the tumor and control samples, respectively, were tested for their MHC-I binding potential with netMHCpan-4.0. RESULTS: A novel pipeline for TSA prediction from RNAseq was used to predict all possible unique peptides size 8-11 on previously published murine and human lung and lymphoma tumors and validated on matched tumor and control lung adenocarcinoma (LUAD) samples. We show that neoantigens predicted by exomeSeq are typically poorly expressed at the RNA level, and a fraction is expressed in matched normal samples. TSAs presented in the proteomics data have higher RNA abundance and lower MHC-I binding percentile, and these attributes are used to discover high confidence TSAs within the validation cohort. Finally, a subset of these high confidence TSAs is expressed in a majority of LUAD tumors and represents attractive vaccine targets. AVAILABILITY AND IMPLEMENTATION: The datasets were derived from sources in the public domain as follows: TSAFinder is open-source software written in python and R. It is licensed under CC-BY-NC-SA and can be downloaded at https://github.com/RNAseqTSA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael F. Sharpnack, Travis S. Johnson, Robert Chalkley, Zhi Han, David Carbone, Kun Huang 0001 |
Bioinform. | 6 |
| 2022 | Application of unsupervised deep learning algorithms for identification of specific clusters of chronic cough patients from EMR dataabstractBACKGROUND: Chronic cough affects approximately 10% of adults. The lack of ICD codes for chronic cough makes it challenging to apply supervised learning methods to predict the characteristics of chronic cough patients, thereby requiring the identification of chronic cough patients by other mechanisms. We developed a deep clustering algorithm with auto-encoder embedding (DCAE) to identify clusters of chronic cough patients based on data from a large cohort of 264,146 patients from the Electronic Medical Records (EMR) system. We constructed features using the diagnosis within the EMR, then built a clustering-oriented loss function directly on embedded features of the deep autoencoder to jointly perform feature refinement and cluster assignment. Lastly, we performed statistical analysis on the identified clusters to characterize the chronic cough patients compared to the non-chronic cough patients. RESULTS: The experimental results show that the DCAE model generated three chronic cough clusters and one non-chronic cough patient cluster. We found various diagnoses, medications, and lab tests highly associated with chronic cough patients by comparing the chronic cough cluster with the non-chronic cough cluster. Comparison of chronic cough clusters demonstrated that certain combinations of medications and diagnoses characterize some chronic cough clusters. CONCLUSIONS: To the best of our knowledge, this study is the first to test the potential of unsupervised deep learning methods for chronic cough investigation, which also shows a great advantage over existing algorithms for patient data clustering. Wei Shao 0005, Xiao Luo 0002, Zuoyi Zhang, Zhi Han, Vasu Chandrasekaran, Vladimir Turzhitsky, Vishal Bali, Anna R. Roberts, Megan Metzger, Jarod Baker, Carmen La Rosa, Jessica Weaver, Paul Richard Dexter, Kun Huang 0001 |
BMC Bioinform. | 14 |
| 2022 | A Deep Language Model for Symptom Extraction From Clinical Text and its Application to Extract COVID-19 Symptoms From Social MediaabstractPatients experience various symptoms when they haveeither acute or chronic diseases or undergo some treatments for diseases. Symptoms are often indicators of the severity of the disease and the need for hospitalization. Symptoms are often described in free text written as clinical notes in the Electronic Health Records (EHR) and are not integrated with other clinical factors for disease prediction and healthcare outcome management. In this research, we propose a novel deep language model to extract patient-reported symptoms from clinical text. The deep language model integrates syntactic and semantic analysis for symptom extraction and identifies the actual symptoms reported by patients and conditional or negation symptoms. The deep language model can extract both complex and straightforward symptom expressions. We used a real-world clinical notes dataset to evaluate our model and demonstrated that our model achieves superior performance compared to three other state-of-the-art symptom extraction models. We extensively analyzed our model to illustrate its effectiveness by examining each component's contribution to the model. Finally, we applied our model on a COVID-19 tweets data set to extract COVID-19 symptoms. The results show that our model can identify all the symptoms suggested by the Center for Disease Control (CDC) ahead of their timeline and many rare symptoms. Xiao Luo 0002, Priyanka Gandhi, Susan Storey, Kun Huang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | PASC CKD: revealing the progression trajectories of sustained COVID-19-related renal injury using real-world evidence
Jing Su 0003, Pengyue Zhang, Zuoyi Zhang, Michael Eadon, Xiaochun Li 0003, Stanley Taylor, Travis Johnson, Zhaorui Liu, Ziyang Tang, Baijian Yang 0001, Qianqian Song 0002, Kun Huang 0001 |
AMIA | 13 |
| 2021 | Transfer Learning via Optimal Transportation for Integrative Cancer Patient StratificationabstractThe Stratification of early-stage cancer patients for the prediction of clinical outcome is a challenging task since cancer is associated with various molecular aberrations. A single biomarker often cannot provide sufficient information to stratify early-stage patients effectively. Understanding the complex mechanism behind cancer development calls for exploiting biomarkers from multiple modalities of data such as histopathology images and genomic data. The integrative analysis of these biomarkers sheds light on cancer diagnosis, subtyping, and prognosis. Another difficulty is that labels for early-stage cancer patients are scarce and not reliable enough for predicting survival times. Given the fact that different cancer types share some commonalities, we explore if the knowledge learned from one cancer type can be utilized to improve prognosis accuracy for another cancer type. We propose a novel unsupervised multi-view transfer learning algorithm to simultaneously analyze multiple biomarkers in different cancer types. We integrate multiple views using non-negative matrix factorization and formulate the transfer learning model based on the Optimal Transport theory to align features of different cancer types. We evaluate the stratification performance on three early-stage cancers from the Cancer Genome Atlas (TCGA) project. Comparing with other benchmark methods, our framework achieves superior accuracy for patient outcome prediction. Wei Shao 0005, Jie Zhang 0010, Kun Huang 0001 |
IJCAI | 5 |
| 2021 | Towards Fair Cross-Domain Adaptation via Generative LearningabstractDomain Adaptation (DA) targets at adapting a model trained over the well-labeled source domain to the unlabeled target domain lying in different distributions. Existing DA normally assumes the well-labeled source domain is class-wise balanced, which means the size per source class is relatively similar. However, in real-world applications, labeled samples for some categories in the source domain could be extremely few due to the difficulty of data collection and annotation, which leads to decreasing performance over target domain on those few-shot categories. To perform fair cross-domain adaptation and boost the performance on these minority categories, we develop a novel Generative Few-shot Cross-domain Adaptation (GFCA) algorithm for fair cross-domain classification. Specifically, generative feature augmentation is explored to synthesize effective training data for few-shot source classes, while effective cross-domain alignment aims to adapt knowledge from source to facilitate the target learning. Experimental results on two large cross-domain visual datasets demonstrate the effectiveness of our proposed method on improving both few-shot and overall classification accuracy comparing with the state-of-the-art DA approaches. Tongxin Wang, Zhengming Ding, Wei Shao 0005, Haixu Tang, Kun Huang 0001 |
WACV | 5 |
| 2021 | TPSC: a module detection method based on topology potential and spectral clustering in weighted networks and its application in gene co-expression module discoveryabstractBACKGROUND: Gene co-expression networks are widely studied in the biomedical field, with algorithms such as WGCNA and lmQCM having been developed to detect co-expressed modules. However, these algorithms have limitations such as insufficient granularity and unbalanced module size, which prevent full acquisition of knowledge from data mining. In addition, it is difficult to incorporate prior knowledge in current co-expression module detection algorithms. RESULTS: In this paper, we propose a novel module detection algorithm based on topology potential and spectral clustering algorithm to detect co-expressed modules in gene co-expression networks. By testing on TCGA data, our novel method can provide more complete coverage of genes, more balanced module size and finer granularity than current methods in detecting modules with significant overall survival difference. In addition, the proposed algorithm can identify modules by incorporating prior knowledge. CONCLUSION: In summary, we developed a method to obtain as much as possible information from networks with increased input coverage and the ability to detect more size-balanced and granular modules. In addition, our method can integrate data from different sources. Our proposed method performs better than current methods with complete coverage of input genes and finer granularity. Moreover, this method is designed not only for gene co-expression networks but can also be applied to any general fully connected weighted network. Yusong Liu, Xiufen Ye, Christina Y. Yu, Wei Shao 0005, Weixing Feng, Jie Zhang 0010, Kun Huang 0001 |
BMC Bioinform. | 8 |
| 2021 | A Computational Framework to Analyze the Associations Between Symptoms and Cancer Patient Attributes Post Chemotherapy Using EHR DataabstractPatients with cancer, such as breast and colorectal cancer, often experience different symptoms post-chemotherapy. The symptoms could be fatigue, gastrointestinal (nausea, vomiting, lack of appetite), psychoneurological symptoms (depressive symptoms, anxiety), or other types. Previous research focused on understanding the symptoms using survey data. In this research, we propose to utilize the data within the Electronic Health Record (EHR). A computational framework is developed to use a natural language processing (NLP) pipeline to extract the clinician-documented symptoms from clinical notes. Then, a patient clustering method is based on the symptom severity levels to group the patient in clusters. The association rule mining is used to analyze the associations between symptoms and patient attributes (smoking history, number of comorbidities, diabetes status, age at diagnosis) in the patient clusters. The results show that the various symptom types and severity levels have different associations between breast and colorectal cancers and different timeframes post-chemotherapy. The results also show that patients with breast or colorectal cancers, who smoke and have severe fatigue, likely have severe gastrointestinal symptoms six months after the chemotherapy. Our framework can be generalized to analyze symptoms or symptom clusters of other chronic diseases where symptom management is critical. Xiao Luo 0002, Priyanka Gandhi, Susan Storey, Zuoyi Zhang, Zhi Han, Kun Huang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Weakly Supervised Deep Ordinal Cox Model for Survival Prediction From Whole-Slide Pathological ImagesabstractWhole-Slide Histopathology Image (WSI) is generally considered the gold standard for cancer diagnosis and prognosis. Given the large inter-operator variation among pathologists, there is an imperative need to develop machine learning models based on WSIs for consistently predicting patient prognosis. The existing WSI-based prediction methods do not utilize the ordinal ranking loss to train the prognosis model, and thus cannot model the strong ordinal information among different patients in an efficient way. Another challenge is that a WSI is of large size (e.g., 100,000-by-100,000 pixels) with heterogeneous patterns but often only annotated with a single WSI-level label, which further complicates the training process. To address these challenges, we consider the ordinal characteristic of the survival process by adding a ranking-based regularization term on the Cox model and propose a weakly supervised deep ordinal Cox model (BDOCOX) for survival prediction from WSIs. Here, we generate amounts of bags from WSIs, and each bag is comprised of the image patches representing the heterogeneous patterns of WSIs, which is assumed to match the WSI-level labels for training the proposed model. The effectiveness of the proposed method is well validated by theoretical analysis as well as the prognosis and patient stratification results on three cancer datasets from The Cancer Genome Atlas (TCGA). Wei Shao 0005, Tongxin Wang, Zhi Han, Jie Zhang 0010, Kun Huang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Multi-task multi-modal learning for joint diagnosis and prognosis of human cancers
Wei Shao 0005, Tongxin Wang, Liang Sun 0009, Tianhan Dong, Zhi Han, Jie Zhang 0010, Daoqiang Zhang, Kun Huang 0001 |
Medical Image Anal. | 9 |
| 2020 | AI in Medical Imaging Informatics: Current Challenges and Future DirectionsabstractThis paper reviews state-of-the-art research solutions across the spectrum of medical imaging informatics, discusses clinical translation, and provides future directions for advancing clinical practice. More specifically, it summarizes advances in medical imaging acquisition technologies for different modalities, highlighting the necessity for efficient medical data management strategies in the context of AI in big healthcare data analytics. It then provides a synopsis of contemporary and emerging algorithmic methods for disease classification and organ/ tissue segmentation, focusing on AI and deep learning architectures that have already become the de facto approach. The clinical benefits of in-silico modelling advances linked with evolving 3D reconstruction and visualization applications are further documented. Concluding, integrative analytics approaches driven by associate research branches highlighted in this study promise to revolutionize imaging informatics as known today across the healthcare continuum for both radiology and digital pathology applications. The latter, is projected to enable informed, more accurate diagnosis, timely prognosis, and effective treatment planning, underpinning precision medicine. Andreas Panayides, Amir A. Amini, Nenad Filipovic, Ashish Sharma 0001, Sotirios A. Tsaftaris, Alistair A. Young, David J. Foran, Nhan Do, Spyretta Golemati, Tahsin M. Kurç, Kun Huang 0001, Konstantina S. Nikita, Benjamin Veasey, Michalis E. Zervakis, Joel H. Saltz, Constantinos S. Pattichis |
IEEE J. Biomed. Health Informatics | 11 |
| 2020 | Integrative Analysis of Pathological Images and Multi-Dimensional Genomic Data for Early-Stage Cancer PrognosisabstractThe integrative analysis of histopathological images and genomic data has received increasing attention for studying the complex mechanisms of driving cancers. However, most image-genomic studies have been restricted to combining histopathological images with the single modality of genomic data (e.g., mRNA transcription or genetic mutation), and thus neglect the fact that the molecular architecture of cancer is manifested at multiple levels, including genetic, epigenetic, transcriptional, and post-transcriptional events. To address this issue, we propose a novel ordinal multi-modal feature selection (OMMFS) framework that can simultaneously identify important features from both pathological images and multi-modal genomic data (i.e., mRNA transcription, copy number variation, and DNA methylation data) for the prognosis of cancer patients. Our model is based on a generalized sparse canonical correlation analysis framework, by which we also take advantage of the ordinal survival information among different patients for survival outcome prediction. We evaluate our method on three early-stage cancer datasets derived from The Cancer Genome Atlas (TCGA) project, and the experimental results demonstrated that both the selected image and multi-modal genomic markers are strongly correlated with survival enabling effective stratification of patients with distinct survival than the comparing methods, which is often difficult for early-stage cancer patients. Wei Shao 0005, Kun Huang 0001, Zhi Han, Jun Cheng 0006, Tongxin Wang, Liang Sun 0009, Zixiao Lu, Jie Zhang 0010, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Image saliency detection via multi-scale iterative CNN
Kun Huang 0001, Shenghua Gao |
Vis. Comput. | 1 |
| 2019 | PPGNet: Learning Point-Pair Graph for Line Segment DetectionabstractIn this paper, we present a novel framework to detect line segments in man-made environments. Specifically, we propose to describe junctions, line segments and relationships between them with a simple graph, which is more structured and informative than end-point representation used in existing line segment detection methods. In order to extract a line segment graph from an image, we further introduce the PPGNet, a convolutional neural network that directly infers a graph from an image. We evaluate our method on published benchmarks including York Urban and Wireframe datasets. The results demonstrate that our method achieves satisfactory performance and generalizes well on all the benchmarks. The source code of our work is available at https://github.com/svip-lab/PPGNet. Ning Bi, Jia Zheng 0002, Kun Huang 0001, Weixin Luo, Yanyu Xu 0001, Shenghua Gao |
CVPR | 6 |
| 2019 | Diagnosis-Guided Multi-modal Feature Selection for Prognosis Prediction of Lung Squamous Cell Carcinoma
Wei Shao 0005, Tongxin Wang, Jun Cheng 0006, Zhi Han, Daoqiang Zhang, Kun Huang 0001 |
MICCAI (4) | 7 |
| 2019 | Integrative cancer patient stratification via subspace mergingabstractMOTIVATION: Technologies that generate high-throughput omics data are flourishing, creating enormous, publicly available repositories of multi-omics data. As many data repositories continue to grow, there is an urgent need for computational methods that can leverage these data to create comprehensive clusters of patients with a given disease. RESULTS: Our proposed approach creates a patient-to-patient similarity graph for each data type as an intermediate representation of each omics data type and merges the graphs through subspace analysis on a Grassmann manifold. We hypothesize that this approach generates more informative clusters by preserving the complementary information from each level of omics data. We applied our approach to The Cancer Genome Atlas (TCGA) breast cancer dataset and show that by integrating gene expression, microRNA and DNA methylation data, our proposed method can produce clinically useful subtypes of breast cancer. We then investigate the molecular characteristics underlying these subtypes. We discover a highly expressed cluster of genes on chromosome 19p13 that strongly correlates with survival in TCGA breast cancer patients and validate these results in three additional breast cancer datasets. We also compare our approach with previous integrative clustering approaches and obtain comparable or superior results. AVAILABILITY AND IMPLEMENTATION: https://github.com/michaelsharpnack/GrassmannCluster. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael F. Sharpnack, Chao Wang 0081, Kun Huang 0001, Raghu Machiraju |
Bioinform. | 4 |
| 2019 | LAmbDA: label ambiguous domain adaptation dataset integration reduces batch effects and improves subtype detectionabstractMOTIVATION: Rapid advances in single cell RNA sequencing (scRNA-seq) have produced higher-resolution cellular subtypes in multiple tissues and species. Methods are increasingly needed across datasets and species to (i) remove systematic biases, (ii) model multiple datasets with ambiguous labels and (iii) classify cells and map cell type labels. However, most methods only address one of these problems on broad cell types or simulated data using a single model type. It is also important to address higher-resolution cellular subtypes, subtype labels from multiple datasets, models trained on multiple datasets simultaneously and generalizability beyond a single model type. RESULTS: We developed a species- and dataset-independent transfer learning framework (LAmbDA) to train models on multiple datasets (even from different species) and applied our framework on simulated, pancreas and brain scRNA-seq experiments. These models mapped corresponding cell types between datasets with inconsistent cell subtype labels while simultaneously reducing batch effects. We achieved high accuracy in labeling cellular subtypes (weighted accuracy simulated 1 datasets: 90%; simulated 2 datasets: 94%; pancreas datasets: 88% and brain datasets: 66%) using LAmbDA Feedforward 1 Layer Neural Network with bagging. This method achieved higher weighted accuracy in labeling cellular subtypes than two other state-of-the-art methods, scmap and CaSTLe in brain (66% versus 60% and 32%). Furthermore, it achieved better performance in correctly predicting ambiguous cellular subtype labels across datasets in 88% of test cases compared with CaSTLe (63%), scmap (50%) and MetaNeighbor (50%). LAmbDA is model- and dataset-independent and generalizable to diverse data types representing an advance in biocomputing. AVAILABILITY AND IMPLEMENTATION: github.com/tsteelejohnson91/LAmbDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Travis S. Johnson, Tongxin Wang, Christina Y. Yu, Yatong Han, Kun Huang 0001, Jie Zhang 0010 |
Bioinform. | 8 |
| 2019 | A protocol to evaluate RNA sequencing normalization methodsabstractBACKGROUND: RNA sequencing technologies have allowed researchers to gain a better understanding of how the transcriptome affects disease. However, sequencing technologies often unintentionally introduce experimental error into RNA sequencing data. To counteract this, normalization methods are standardly applied with the intent of reducing the non-biologically derived variability inherent in transcriptomic measurements. However, the comparative efficacy of the various normalization techniques has not been tested in a standardized manner. Here we propose tests that evaluate numerous normalization techniques and applied them to a large-scale standard data set. These tests comprise a protocol that allows researchers to measure the amount of non-biological variability which is present in any data set after normalization has been performed, a crucial step to assessing the biological validity of data following normalization. RESULTS: In this study we present two tests to assess the validity of normalization methods applied to a large-scale data set collected for systematic evaluation purposes. We tested various RNASeq normalization procedures and concluded that transcripts per million (TPM) was the best performing normalization method based on its preservation of biological signal as compared to the other methods tested. CONCLUSION: Normalization is of vital importance to accurately interpret the results of genomic and transcriptomic experiments. More work, however, needs to be performed to optimize normalization methods for RNASeq data. The present effort helps pave the way for more systematic evaluations of normalization methods across different platforms. With our proposed schema researchers can evaluate their own or future normalization methods to further improve the field of RNASeq normalization. Zachary B. Abrams, Travis S. Johnson, Kun Huang 0001, Philip R. O. Payne, Kevin R. Coombes |
BMC Bioinform. | 3 |
| 2019 | Generalized gene co-expression analysis via subspace clustering using low-rank representationabstractBACKGROUND: Gene Co-expression Network Analysis (GCNA) helps identify gene modules with potential biological functions and has become a popular method in bioinformatics and biomedical research. However, most current GCNA algorithms use correlation to build gene co-expression networks and identify modules with highly correlated genes. There is a need to look beyond correlation and identify gene modules using other similarity measures for finding novel biologically meaningful modules. RESULTS: We propose a new generalized gene co-expression analysis algorithm via subspace clustering that can identify biologically meaningful gene co-expression modules with genes that are not all highly correlated. We use low-rank representation to construct gene co-expression networks and local maximal quasi-clique merger to identify gene co-expression modules. We applied our method on three large microarray datasets and a single-cell RNA sequencing dataset. We demonstrate that our method can identify gene modules with different biological functions than current GCNA methods and find gene modules with prognostic values. CONCLUSIONS: The presented method takes advantage of subspace clustering to generate gene co-expression networks rather than using correlation as the similarity measure between genes. Our generalized GCNA method can provide new insights from gene expression datasets and serve as a complement to current GCNA algorithms. Tongxin Wang, Jie Zhang 0010, Kun Huang 0001 |
BMC Bioinform. | 3 |
| 2018 | Genetic Mutations Associated with Histopathology Changes in Kidney Cancer
Jun Cheng 0006, Zhi Han, Qianjin Feng 0002, Jie Zhang 0010, Kun Huang 0001 |
AMIA | 6 |
| 2018 | Learning to Parse Wireframes in Images of Man-Made EnvironmentsabstractIn this paper, we propose a learning-based approach to the task of automatically extracting a "wireframe" representation for images of cluttered man-made environments. The wireframe (see Fig. 1) contains all salient straight lines and their junctions of the scene that encode efficiently and accurately large-scale geometry and object shapes. To this end, we have built a very large new dataset of over 5,000 images with wireframes thoroughly labelled by humans. We have proposed two convolutional neural networks that are suitable for extracting junctions and lines with large spatial support, respectively. The networks trained on our dataset have achieved significantly better performance than state-of-the-art methods for junction detection and line segment detection, respectively. We have conducted extensive experiments to evaluate quantitatively and qualitatively the wireframes obtained by our method, and have convincingly shown that effectively and efficiently parsing wireframes for images of man-made environments is a feasible goal within reach. Such wireframes could benefit many important visual tasks such as feature correspondence, 3D reconstruction, vision-based mapping, localization, and navigation. The data and source code are available at https://github.com/huangkuns/wireframe. Kun Huang 0001, Zihan Zhou 0001, Tianjiao Ding, Shenghua Gao, Yi Ma 0001 |
CVPR | 1 |
| 2018 | Ordinal Multi-modal Feature Selection for Survival Analysis of Early-Stage Renal Cancer
Wei Shao 0005, Jun Cheng 0006, Liang Sun 0009, Zhi Han, Qianjin Feng 0002, Daoqiang Zhang, Kun Huang 0001 |
MICCAI (2) | 7 |
| 2018 | Identification of topological features in renal tumor microenvironment associated with patient survivalabstractMotivation: As a highly heterogeneous disease, the progression of tumor is not only achieved by unlimited growth of the tumor cells, but also supported, stimulated, and nurtured by the microenvironment around it. However, traditional qualitative and/or semi-quantitative parameters obtained by pathologist's visual examination have very limited capability to capture this interaction between tumor and its microenvironment. With the advent of digital pathology, computerized image analysis may provide a better tumor characterization and give new insights into this problem. Results: We propose a novel bioimage informatics pipeline for automatically characterizing the topological organization of different cell patterns in the tumor microenvironment. We apply this pipeline to the only publicly available large histopathology image dataset for a cohort of 190 patients with papillary renal cell carcinoma obtained from The Cancer Genome Atlas project. Experimental results show that the proposed topological features can successfully stratify early- and middle-stage patients with distinct survival, and show superior performance to traditional clinical features and cellular morphological and intensity features. The proposed features not only provide new insights into the topological organizations of cancers, but also can be integrated with genomic data in future studies to develop new integrative biomarkers. Availability and implementation: https://github.com/chengjun583/KIRP-topological-features. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jun Cheng 0006, Xiaokui Mo, Anil V. Parwani, Qianjin Feng 0002, Kun Huang 0001 |
Bioinform. | 6 |
| 2018 | annoPeak: a web application to annotate and visualize peaks from ChIP-seq/ChIP-exo-seq
Arunima Srivastava, Huayang Liu, Raghu Machiraju, Kun Huang 0001, Gustavo Leone |
Bioinform. | 5 |
| 2017 | annoPeak: a web application to annotate and visualize peaks from ChIP-seq/ChIP-exo-seqabstractSUMMARY: We developed annoPeak, a web application to annotate, visualize and compare predicted protein-binding regions derived from ChIP-seq/ChIP-exo-seq experiments using human and mouse cells. Users can upload peak regions from multiple experiments onto the annoPeak server to annotate them with biological context, identify associated target genes and categorize binding sites with respect to gene structure. Users can also compare multiple binding profiles intuitively with the help of visualization tools and tables provided by annoPeak. In general, annoPeak will help users identify patterns of genome wide transcription factor binding profiles, assess binding profiles in different biological contexts and generate new hypotheses. AVAILABILITY AND IMPLEMENTATION: The web service is freely accessible through URL: http://ccc-annopeak.osumc.edu/annoPeak . Source code is available at https://github.com/XingTang2014/annoPeak . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Arunima Srivastava, Huayang Liu, Raghu Machiraju, Kun Huang 0001, Gustavo Leone |
Bioinform. | 5 |
| 2016 | Identification of recurrent combinatorial patterns of chromatin modifications at promoters across various tissue typesabstractBACKGROUND: Identification and analysis of recurrent combinatorial patterns of multiple chromatin modifications provide invaluable information for understanding epigenetic regulations. Furthermore, as more data becomes available, it is computationally expensive and unnecessary to study combinatorial patterns of all modifications. METHODS: A novel framework is proposed to investigate recurrent combinatorial patterns of a subset of quantitatively selected chromatin modifications. The framework is based on heirarchical clustering and selects subsets of chromatin modifications that form distinct recurrent patterns at regulatory regions. The identified recurrent combinatorial patterns can be further utilized to discover novel regulatory regions. Data is in the form of genome wide maps of histone acetylations, methylations, and histone variant of human skeletal muscular and B-lymphocyte cells both derived from the ENCODE project. RESULTS: A case study conducted at promoter regions is presented: four out of twelve chromatin modifications were selected, eight different promoter states were identified and the identified patterns of active promoters were further utilized to discover novel promoter regions. Several previously un-annotated promoters were discovered, further investigations confirm their promoter functions. CONCLUSIONS: This framework is approproiately general and could lead to better understanding of epigenetic regulations by discovering previously unknown regulatory regions. Nan Meng, Raghu Machiraju, Kun Huang 0001 |
BMC Bioinform. | 3 |
| 2016 | Single-Cell Co-expression Analysis Reveals Distinct Functional Modules, Co-regulation Mechanisms and Clinical OutcomesabstractCo-expression analysis has been employed to predict gene function, identify functional modules, and determine tumor subtypes. Previous co-expression analysis was mainly conducted at bulk tissue level. It is unclear whether co-expression analysis at the single-cell level will provide novel insights into transcriptional regulation. Here we developed a computational approach to compare glioblastoma expression profiles at the single-cell level with those obtained from bulk tumors. We found that the co-expressed genes observed in single cells and bulk tumors have little overlap and show distinct characteristics. The co-expressed genes identified in bulk tumors tend to have similar biological functions, and are enriched for intrachromosomal interactions with synchronized promoter activity. In contrast, single-cell co-expressed genes are enriched for known protein-protein interactions, and are regulated through interchromosomal interactions. Moreover, gene members of some protein complexes are co-expressed only at the bulk level, while those of other complexes are co-expressed at both single-cell and bulk levels. Finally, we identified a set of co-expressed genes that can predict the survival of glioblastoma patients. Our study highlights that comparative analyses of single-cell and bulk gene expression profiles enable us to identify functional modules that are regulated at different levels and hold great translational potential. Jie Wang 0066, Shuli Xia, Brian Arand, Raghu Machiraju, Kun Huang 0001, Hongkai Ji |
PLoS Comput. Biol. | 6 |
| 2015 | GRAPHIE: graph based histology image explorerabstractBACKGROUND: Histology images comprise one of the important sources of knowledge for phenotyping studies in systems biology. However, the annotation and analyses of histological data have remained a manual, subjective and relatively low-throughput process. RESULTS: We introduce Graph based Histology Image Explorer (GRAPHIE)-a visual analytics tool to explore, annotate and discover potential relationships in histology image collections within a biologically relevant context. The design of GRAPHIE is guided by domain experts' requirements and well-known InfoVis mantras. By representing each image with informative features and then subsequently visualizing the image collection with a graph, GRAPHIE allows users to effectively explore the image collection. The features were designed to capture localized morphological properties in the given tissue specimen. More importantly, users can perform feature selection in an interactive way to improve the visualization of the image collection and the overall annotation process. Finally, the annotation allows for a better prospective examination of datasets as demonstrated in the users study. Thus, our design of GRAPHIE allows for the users to navigate and explore large collections of histology image datasets. CONCLUSIONS: We demonstrated the usefulness of our visual analytics approach through two case studies. Both of the cases showed efficient annotation and analysis of histology image collection. Chao Wang 0081, Kun Huang 0001, Raghu Machiraju |
BMC Bioinform. | 3 |
| 2015 | Identify Critical Genes in Development with Consistent H3K4me2 Patterns across Multiple TissuesabstractHistone modification is an important epigenetic event which plays essential roles in cell differentiation and tissue development. Recent studies show that a unique dimethylation of lysine 4 residue on histone 3 (H3K4me2) distribution pattern around transcription starting sites (TSS) of genes marks tissue specific genes in human CD4 þ T cells and mouse nervous tissue cells. However, existence of this pattern has not been widely tested and its implication remains unclear. In this paper, we study the H3K4me2 distribution patterns across six different cell lines from five major tissue types (including muscular tissue, nervous tissue, non-blood connective tissue, blood, and epithelial tissue) as well as embryonic stem cells. We define a metric ‘tail length’ to quantitatively describe H3K4me2 distribution patterns around the TSS. While confirming the previous observations, we also identified a group of 217 genes with ubiquitous long-tail H3K4me2 patterns in all the tested tissues and the embryonic stem cells (ESC). Further analyses confirmed that these genes are critical for development, and highly interactive with other tissue specific genes as evinced by protein-protein interaction networks, suggesting their critical regulatory functions. Our results suggest that rich information on gene functions and epigenetic events can be revealed using pattern recognition methods. Nan Meng, Raghu Machiraju, Kun Huang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2015 | Guest Editors Introduction to the Special Issue on Software and DatabasesabstractThe papers in this special section focus on software and databases that are central in bioinformatics and computational biology.. These programs are playing more and more important roles in biology and medical research. These papers cover a broad range of topics, including computational genomics and transcriptomics, analysis of biological networks and interactions, drug design, biomedical signal/image analysis, biomedical text mining and ontologies, biological data mining, visualization and integration, and high performance computing application in bioinformatics. Dong Xu 0002, Kun Huang 0001, Jeanette Schmidt |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | iGPSe: A visual analytic system for integrative genomic based cancer patient stratificationabstractBACKGROUND: Cancers are highly heterogeneous with different subtypes. These subtypes often possess different genetic variants, present different pathological phenotypes, and most importantly, show various clinical outcomes such as varied prognosis and response to treatment and likelihood for recurrence and metastasis. Recently, integrative genomics (or panomics) approaches are often adopted with the goal of combining multiple types of omics data to identify integrative biomarkers for stratification of patients into groups with different clinical outcomes. RESULTS: In this paper we present a visual analytic system called Interactive Genomics Patient Stratification explorer (iGPSe) which significantly reduces the computing burden for biomedical researchers in the process of exploring complicated integrative genomics data. Our system integrates unsupervised clustering with graph and parallel sets visualization and allows direct comparison of clinical outcomes via survival analysis. Using a breast cancer dataset obtained from the The Cancer Genome Atlas (TCGA) project, we are able to quickly explore different combinations of gene expression (mRNA) and microRNA features and identify potential combined markers for survival prediction. CONCLUSIONS: Visualization plays an important role in the process of stratifying given population patients. Visual tools allowed for the selection of possibly features across various datasets for the given patient population. We essentially made a case for visualization for a very important problem in translational informatics. Chao Wang 0081, Kun Huang 0001, Raghu Machiraju |
BMC Bioinform. | 3 |
| 2014 | Visualizing Multidimensional Data with Glyph SPLOMsabstractAbstract Scatterplot matrices or SPLOMs provide a feasible method of visualizing and representing multi‐dimensional data especially for a small number of dimensions. For very high dimensional data, we introduce a novel technique to summarize a SPLOM, as a clustered matrix of glyphs, or a Glyph SPLOM. Each glyph visually encodes a general measure of dependency strength, distance correlation, and a logical dependency class based on the occupancy of the scatterplot quadrants. We present the Glyph SPLOM as a general alternative to the traditional correlation based heatmap and the scatterplot matrix in two examples: demography data from the World Health Organization (WHO), and gene expression data from developmental biology. By using both, dependency class and strength, the Glyph SPLOM illustrates high dimensional data in more detail than a heatmap but with more summarization than a SPLOM. More importantly, the summarization capabilities of Glyph SPLOM allow for the assertion of “necessity” causal relationships in the data and the reconstruction of interaction networks in various dynamic systems. A. Yates, Amy Webb, Michael F. Sharpnack, H. Chamberlin, Kun Huang 0001, Raghu Machiraju |
Comput. Graph. Forum | 5 |
| 2013 | Research and applications: Identifying survival associated morphological features of triple negative breast cancer using multiple datasetsabstractBACKGROUND AND OBJECTIVE: Biomarkers for subtyping triple negative breast cancer (TNBC) are needed given the absence of responsive therapy and relatively poor prediction of survival. Morphology of cancer tissues is widely used in clinical practice for stratifying cancer patients, while genomic data are highly effective to classify cancer patients into subgroups. Thus integration of both morphological and genomic data is a promising approach in discovering new biomarkers for cancer outcome prediction. Here we propose a workflow for analyzing histopathological images and integrate them with genomic data for discovering biomarkers for TNBC. MATERIALS AND METHODS: We developed an image analysis workflow for extracting a large collection of morphological features and deployed the same on histological images from The Cancer Genome Atlas (TCGA) TNBC samples during the discovery phase (n=44). Strong correlations between salient morphological features and gene expression profiles from the same patients were identified. We then evaluated the same morphological features in predicting survival using a local TNBC cohort (n=143). We further tested the predictive power on patient prognosis of correlated gene clusters using two other public gene expression datasets. RESULTS AND CONCLUSION: Using TCGA data, we identified 48 pairs of significantly correlated morphological features and gene clusters; four morphological features were able to separate the local cohort with significantly different survival outcomes. Gene clusters correlated with these four morphological features further proved to be effective in predicting patient survival using multiple public gene expression datasets. These results suggest the efficacy of our workflow and demonstrate that integrative analysis holds promise for discovering biomarkers of complex diseases. Chao Wang 0081, Thierry Pécot, Debra L. Zynger, Raghu Machiraju, Charles L. Shapiro, Kun Huang 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2012 | Visualizing Clusters in Parallel Coordinates for Visual Knowledge Discovery
Yang Xiang 0007, David Fuhry, Ruoming Jin, Ye Zhao 0003, Kun Huang 0001 |
PAKDD (1) | 5 |
| 2012 | A signal processing approach for enriched region detection in RNA polymerase II ChIP-seq dataabstractBACKGROUND: RNA polymerase II (PolII) is essential in gene transcription and ChIP-seq experiments have been used to study PolII binding patterns over the entire genome. However, since PolII enriched regions in the genome can be very long, existing peak finding algorithms for ChIP-seq data are not adequate for identifying such long regions. METHODS: Here we propose an enriched region detection method for ChIP-seq data to identify long enriched regions by combining a signal denoising algorithm with a false discovery rate (FDR) approach. The binned ChIP-seq data for PolII are first processed using a non-local means (NL-means) algorithm for purposes of denoising. Then, a FDR approach is developed to determine the threshold for marking enriched regions in the binned histogram. RESULTS: We first test our method using a public PolII ChIP-seq dataset and compare our results with published results obtained using the published algorithm HPeak. Our results show a high consistency with the published results (80-100%). Then, we apply our proposed method on PolII ChIP-seq data generated in our own study on the effects of hormone on the breast cancer cell line MCF7. The results demonstrate that our method can effectively identify long enriched regions in ChIP-seq datasets. Specifically, pertaining to MCF7 control samples we identified 5,911 segments with length of at least 4 Kbp (maximum 233,000 bp); and in MCF7 treated with E2 samples, we identified 6,200 such segments (maximum 325,000 bp). CONCLUSIONS: We demonstrated the effectiveness of this method in studying binding patterns of PolII in cancer cells which enables further deep analysis in transcription regulation and epigenetics. Our method complements existing peak detection algorithms for ChIP-seq experiments. Zhi Han, Thierry Pécot, Tim Hui-Ming Huang, Raghu Machiraju, Kun Huang 0001 |
BMC Bioinform. | 6 |
| 2012 | Predicting glioblastoma prognosis networks using weighted gene co-expression network analysis on TCGA dataabstractBACKGROUND: Using gene co-expression analysis, researchers were able to predict clusters of genes with consistent functions that are relevant to cancer development and prognosis. We applied a weighted gene co-expression network (WGCN) analysis algorithm on glioblastoma multiforme (GBM) data obtained from the TCGA project and predicted a set of gene co-expression networks which are related to GBM prognosis. METHODS: We modified the Quasi-Clique Merger algorithm (QCM algorithm) into edge-covering Quasi-Clique Merger algorithm (eQCM) for mining weighted sub-network in WGCN. Each sub-network is considered a set of features to separate patients into two groups using K-means algorithm. Survival times of the two groups are compared using log-rank test and Kaplan-Meier curves. Simulations using random sets of genes are carried out to determine the thresholds for log-rank test p-values for network selection. Sub-networks with p-values less than their corresponding thresholds were further merged into clusters based on overlap ratios (>50%). The functions for each cluster are analyzed using gene ontology enrichment analysis. RESULTS: Using the eQCM algorithm, we identified 8,124 sub-networks in the WGCN, out of which 170 sub-networks show p-values less than their corresponding thresholds. They were then merged into 16 clusters. CONCLUSIONS: We identified 16 gene clusters associated with GBM prognosis using the eQCM algorithm. Our results not only confirmed previous findings including the importance of cell cycle and immune response in GBM, but also suggested important epigenetic events in GBM development and prognosis. Yang Xiang 0007, Cun-Quan Zhang, Kun Huang 0001 |
BMC Bioinform. | 3 |
| 2012 | k-Neighborhood decentralization: A comprehensive solution to index the UMLS for large scale knowledge discovery
Yang Xiang 0007, Kewei Lu, Stephen L. James, Tara Borlawsky, Kun Huang 0001, Philip R. O. Payne |
J. Biomed. Informatics | 5 |
| 2012 | Weighted Frequent Gene Co-expression Network Mining to Identify Genes Involved in Genome StabilityabstractGene co-expression network analysis is an effective method for predicting gene functions and disease biomarkers. However, few studies have systematically identified co-expressed genes involved in the molecular origin and development of various types of tumors. In this study, we used a network mining algorithm to identify tightly connected gene co-expression networks that are frequently present in microarray datasets from 33 types of cancer which were derived from 16 organs/tissues. We compared the results with networks found in multiple normal tissue types and discovered 18 tightly connected frequent networks in cancers, with highly enriched functions on cancer-related activities. Most networks identified also formed physically interacting networks. In contrast, only 6 networks were found in normal tissues, which were highly enriched for housekeeping functions. The largest cancer network contained many genes with genome stability maintenance functions. We tested 13 selected genes from this network for their involvement in genome maintenance using two cell-based assays. Among them, 10 were shown to be involved in either homology-directed DNA repair or centrosome duplication control including the well-known cancer marker MKI67. Our results suggest that the commonly recognized characteristics of cancers are supported by highly coordinated transcriptomic activities. This study also demonstrated that the co-expression network directed approach provides a powerful tool for understanding cancer physiology, predicting new gene functions, as well as providing new target candidates for cancer therapeutics. Jie Zhang 0010, Kewei Lu, Yang Xiang 0007, Muhtadi Islam, Shweta Kotian, Zeina Kais, Cindy Lee, Mansi Arora, Hui-wen Liu, Jeffrey D. Parvin, Kun Huang 0001 |
PLoS Comput. Biol. | 11 |
| 2012 | Transactional Database Transformation and Its Application in Prioritizing Human Disease GenesabstractBinary (0,1) matrices, commonly known as transactional databases, can represent many application data, including genephenotype data where “1” represents a confirmed gene-phenotype relation and “0” represents an unknown relation. It is natural to ask what information is hidden behind these “0”s and “1”s. Unfortunately, recent matrix completion methods, though very effective in many cases, are less likely to infer something interesting from these (0,1)-matrices. To answer this challenge, we propose INDEVI, a very succinct and effective algorithm to perform independent-evidence-based transactional database transformation. Each entry of a (0,1)-matrix is evaluated by “independent evidence” (maximal supporting patterns) extracted from the whole matrix for this entry. The value of an entry, regardless of its value as 0 or 1, has completely no effect for its independent evidence. The experiment on a genephenotype database shows that our method is highly promising in ranking candidate genes and predicting unknown disease genes. Yang Xiang 0007, Philip R. O. Payne, Kun Huang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2011 | Non-parametric Population Analysis of Cellular Phenotypes
Shantanu Singh, Firdaus Janoos, Thierry Pécot, Enrico Caserta, Kun Huang 0001, Jens Rittscher, Gustavo Leone, Raghu Machiraju |
MICCAI (2) | 5 |
| 2010 | Multi-dimensional discovery of biomarker and phenotype complexesabstractBACKGROUND: Given the rapid growth of translational research and personalized healthcare paradigms, the ability to relate and reason upon networks of bio-molecular and phenotypic variables at various levels of granularity in order to diagnose, stage and plan treatments for disease states is highly desirable. Numerous techniques exist that can be used to develop networks of co-expressed or otherwise related genes and clinical features. Such techniques can also be used to create formalized knowledge collections based upon the information incumbent to ontologies and domain literature. However, reports of integrative approaches that bridge such networks to create systems-level models of disease or wellness are notably lacking in the contemporary literature. RESULTS: In response to the preceding gap in knowledge and practice, we report upon a prototypical series of experiments that utilize multi-modal approaches to network induction. These experiments are intended to elicit meaningful and significant biomarker-phenotype complexes spanning multiple levels of granularity. This work has been performed in the experimental context of a large-scale clinical and basic science data repository maintained by the National Cancer Institute (NCI) funded Chronic Lymphocytic Leukemia Research Consortium. CONCLUSIONS: Our results indicate that it is computationally tractable to link orthogonal networks of genes, clinical features, and conceptual knowledge to create multi-dimensional models of interrelated biomarkers and phenotypes. Further, our results indicate that such systems-level models contain interrelated bio-molecular and clinical markers capable of supporting hypothesis discovery and testing. Based on such findings, we propose a conceptual model intended to inform the cross-linkage of the results of such methods. This model has as its aim the identification of novel and knowledge-anchored biomarker-phenotype complexes. Philip R. O. Payne, Kun Huang 0001, Kristin Keen-Circle, Abhisek Kundu, Jie Zhang 0010, Tara Borlawsky |
BMC Bioinform. | 2 |
| 2010 | Using gene co-expression network analysis to predict biomarkers for chronic lymphocytic leukemiaabstractBACKGROUND: Chronic lymphocytic leukemia (CLL) is the most common adult leukemia. It is a highly heterogeneous disease, and can be divided roughly into indolent and progressive stages based on classic clinical markers. Immunoglobin heavy chain variable region (IgVH) mutational status was found to be associated with patient survival outcome, and biomarkers linked to the IgVH status has been a focus in the CLL prognosis research field. However, biomarkers highly correlated with IgVH mutational status which can accurately predict the survival outcome are yet to be discovered. RESULTS: In this paper, we investigate the use of gene co-expression network analysis to identify potential biomarkers for CLL. Specifically we focused on the co-expression network involving ZAP70, a well characterized biomarker for CLL. We selected 23 microarray datasets corresponding to multiple types of cancer from the Gene Expression Omnibus (GEO) and used the frequent network mining algorithm CODENSE to identify highly connected gene co-expression networks spanning the entire genome, then evaluated the genes in the co-expression network in which ZAP70 is involved. We then applied a set of feature selection methods to further select genes which are capable of predicting IgVH mutation status from the ZAP70 co-expression network. CONCLUSIONS: We have identified a set of genes that are potential CLL prognostic biomarkers IL2RB, CD8A, CD247, LAG3 and KLRK1, which can predict CLL patient IgVH mutational status with high accuracies. Their prognostic capabilities were cross-validated by applying these biomarker candidates to classify patients into different outcome groups using a CLL microarray datasets with clinical information. Jie Zhang 0010, Yang Xiang 0007, Liya Ding 0001, Kristin Keen-Circle, Tara Borlawsky, Hatice Gulcin Ozer, Ruoming Jin, Philip R. O. Payne, Kun Huang 0001 |
BMC Bioinform. | 9 |
| 2009 | Enabling Data Analysis on High-Throughput Data in Large Data Depository Using Web-Based Analysis Platform - A Case Study on Integrating QUEST with GenePattern in Epigenetics ResearchabstractEnabling data analysis in large data depositories for high throughput experimental data such as gene microarrays and ChIP-seq is challenging. In this paper, we discuss three methods for integrating QUEST, a data depository for epigenetic experiments, with a web-based data analysis platform GenePattern. These methods are universal and can serve as an exemplary implementation resolving the dilemma facing many similar database systems in integrating data analysis tools. Terry Camerlengo, Hatice Gulcin Ozer, Pearlly Yan, Jeffrey D. Parvin, Tim Hui-Ming Huang, Kun Huang 0001, Mingxiang Teng, Lang Li 0001, Francisco Perez, Tahsin M. Kurç |
BIBM | 6 |
| 2009 | Clinical Attribute Network for Chronic Lymphocytic LeukemiaabstractIn this paper, we present a study on the relationships between Chronic Lymphocytic Leukemia (CLL) related clinical attributes using network visualization and analysis techniques. This work is the first step of our long-term project to identify novel biomarkers for CLL, working in coordination with the NCI-funded CLL Research Consortium (cll.ucsd.edu). By computing Spearman correlation coefficients for 125 clinical attributes, we established an attribute network for CLL. Using network visualization techniques we identified a core network with peripheral nodes around it. An important observation is that many cytogenetic attributes are in the core network, indicating the connection between karyotypes (e.g., abnormality in Chromosome 17) and other clinical attributes. Since these karyotypes are linked to CLL etiology in many ways, the corresponding clinical attributes may be further selected and tested as markers for disease screening. The correlation analysis results are also consistent with previous studies using a knowledge engineering based approach for establishing relationships between such attributes. Abhisek Kundu, Hatice Gulcin Ozer, Tara Borlawsky, Kristin C. Circle, Kun Huang 0001, Philip R. O. Payne |
BIBM | 5 |
| 2009 | Comparative study on ChIP-seq data: normalization and binding pattern characterizationabstractMOTIVATION: Antibody-based Chromatin Immunoprecipitation assay followed by high-throughput sequencing technology (ChIP-seq) is a relatively new method to study the binding patterns of specific protein molecules over the entire genome. ChIP-seq technology allows scientist to get more comprehensive results in shorter time. Here, we present a non-linear normalization algorithm and a mixture modeling method for comparing ChIP-seq data from multiple samples and characterizing genes based on their RNA polymerase II (Pol II) binding patterns. RESULTS: We apply a two-step non-linear normalization method based on locally weighted regression (LOESS) approach to compare ChIP-seq data across multiple samples and model the difference using an Exponential-Normal(K) mixture model. Fitted model is used to identify genes associated with differential binding sites based on local false discovery rate (fdr). These genes are then standardized and hierarchically clustered to characterize their Pol II binding patterns. As a case study, we apply the analysis procedure comparing normal breast cancer (MCF7) to tamoxifen-resistant (OHT) cell line. We find enriched regions that are associated with cancer (P < 0.0001). Our findings also imply that there may be a dysregulation of cell cycle and gene expression control pathways in the tamoxifen-resistant cells. These results show that the non-linear normalization method can be used to analyze ChIP-seq data across multiple samples. AVAILABILITY: Data are available at http://www.bmi.osu.edu/~khuang/Data/ChIP/RNAPII/. Cenny Taslim, Jiejun Wu, Pearlly Yan, Greg Singer, Jeffrey D. Parvin, Tim Hui-Ming Huang, Shili Lin, Kun Huang 0001 |
Bioinform. | 8 |
| 2009 | Robust 3D reconstruction and identification of dendritic spines from optical microscopy imaging
Firdaus Janoos, Kishore Mosaliganti, Xiaoyin Xu, Raghu Machiraju, Kun Huang 0001, Stephen T. C. Wong |
Medical Image Anal. | 5 |
| 2009 | Tensor classification of N-point correlation function features for histology tissue segmentation
Kishore Mosaliganti, Firdaus Janoos, M. Okan Irfanoglu, Randall Ridgway, Raghu Machiraju, Kun Huang 0001, Joel H. Saltz, Gustavo Leone, Michael C. Ostrowski |
Medical Image Anal. | 6 |
| 2008 | Geometry-driven Visualization of Microscopic Structures in BiologyabstractAbstract At a microscopic resolution, biological structures are composed of cells, red blood corpuscles (RBCs), cytoplasm and other microstructural components. There is a natural pattern in terms of distribution, arrangement and packing density of these components in biological organization. In this work, we propose to use N‐point correlation functions to guide the analysis and exploration process in microscopic datasets. These functions provide useful feature spaces to aid segmentation and visualization tasks. We show3D visualizations of mouse placenta tissue layers and mouse mammary ducts as well as2D segmentation/tracking of clonal populations. Further confidence in our results stems from validation studies that were performed with manual ground‐truth for segmentation. Kishore Mosaliganti, Raghu Machiraju, Kun Huang 0001, Gustavo Leone |
Comput. Graph. Forum | 3 |
| 2008 | An imaging workflow for characterizing phenotypical change in large histological mouse model datasets
Kishore Mosaliganti, Tony Pan, Randall Ridgway, Richard Sharp, Lee A. D. Cooper, Alexandra Gulacy, Ashish Sharma 0001, M. Okan Irfanoglu, Raghu Machiraju, Tahsin M. Kurç, Alain de Bruin, Pamela Wenzel, Gustavo Leone, Joel H. Saltz, Kun Huang 0001 |
J. Biomed. Informatics | 15 |
| 2008 | Reconstruction of Cellular Biological Structures from Optical Microscopy DataabstractDevelopments in optical microscopy imaging have generated large high-resolution data sets that have spurred medical researchers to conduct investigations into mechanisms of disease, including cancer at cellular and subcellular levels. The work reported here demonstrates that a suitable methodology can be conceived that isolates modality-dependent effects from the larger segmentation task and that 3D reconstructions can be cognizant of shapes as evident in the available 2D planar images. In the current realization, a method based on active geodesic contours is first deployed to counter the ambiguity that exists in separating overlapping cells on the image plane. Later, another segmentation effort based on a variant of Voronoi tessellations improves the delineation of the cell boundaries using a Bayesian formulation. In the next stage, the cells are interpolated across the third dimension thereby mitigating the poor structural correlation that exists in that dimension. We deploy our methods on three separate data sets obtained from light, confocal, and phase-contrast microscopy and validate the results appropriately. Kishore Mosaliganti, Lee A. D. Cooper, Richard Sharp, Raghu Machiraju, Gustavo Leone, Kun Huang 0001, Joel H. Saltz |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2007 | The GPU on biomedical image processing for color and phenotype analysisabstractThe computational power and memory bandwidth of graphics processing units (GPUs) have turned them into attractive platforms for general-purpose applications. In this paper, we exploit this power in the context of biomedical image processing by establishing a cooperative environment between the CPU and the GPU. We deal with phenotype and color analysis on a wide variety of microscopic images from studies of cartilage and bone tissue regeneration using stem cells and genetics involving cancer pathology. Both processors are used in parallel to map algorithms for computing color histograms, contour detection using the Canny filter and pattern recognition based on the Hough transform. Task, data and instruction parallelism are exploited in the GPU to accomplish performance gains between 4x and 100x more than the typical CPU code. Antonio Ruiz 0001, Manuel Ujaldon, Jose Antonio Andrades, Jose Becerra, Kun Huang 0001, Tony Pan, Joel H. Saltz |
BIBE | 5 |
| 2007 | Detection and Visualization of Surface-Pockets to Enable Phenotyping StudiesabstractIn this paper, we propose a technique for detecting pockets on a surface-of-interest. A sequence of propagating fronts converging to the target surface is used as the basis for inspection. We compute a correspondence function between the initial and the target surface. This leads to a natural definition of the local feature size measured as the evolution distance between mapped points. Surface pockets are then extracted as salient clusters embedded in the feature space. The level-set initialization also determines the scale-space of the extracted pockets. Results are presented on a case-study in which the focus is to chronicle the phenotyping differences in genetically modified mouse placenta. Our results are validated based on manually verified ground-truth. Kishore Mosaliganti, Firdaus Janoos, Richard Sharp, Randall Ridgway, Raghu Machiraju, Kun Huang 0001, Pamela Wenzel, Alain de Bruin, Gustavo Leone, Joel H. Saltz |
IEEE Trans. Medical Imaging | 6 |
| 2006 | Multiscale Hybrid Linear Models for Lossy Image RepresentationabstractIn this paper, we introduce a simple and efficient representation for natural images. We view an image (in either the spatial domain or the wavelet domain) as a collection of vectors in a high-dimensional space. We then fit a piece-wise linear model (i.e., a union of affine subspaces) to the vectors at each downsampling scale. We call this a multiscale hybrid linear model for the image. The model can be effectively estimated via a new algebraic method known as generalized principal component analysis (GPCA). The hybrid and hierarchical structure of this model allows us to effectively extract and exploit multimodal correlations among the imagery data at different scales. It conceptually and computationally remedies limitations of many existing image representation methods that are based on either a fixed linear transformation (e.g., DCT, wavelets), or an adaptive uni-modal linear transformation (e.g., PCA), or a multimodal model that uses only cluster means (e.g., VQ). We will justify both quantitatively and experimentally why and how such a simple multiscale hybrid model is able to reduce simultaneously the model complexity and computational cost. Despite a small overhead of the model, our careful and extensive experimental results show that this new model gives more compact representations for a wide variety of natural images under a wide range of signal-to-noise ratios than many existing methods, including wavelets. We also briefly address how the same (hybrid linear) modeling paradigm can be extended to be potentially useful for other applications, such as image segmentation. Wei Hong 0003, John Wright 0001, Kun Huang 0001, Yi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2005 | A Multi-Scale Hybrid Linear Model for Lossy Image RepresentationabstractThis paper introduces a simple and efficient representation for natural images. We partition an image into blocks and treat the blocks as vectors in a high-dimensional space. We then fit a piecewise linear model (i.e. a union of affine subspaces) to the vectors at each down-sampling scale. We call this a multiscale hybrid linear model of the image. The hybrid and hierarchical structure of this model allows us effectively to extract and exploit multimodal correlations among the imagery data at different scales. It conceptually and computationally remedies limitations of many existing image representation methods that are based on either a fixed linear transformation (e.g. DCT, wavelets), an adaptive unimodal linear transformation (e.g. PCA), or a multi-modal model at a single scale. We will justify both analytically and experimentally why and how such a simple multiscale hybrid model is able to reduce simultaneously the model complexity and computational cost. Despite a small overhead for the model, our results show that this new model gives more compact representations for a wide variety of natural images under a wide range of signal-to-noise ratio than many existing methods, including wavelets. Wei Hong 0003, John Wright 0001, Kun Huang 0001, Yi Ma 0001 |
ICCV | 3 |
| 2005 | Symmetry-based 3-D reconstruction from perspective images
Allen Y. Yang, Kun Huang 0001, Shankar R. Rao, Wei Hong 0003, Yi Ma 0001 |
Comput. Vis. Image Underst. | 2 |
| 2005 | Symmetry-based photo-editing
Kun Huang 0001, Wei Hong 0003, Yi Ma 0001 |
Pattern Recognit. | 1 |
| 2004 | Minimum Effective Dimension for Mixtures of Subspaces: A Robust GPCA Algorithm and Its Applications
Kun Huang 0001, Yi Ma 0001, René Vidal |
CVPR (2) | 1 |
| 2004 | Sparse representation of images with hybrid linear modelsabstractWe propose a mixture of multiple linear models, also known as hybrid linear model, for a sparse representation of an image. This is a generalization of the conventional Karhunen-Loeve transform (KLT) or principal component analysis (PCA). We provide an algebraic algorithm based on generalized principal component analysis (GPCA) that gives a global and noniterative solution to the identification of a hybrid linear model for any given image. We demonstrate the efficiency of the proposed hybrid linear model by experiments and comparison with other transforms such as the KLT, DCT and wavelet transforms. Such an efficient representation can be very useful for later stages of image processing, especially in applications such as image segmentation and image compression. Kun Huang 0001, Allen Y. Yang, Yi Ma 0001 |
ICIP | 1 |
| 2004 | Large-baseline Matching and Reconstruction from Symmetry CellsabstractIn this paper, we study how the presence of symmetry in man-made environments may significantly facilitate the task of automatic matching features and recovering 3-D camera pose and scene structure from multiple perspective images. While conventional methods typically rely on small-motion tracking or robust statistic techniques to resolve the coupling between feature matching and 3-D recovery, we here propose a new symmetry-based approach which allows automatic feature matching between images taken with arbitrary (both large and small) camera motions. To this end, we develop the multiple-view geometry of symmetry cells. To resolve possible ambiguities that may arise in matching symmetry cells and camera pose recovery, we find a consistent solution by finding the maximal complete subgraph of a matching graph; we also use a topological check to avoid mismatches. As our experiments shows, the resulting algorithms are simple, accurate and easy to implement. Kun Huang 0001, Allen Y. Yang, Wei Hong 0003, Yi Ma 0001 |
ICRA | 1 |
| 2004 | On Symmetry and Multiple-View Geometry: Structure, Pose, and Calibration from a Single Image
Wei Hong 0003, Allen Y. Yang, Kun Huang 0001, Yi Ma 0001 |
Int. J. Comput. Vis. | 3 |
| 2004 | Rank Conditions on the Multiple-View Matrix
Yi Ma 0001, Kun Huang 0001, René Vidal, Jana Kosecka, S. Shankar Sastry |
Int. J. Comput. Vis. | 2 |
| 2003 | Geometric Segmentation of Perspective Images Based on Symmetry GroupsabstractSymmetry is an effective geometric cue to facilitate conventional segmentation techniques on images of man-made environment. Based on three fundamental principles that summarize the relations between symmetry and perspective imaging, namely, structure from symmetry, symmetry hypothesis testing, and global symmetry testing, we develop a prototype system which is able to automatically segment symmetric objects in space from single 2D perspective images. The result of such a segmentation is a hierarchy of geometric primitives, called symmetry cells and complexes, whose 3D structure and pose are fully recovered. Such a geometrically meaningful segmentation may greatly facilitate applications such as feature matching and robot navigation. Allen Y. Yang, Shankar R. Rao, Kun Huang 0001, Wei Hong 0003, Yi Ma 0001 |
ICCV | 3 |
| 2002 | Generalized Rank Conditions in Multiple View Geometry with Applications to Dynamical Scenes
Kun Huang 0001, Robert M. Fossum, Yi Ma 0001 |
ECCV (2) | 1 |