EDBT 2026 Demo / reviewers in the wild / expert
Guanglan Zhang
dblp:16/2037
· DBLP profile ↗
23ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0001-6010-490XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PREDBL6: a system for predicting C57BL/6 mouse T-cell epitopesabstractThe MHC class I antigen processing pathway plays a critical role in the adaptive immune system by presenting peptides for recognition by CD8+ T cells. While most prediction tools focus on MHC binding, accurately identifying immunogenic T-cell epitopes requires accounting for additional factors such as antigen processing and peptide-MHC binding thermostability. We developed a bioinformatics tool that integrates MHC binding predictions with thermostability assessment and antigen processing steps to enhance T cell epitope identification accuracy for C57BL/6 mice. Our machine learning models, trained on a comprehensive dataset of eluted H2-Kband H2-Dbligands and thermostability data across a range of physiologically relevant temperatures (37°C, 50°C, 70°C), were rigorously validated. These models showed improved overall accuracy on an external validation dataset compared to the widely used NetMHCPan-4.1. We consolidated the models into a user-friendly web-based application named PREDBL6 to facilitate accurate predictions of immunogenic peptides that stably bind H2b molecules and stimulate immune responses in C57BL/6 mice. PREDBL6 is accessible at http://met-hilab.org:3001/. Zitian Zhen, Guancheng Huang, Lubomir T. Chitkushev, Vladimir Brusic, Derin B. Keskin, Guanglan Zhang |
BIBM | 7 |
| 2021 | Correctness of Cell Labels in Public Single Cell Transcriptomics DatasetsabstractThe number of single-cell transcriptomic (SCT) studies is rapidly increasing. More than 15000 single cell gene expression data sets are available in public repositories. More than 2400 of these sets involve Peripheral Blood Mononuclear Cells (PBMC) data sets. Main cell types of PBMC are B cells, dendritic cells, monocytes, natural killer cells, and T cells. Labels of individual PBMC are usually provided in metadata accompanying the data sets or are implicit as data set partitions for sorted cells. We analyzed the correctness of labels assigned to individual cells from PBMC in primary reports. The correctness of primary labels was assessed by using Artificial Neural Network (ANN) classifier and Confident Learning (CL) approach. We assessed that the number of mislabels on average in our data sets is about2%. The label accuracy varied broadly between data sets, particularly among those generated by experimental cell sorting followed by SCT. Minjie Lyu, Yihan Zhang 0003, Derin B. Keskin, Lubomir T. Chitkushev, Guanglan Zhang, Vladimir Brusic |
BIBM | 6 |
| 2021 | Examining Mental Illness Trends in the United States From 2006 to 2019abstractWe investigate the characteristics of medical expenditures associated with mental illness hospitalizations using the Truven Health MarketScan Database. We focus on the inpatient admissions due to mental illness of adults aged 1S to 64 between 2006 to 2019. We aim to answer the following questions: (1) Did the financial crisis of 2008 impact mental health in the U.S.?(2) What are the other macro-level (socioeconomic and regulartory) and micro-level (individualpatient related) factors that affect the cost of inpatient care due to mental illness; (3) Did mental illness affect men and women differently? (4) How were different regions within the U.S. affected by mental illness? Thomas Olson, Irena Vodenska, Guanglan Zhang, Marislei Nishijima, Lubomir T. Chitkushev |
BIBM | 3 |
| 2021 | Classification of Single Cell Types using Small Sets of Expressed Genes: Comparative Analysis of Supervised Machine Learning MethodsabstractSingle cell transcriptomics measures gene expression data of large number of genes, concurrently, from tens of thousands of cells present in a studied biological sample. It is difficult to obtain good classification results due to high data dimensionality and variability of biological states. We performed a preliminary study to assess the feasibility of using supervised machine learning methods to classify peripheral blood mononuclear cell (PBMC) types from single cell gene expression data. We analyzed a large PBMC data set $(\sim 120,000$ PBMC cells), selected 47 genes (from 30698 features) suitable as SML classification features, and performed classification using 20 machine learning algorithms. Data sets represented three sample processing strategies: PBMC separation (two data sets), and experimental cell sorting by (two data sets). The accuracy in 5-class classification among 20 methods was 91-97% (PBMC separation), 97-100% (magnetic-activated cell sorting), and 82-99% (fluorescence-activated cell sorting). Our results indicate the feasibility of supervised machine learning for classification of cells into major PBMC cell types using a small number of classification features from single cell gene expression data. Aleksandar Veljkovic, Mirjana M. Maljkovic, Nenad S. Mitic, Sasa N. Malkov, Minjie Lyu, Marek T. Michalewicz, Guanglan Zhang, Vladimir Brusic |
BIBM | 8 |
| 2021 | Applications of single cell profiles of PBMC: Improvements of cell type classificationabstractSingle-cell-derived-class (SCDC) profiles capture characteristic gene expression from single cells representing types and subtypes and their conditions. SCDC profiles show high reproducibility across similar single-cell types processed under the same conditions. We have demonstrated two applications of SCDC profiles-classification of single cells from PBMC into six main classes (B cells, cDC, pDC, monocytes, NK cells, and T cells) and into three super-classes (BC+pDC,MC+cDC, and TC+NK). The minimum number of individual cells required for building an effective reference SCDC profile has been assessed to be between 160 and 640 cells. The variability of SCDC gradually decreases as the number of cells used to derive the profile increases from 10-cells to 640-cells. The classification accuracy of PBMC extracted by PBMC separation by SCDC profiles was 85-100% and 95-100% for supertypes depending on the cell type or supertype. Classification accuracy for PBMC cell types is lower for samples that are processed by cell sorting, or other sample processing steps. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Richard Rankin, Vladimir Brusic |
BIBM | 5 |
| 2021 | Blockchain Technology in Healthcare: A Scientific and Technological Driving ForceabstractBlockchain is a technology to enable decentralized collaboration among un-trusted entities. Academia and industry are rushing to uncover its potential for their field of interest. Due to the novelty of the technology and its diverse applications, there are some ambiguities in approaches and trends. In this paper, we analyze the scientific publications and patents from the past five years to identify the trends for blockchain integration with healthcare. For this purpose, we have adopted a quantitative (clustering) and qualitative (theme extraction) approach to discover themes and temporal dynamics in academia and industry. Our results shed light on the potential challenges and vision for future works. Irena Vodenska, Lubomir T. Chitkushev, Guanglan Zhang, Shahin Gheitanchi, Reza Rawassizadeh |
CBMS | 5 |
| 2021 | TANTIGEN 2.0: a knowledge base of tumor T cell antigens and epitopesabstractWe previously developed TANTIGEN, a comprehensive online database cataloging more than 1000 T cell epitopes and HLA ligands from 292 tumor antigens. In TANTIGEN 2.0, we significantly expanded coverage in both immune response targets (T cell epitopes and HLA ligands) and tumor antigens. It catalogs 4,296 antigen variants from 403 unique tumor antigens and more than 1500 T cell epitopes and HLA ligands. We also included neoantigens, a class of tumor antigens generated through mutations resulting in new amino acid sequences in tumor antigens. TANTIGEN 2.0 contains validated TCR sequences specific for cognate T cell epitopes and tumor antigen gene/mRNA/protein expression information in major human cancers extracted by Human Pathology Atlas. TANTIGEN 2.0 is a rich data resource for tumor antigens and their associated epitopes and neoepitopes. It hosts a set of tailored data analytics tools tightly integrated with the data to form meaningful analysis workflows. It is freely available at http://projects.met-hilab.org/tadb . Guanglan Zhang, Lubomir T. Chitkushev, Lars Rønn Olsen, Derin B. Keskin, Vladimir Brusic |
BMC Bioinform. | 1 |
| 2020 | Artificial Neural Network System for Cell Classification using Single Cell RNA ExpressionabstractWe implemented an automated system for single-cell classification using artificial neural networks (ANN). Our system takes single-cell gene expression sparse matrices and trains ANN to classify cell types and subtypes. The assemblies of ANNs predict cell classes by voting. We tested the system in a case study where we trained ANNs with a dataset containing approximately 120,000 single cells and tested the resulting model using an independent data set of 13,000 single cells. The overall accuracy of the 5-class classification was 95%. We trained and tested a total of 100 ANNs in 10 cycles. The prediction system demonstrated excellent reproducibility. The analysis of misclassifications indicated that 2% were likely classification errors, while the remaining 3% were likely due to mislabeled types and subtypes in the test set. Jiahui Zhong, Minjie Lyu, Derin B. Keskin, Guanglan Zhang, Vladimir Brusic, Lubomir T. Chitkushev |
BIBM | 6 |
| 2020 | A Review of Telemedicine in time of COVID-19abstractTelemedicine plays an increasingly important role in global healthcare. In this study, we summarized the latest developments related to telemedicine and discussed the obstacles and challenges to its wide adoption with a focus on the impact of COVID-19. Zhidong Wu, Lubomir T. Chitkushev, Guanglan Zhang |
BIBM | 3 |
| 2020 | Single-cell mRNA Profiles in PBMCabstractWe developed a method for building gene expression profiles from single-cell gene expression matrices. We named these profiles the “single-cell-derived-class” or SCDC profiles. They represent characteristic patterns of gene expressions of the types and subtypes of cells derived from single-cell transcriptome experiments. We deployed this method on classes and subclasses of peripheral blood mononuclear cells (PBMC). We used 47 human single-cell transcriptomics (SCT) data sets representing various classes, subclasses, and sample processing conditions. From comparisons of these profiles we found that they are highly reproducible, even when derived from unrelated studies as long as the processing steps are identical. The most similar profiles are those that are minimally processed. Cell sorting using FACS, cell enrichment, or fixing in methanol make profiles distinct from those derived from normal healthy samples. Our results suggest that approximately 50-200 cells are sufficient for building a useful SCDC profile. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 5 |
| 2020 | Prediction of PBMC Cell Types Using scRNAseq Reference ProfilesabstractSingle cell transcriptomics enables a high-resolution concurrent measurement of gene expression from tens of thousands of cells. We developed a method for determining standardized profiles from SCT data. We defined 48 data sets from 13 different studies and developed single-cellderived-class” (SCDC) profiles representing multiple classes and subclasses of peripheral blood mononuclear cells (PBMC). We applied pattern recognition analysis by calculating the distance from each query cell to the SCDC profiles (excluding the profiles of the query cells). Classification of cells by pattern recognition showed excellent performance for PBMC that were isolated, but not further processed by cell sorting. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 5 |
| 2020 | Classification of PBMC cell types using scRNAseq, ANN, and incremental learningabstractSingle cell transcriptomics (SCT) technology reveals gene expression of individual cells. Peripheral blood mononuclear cells (PBMC) are important diagnostic targets in immunology. In this study, we obtained and standardized 27 SCT data sets, derived from healthy PBMC samples using 10x SCT. We used artificial neural networks (ANN) to assess the ability of ANN to classify main PBMC cell types. Incremental learning by the gradual addition of new data sets to ANN training improved classification. The overall prediction accuracy of the final step of incremental learning reached 93% in 4-class classification. Jiahui Zhong, Razin A. Shaikh, Haoguo Wu, Lubomir T. Chitkushev, Guanglan Zhang, Derin B. Keskin, Vladimir Brusic |
BIBM | 7 |
| 2019 | Classification of Five Cell Types from PBMC Samples using Single Cell Transcriptomics and Artificial Neural NetworksabstractWe used 27 human single cell transcriptomics (SCT) data sets to develop an artificial neural network (ANN) model for classification of Peripheral Blood Mononuclear Cells (PBMC). We demonstrated that highly accurate models for the classification of PBMC subtypes can be developed by combining multiple independent data sets to form training data sets. A significant data preparation effort was needed for building predictive models. Using a data set of ~120,000 single cell instances we showed the accuracy of classification of PBMC call of ~ 90%. Optimization techniques and the addition of new high-quality data sets for model training are expected to improve PBMC subtype classification accuracy. Razin A. Shaikh, Jiahui Zhong, Minjie Lyu, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 6 |
| 2019 | TANTIGEN 2.0: an online database and analysis platform for tumor T cell antigensabstractWe previously developed TANTIGEN, a comprehensive web-based database cataloging more than 1,000 T cell epitopes and HLA ligands from 292 tumor antigens. TANTIGEN 2.0 is significantly expanded the number and coverage of immune response targets (T cell epitopes and HLA ligands) of previously cataloged tumor antigens. We expanded the number of cataloged tumor antigens to more than 4,000 and have added their reported targets of immune responses. We also included neoantigens, a new class of tumor antigens generated through mutations that results in a new amino acid sequence in tumor antigens. TANTIGEN 2.0 contains validated TCR sequences specific for cognate T cell epitopes. Gene expression information was extracted from tumor antigen gene/mRNA/protein expression information in major human cancers provided by Human Pathology Atlas. TANTIGEN 2.0 provides a rich data resource for tumor-associated epitope and neoepitope discovery studies. It is freely available at http://projects.met-hilab.org/tadb. Guanglan Zhang, Lubomir T. Chitkushev, Derin B. Keskin, Vladimir Brusic |
BIBM | 1 |
| 2018 | Computational modeling approach to suggest possible therapeutic interventions in spinal muscular atrophy
Giulia Russo, Guanglan Zhang |
BIBM | 2 |
| 2017 | MCVdb: A database for knowledge discovery in Merkel cell polyomavirus with applications in T cell immunology and vaccinologyabstractMerkel Cell Polyomavirus (MCV) is associated with more than 80% of Merkel cell carcinoma (MCC), a rare but highly lethal form of skin cancer. We made use of the immunological data on MCV available through publications and databases and constructed MCV T cell Antigen Database (MCVdb). MCVdb contains 734 curated antigen entries of MCV antigenic proteins and 30 experimentally verified T cell epitopes. The data were subject to extensive quality control (redundancy elimination, error detection, and vocabulary consolidation). A set of computational tools for in-depth analysis, such as sequence comparison using BLAST search, multiple alignments of antigens, and T cell epitope conservation analysis have been integrated within the MCVdb. Predicted Class I and Class II HLA-binding peptides for 15 common HLA alleles are included in this database as putative targets. MCVdb is a unique data source providing a comprehensive list of MCV antigens and peptides. MCVdb is publicly available at http://projects.met-hilab.org/mcv/. Guanglan Zhang, Derin B. Keskin, James A. DeCaprio, Catherine J. Wu, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 1 |
| 2008 | Evaluation of MHC-II peptide binding prediction servers: applications for vaccine researchabstractBACKGROUND: Initiation and regulation of immune responses in humans involves recognition of peptides presented by human leukocyte antigen class II (HLA-II) molecules. These peptides (HLA-II T-cell epitopes) are increasingly important as research targets for the development of vaccines and immunotherapies. HLA-II peptide binding studies involve multiple overlapping peptides spanning individual antigens, as well as complete viral proteomes. Antigen variation in pathogens and tumor antigens, and extensive polymorphism of HLA molecules increase the number of targets for screening studies. Experimental screening methods are expensive and time consuming and reagents are not readily available for many of the HLA class II molecules. Computational prediction methods complement experimental studies, minimize the number of validation experiments, and significantly speed up the epitope mapping process. We collected test data from four independent studies that involved 721 peptide binding assays. Full overlapping studies of four antigens identified binding affinity of 103 peptides to seven common HLA-DR molecules (DRB1*0101, 0301, 0401, 0701, 1101, 1301, and 1501). We used these data to analyze performance of 21 HLA-II binding prediction servers accessible through the WWW. RESULTS: Because not all servers have predictors for all tested HLA-II molecules, we assessed a total of 113 predictors. The length of test peptides ranged from 15 to 19 amino acids. We tried three prediction strategies - the best 9-mer within the longer peptide, the average of best three 9-mer predictions, and the average of all 9-mer predictions within the longer peptide. The best strategy was the identification of a single best 9-mer within the longer peptide. Overall, measured by the receiver operating characteristic method (AROC), 17 predictors showed good (AROC > 0.8), 41 showed marginal (AROC > 0.7), and 55 showed poor performance (AROC < 0.7). Good performance predictors included HLA-DRB1*0101 (seven), 1101 (six), 0401 (three), and 0701 (one). The best individual predictor was NETMHCIIPAN, closely followed by PROPRED, IEDB (Consensus), and MULTIPRED (SVM). None of the individual predictors was shown to be suitable for prediction of promiscuous peptides. Current predictive capabilities allow prediction of only 50% of actual T-cell epitopes using practical thresholds. CONCLUSION: The available HLA-II servers do not match prediction capabilities of HLA-I predictors. Currently available HLA-II prediction servers offer only a limited prediction accuracy and the development of improved predictors is needed for large-scale studies, such as proteome-wide epitope mapping. The requirements for accuracy of HLA-II binding predictions are stringent because of the substantial effect of false positives. Honghuang Lin, Guanglan Zhang, Songsak Tongchusak, Ellis L. Reinherz, Vladimir Brusic |
BMC Bioinform. | 2 |
| 2008 | Hotspot Hunter: a computational system for large-scale screening and selection of candidate immunological hotspots in pathogen proteomesabstractBACKGROUND: T-cell epitopes that promiscuously bind to multiple alleles of a human leukocyte antigen (HLA) supertype are prime targets for development of vaccines and immunotherapies because they are relevant to a large proportion of the human population. The presence of clusters of promiscuous T-cell epitopes, immunological hotspots, has been observed in several antigens. These clusters may be exploited to facilitate the development of epitope-based vaccines by selecting a small number of hotspots that can elicit all of the required T-cell activation functions. Given the large size of pathogen proteomes, including of variant strains, computational tools are necessary for automated screening and selection of immunological hotspots. RESULTS: Hotspot Hunter is a web-based computational system for large-scale screening and selection of candidate immunological hotspots in pathogen proteomes through analysis of antigenic diversity. It allows screening and selection of hotspots specific to four common HLA supertypes, namely HLA class I A2, A3, B7 and class II DR. The system uses Artificial Neural Network and Support Vector Machine methods as predictive engines. Soft computing principles were employed to integrate the prediction results produced by both methods for robust prediction performance. Experimental validation of the predictions showed that Hotspot Hunter can successfully identify majority of the real hotspots. Users can predict hotspots from a single protein sequence, or from a set of aligned protein sequences representing pathogen proteome. The latter feature provides a global view of the localizations of the hotspots in the proteome set, enabling analysis of antigenic diversity and shift of hotspots across protein variants. The system also allows the integration of prediction results of the four supertypes for identification of hotspots common across multiple supertypes. The target selection feature of the system shortlists candidate peptide hotspots for the formulation of an epitope-based vaccine that could be effective against multiple variants of the pathogen and applicable to a large proportion of the human population. CONCLUSION: Hotspot Hunter is publicly accessible at http://antigen.i2r.a-star.edu.sg/hh/. It is a new generation computational tool aiding in epitope-based vaccine design. Guanglan Zhang, Asif M. Khan, Kellathur N. Srinivasan, A. T. Heiny, Kenneth X. Lee, Chee Keong Kwoh 0001, J. Thomas August, Vladimir Brusic |
BMC Bioinform. | 1 |
| 2007 | AllerTool: a web server for predicting allergenicity and allergic cross-reactivity in proteinsabstractUNLABELLED: Assessment of potential allergenicity and patterns of cross-reactivity is necessary whenever novel proteins are introduced into human food chain. Current bioinformatic methods in allergology focus mainly on the prediction of allergenic proteins, with no information on cross-reactivity patterns among known allergens. In this study, we present AllerTool, a web server with essential tools for the assessment of predicted as well as published cross-reactivity patterns of allergens. The analysis tools include graphical representation of allergen cross-reactivity information; a local sequence comparison tool that displays information of known cross-reactive allergens; a sequence similarity search tool for assessment of cross-reactivity in accordance to FAO/WHO Codex alimentarius guidelines; and a method based on support vector machine (SVM). A 10-fold cross-validation results showed that the area under the receiver operating curve (A(ROC)) of SVM models is 0.90 with 86.00% sensitivity (SE) at specificity (SP) of 86.00%. AVAILABILITY: AllerTool is freely available at http://research.i2r.a-star.edu.sg/AllerTool/. Zong Hong Zhang, Judice L. Y. Koh, Guanglan Zhang, Khar Heng Choo, Martti T. Tammi, Joo Chuan Tong |
Bioinform. | 3 |
| 2006 | Extreme Learning Machine for Predicting HLA-Peptide Binding
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong, Guanglan Zhang, Vladimir Brusic |
ISNN (2) | 4 |
| 2006 | Prediction of HLA-DQ3.2ß Ligands: evidence of multiple registers in class II binding peptidesabstractMOTIVATION: While processing of MHC class II antigens for presentation to helper T-cells is essential for normal immune response, it is also implicated in the pathogenesis of autoimmune disorders and hypersensitivity reactions. Sequence-based computational techniques for predicting HLA-DQ binding peptides have encountered limited success, with few prediction techniques developed using three-dimensional models. METHODS: We describe a structure-based prediction model for modeling peptide-DQ3.2beta complexes. We have developed a rapid and accurate protocol for docking candidate peptides into the DQ3.2beta receptor and a scoring function to discriminate binders from the background. The scoring function was rigorously trained, tested and validated using experimentally verified DQ3.2beta binding and non-binding peptides obtained from biochemical and functional studies. RESULTS: Our model predicts DQ3.2beta binding peptides with high accuracy [area under the receiver operating characteristic (ROC) curve A(ROC) > 0.90], compared with experimental data. We investigated the binding patterns of DQ3.2beta peptides and illustrate that several registers exist within a candidate binding peptide. Further analysis reveals that peptides with multiple registers occur predominantly for high-affinity binders. Joo Chuan Tong, Guanglan Zhang, Tin Wee Tan, J. Thomas August, Vladimir Brusic, Shoba Ranganathan |
Bioinform. | 2 |
| 2005 | Predictive Vaccinology: Optimisation of Predictions Using Support Vector Machine Classifiers
Ivana Bozic, Guanglan Zhang, Vladimir Brusic |
IDEAL | 2 |
| 2002 | Dragon Promoter Finder: recognition of vertebrate RNA polymerase II promotersabstractAbstract Summary: Dragon Promoter Finder (DPF) locates RNA polymerase II promoters in DNA sequences of vertebrates by predicting Transcription Start Site (TSS) positions. DPF’s algorithm uses sensors for three functional regions (promoters, exons and introns) and an Artificial Neural Network (ANN). Results on a large and diverse evaluation set indicate that DPF exhibits a superior predicting ability for TSS location compared to three other promoter-finding programs. Availability: http://sdmc.krdl.org.sg:8080/promoter/ Contact: [email protected] * To whom correspondence should be addressed. Vladimir B. Bajic, Seng Hong Seah, Allen Chong, Guanglan Zhang, Judice L. Y. Koh, Vladimir Brusic |
Bioinform. | 4 |