VLDB 2026 Research / reviewers in the wild / expert
Derin B. Keskin
dblp:24/10786
· DBLP profile ↗
13ranked-venue papers
0as first author
4since 2021 · last 2024
0000-0002-8496-6181ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PREDBL6: a system for predicting C57BL/6 mouse T-cell epitopesabstractThe MHC class I antigen processing pathway plays a critical role in the adaptive immune system by presenting peptides for recognition by CD8+ T cells. While most prediction tools focus on MHC binding, accurately identifying immunogenic T-cell epitopes requires accounting for additional factors such as antigen processing and peptide-MHC binding thermostability. We developed a bioinformatics tool that integrates MHC binding predictions with thermostability assessment and antigen processing steps to enhance T cell epitope identification accuracy for C57BL/6 mice. Our machine learning models, trained on a comprehensive dataset of eluted H2-Kband H2-Dbligands and thermostability data across a range of physiologically relevant temperatures (37°C, 50°C, 70°C), were rigorously validated. These models showed improved overall accuracy on an external validation dataset compared to the widely used NetMHCPan-4.1. We consolidated the models into a user-friendly web-based application named PREDBL6 to facilitate accurate predictions of immunogenic peptides that stably bind H2b molecules and stimulate immune responses in C57BL/6 mice. PREDBL6 is accessible at http://met-hilab.org:3001/. Zitian Zhen, Guancheng Huang, Lubomir T. Chitkushev, Vladimir Brusic, Derin B. Keskin, Guanglan Zhang |
BIBM | 6 |
| 2021 | Correctness of Cell Labels in Public Single Cell Transcriptomics DatasetsabstractThe number of single-cell transcriptomic (SCT) studies is rapidly increasing. More than 15000 single cell gene expression data sets are available in public repositories. More than 2400 of these sets involve Peripheral Blood Mononuclear Cells (PBMC) data sets. Main cell types of PBMC are B cells, dendritic cells, monocytes, natural killer cells, and T cells. Labels of individual PBMC are usually provided in metadata accompanying the data sets or are implicit as data set partitions for sorted cells. We analyzed the correctness of labels assigned to individual cells from PBMC in primary reports. The correctness of primary labels was assessed by using Artificial Neural Network (ANN) classifier and Confident Learning (CL) approach. We assessed that the number of mislabels on average in our data sets is about2%. The label accuracy varied broadly between data sets, particularly among those generated by experimental cell sorting followed by SCT. Minjie Lyu, Yihan Zhang 0003, Derin B. Keskin, Lubomir T. Chitkushev, Guanglan Zhang, Vladimir Brusic |
BIBM | 4 |
| 2021 | Applications of single cell profiles of PBMC: Improvements of cell type classificationabstractSingle-cell-derived-class (SCDC) profiles capture characteristic gene expression from single cells representing types and subtypes and their conditions. SCDC profiles show high reproducibility across similar single-cell types processed under the same conditions. We have demonstrated two applications of SCDC profiles-classification of single cells from PBMC into six main classes (B cells, cDC, pDC, monocytes, NK cells, and T cells) and into three super-classes (BC+pDC,MC+cDC, and TC+NK). The minimum number of individual cells required for building an effective reference SCDC profile has been assessed to be between 160 and 640 cells. The variability of SCDC gradually decreases as the number of cells used to derive the profile increases from 10-cells to 640-cells. The classification accuracy of PBMC extracted by PBMC separation by SCDC profiles was 85-100% and 95-100% for supertypes depending on the cell type or supertype. Classification accuracy for PBMC cell types is lower for samples that are processed by cell sorting, or other sample processing steps. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Richard Rankin, Vladimir Brusic |
BIBM | 4 |
| 2021 | TANTIGEN 2.0: a knowledge base of tumor T cell antigens and epitopesabstractWe previously developed TANTIGEN, a comprehensive online database cataloging more than 1000 T cell epitopes and HLA ligands from 292 tumor antigens. In TANTIGEN 2.0, we significantly expanded coverage in both immune response targets (T cell epitopes and HLA ligands) and tumor antigens. It catalogs 4,296 antigen variants from 403 unique tumor antigens and more than 1500 T cell epitopes and HLA ligands. We also included neoantigens, a class of tumor antigens generated through mutations resulting in new amino acid sequences in tumor antigens. TANTIGEN 2.0 contains validated TCR sequences specific for cognate T cell epitopes and tumor antigen gene/mRNA/protein expression information in major human cancers extracted by Human Pathology Atlas. TANTIGEN 2.0 is a rich data resource for tumor antigens and their associated epitopes and neoepitopes. It hosts a set of tailored data analytics tools tightly integrated with the data to form meaningful analysis workflows. It is freely available at http://projects.met-hilab.org/tadb . Guanglan Zhang, Lubomir T. Chitkushev, Lars Rønn Olsen, Derin B. Keskin, Vladimir Brusic |
BMC Bioinform. | 4 |
| 2020 | Artificial Neural Network System for Cell Classification using Single Cell RNA ExpressionabstractWe implemented an automated system for single-cell classification using artificial neural networks (ANN). Our system takes single-cell gene expression sparse matrices and trains ANN to classify cell types and subtypes. The assemblies of ANNs predict cell classes by voting. We tested the system in a case study where we trained ANNs with a dataset containing approximately 120,000 single cells and tested the resulting model using an independent data set of 13,000 single cells. The overall accuracy of the 5-class classification was 95%. We trained and tested a total of 100 ANNs in 10 cycles. The prediction system demonstrated excellent reproducibility. The analysis of misclassifications indicated that 2% were likely classification errors, while the remaining 3% were likely due to mislabeled types and subtypes in the test set. Jiahui Zhong, Minjie Lyu, Derin B. Keskin, Guanglan Zhang, Vladimir Brusic, Lubomir T. Chitkushev |
BIBM | 5 |
| 2020 | Classification of Single Cell Types During Leukemia Therapy using Artificial Neural NetworksabstractWe trained artificial neural network (ANN) models to classify peripheral blood mononuclear cells (PBMC) in chronic lymphoid leukemia (CLL) patients. The classification task was to determine differences in gene expression profiles in PBMC pre-treatment (with ibrutinib) and on days 30, 120, 150, and 280 after the start of treatment. Twelve datasets represented clinical samples containing a total 48,016 single cell profiles were used to train and test ANN models to classify the progress of therapy by gene expression changes. The accuracy of ANN classification was $ \gt 92$% in internal cross-validation. External cross-validation, using independent data sets for training and testing, showed the accuracy of classification of post-treatment PBMCs to more than 80%. To the best of our knowledge, this is the first study that has demonstrated the potential of ANNs with 10x single cell gene expression data for detecting the changes during treatment of CLL. Minjie Lyu, Milena Radenkovic 0001, Derin B. Keskin, Vladimir Brusic |
BIBM | 3 |
| 2020 | Single-cell mRNA Profiles in PBMCabstractWe developed a method for building gene expression profiles from single-cell gene expression matrices. We named these profiles the “single-cell-derived-class” or SCDC profiles. They represent characteristic patterns of gene expressions of the types and subtypes of cells derived from single-cell transcriptome experiments. We deployed this method on classes and subclasses of peripheral blood mononuclear cells (PBMC). We used 47 human single-cell transcriptomics (SCT) data sets representing various classes, subclasses, and sample processing conditions. From comparisons of these profiles we found that they are highly reproducible, even when derived from unrelated studies as long as the processing steps are identical. The most similar profiles are those that are minimally processed. Cell sorting using FACS, cell enrichment, or fixing in methanol make profiles distinct from those derived from normal healthy samples. Our results suggest that approximately 50-200 cells are sufficient for building a useful SCDC profile. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 4 |
| 2020 | Prediction of PBMC Cell Types Using scRNAseq Reference ProfilesabstractSingle cell transcriptomics enables a high-resolution concurrent measurement of gene expression from tens of thousands of cells. We developed a method for determining standardized profiles from SCT data. We defined 48 data sets from 13 different studies and developed single-cellderived-class” (SCDC) profiles representing multiple classes and subclasses of peripheral blood mononuclear cells (PBMC). We applied pattern recognition analysis by calculating the distance from each query cell to the SCDC profiles (excluding the profiles of the query cells). Classification of cells by pattern recognition showed excellent performance for PBMC that were isolated, but not further processed by cell sorting. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 4 |
| 2020 | Classification of PBMC cell types using scRNAseq, ANN, and incremental learningabstractSingle cell transcriptomics (SCT) technology reveals gene expression of individual cells. Peripheral blood mononuclear cells (PBMC) are important diagnostic targets in immunology. In this study, we obtained and standardized 27 SCT data sets, derived from healthy PBMC samples using 10x SCT. We used artificial neural networks (ANN) to assess the ability of ANN to classify main PBMC cell types. Incremental learning by the gradual addition of new data sets to ANN training improved classification. The overall prediction accuracy of the final step of incremental learning reached 93% in 4-class classification. Jiahui Zhong, Razin A. Shaikh, Haoguo Wu, Lubomir T. Chitkushev, Guanglan Zhang, Derin B. Keskin, Vladimir Brusic |
BIBM | 8 |
| 2019 | Classification of Five Cell Types from PBMC Samples using Single Cell Transcriptomics and Artificial Neural NetworksabstractWe used 27 human single cell transcriptomics (SCT) data sets to develop an artificial neural network (ANN) model for classification of Peripheral Blood Mononuclear Cells (PBMC). We demonstrated that highly accurate models for the classification of PBMC subtypes can be developed by combining multiple independent data sets to form training data sets. A significant data preparation effort was needed for building predictive models. Using a data set of ~120,000 single cell instances we showed the accuracy of classification of PBMC call of ~ 90%. Optimization techniques and the addition of new high-quality data sets for model training are expected to improve PBMC subtype classification accuracy. Razin A. Shaikh, Jiahui Zhong, Minjie Lyu, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 5 |
| 2019 | TANTIGEN 2.0: an online database and analysis platform for tumor T cell antigensabstractWe previously developed TANTIGEN, a comprehensive web-based database cataloging more than 1,000 T cell epitopes and HLA ligands from 292 tumor antigens. TANTIGEN 2.0 is significantly expanded the number and coverage of immune response targets (T cell epitopes and HLA ligands) of previously cataloged tumor antigens. We expanded the number of cataloged tumor antigens to more than 4,000 and have added their reported targets of immune responses. We also included neoantigens, a new class of tumor antigens generated through mutations that results in a new amino acid sequence in tumor antigens. TANTIGEN 2.0 contains validated TCR sequences specific for cognate T cell epitopes. Gene expression information was extracted from tumor antigen gene/mRNA/protein expression information in major human cancers provided by Human Pathology Atlas. TANTIGEN 2.0 provides a rich data resource for tumor-associated epitope and neoepitope discovery studies. It is freely available at http://projects.met-hilab.org/tadb. Guanglan Zhang, Lubomir T. Chitkushev, Derin B. Keskin, Vladimir Brusic |
BIBM | 3 |
| 2017 | MCVdb: A database for knowledge discovery in Merkel cell polyomavirus with applications in T cell immunology and vaccinologyabstractMerkel Cell Polyomavirus (MCV) is associated with more than 80% of Merkel cell carcinoma (MCC), a rare but highly lethal form of skin cancer. We made use of the immunological data on MCV available through publications and databases and constructed MCV T cell Antigen Database (MCVdb). MCVdb contains 734 curated antigen entries of MCV antigenic proteins and 30 experimentally verified T cell epitopes. The data were subject to extensive quality control (redundancy elimination, error detection, and vocabulary consolidation). A set of computational tools for in-depth analysis, such as sequence comparison using BLAST search, multiple alignments of antigens, and T cell epitope conservation analysis have been integrated within the MCVdb. Predicted Class I and Class II HLA-binding peptides for 15 common HLA alleles are included in this database as putative targets. MCVdb is a unique data source providing a comprehensive list of MCV antigens and peptides. MCVdb is publicly available at http://projects.met-hilab.org/mcv/. Guanglan Zhang, Derin B. Keskin, James A. DeCaprio, Catherine J. Wu, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 2 |
| 2011 | PB1-F2 Finder: scanning influenza sequences for PB1-F2 encoding RNA segmentsabstractBACKGROUND: PB1-F2 is a major virulence factor of influenza A. This protein is a product of an alternative reading frame in the PB1-encoding RNA segment 2. Its presence of is dictated by the presence or absence of premature stop codons. This virulence factor is present in every influenza pandemic and major epidemic of the 20th century. Absence of PB1-F2 is associated with mild disease, such as the 2009 H1N1 ("swine flu"). RESULTS: The analysis of 8608 segment 2 sequences showed that only 8.5% have been annotated for the presence of PB1-F2. Our analysis indicates that 75% of segment 2 sequences are likely to encode PB1-F2. Two major populations of PB1-F2 are of lengths 90 and 57 while minor populations include lengths 52, 63, 79, 81, 87, and 101. Additional possible populations include the lengths of 59, 69, 81, 95, and 106. Previously described sequences include only lengths 57, 87, and 90. We observed substantial variation in PB1-F2 sequences where certain variants show up to 35% difference to well-defined reference sequences. Therefore this dataset indicates that there are many more variants that need to be functionally characterized. CONCLUSIONS: Our web-accessible tool PB1-F2 Finder enables scanning of influenza sequences for potential PB1-F2 protein products. It provides an initial screen and annotation of PB1-F2 products. It is accessible at http://cvc.dfci.harvard.edu/pb1-f2. David S. DeLuca, Derin B. Keskin, Ellis L. Reinherz, Vladimir Brusic |
BMC Bioinform. | 2 |