EDBT 2026 Demo / reviewers in the wild / expert
Vladimir Brusic
dblp:00/6815
· DBLP profile ↗
54ranked-venue papers
5as first author
11since 2021 · last 2024
0000-0003-0523-5266ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PREDBL6: a system for predicting C57BL/6 mouse T-cell epitopesabstractThe MHC class I antigen processing pathway plays a critical role in the adaptive immune system by presenting peptides for recognition by CD8+ T cells. While most prediction tools focus on MHC binding, accurately identifying immunogenic T-cell epitopes requires accounting for additional factors such as antigen processing and peptide-MHC binding thermostability. We developed a bioinformatics tool that integrates MHC binding predictions with thermostability assessment and antigen processing steps to enhance T cell epitope identification accuracy for C57BL/6 mice. Our machine learning models, trained on a comprehensive dataset of eluted H2-Kband H2-Dbligands and thermostability data across a range of physiologically relevant temperatures (37°C, 50°C, 70°C), were rigorously validated. These models showed improved overall accuracy on an external validation dataset compared to the widely used NetMHCPan-4.1. We consolidated the models into a user-friendly web-based application named PREDBL6 to facilitate accurate predictions of immunogenic peptides that stably bind H2b molecules and stimulate immune responses in C57BL/6 mice. PREDBL6 is accessible at http://met-hilab.org:3001/. Zitian Zhen, Guancheng Huang, Lubomir T. Chitkushev, Vladimir Brusic, Derin B. Keskin, Guanglan Zhang |
BIBM | 5 |
| 2023 | A privacy preserving framework for federated learning in smart healthcare systems
Wenshuo Wang 0006, Xiuqin Qiu, Jindong Zhao, Vladimir Brusic |
Inf. Process. Manag. | 6 |
| 2022 | Ethically Informed Software Process for Smart Health HomeabstractSmart health homes (SHHs) integrate wearable sensors and various interconnected devices using the Internet of Things (IoT) technologies. SHHs combine IoT, data communication, and health-related applications to deliver healthcare services at home. The existing regulations and standards for SHH design are insufficient for home health care. Technical and device standards are available for guiding SHH design and implementation, but ethical standards are lacking. We identified six ethical requirements important for SHH: safety/trust, privacy/data security, vulnerable groups, individual autonomy, transparency/explainability/fairness, and social responsibility/ morality. We identified a set of questions useful for software engineering (SE) process for ethically informed software in SHH design and mapped them to the steps of software process. We mapped related guidelines from relevant professional codes of conduct. These questions can guide ethically informed software process of SHH. Matthew Pike, Nasser Mustafa, Vladimir Brusic |
CBMS | 4 |
| 2022 | Social Impact of Smart Environments: Software Engineering Perspectives and Challenges
Stuart McDonald, Dave Towey, Vladimir Brusic |
COMPSAC | 3 |
| 2022 | The development of ethically informed standards for intelligent monitoring systems of electric machinesabstractThis paper presents a development process of ethically-informed standards for intelligent monitoring systems (IMS). It considers the gap of socio-ethical standards for IMS over global industry, in both general terms and field-specific cases. Ethical issues of IMS are defined to include safety, data security, transparency, privacy, justice, equality, sustainability, and beneficence. We argue that they can be investigated and assessed by answering a set of questions for each ethical issue. This ethical evaluation questionnaire can be considered as an effective assessment tool for standard constructions. Additionally, field-specific industrial cases for the development of IMS ethical standards were demonstrated, including IMS for wind power systems, more electric aircrafts, and agricultural industry. Eventually, the summary of necessary items for building the ethical standards of IMS are presented. Kun Shang 0001, Stuart McDonald, Giampaolo Buticchi, Vladimir Brusic |
COMPSAC | 4 |
| 2021 | Correctness of Cell Labels in Public Single Cell Transcriptomics DatasetsabstractThe number of single-cell transcriptomic (SCT) studies is rapidly increasing. More than 15000 single cell gene expression data sets are available in public repositories. More than 2400 of these sets involve Peripheral Blood Mononuclear Cells (PBMC) data sets. Main cell types of PBMC are B cells, dendritic cells, monocytes, natural killer cells, and T cells. Labels of individual PBMC are usually provided in metadata accompanying the data sets or are implicit as data set partitions for sorted cells. We analyzed the correctness of labels assigned to individual cells from PBMC in primary reports. The correctness of primary labels was assessed by using Artificial Neural Network (ANN) classifier and Confident Learning (CL) approach. We assessed that the number of mislabels on average in our data sets is about2%. The label accuracy varied broadly between data sets, particularly among those generated by experimental cell sorting followed by SCT. Minjie Lyu, Yihan Zhang 0003, Derin B. Keskin, Lubomir T. Chitkushev, Guanglan Zhang, Vladimir Brusic |
BIBM | 7 |
| 2021 | PBMC Cell Classification from Single Cell mRNA Expression by Artificial Neural Networks, Profiles, Gene Markers, and Protein MarkersabstractWe performed classification of healthy Peripheral Blood Mononuclear Cells cell types using four methods Artificial Neural Network (ANN), Profiles, Protein Markers (PMs), and RNA markers (RNAMs). Profiles represent patterns of gene expressions characteristic of the subtypes of cells. PMs are protein found exclusively in certain types or subtypes of cells, or represent particular cell states, RNAMs are genes which demonstrate significant differential expressions between cell types. A total of 109 datasets from four different sources containing $\sim$ 120,000 single cells gene expression were used to train and test prediction models. We combined the methods which perform prediction using the whole set of gene features (ANN and Profiles), and those that used specific gene features (PMs and RNAMs) to predict the cell type. The overall classification accuracy was 94.8% for ANN, 94.5% for Profiles, 90.7% for PMs, 67.9% for RNAMs. The combination of four methods showed accuracy of 90.9% with high confidence of positive predictions. The combination of four methods allowed identification of mislabeled cell types in test data sets. Minjie Lyu, Yihan Zhang 0003, Luning Yang, Huan Jin, Anthony Bellotti, Nenad S. Mitic, Vladimir Brusic |
BIBM | 9 |
| 2021 | Classification of Single Cell Types using Small Sets of Expressed Genes: Comparative Analysis of Supervised Machine Learning MethodsabstractSingle cell transcriptomics measures gene expression data of large number of genes, concurrently, from tens of thousands of cells present in a studied biological sample. It is difficult to obtain good classification results due to high data dimensionality and variability of biological states. We performed a preliminary study to assess the feasibility of using supervised machine learning methods to classify peripheral blood mononuclear cell (PBMC) types from single cell gene expression data. We analyzed a large PBMC data set $(\sim 120,000$ PBMC cells), selected 47 genes (from 30698 features) suitable as SML classification features, and performed classification using 20 machine learning algorithms. Data sets represented three sample processing strategies: PBMC separation (two data sets), and experimental cell sorting by (two data sets). The accuracy in 5-class classification among 20 methods was 91-97% (PBMC separation), 97-100% (magnetic-activated cell sorting), and 82-99% (fluorescence-activated cell sorting). Our results indicate the feasibility of supervised machine learning for classification of cells into major PBMC cell types using a small number of classification features from single cell gene expression data. Aleksandar Veljkovic, Mirjana M. Maljkovic, Nenad S. Mitic, Sasa N. Malkov, Minjie Lyu, Marek T. Michalewicz, Guanglan Zhang, Vladimir Brusic |
BIBM | 9 |
| 2021 | Applications of single cell profiles of PBMC: Improvements of cell type classificationabstractSingle-cell-derived-class (SCDC) profiles capture characteristic gene expression from single cells representing types and subtypes and their conditions. SCDC profiles show high reproducibility across similar single-cell types processed under the same conditions. We have demonstrated two applications of SCDC profiles-classification of single cells from PBMC into six main classes (B cells, cDC, pDC, monocytes, NK cells, and T cells) and into three super-classes (BC+pDC,MC+cDC, and TC+NK). The minimum number of individual cells required for building an effective reference SCDC profile has been assessed to be between 160 and 640 cells. The variability of SCDC gradually decreases as the number of cells used to derive the profile increases from 10-cells to 640-cells. The classification accuracy of PBMC extracted by PBMC separation by SCDC profiles was 85-100% and 95-100% for supertypes depending on the cell type or supertype. Classification accuracy for PBMC cell types is lower for samples that are processed by cell sorting, or other sample processing steps. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Richard Rankin, Vladimir Brusic |
BIBM | 8 |
| 2021 | Multi-Step Optimization of Indoor Localization Accuracy Using Commodity WiFiabstractWe present a Multi-step optimization Localization Algorithm (MoLA) for improved accuracy of indoor localization by WiFi. Using consumer-grade WiFi devices, we can determine the location of an object carrying WiFi device with higher accuracy than the comparable popular systems. MoLA removes the phase error by using calibration, then estimates the angle-of-arrival (AoA) of the signal by combining I-MUSIC algorithm and the minimum description length (MDL) equation. The line-of-sight (LoS) path is identified using our novel estimator function. In the last step, the object location is estimated. Extensive measurement experiments have shown that MoLA can achieve improvement in both median localization and point localization error in a 290 m2environment having adequate LoS with a single receiver. Sherif Welsen, Vladimir Brusic |
PIMRC | 3 |
| 2021 | TANTIGEN 2.0: a knowledge base of tumor T cell antigens and epitopesabstractWe previously developed TANTIGEN, a comprehensive online database cataloging more than 1000 T cell epitopes and HLA ligands from 292 tumor antigens. In TANTIGEN 2.0, we significantly expanded coverage in both immune response targets (T cell epitopes and HLA ligands) and tumor antigens. It catalogs 4,296 antigen variants from 403 unique tumor antigens and more than 1500 T cell epitopes and HLA ligands. We also included neoantigens, a class of tumor antigens generated through mutations resulting in new amino acid sequences in tumor antigens. TANTIGEN 2.0 contains validated TCR sequences specific for cognate T cell epitopes and tumor antigen gene/mRNA/protein expression information in major human cancers extracted by Human Pathology Atlas. TANTIGEN 2.0 is a rich data resource for tumor antigens and their associated epitopes and neoepitopes. It hosts a set of tailored data analytics tools tightly integrated with the data to form meaningful analysis workflows. It is freely available at http://projects.met-hilab.org/tadb . Guanglan Zhang, Lubomir T. Chitkushev, Lars Rønn Olsen, Derin B. Keskin, Vladimir Brusic |
BMC Bioinform. | 5 |
| 2020 | Artificial Neural Network System for Cell Classification using Single Cell RNA ExpressionabstractWe implemented an automated system for single-cell classification using artificial neural networks (ANN). Our system takes single-cell gene expression sparse matrices and trains ANN to classify cell types and subtypes. The assemblies of ANNs predict cell classes by voting. We tested the system in a case study where we trained ANNs with a dataset containing approximately 120,000 single cells and tested the resulting model using an independent data set of 13,000 single cells. The overall accuracy of the 5-class classification was 95%. We trained and tested a total of 100 ANNs in 10 cycles. The prediction system demonstrated excellent reproducibility. The analysis of misclassifications indicated that 2% were likely classification errors, while the remaining 3% were likely due to mislabeled types and subtypes in the test set. Jiahui Zhong, Minjie Lyu, Derin B. Keskin, Guanglan Zhang, Vladimir Brusic, Lubomir T. Chitkushev |
BIBM | 7 |
| 2020 | Classification of Single Cell Types During Leukemia Therapy using Artificial Neural NetworksabstractWe trained artificial neural network (ANN) models to classify peripheral blood mononuclear cells (PBMC) in chronic lymphoid leukemia (CLL) patients. The classification task was to determine differences in gene expression profiles in PBMC pre-treatment (with ibrutinib) and on days 30, 120, 150, and 280 after the start of treatment. Twelve datasets represented clinical samples containing a total 48,016 single cell profiles were used to train and test ANN models to classify the progress of therapy by gene expression changes. The accuracy of ANN classification was $ \gt 92$% in internal cross-validation. External cross-validation, using independent data sets for training and testing, showed the accuracy of classification of post-treatment PBMCs to more than 80%. To the best of our knowledge, this is the first study that has demonstrated the potential of ANNs with 10x single cell gene expression data for detecting the changes during treatment of CLL. Minjie Lyu, Milena Radenkovic 0001, Derin B. Keskin, Vladimir Brusic |
BIBM | 4 |
| 2020 | Single-cell mRNA Profiles in PBMCabstractWe developed a method for building gene expression profiles from single-cell gene expression matrices. We named these profiles the “single-cell-derived-class” or SCDC profiles. They represent characteristic patterns of gene expressions of the types and subtypes of cells derived from single-cell transcriptome experiments. We deployed this method on classes and subclasses of peripheral blood mononuclear cells (PBMC). We used 47 human single-cell transcriptomics (SCT) data sets representing various classes, subclasses, and sample processing conditions. From comparisons of these profiles we found that they are highly reproducible, even when derived from unrelated studies as long as the processing steps are identical. The most similar profiles are those that are minimally processed. Cell sorting using FACS, cell enrichment, or fixing in methanol make profiles distinct from those derived from normal healthy samples. Our results suggest that approximately 50-200 cells are sufficient for building a useful SCDC profile. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 7 |
| 2020 | Prediction of PBMC Cell Types Using scRNAseq Reference ProfilesabstractSingle cell transcriptomics enables a high-resolution concurrent measurement of gene expression from tens of thousands of cells. We developed a method for determining standardized profiles from SCT data. We defined 48 data sets from 13 different studies and developed single-cellderived-class” (SCDC) profiles representing multiple classes and subclasses of peripheral blood mononuclear cells (PBMC). We applied pattern recognition analysis by calculating the distance from each query cell to the SCDC profiles (excluding the profiles of the query cells). Classification of cells by pattern recognition showed excellent performance for PBMC that were isolated, but not further processed by cell sorting. Luning Yang, Yihan Zhang 0003, Nenad S. Mitic, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 7 |
| 2020 | Automation of Gene Expression Profile Analysis in Single Cell DataabstractWe designed and implemented an automated system, named Pierse for pattern recognition of single cell transcriptomics (SCT) data. The Pierse system takes sparse matrices and corresponding metadata as input to generate SCDC profiles (SCT gene expression profiles characteristic of types or subtypes of cells). These profiles can be used for profile comparison, feature extraction, and differential gene expression analysis. Hierarchical clustering is used for similarity analysis between SCDC profiles and resulting heatmaps are produced. We performed a demonstration study to test functional modules in the Pierse system. To improve efficiency, we deployed parallel programming scripts and implemented efficient matrix analysis functions in the demonstration study. Yihan Zhang 0003, Luning Yang, Vladimir Brusic |
BIBM | 3 |
| 2020 | Tissue of origin classification from single cell mRNA expression by Artificial Neural NetworksabstractSingle cell transcriptomics (SCT) enables high-throughput measurement of mRNA expression concurrently from tens of thousands of single cells. Gene expression profiles in single cells cover only a small fraction of expressed genes and these data are inherently noisy. We developed a method that utilizes artificial neural networks (ANN) for classification of single cells by their tissue of origin. Data sets representing 10 different organs and tissues from C57BL/6 laboratory mice were standardized and used for training and testing ANN models. Each organ was represented by at least two datasets derived from different mice. We achieved 80% accuracy in 10-class classification. After combining data sets from spleen, bone marrow, and lung into one super-class and mammary tissue and muscle into another, we achieved overall cell classification accuracy of 98% across two tissue super-classes and five organs. Bangrui Zheng, Minjie Lyu, Vladimir Brusic |
BIBM | 4 |
| 2020 | Classification of PBMC cell types using scRNAseq, ANN, and incremental learningabstractSingle cell transcriptomics (SCT) technology reveals gene expression of individual cells. Peripheral blood mononuclear cells (PBMC) are important diagnostic targets in immunology. In this study, we obtained and standardized 27 SCT data sets, derived from healthy PBMC samples using 10x SCT. We used artificial neural networks (ANN) to assess the ability of ANN to classify main PBMC cell types. Incremental learning by the gradual addition of new data sets to ANN training improved classification. The overall prediction accuracy of the final step of incremental learning reached 93% in 4-class classification. Jiahui Zhong, Razin A. Shaikh, Haoguo Wu, Lubomir T. Chitkushev, Guanglan Zhang, Derin B. Keskin, Vladimir Brusic |
BIBM | 9 |
| 2019 | Classification of Five Cell Types from PBMC Samples using Single Cell Transcriptomics and Artificial Neural NetworksabstractWe used 27 human single cell transcriptomics (SCT) data sets to develop an artificial neural network (ANN) model for classification of Peripheral Blood Mononuclear Cells (PBMC). We demonstrated that highly accurate models for the classification of PBMC subtypes can be developed by combining multiple independent data sets to form training data sets. A significant data preparation effort was needed for building predictive models. Using a data set of ~120,000 single cell instances we showed the accuracy of classification of PBMC call of ~ 90%. Optimization techniques and the addition of new high-quality data sets for model training are expected to improve PBMC subtype classification accuracy. Razin A. Shaikh, Jiahui Zhong, Minjie Lyu, Derin B. Keskin, Guanglan Zhang, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 8 |
| 2019 | TANTIGEN 2.0: an online database and analysis platform for tumor T cell antigensabstractWe previously developed TANTIGEN, a comprehensive web-based database cataloging more than 1,000 T cell epitopes and HLA ligands from 292 tumor antigens. TANTIGEN 2.0 is significantly expanded the number and coverage of immune response targets (T cell epitopes and HLA ligands) of previously cataloged tumor antigens. We expanded the number of cataloged tumor antigens to more than 4,000 and have added their reported targets of immune responses. We also included neoantigens, a new class of tumor antigens generated through mutations that results in a new amino acid sequence in tumor antigens. TANTIGEN 2.0 contains validated TCR sequences specific for cognate T cell epitopes. Gene expression information was extracted from tumor antigen gene/mRNA/protein expression information in major human cancers provided by Human Pathology Atlas. TANTIGEN 2.0 provides a rich data resource for tumor-associated epitope and neoepitope discovery studies. It is freely available at http://projects.met-hilab.org/tadb. Guanglan Zhang, Lubomir T. Chitkushev, Derin B. Keskin, Vladimir Brusic |
BIBM | 4 |
| 2019 | Sensor Networks and Data Management in Healthcare: Emerging Technologies and New ChallengesabstractSmart pervasive sensor networks are becoming an important part of our daily lives. Low-power, high-availability and high-throughput 5G mobile networks provide the necessary communication means for highly pervasive sensor networks, introducing a technological disruption to health monitoring. The meaningful use of large concurrent sensor networks in healthcare requires multi-level health knowledge integration with sensor data streams. In this paper, we highlight some software engineering and data-processing issues that can be addressed by metamorphic testing. The proposed solution combines data streaming with filtering and cross-calibration, use of medical knowledge for system operation and data interpretation, and IoT-based calibration using certified linked diagnostic devices. Matthew Pike, Nasser Mustafa, Dave Towey, Vladimir Brusic |
COMPSAC (1) | 4 |
| 2018 | Single Cell Transcriptomics Reveals Summary Patterns Specific for PBMCs and Other Cell Types
Jingjie Xu, Razin A. Shaikh, Vladimir Brusic |
BIBM | 3 |
| 2018 | Elemental metabolomicsabstractElemental metabolomics is quantification and characterization of total concentration of chemical elements in biological samples and monitoring of their changes. Recent advances in inductively coupled plasma mass spectrometry have enabled simultaneous measurement of concentrations of > 70 elements in biological samples. In living organisms, elements interact and compete with each other for absorption and molecular interactions. They also interact with proteins and nucleotide sequences. These interactions modulate enzymatic activities and are critical for many molecular and cellular functions. Testing for concentration of > 40 elements in blood, other bodily fluids and tissues is now in routine use in advanced medical laboratories. In this article, we define the basic concepts of elemental metabolomics, summarize standards and workflows, and propose minimum information for reporting the results of an elemental metabolomics experiment. Major statistical and informatics tools for elemental metabolomics are reviewed, and examples of applications are discussed. Elemental metabolomics is emerging as an important new technology with applications in medical diagnostics, nutrition, agriculture, food science, environmental science and multiplicity of other areas. Ping Zhang 0008, Constantinos A. Georgiou, Vladimir Brusic |
Briefings Bioinform. | 3 |
| 2017 | MCVdb: A database for knowledge discovery in Merkel cell polyomavirus with applications in T cell immunology and vaccinologyabstractMerkel Cell Polyomavirus (MCV) is associated with more than 80% of Merkel cell carcinoma (MCC), a rare but highly lethal form of skin cancer. We made use of the immunological data on MCV available through publications and databases and constructed MCV T cell Antigen Database (MCVdb). MCVdb contains 734 curated antigen entries of MCV antigenic proteins and 30 experimentally verified T cell epitopes. The data were subject to extensive quality control (redundancy elimination, error detection, and vocabulary consolidation). A set of computational tools for in-depth analysis, such as sequence comparison using BLAST search, multiple alignments of antigens, and T cell epitope conservation analysis have been integrated within the MCVdb. Predicted Class I and Class II HLA-binding peptides for 15 common HLA alleles are included in this database as putative targets. MCVdb is a unique data source providing a comprehensive list of MCV antigens and peptides. MCVdb is publicly available at http://projects.met-hilab.org/mcv/. Guanglan Zhang, Derin B. Keskin, James A. DeCaprio, Catherine J. Wu, Lubomir T. Chitkushev, Vladimir Brusic |
BIBM | 6 |
| 2015 | An adaptive genetic algorithm for selection of blood-based biomarkers for prediction of Alzheimer's disease progressionabstractBACKGROUND: Alzheimer's disease is a multifactorial disorder that may be diagnosed earlier using a combination of tests rather than any single test. Search algorithms and optimization techniques in combination with model evaluation techniques have been used previously to perform the selection of suitable feature sets. Previously we successfully applied GA with LR to neuropsychological data contained within the The Australian Imaging, Biomarkers and Lifestyle (AIBL) study of aging, to select cognitive tests for prediction of progression of AD. This research addresses an Adaptive Genetic Algorithm (AGA) in combination with LR for identifying the best biomarker combination for prediction of the progression to AD. RESULTS: The model has been explored in terms of parameter optimization to predict conversion from healthy stage to AD with high accuracy. Several feature sets were selected - the resulting prediction moddels showed higher area under the ROC values (0.83-0.89). The results has shown consistency with some of the medical research reported in literature. CONCLUSION: The AGA has proven useful in selecting the best combination of biomarkers for prediction of AD progression. The algorithm presented here is generic and can be extended to other data sets generated in projects that seek to identify combination of biomarkers or other features that are predictive of disease onset or progression. Luke Vandewater, Vladimir Brusic, William J. Wilson, Lance S. Macaulay, Ping Zhang 0008 |
BMC Bioinform. | 2 |
| 2013 | Computational vaccinology and the ICoVax 2012 workshopabstractComputational vaccinology or vaccine informatics is an interdisciplinary field that addresses scientific and clinical questions in vaccinology using computational and informatics approaches. Computational vaccinology overlaps with many other fields such as immunoinformatics, reverse vaccinology, postlicensure vaccine research, vaccinomics, literature mining, and systems vaccinology. The second ISV Pre-conference Computational Vaccinology Workshop (ICoVax 2012) was held on October 13, 2013 in Shanghai, China. A number of topics were presented in the workshop, including allergen predictions, prediction of linear T cell epitopes and functional conformational epitopes, prediction of protein-ligand binding regions, vaccine design using reverse vaccinology, and case studies in computational vaccinology. Although a significant progress has been made to date, a number of challenges still exist in the field. This Editorial provides a list of major challenges for the future of computational vaccinology and identifies developing themes that will expand and evolve over the next few years. Yongqun He, Anne S. De Groot, Vladimir Brusic, Christian Schönbach, Nikolai Petrovsky |
BMC Bioinform. | 4 |
| 2012 | InCoB2012 Conference: from biological data to knowledge to technological breakthroughsabstractTen years ago when Asia-Pacific Bioinformatics Network held the first International Conference on Bioinformatics (InCoB) in Bangkok its theme was North-South Networking. At that time InCoB aimed to provide biologists and bioinformatics researchers in the Asia-Pacific region a forum to meet, interact with, and disseminate knowledge about the burgeoning field of bioinformatics. Meanwhile InCoB has evolved into a major regional bioinformatics conference that attracts not only talented and established scientists from the region but increasingly also from East Asia, North America and Europe. Since 2006 InCoB yielded 114 articles in BMC Bioinformatics supplement issues that have been cited nearly 1,000 times to date. In part, these developments reflect the success of bioinformatics education and continuous efforts to integrate and utilize bioinformatics in biotechnology and biosciences in the Asia-Pacific region. A cross-section of research leading from biological data to knowledge and to technological applications, the InCoB2012 theme, is introduced in this editorial. Other highlights included sessions organized by the Pan-Asian Pacific Genome Initiative and a Machine Learning in Immunology competition. InCoB2013 is scheduled for September 18-21, 2013 at Suzhou, China. Christian Schönbach, Sissades Tongsima, Jonathan H. Chan, Vladimir Brusic, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 4 |
| 2011 | PB1-F2 Finder: scanning influenza sequences for PB1-F2 encoding RNA segmentsabstractBACKGROUND: PB1-F2 is a major virulence factor of influenza A. This protein is a product of an alternative reading frame in the PB1-encoding RNA segment 2. Its presence of is dictated by the presence or absence of premature stop codons. This virulence factor is present in every influenza pandemic and major epidemic of the 20th century. Absence of PB1-F2 is associated with mild disease, such as the 2009 H1N1 ("swine flu"). RESULTS: The analysis of 8608 segment 2 sequences showed that only 8.5% have been annotated for the presence of PB1-F2. Our analysis indicates that 75% of segment 2 sequences are likely to encode PB1-F2. Two major populations of PB1-F2 are of lengths 90 and 57 while minor populations include lengths 52, 63, 79, 81, 87, and 101. Additional possible populations include the lengths of 59, 69, 81, 95, and 106. Previously described sequences include only lengths 57, 87, and 90. We observed substantial variation in PB1-F2 sequences where certain variants show up to 35% difference to well-defined reference sequences. Therefore this dataset indicates that there are many more variants that need to be functionally characterized. CONCLUSIONS: Our web-accessible tool PB1-F2 Finder enables scanning of influenza sequences for potential PB1-F2 protein products. It provides an initial screen and annotation of PB1-F2 products. It is accessible at http://cvc.dfci.harvard.edu/pb1-f2. David S. DeLuca, Derin B. Keskin, Ellis L. Reinherz, Vladimir Brusic |
BMC Bioinform. | 5 |
| 2010 | Using Gaussian Process with Test Rejection to Detect T-Cell Epitopes in Pathogen GenomesabstractA major challenge in the development of peptide-based vaccines is finding the right immunogenic element, with efficient and long-lasting immunization effects, from large potential targets encoded by pathogen genomes. Computer models are convenient tools for scanning pathogen genomes to preselect candidate immunogenic peptides for experimental validation. Current methods predict many false positives resulting from a low prevalence of true positives. We develop a test reject method based on the prediction uncertainty estimates determined by Gaussian process regression. This method filters false positives among predicted epitopes from a pathogen genome. The performance of stand-alone Gaussian process regression is compared to other state-of-the-art methods using cross validation on 11 benchmark data sets. The results show that the Gaussian process method has the same accuracy as the top performing algorithms. The combination of Gaussian process regression with the proposed test reject method is used to detect true epitopes from the Vaccinia virus genome. The test rejection increases the prediction accuracy by reducing the number of false positives without sacrificing the method's sensitivity. We show that the Gaussian process in combination with test rejection is an effective method for prediction of T-cell epitopes in large and diverse pathogen genomes, where false positives are of concern. Liwen You, Vladimir Brusic, Marcus Gallagher, Mikael Bodén |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2009 | ImmunoGrid, an integrative environment for large-scale simulation of the immune system for vaccine discovery, design and optimizationabstractVaccine research is a combinatorial science requiring computational analysis of vaccine components, formulations and optimization. We have developed a framework that combines computational tools for the study of immune function and vaccine development. This framework, named ImmunoGrid combines conceptual models of the immune system, models of antigen processing and presentation, system-level models of the immune system, Grid computing, and database technology to facilitate discovery, formulation and optimization of vaccines. ImmunoGrid modules share common conceptual models and ontologies. The ImmunoGrid portal offers access to educational simulators where previously defined cases can be displayed, and to research simulators that allow the development of new, or tuning of existing, computational models. The portal is accessible at . Francesco Pappalardo 0001, Mark D. Halling-Brown, Nicolas Rapin, Ping Zhang 0008, Davide Alemani, Andrew P. J. Emerson, Paola Paci, Patrice Duroux, Marzio Pennisi, Arianna Palladini, Olivo Miotto, Daniel Churchill, Elda Rossi, Adrian J. Shepherd, David S. Moss, Filippo Castiglione, Massimo Bernaschi, Marie-Paule Lefranc, Søren Brunak, Santo Motta, Pierluigi Lollini, Kaye E. Basford, Vladimir Brusic |
Briefings Bioinform. | 23 |
| 2008 | A Hybrid Model for Prediction of Peptide Binding to MHC Molecules
Ping Zhang 0008, Vladimir Brusic, Kaye E. Basford |
ICONIP (1) | 2 |
| 2008 | Critical technologies for bioinformaticsabstractScientific advances over the last 50 years have provided a basis for parallel revolutions in engineering and in biomedicine. Advances in automation and miniaturization have enabled the development of modern instrumentation that supports large-scale measurement of biological entities. The burgeoning fields of genomics and proteomics, as well as other ‘omics’, keep generating large amounts of molecular expression and interaction data. Advances in cytometry enable quantification of cellular states and related functional properties. The latest scanning and imaging technologies have made it possible to scan and make detailed measurements of whole organisms. Clinical data complement biological data, enabling detailed descriptions of various healthy and diseased states, progression and responses to therapies. The availability of data representing various biological states, processes, and their time dependencies enable the study of biological systems at various levels of organization, from molecule to organism, and even population levels. Multiple sources of data support a rapidly growing body of biomedical knowledge; however our ability to analyze and interpret these data lags far behind data generation and storage capacity. The field of bioinformatics is experiencing rapid growth, principally in three directions: management, analysis and modeling of biological data. Data management refers to acquisition, management, dissemination and basic interpretation of biological data. Analysis refers to integrative analysis, as well as complex analysis of biological data. Modeling of biological data refers to in silico approaches such as mathematical and computational modeling, simulation and prediction of biological systems and processes. Bioinformatics uses information technologies to gather data from biomedicine and translate it into information and various forms of knowledge. Information technologies support this quest: hardware, networking, databases, algorithms and computational models are used for tasks ranging from simple retrieval of publication abstracts or data from molecular databases, to complex simulations that complement multi-step experimentation. Early bioinformatics infrastructure of the 1980s comprised simple molecular databases containing several thousand molecular entries, basic sequence comparison algorithms and simple mathematical models. Current infrastructure comprises a network of databases, web-accessible analytical tools, computer networks and supercomputers that connect users to bioinformatics resources across the world. The revolutions in information technology and biotechnology and related advances in instrumentation produce huge amounts of data. The interpretation of these data and extraction of new knowledge requires bioinformatics approaches of ever increasing sophistication. This special issue of Briefings in Bioinformatics covers eight bioinformatics technologies which have made a significant impact on both the theory and practice of biomedical research. These are representative of important bioinformatics topics that can be described as enabling technologies for generating and sharing of biological knowledge. Lefranc et al. describe IMGT, a system that integrates complex immunogenetic data, ontologies, analysis tools and web resources. This system offers a formal description of objects, processes and relations necessary for conceptualization of knowledge in molecular immunogenetics and provides tools for multi-scale approaches at various levels of hierarchy from molecule to organism and, possibly, population level. In particular, this review offers an insight on merging the conceptual models and ontologies with sequence databases and their combination with related analytical tools. The IMGT technology demonstrates the conversion of precisely defined biological concepts into an integrated bioinformatics environment. The article by Standley et al. describes the latest developments in molecular structure databases, in particular the enhancement of structural information with biochemical, functional and experimental information. A new generation of tools enable querying by keywords/text, sequence similarity and structure similarity. Advanced visualization tools have been developed to assist the analysis and interpretation of molecular structures. The new generation of tools improves identification of distant functional relationship relative to traditional sequence comparison-based methods. Multiple sequence alignment (MSA) methods are among the most important bioinformatics tools whereby biological sequences are compared to identify the regions of similarity that confer evolutionary, structural or functional relatedness. Because the MSA problem is NP-hard, the optimal solution is not possible unless MSA involves a small number of sequences. The rapid growth of biological sequence data, both in length and number of sequences, the need for identification of distant homologs and for web accessibility require increasingly sophisticated algorithms. Katoh and Toh have described the latest version of the MAFFT program for MSA. MAFFT is an example of tools where high-quality MSAs can be performed rapidly over the Internet. Kumar et al. have described MEGA, a tool for evolutionary analysis of biological sequences. MEGA is designed for comparative analysis of genes and proteins and investigation of their evolutionary relationships. It offers a convenient means for managing data from local files and web repositories, for statistical analysis of data, and visualization of results. In their article, the authors describe the design features that make MEGA a biologist-centric tool, that enable complex analyses where the results are clearly, unambiguously presented to the user through advanced visualization tools. The analysis of complex biological data brings formidable challenges to the developers of advanced bioinformatics tools. Formal mathematical and statistical approaches are complemented with tools of computational intelligence that employ heuristic algorithms. Computational intelligence methods employ fuzzy logic, search and classification to study systems than learn, evolve and adapt. Fogel provides a primer of computational intelligence methods and their application to selected biological problems. Systems biology is concerned with the study of complex interactions in biological systems. It uses integrative rather than reductionist approach to discover properties arising from combination of multiple biological entities and their effects on multiple levels of biological organization. Hu et al. have described a visual data mining system VisANT that integrates multiple data types, performs the analysis of complex biological networks, and displays the results through visualization tools. Transcriptional regulation is a key area of systems biology. Wingender's article describes the TRANSFAC database and analysis of system whose original aim was to provide a genome-wide map of interaction sites for transcription factors. Over the years, TRANSFAC became the industry standard. However, with the growth of genomic information and genome mapping of model organisms, the original development evolved into a system that provides classifications and ontologies for the domain, and also integrates a catalog of transcription networks and their parts with a set of prediction tools. Hunter et al. have described the Physiome project that focuses on multi-scale modeling necessary for relating molecular and cellular processes to the events observed at the organ or whole organism. Such multi-scale modeling requires advanced techniques of mechanistic modeling, visualization and numerical techniques. Advances in bioinformatics are reshaping biomedical research and applications in biotechnology. Integration of large quantities of data, instant access to information, and the ability to perform complex analyses and simulations provide the capacity to amplify the results of biomedical research many-fold. Biomedicine is increasingly becoming an information-rich field. Deciphering qualitative relationships between components of living organisms, and quantification of interactions will depend on further development of bioinformatics technologies. The combination of large-scale screening experiments with large-scale simulations in silico is already used for selection of a limited number of key experiments, resulting in decreased cost and time required for biomedical research and discovery. Vladimir Brusic, Shoba Ranganathan |
Briefings Bioinform. | 1 |
| 2008 | Evaluation of MHC-II peptide binding prediction servers: applications for vaccine researchabstractBACKGROUND: Initiation and regulation of immune responses in humans involves recognition of peptides presented by human leukocyte antigen class II (HLA-II) molecules. These peptides (HLA-II T-cell epitopes) are increasingly important as research targets for the development of vaccines and immunotherapies. HLA-II peptide binding studies involve multiple overlapping peptides spanning individual antigens, as well as complete viral proteomes. Antigen variation in pathogens and tumor antigens, and extensive polymorphism of HLA molecules increase the number of targets for screening studies. Experimental screening methods are expensive and time consuming and reagents are not readily available for many of the HLA class II molecules. Computational prediction methods complement experimental studies, minimize the number of validation experiments, and significantly speed up the epitope mapping process. We collected test data from four independent studies that involved 721 peptide binding assays. Full overlapping studies of four antigens identified binding affinity of 103 peptides to seven common HLA-DR molecules (DRB1*0101, 0301, 0401, 0701, 1101, 1301, and 1501). We used these data to analyze performance of 21 HLA-II binding prediction servers accessible through the WWW. RESULTS: Because not all servers have predictors for all tested HLA-II molecules, we assessed a total of 113 predictors. The length of test peptides ranged from 15 to 19 amino acids. We tried three prediction strategies - the best 9-mer within the longer peptide, the average of best three 9-mer predictions, and the average of all 9-mer predictions within the longer peptide. The best strategy was the identification of a single best 9-mer within the longer peptide. Overall, measured by the receiver operating characteristic method (AROC), 17 predictors showed good (AROC > 0.8), 41 showed marginal (AROC > 0.7), and 55 showed poor performance (AROC < 0.7). Good performance predictors included HLA-DRB1*0101 (seven), 1101 (six), 0401 (three), and 0701 (one). The best individual predictor was NETMHCIIPAN, closely followed by PROPRED, IEDB (Consensus), and MULTIPRED (SVM). None of the individual predictors was shown to be suitable for prediction of promiscuous peptides. Current predictive capabilities allow prediction of only 50% of actual T-cell epitopes using practical thresholds. CONCLUSION: The available HLA-II servers do not match prediction capabilities of HLA-I predictors. Currently available HLA-II prediction servers offer only a limited prediction accuracy and the development of improved predictors is needed for large-scale studies, such as proteome-wide epitope mapping. The requirements for accuracy of HLA-II binding predictions are stringent because of the substantial effect of false positives. Honghuang Lin, Guanglan Zhang, Songsak Tongchusak, Ellis L. Reinherz, Vladimir Brusic |
BMC Bioinform. | 5 |
| 2008 | Identification of human-to-human transmissibility factors in PB2 proteins of influenza A by large-scale mutual information analysisabstractBACKGROUND: The identification of mutations that confer unique properties to a pathogen, such as host range, is of fundamental importance in the fight against disease. This paper describes a novel method for identifying amino acid sites that distinguish specific sets of protein sequences, by comparative analysis of matched alignments. The use of mutual information to identify distinctive residues responsible for functional variants makes this approach highly suitable for analyzing large sets of sequences. To support mutual information analysis, we developed the AVANA software, which utilizes sequence annotations to select sets for comparison, according to user-specified criteria. The method presented was applied to an analysis of influenza A PB2 protein sequences, with the objective of identifying the components of adaptation to human-to-human transmission, and reconstructing the mutation history of these components. RESULTS: We compared over 3,000 PB2 protein sequences of human-transmissible and avian isolates, to produce a catalogue of sites involved in adaptation to human-to-human transmission. This analysis identified 17 characteristic sites, five of which have been present in human-transmissible strains since the 1918 Spanish flu pandemic. Sixteen of these sites are located in functional domains, suggesting they may play functional roles in host-range specificity. The catalogue of characteristic sites was used to derive sequence signatures from historical isolates. These signatures, arranged in chronological order, reveal an evolutionary timeline for the adaptation of the PB2 protein to human hosts. CONCLUSION: By providing the most complete elucidation to date of the functional components participating in PB2 protein adaptation to humans, this study demonstrates that mutual information is a powerful tool for comparative characterization of sequence sets. In addition to confirming previously reported findings, several novel characteristic sites within PB2 are reported. Sequence signatures generated using the characteristic sites catalogue characterize concisely the adaptation characteristics of individual isolates. Evolutionary timelines derived from signatures of early human influenza isolates suggest that characteristic variants emerged rapidly, and remained remarkably stable through subsequent pandemics. In addition, the signatures of human-infecting H5N1 isolates suggest that this avian subtype has low pandemic potential at present, although it presents more human adaptation components than most avian subtypes. Olivo Miotto, A. T. Heiny, Tin Wee Tan, J. Thomas August, Vladimir Brusic |
BMC Bioinform. | 5 |
| 2008 | Rule-based knowledge aggregation for large-scale protein sequence analysis of influenza A virusesabstractBACKGROUND: The explosive growth of biological data provides opportunities for new statistical and comparative analyses of large information sets, such as alignments comprising tens of thousands of sequences. In such studies, sequence annotations frequently play an essential role, and reliable results depend on metadata quality. However, the semantic heterogeneity and annotation inconsistencies in biological databases greatly increase the complexity of aggregating and cleaning metadata. Manual curation of datasets, traditionally favoured by life scientists, is impractical for studies involving thousands of records. In this study, we investigate quality issues that affect major public databases, and quantify the effectiveness of an automated metadata extraction approach that combines structural and semantic rules. We applied this approach to more than 90,000 influenza A records, to annotate sequences with protein name, virus subtype, isolate, host, geographic origin, and year of isolation. RESULTS: Over 40,000 annotated Influenza A protein sequences were collected by combining information from more than 90,000 documents from NCBI public databases. Metadata values were automatically extracted, aggregated and reconciled from several document fields by applying user-defined structural rules. For each property, values were recovered from >/=88.8% of records, with accuracy exceeding 96% in most cases. Because of semantic heterogeneity, each property required up to six different structural rules to be combined. Significant quality differences between databases were found: GenBank documents yield values more reliably than documents extracted from GenPept. Using a simple set of semantic rules and a reasoner, we reconstructed relationships between sequences from the same isolate, thus identifying 7640 isolates. Validation of isolate metadata against a simple ontology highlighted more than 400 inconsistencies, leading to over 3,000 property value corrections. CONCLUSION: To overcome the quality issues inherent in public databases, automated knowledge aggregation with embedded intelligence is needed for large-scale analyses. Our results show that user-controlled intuitive approaches, based on combination of simple rules, can reliably automate various curation tasks, reducing the need for manual corrections to approximately 5% of the records. Emerging semantic technologies possess desirable features to support today's knowledge aggregation tasks, with a potential to bring immediate benefits to this field. Olivo Miotto, Tin Wee Tan, Vladimir Brusic |
BMC Bioinform. | 3 |
| 2008 | Hotspot Hunter: a computational system for large-scale screening and selection of candidate immunological hotspots in pathogen proteomesabstractBACKGROUND: T-cell epitopes that promiscuously bind to multiple alleles of a human leukocyte antigen (HLA) supertype are prime targets for development of vaccines and immunotherapies because they are relevant to a large proportion of the human population. The presence of clusters of promiscuous T-cell epitopes, immunological hotspots, has been observed in several antigens. These clusters may be exploited to facilitate the development of epitope-based vaccines by selecting a small number of hotspots that can elicit all of the required T-cell activation functions. Given the large size of pathogen proteomes, including of variant strains, computational tools are necessary for automated screening and selection of immunological hotspots. RESULTS: Hotspot Hunter is a web-based computational system for large-scale screening and selection of candidate immunological hotspots in pathogen proteomes through analysis of antigenic diversity. It allows screening and selection of hotspots specific to four common HLA supertypes, namely HLA class I A2, A3, B7 and class II DR. The system uses Artificial Neural Network and Support Vector Machine methods as predictive engines. Soft computing principles were employed to integrate the prediction results produced by both methods for robust prediction performance. Experimental validation of the predictions showed that Hotspot Hunter can successfully identify majority of the real hotspots. Users can predict hotspots from a single protein sequence, or from a set of aligned protein sequences representing pathogen proteome. The latter feature provides a global view of the localizations of the hotspots in the proteome set, enabling analysis of antigenic diversity and shift of hotspots across protein variants. The system also allows the integration of prediction results of the four supertypes for identification of hotspots common across multiple supertypes. The target selection feature of the system shortlists candidate peptide hotspots for the formulation of an epitope-based vaccine that could be effective against multiple variants of the pathogen and applicable to a large proportion of the human population. CONCLUSION: Hotspot Hunter is publicly accessible at http://antigen.i2r.a-star.edu.sg/hh/. It is a new generation computational tool aiding in epitope-based vaccine design. Guanglan Zhang, Asif M. Khan, Kellathur N. Srinivasan, A. T. Heiny, Kenneth X. Lee, Chee Keong Kwoh 0001, J. Thomas August, Vladimir Brusic |
BMC Bioinform. | 8 |
| 2007 | The growth of bioinformaticsabstractBriefings in Bioinformatics, or BiB for short, will celebrate its 7th anniversary this year. The journal's mission has remained unchanged throughout this period: we are committed to disseminating knowledge on databases and computational tools for life sciences through review articles. During its history, some 250 articles have been published and they have been cited more than 2700 times, making BiB the premier bioinformatics journal using the per-article citation impact measure. The common theme for this issue is the cross-disciplinary nature of bioinformatics and the proliferation of bioinformatics into new areas of life sciences. This issue brings six outstanding reviews that collectively demonstrate the broad outreach of bioinformatics. Perez-Iratxeta, Andrade-Navarro and Wren performed a meta-analysis of abstracts published in MEDLINE and abstracts of NIH-funded project grants to determine the growth and spread of computational approaches across the various subfields of biomedicine during the past 30 years. They explore three major bioinformatics concepts: computation, the Internet and databases. Their analysis of MeSH terms indicate the major areas of focus within bioinformatics are protein, gene and nucleic acid databases, computational biology, computing methodologies and programming languages. Software and software design, database management systems and principal component analysis, are found to be of high importance, followed by several other sub-areas of biocomputing. The areas with highest growth during the period of 2000–03 include bioinformatics-dependent areas of genomics, genetic databases, gene expression profiling and oligonucleotide array sequence analysis. Computational biology alone showed a 3-fold increase during this period, while bioinformatics showed a 15-fold increase. Bioinformatics has spawned into sub-disciplines such as cheminformatics, neuroinformatics and immunoinformatics, and the boundaries between bioinformatics and biomedical disciplines are increasingly blurred. Tong, Tan and Ranganathan have summarized the latest developments of methods and protocols for predicting immunogenic epitopes, a major topic within the rapidly growing field of immunoinformatics. They present a clear case of how bioinformatics-driven methods for the selection of key experiments resulted in a significant increase in the speed and economy of mapping of vaccine targets. These methods enable the formulation of new testable hypotheses through the in-depth analysis of complex immunological data that could not have been developed by traditional experimental approaches alone. Bioinformatics keeps proliferating into diverse biomedical disciplines. Law enforcement increasingly uses biomolecular data and databases. Forensic DNA databases, for example, have been established in a large number of countries. Bianchi and Liò discuss the state of the art in forensic DNA science and the potential of bioinformatics in developing this field. They point out that the bioinformatics analysis of forensic DNA has important implications for the organization of forensic evidence and the integration of crime databases with public health and population genetics databases. Any information of such nature has potential for misuse. The authors also discuss privacy rights and the role of bioinformatics in protection of these rights. Genomics is the fastest growing area at the intersection of biomedicine and bioinformatics. A large number of statistical methods and software solutions are appearing in support of genomic studies. Gold and co-authors address the issues of statistical testing and of the use of a priori knowledge for the assignment of biological properties to genes and have a number of recommendations for the users. Their results support the assumption of gene independence for the analysis of genomic data, thus allowing the use of a range of statistical techniques. They also offer insights into the practical use of software packages for these tasks. Their final words of wisdom encourage the bioinformatics community to remain wary of the implicit assumptions present in software packages. Bayesian statistics is used as inference engine and for information extraction. It is particularly suitable for the analysis of data produced by complex systems and which are subject to high level of noise. D. J. Wilkinson provides the primer on the use of Bayesian approaches in bioinformatics. These applications include biological sequence analysis using (hidden Markov models), analysis of microarray data, protein informatics and expression networks. The author gives a number of examples, such as stochastic kinetic models of biological processes, modeling of synchrony in yeast populations, multiple sequence analysis, motif detection and prediction of transcription factor binding sites. Further examples of use of Bayesian methods are given in other reviews within this issue: Bianchi and Liò (forensics), Tong and co-authors (prediction of immune epitopes) and Gold (gene annotation). The expansion of bioinformatics requires the refinement of well-established tools and methods. Although protein modeling has a long history, some aspects are only now being addressed due to the large quantities of proteomic data and the sophistication of technique required. Fariselli and co-authors address the issue of how to detect remote structural homologs among proteins. They introduce the concept of WWWH (When, Why, Where and How) of structural homology modeling and review popular computational approaches. Perez-Iratxeta and co-authors note the need for commercial, governmental and educational institutions to make long-term and short-term strategic decisions about bioinformatics-based resource allocation, training and workforce education. Briefings in Bioinformatics provides reference materials and standards to help decision makers and experts toward achieving these goals. Vladimir Brusic |
Briefings Bioinform. | 1 |
| 2007 | Predicting peptides binding to MHC class II molecules using multi-objective evolutionary algorithmsabstractBACKGROUND: Peptides binding to Major Histocompatibility Complex (MHC) class II molecules are crucial for initiation and regulation of immune responses. Predicting peptides that bind to a specific MHC molecule plays an important role in determining potential candidates for vaccines. The binding groove in class II MHC is open at both ends, allowing peptides longer than 9-mer to bind. Finding the consensus motif facilitating the binding of peptides to a MHC class II molecule is difficult because of different lengths of binding peptides and varying location of 9-mer binding core. The level of difficulty increases when the molecule is promiscuous and binds to a large number of low affinity peptides. In this paper, we propose two approaches using multi-objective evolutionary algorithms (MOEA) for predicting peptides binding to MHC class II molecules. One uses the information from both binders and non-binders for self-discovery of motifs. The other, in addition, uses information from experimentally determined motifs for guided-discovery of motifs. RESULTS: The proposed methods are intended for finding peptides binding to MHC class II I-Ag7 molecule - a promiscuous binder to a large number of low affinity peptides. Cross-validation results across experiments on two motifs derived for I-Ag7 datasets demonstrate better generalization abilities and accuracies of the present method over earlier approaches. Further, the proposed method was validated and compared on two publicly available benchmark datasets: (1) an ensemble of qualitative HLA-DRB1*0401 peptide data obtained from five different sources, and (2) quantitative peptide data obtained for sixteen different alleles comprising of three mouse alleles and thirteen HLA alleles. The proposed method outperformed earlier methods on most datasets, indicating that it is well suited for finding peptides binding to MHC class II molecules. CONCLUSION: We present two MOEA-based algorithms for finding motifs, one for self-discovery and the other for guided-discovery by experimentally determined motifs, and thereby predicting binding peptides to I-Ag7 molecule. Our experiments show that the proposed MOEA-based algorithms are better than earlier methods in predicting binding sites not only on I-Ag7 but also on most alleles of class II MHC benchmark datasets. This shows that our methods could be applicable to find binding motifs in a wide range of alleles. Menaka Rajapakse, Bertil Schmidt, Feng Lin 0002, Vladimir Brusic |
BMC Bioinform. | 4 |
| 2006 | Functional Prediction of Snake NeurotoxinsabstractSnake neurotoxins are important experimental tool in pharmacological research. Over the years, the number of snake neurotoxin sequences identified is increasing at a very fast pace. However, only a small portion of them are experimentally characterized from more than 200,000 variants estimated to exist in nature. In this paper, we report a systematic functional analysis on snake neurotoxins using a statistical machine learning method - nearest neighbour approach for functional prediction together with a set of rules. Based on this method we built a highly accurate functional prediction tool for putative annotation for snake neurotoxins Seng Hong Seah, Chee Keong Kwoh 0001, Vladimir Brusic, Meena Kishore Sakharkar, Geok See Ng |
ICARCV | 3 |
| 2006 | Extreme Learning Machine for Predicting HLA-Peptide Binding
Stephanus Daniel Handoko, Chee Keong Kwoh 0001, Yew-Soon Ong, Guanglan Zhang, Vladimir Brusic |
ISNN (2) | 5 |
| 2006 | Prediction of HLA-DQ3.2ß Ligands: evidence of multiple registers in class II binding peptidesabstractMOTIVATION: While processing of MHC class II antigens for presentation to helper T-cells is essential for normal immune response, it is also implicated in the pathogenesis of autoimmune disorders and hypersensitivity reactions. Sequence-based computational techniques for predicting HLA-DQ binding peptides have encountered limited success, with few prediction techniques developed using three-dimensional models. METHODS: We describe a structure-based prediction model for modeling peptide-DQ3.2beta complexes. We have developed a rapid and accurate protocol for docking candidate peptides into the DQ3.2beta receptor and a scoring function to discriminate binders from the background. The scoring function was rigorously trained, tested and validated using experimentally verified DQ3.2beta binding and non-binding peptides obtained from biochemical and functional studies. RESULTS: Our model predicts DQ3.2beta binding peptides with high accuracy [area under the receiver operating characteristic (ROC) curve A(ROC) > 0.90], compared with experimental data. We investigated the binding patterns of DQ3.2beta peptides and illustrate that several registers exist within a candidate binding peptide. Further analysis reveals that peptides with multiple registers occur predominantly for high-affinity binders. Joo Chuan Tong, Guanglan Zhang, Tin Wee Tan, J. Thomas August, Vladimir Brusic, Shoba Ranganathan |
Bioinform. | 5 |
| 2006 | Large-scale analysis of antigenic diversity of T-cell epitopes in dengue virusabstractBACKGROUND: Antigenic diversity in dengue virus strains has been studied, but large-scale and detailed systematic analyses have not been reported. In this study, we report a bioinformatics method for analyzing viral antigenic diversity in the context of T-cell mediated immune responses. We applied this method to study the relationship between short-peptide antigenic diversity and protein sequence diversity of dengue virus. We also studied the effects of sequence determinants on viral antigenic diversity. Short peptides, principally 9-mers were studied because they represent the predominant length of binding cores of T-cell epitopes, which are important for formulation of vaccines. RESULTS: Our analysis showed that the number of unique protein sequences required to represent complete antigenic diversity of short peptides in dengue virus is significantly smaller than that required to represent complete protein sequence diversity. Short-peptide antigenic diversity shows an asymptotic relationship to the number of unique protein sequences, indicating that for large sequence sets (approximately 200) the addition of new protein sequences has marginal effect to increasing antigenic diversity. A near-linear relationship was observed between the extent of antigenic diversity and the length of protein sequences, suggesting that, for the practical purpose of vaccine development, antigenic diversity of short peptides from dengue virus can be represented by short regions of sequences (approximately <100 aa) within viral antigens that are specific targets of immune responses (such as T-cell epitopes specific to particular human leukocyte antigen alleles). CONCLUSION: This study provides evidence that there are limited numbers of antigenic combinations in protein sequence variants of a viral species and that short regions of the viral protein are sufficient to capture antigenic diversity of T-cell epitopes. The approach described herein has direct application to the analysis of other viruses, in particular those that show high diversity and/or rapid evolution, such as influenza A virus and human immunodeficiency virus (HIV). Asif M. Khan, A. T. Heiny, Kenneth X. Lee, Kellathur N. Srinivasan, Tin Wee Tan, J. Thomas August, Vladimir Brusic |
BMC Bioinform. | 7 |
| 2005 | Predictive Vaccinology: Optimisation of Predictions Using Support Vector Machine Classifiers
Ivana Bozic, Guanglan Zhang, Vladimir Brusic |
IDEAL | 3 |
| 2005 | Extraction by Example: Induction of Structural Rules for the Analysis of Molecular Sequence Data from Heterogeneous Sources
Olivo Miotto, Tin Wee Tan, Vladimir Brusic |
IDEAL | 3 |
| 2005 | Deriving Matrix of Peptide-MHC Interactions in Diabetic Mouse by Genetic Algorithm
Menaka Rajapakse, Lonce L. Wyse, Bertil Schmidt, Vladimir Brusic |
IDEAL | 4 |
| 2004 | Systematic analysis of snake neurotoxins' functional classification using a data warehousing approachabstractMOTIVATION: Sequence annotations, functional and structural data on snake venom neurotoxins (svNTXs) are scattered across multiple databases and literature sources. Sequence annotations and structural data are available in the public molecular databases, while functional data are almost exclusively available in the published articles. There is a need for a specialized svNTXs database that contains NTX entries, which are organized, well annotated and classified in a systematic manner. RESULTS: We have systematically analyzed svNTXs and classified them using structure-function groups based on their structural, functional and phylogenetic properties. Using conserved motifs in each phylogenetic group, we built an intelligent module for the prediction of structural and functional properties of unknown NTXs. We also developed an annotation tool to aid the functional prediction of newly identified NTXs as an additional resource for the venom research community. AVAILABILITY: We created a searchable online database of NTX proteins sequences (http://research.i2r.a-star.edu.sg/Templar/DB/snake_neurotoxin). This database can also be found under Swiss-Prot Toxin Annotation Project website (http://www.expasy.org/sprot/). Joyce Phui Yee Siew, Asif M. Khan, Paul T. J. Tan, Judice L. Y. Koh, Seng Hong Seah, Chuay Yeng Koo, Siaw Ching Chai, Arunmozhiarasi Armugam, Vladimir Brusic, Kandiah Jeyaseelan |
Bioinform. | 9 |
| 2003 | From Informatics to Bioinformatics
Vladimir B. Bajic, Vladimir Brusic, Jinyan Li 0001, See-Kiong Ng, Limsoon Wong |
APBC | 2 |
| 2003 | Bioinformatics for Venom and Toxin SciencesabstractVenomous animals produce a myriad of important pharmacological components. The individual components, or venoms (toxins), are used in ion channel and receptor studies, drug discovery, and formulation of insecticides. The toxin data are scattered across public databases which provide sequence and structural descriptions, but very limited functional annotation. The exponential growth of newly identified toxin data has created a need for better data management. Venominformatics is a systematic bioinformatics approach in which classified, consolidated and cleaned venom data are stored into repositories and integrated with advanced bioinformatics tools for the analysis of structure and function of toxins. Venominformatics complements experimental studies and helps reduce the number of essential experiments. Paul T. J. Tan, Asif M. Khan, Vladimir Brusic |
Briefings Bioinform. | 3 |
| 2002 | Dragon Promoter Finder: recognition of vertebrate RNA polymerase II promotersabstractAbstract Summary: Dragon Promoter Finder (DPF) locates RNA polymerase II promoters in DNA sequences of vertebrates by predicting Transcription Start Site (TSS) positions. DPF’s algorithm uses sensors for three functional regions (promoters, exons and introns) and an Artificial Neural Network (ANN). Results on a large and diverse evaluation set indicate that DPF exhibits a superior predicting ability for TSS location compared to three other promoter-finding programs. Availability: http://sdmc.krdl.org.sg:8080/promoter/ Contact: [email protected] * To whom correspondence should be addressed. Vladimir B. Bajic, Seng Hong Seah, Allen Chong, Guanglan Zhang, Judice L. Y. Koh, Vladimir Brusic |
Bioinform. | 6 |
| 2000 | Data Warehousing in Molecular BiologyabstractIn the business and healthcare sectors data warehousing has provided effective solutions for information usage and knowledge discovery from databases. However, data warehousing applications in the biological research and development (R&D) sector are lagging far behind. The fuzziness and complexity of biological data represent a major challenge in data warehousing for molecular biology. By combining experiences in other domains with our findings from building a model database, we have defined the requirements for data warehousing in molecular biology. Christian Schönbach, Peter Kowalski-Saunders, Vladimir Brusic |
Briefings Bioinform. | 3 |
| 1999 | Artificial neural network applications in immunologyabstractArtificial neural network (ANN) applications in immunology include simulations of peptide binding to histocompatibility complex molecules, which present peptides for recognition by the immune system. These peptides are derived from protein antigens and represent prime targets for vaccine discovery. ANN models have proven superior when compared to the alternative models. Applications of ANN models help minimise the number of necessary wet-lab experiments. In this article we describe three specific applications in which targets of immune recognition have been determined from diabetes-, melanoma-, and malaria-related antigens. Vladimir Brusic, John Zeleznikow |
IJCNN | 1 |
| 1998 | MHCWeb: converting a WWW database into a knowledge-based collaborative environment
Lawrence S. Hon, Neil F. Abernethy, Vladimir Brusic, Jenny Chai, Russ B. Altman |
AMIA | 3 |
| 1998 | Prediction of MHC class II-binding peptides using an evolutionary algorithm and artificial neural networkabstractMOTIVATION: Prediction methods for identifying binding peptides could minimize the number of peptides required to be synthesized and assayed, and thereby facilitate the identification of potential T-cell epitopes. We developed a bioinformatic method for the prediction of peptide binding to MHC class II molecules. RESULTS: Experimental binding data and expert knowledge of anchor positions and binding motifs were combined with an evolutionary algorithm (EA) and an artificial neural network (ANN): binding data extraction --> peptide alignment --> ANN training and classification . This method, termed PERUN, was implemented for the prediction of peptides that bind to HLA-DR4(B1*0401). The respective positive predictive values of PERUN predictions of high-, moderate-, low- and zero-affinity binders were assessed as 0.8, 0.7, 0.5 and 0.8 by cross-validation, and 1.0, 0.8, 0.3 and 0.7 by experimental binding. This illustrates the synergy between experimentation and computer modeling, and its application to the identification of potential immunotherapeutic peptides. AVAILABILITY: Software and data are available from the authors upon request. CONTACT: [email protected]. au Vladimir Brusic, George B. Rudy, G. Honeyman, Jürgen Hammer, Leonard C. Harrison |
Bioinform. | 1 |
| 1997 | Application of Genetic Search in Derivation of Matrix Models of Peptide Binding to MHC Molecules
Vladimir Brusic, Christian Schönbach, Masafumi Takiguchi, Victor Ciesielski, Leonard C. Harrison |
ISMB | 1 |