EDBT 2026 Demo / reviewers in the wild / expert
Tin Wee Tan
dblp:33/1961
· DBLP profile ↗
39ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0008-7829-3647ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
immunoinformatics |
0.2 | 3 | 2007 | In silico grouping of peptide/HLA class I complexes using structural interaction characteristics · Bioinform. 2007 Prediction of HLA-DQ3.2ß Ligands: evidence of multiple registers in class II binding peptides · Bioinform. 2006 MPID: MHC-Peptide Interaction Database for sequence-structure-function information on peptides binding to MHC molecules · Bioinform. 2003 |
Distributed systems › peer-to-peer systems › file sharing
peer-to-peer file sharing |
0.1 | 1 | 2008 | Automatic synchronization and distribution of biological databases and software over low-bandwidth networks among developing countries · Bioinform. 2008 |
Bioinformatics and computational biology › protein function prediction
caspase substrate cleavage site prediction |
0.1 | 1 | 2007 | CASVM: web server for SVM-based prediction of caspase substrates cleavage sites · Bioinform. 2007 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2007 | CASVM: web server for SVM-based prediction of caspase substrates cleavage sites · Bioinform. 2007 |
Bioinformatics and computational biology › immunoinformatics
MHC class II binding prediction |
0.1 | 1 | 2006 | Prediction of HLA-DQ3.2ß Ligands: evidence of multiple registers in class II binding peptides · Bioinform. 2006 |
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure visualization |
0.0 | 1 | 2003 | XdomView: protein domain and exon position visualization · Bioinform. 2003 |
Bioinformatics and computational biology
phylogenetics |
0.0 | 1 | 2001 | Cladogramer: incorporating haplotype frequency into cladogram analysis · Bioinform. 2001 |
Bioinformatics and computational biology › biological database
gene database |
0.0 | 1 | 2000 | IE-Kb: intron exon knowledge base · Bioinform. 2000 |
Bioinformatics and computational biology › immunoinformatics
epitope-based vaccine design |
0.0 | 1 | 2007 | In silico grouping of peptide/HLA class I complexes using structural interaction characteristics · Bioinform. 2007 |
Bioinformatics and computational biology
protein function prediction |
0.0 | 1 | 2007 | CASVM: web server for SVM-based prediction of caspase substrates cleavage sites · Bioinform. 2007 |
Bioinformatics and computational biology › kernel methods
support vector machine |
0.0 | 1 | 2007 | CASVM: web server for SVM-based prediction of caspase substrates cleavage sites · Bioinform. 2007 |
Bioinformatics and computational biology › bioinformatics infrastructure
sequence analysis software |
0.0 | 1 | 1993 | A hypertext-like approach to navigating through the GCG sequence analysis package · Comput. Appl. Biosci. 1993 |
Bioinformatics and computational biology
user interface |
0.0 | 1 | 1993 | A hypertext-like approach to navigating through the GCG sequence analysis package · Comput. Appl. Biosci. 1993 |
Methods — techniques the papers use, named apart from their topics
rsync · 0.1data grids · 0.1FTP · 0.1bittorrent · 0.1support vector machine · 0.1structural interaction descriptors · 0.1crystallographic structure analysis · 0.1structure-based prediction · 0.1scoring function · 0.1molecular docking · 0.1interaction parameter computation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Establishing the Asia & Pacific Bioinformatics Joint Congress: a historic milestone in regional bioinformatics collaborationabstractIn response to the need for greater cohesion among regional conferences, the Asia Pacific Bioinformatics Network (APBioNET) set out in 2015 to realize a long-held aspiration-a single, unifying bioinformatics "super conference" for the Asia & Pacific community. Nearly a decade of persistence, coordination, and coalition-building led to the inaugural Asia & Pacific Bioinformatics Joint Congress (APBJC2024) in Okinawa, Japan. Now established as a triennial event, APBJC stands as a testament to the power of collective vision and shared purpose, offering a unifying platform for regional collaboration and scientific exchange. Tagline: Bringing a Region Together: The Making of APBJC. Asif M. Khan, Susumu Goto, Kenta Nakai, Limsoon Wong, Diane E. Kovats, Shinya Ikematsu, Yoshihiro Yamanishi, Nurul Salwanie Che Wahid, Pradeep Eranti, Yi-Ping Phoebe Chen, Tae-Min Kim, Shinn-Ying Ho, Jessica Cara Mar, Wataru Iwasaki 0001, Jayaraman Valadi, Prashanth Suravajhala, Christian Schönbach, Tin Wee Tan, Shoba Ranganathan, Kiyoko F. Aoki-Kinoshita |
Briefings Bioinform. | 20 |
| 2024 | A global initiative on addressing bioinformatics' grand challengesabstractThe Bioinformatics Grand Challenges Consortium (BGCC) is a collaborative effort to address the most pressing challenges in bioinformatics. Initially focusing on education and training, the consortium successfully defined seven key grand challenges and is actively developing actionable solutions for these challenges. Building on this foundation, the BGCC plans to broaden its focus to include additional grand challenges in emerging areas. Asif M. Khan, Esra Büsra Isik, Tin Wee Tan |
Briefings Bioinform. | 3 |
| 2021 | A multi-task CNN learning model for taxonomic assignment of human virusesabstractBACKGROUND: Taxonomic assignment is a key step in the identification of human viral pathogens. Current tools for taxonomic assignment from sequencing reads based on alignment or alignment-free k-mer approaches may not perform optimally in cases where the sequences diverge significantly from the reference sequences. Furthermore, many tools may not incorporate the genomic coverage of assigned reads as part of overall likelihood of a correct taxonomic assignment for a sample. RESULTS: In this paper, we describe the development of a pipeline that incorporates a multi-task learning model based on convolutional neural network (MT-CNN) and a Bayesian ranking approach to identify and rank the most likely human virus from sequence reads. For taxonomic assignment of reads, the MT-CNN model outperformed Kraken 2, Centrifuge, and Bowtie 2 on reads generated from simulated divergent HIV-1 genomes and was more sensitive in identifying SARS as the closest relation in four RNA sequencing datasets for SARS-CoV-2 virus. For genomic region assignment of assigned reads, the MT-CNN model performed competitively compared with Bowtie 2 and the region assignments were used for estimation of genomic coverage that was incorporated into a naïve Bayesian network together with the proportion of taxonomic assignments to rank the likelihood of candidate human viruses from sequence data. CONCLUSIONS: We have developed a pipeline that combines a novel MT-CNN model that is able to identify viruses with divergent sequences together with assignment of the genomic region, with a Bayesian approach to ranking of taxonomic assignments by taking into account both the number of assigned reads and genomic coverage. The pipeline is available at GitHub via https://github.com/MaHaoran627/CNN_Virus . Tin Wee Tan, Kenneth Ban Hon Kim |
BMC Bioinform. | 2 |
| 2017 | Exploring the transcriptome of non-model oleaginous microalga Dunaliella tertiolecta through high-throughput sequencing and high performance computingabstractBACKGROUND: RNA-Seq technology has received a lot of attention in recent years for microalgal global transcriptomic profiling. It is widely used in transcriptome-wide analysis of gene expression., particularly for microalgal strains with potential as biofuel sources. However, insufficient genomic or transcriptomic information of non-model microalgae has limited the understanding of their regulatory mechanisms and hampered genetic manipulation to enhance biofuel production. As such, an optimal microalgal transcriptomic database construction is a subject of urgent investigation. RESULTS: Dunaliella tertiolecta, a non-model oleaginous microalgal species, was sequenced via Illumina MISEQ and HISEQ 4000 in RNA-Seq studies. The high quality high-throughout sequencing data were explored using high performance computing (HPC) in a petascale data center and subjected to de novo assembly and parallelized mpiBLASTX search with multiple species. As a result, a transcriptome database of 17,845 was constructed (~95% completeness). This enlarged database constructed fueled the RNA-Seq data analysis, which was validated by a nitrogen deprivation (ND) study that induces triacylglycerol (TAG) production. CONCLUSIONS: The new paralleled assembly and annotation method under HPC presented here allows the solution of large-scale data processing problems in acceptable computation time. There is significant increase in the number of transcriptomic data achieved and observable heterogeneity in the performance to identify differentially expressed genes in the ND treatment paradigm. The results provide new insights as to how response to ND treatment in microalgae is regulated. ND analyses highlight the advantages of this database generated in this study that could also serve as a useful resource for future gene manipulation and transcriptome-wide analysis. We thus demonstrate the usefulness of exploring the transcriptome as an informative platform for functional studies and genetic manipulations in similar species. Kenneth Wei Min Tan, Tin Wee Tan, Yuan Kun Lee |
BMC Bioinform. | 3 |
| 2015 | GIW and InCoB are advancing bioinformatics in the Asia-PacificabstractGIW/InCoB2015 the joint 26th International Conference on Genome Informatics (GIW) and 14th International Conference on Bioinformatics (InCoB) held in Tokyo, September 9-11, 2015 was attended by over 200 delegates. Fifty-one out of 89 oral presentations were based on research articles accepted for publication in four BMC journal supplements and three other journals. Sixteen articles in this supplement and six articles in the BMC Systems Biology GIW/InCoB2015 Supplement are covered by this introduction. The topics range from genome informatics, protein structure informatics, image analysis to biological networks and biomarker discovery. Christian Schönbach, Paul Horton, Siu-Ming Yiu, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 4 |
| 2014 | InCoB2014: bioinformatics to tackle the data to knowledge challengeabstractSince 2006, the International Conference on Bioinformatics (InCoB) has been publishing selected papers in BMC Bioinformatics. Papers within the scope of the journal from the 13th InCoB July 31-2 August, 2014 in Sydney, Australia have been compiled in this supplement. These span protein and proteome informatics, structural bioinformatics, software development and bioimaging to pharmacoinformatics and disease informatics, representing the breadth of bioinformatics research in the Asia-Pacific. Shoba Ranganathan, Tin Wee Tan, Christian Schönbach |
BMC Bioinform. | 2 |
| 2013 | APBioNet - Transforming Bioinformatics in the Asia-Pacific Regionabstract10.1371/journal.pcbi.1003317 Asif M. Khan, Tin Wee Tan, Christian Schönbach, Shoba Ranganathan |
PLoS Comput. Biol. | 2 |
| 2012 | InCoB2012 Conference: from biological data to knowledge to technological breakthroughsabstractTen years ago when Asia-Pacific Bioinformatics Network held the first International Conference on Bioinformatics (InCoB) in Bangkok its theme was North-South Networking. At that time InCoB aimed to provide biologists and bioinformatics researchers in the Asia-Pacific region a forum to meet, interact with, and disseminate knowledge about the burgeoning field of bioinformatics. Meanwhile InCoB has evolved into a major regional bioinformatics conference that attracts not only talented and established scientists from the region but increasingly also from East Asia, North America and Europe. Since 2006 InCoB yielded 114 articles in BMC Bioinformatics supplement issues that have been cited nearly 1,000 times to date. In part, these developments reflect the success of bioinformatics education and continuous efforts to integrate and utilize bioinformatics in biotechnology and biosciences in the Asia-Pacific region. A cross-section of research leading from biological data to knowledge and to technological applications, the InCoB2012 theme, is introduced in this editorial. Other highlights included sessions organized by the Pan-Asian Pacific Genome Initiative and a Machine Learning in Immunology competition. InCoB2013 is scheduled for September 18-21, 2013 at Suzhou, China. Christian Schönbach, Sissades Tongsima, Jonathan H. Chan, Vladimir Brusic, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 5 |
| 2011 | Towards big data science in the decade ahead from ten years of InCoB and the 1st ISCB-Asia Joint ConferenceabstractThe 2011 International Conference on Bioinformatics (InCoB) conference, which is the annual scientific conference of the Asia-Pacific Bioinformatics Network (APBioNet), is hosted by Kuala Lumpur, Malaysia, is co-organized with the first ISCB-Asia conference of the International Society for Computational Biology (ISCB). InCoB and the sequencing of the human genome are both celebrating their tenth anniversaries and InCoB's goalposts for the next decade, implementing standards in bioinformatics and globally distributed computational networks, will be discussed and adopted at this conference. Of the 49 manuscripts (selected from 104 submissions) accepted to BMC Genomics and BMC Bioinformatics conference supplements, 24 are featured in this issue, covering software tools, genome/proteome analysis, systems biology (networks, pathways, bioimaging) and drug discovery and design. Shoba Ranganathan, Christian Schönbach, Janet Kelso, Burkhard Rost, Sheila Nathan, Tin Wee Tan |
BMC Bioinform. | 6 |
| 2010 | InCoB2010 - 9th International Conference on Bioinformatics at Tokyo, Japan, September 26-28, 2010abstractThe International Conference on Bioinformatics (InCoB), the annual conference of the Asia-Pacific Bioinformatics Network (APBioNet), is hosted in one of countries of the Asia-Pacific region. The 2010 conference was awarded to Japan and has attracted more than one hundred high-quality research paper submissions. Thorough peer reviewing resulted in 47 (43.5%) accepted papers out of 108 submissions. Submissions from Japan, R.O. Korea, P.R. China, Australia, Singapore and U.S.A totaled 43.8% and contributed to 57.4% of accepted papers. Manuscripts originating from Taiwan and India added up to 42.8% of submissions and 28.3% of acceptances. The fifteen articles published in this BMC Bioinformatics supplement cover disease informatics, structural bioinformatics and drug design, biological databases and software tools, signaling pathways, gene regulatory and biochemical networks, evolution and sequence analysis. Christian Schönbach, Kenta Nakai, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 3 |
| 2010 | T3SEdb: data warehousing of virulence effectors secreted by the bacterial Type III Secretion SystemabstractBACKGROUND: Effectors of Type III Secretion System (T3SS) play a pivotal role in establishing and maintaining pathogenicity in the host and therefore the identification of these effectors is important in understanding virulence. However, the effectors display high level of sequence diversity, therefore making the identification a difficult process. There is a need to collate and annotate existing effector sequences in public databases to enable systematic analyses of these sequences for development of models for screening and selection of putative novel effectors from bacterial genomes that can be validated by a smaller number of key experiments. RESULTS: Herein, we present T3SEdb http://effectors.bic.nus.edu.sg/T3SEdb, a specialized database of annotated T3SS effector (T3SE) sequences containing 1089 records from 46 bacterial species compiled from the literature and public protein databases. Procedures have been defined for i) comprehensive annotation of experimental status of effectors, ii) submission and curation review of records by users of the database, and iii) the regular update of T3SEdb existing and new records. Keyword fielded and sequence searches (BLAST, regular expression) are supported for both experimentally verified and hypothetical T3SEs. More than 171 clusters of T3SEs were detected based on sequence identity comparisons (intra-cluster difference up to ~60%). Owing to this high level of sequence diversity of T3SEs, the T3SEdb provides a large number of experimentally known effector sequences with wide species representation for creation of effector predictors. We created a reliable effector prediction tool, integrated into the database, to demonstrate the application of the database for such endeavours. CONCLUSIONS: T3SEdb is the first specialised database reported for T3SS effectors, enriched with manual annotations that facilitated systematic construction of a reliable prediction model for identification of novel effectors. The T3SEdb represents a platform for inclusion of additional annotations of metadata for future developments of sophisticated effector prediction models for screening and selection of putative novel effectors from bacterial genomes/proteomes that can be validated by a small number of key experiments. Daniel Tay, Kunde Ramamoorthy Govindarajan, Asif M. Khan, Terenze Ong, Hanif M. Samad, Wei Soh, Minyan Tong, Tin Wee Tan |
BMC Bioinform. | 9 |
| 2009 | A comprehensive assessment of N-terminal signal peptides prediction methodsabstractBACKGROUND: Amino-terminal signal peptides (SPs) are short regions that guide the targeting of secretory proteins to the correct subcellular compartments in the cell. They are cleaved off upon the passenger protein reaching its destination. The explosive growth in sequencing technologies has led to the deposition of vast numbers of protein sequences necessitating rapid functional annotation techniques, with subcellular localization being a key feature. Of the myriad software prediction tools developed to automate the task of assigning the SP cleavage site of these new sequences, we review here, the performance and reliability of commonly used SP prediction tools. RESULTS: The available signal peptide data has been manually curated and organized into three datasets representing eukaryotes, Gram-positive and Gram-negative bacteria. These datasets are used to evaluate thirteen prediction tools that are publicly available. SignalP (both the HMM and ANN versions) maintains consistency and achieves the best overall accuracy in all three benchmarking experiments, ranging from 0.872 to 0.914 although other prediction tools are narrowing the performance gap. CONCLUSION: The majority of the tools evaluated in this study encounter no difficulty in discriminating between secretory and non-secretory proteins. The challenge clearly remains with pinpointing the correct SP cleavage site. The composite scoring schemes employed by SignalP may help to explain its accuracy. Prediction task is divided into a number of separate steps, thus allowing each score to tackle a particular aspect of the prediction. Khar Heng Choo, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 2 |
| 2009 | The implementation of e-learning tools to enhance undergraduate bioinformatics teaching and learning: a case study in the National University of SingaporeabstractBACKGROUND: The rapid advancement of computer and information technology in recent years has resulted in the rise of e-learning technologies to enhance and complement traditional classroom teaching in many fields, including bioinformatics. This paper records the experience of implementing e-learning technology to support problem-based learning (PBL) in the teaching of two undergraduate bioinformatics classes in the National University of Singapore. RESULTS: Survey results further established the efficiency and suitability of e-learning tools to supplement PBL in bioinformatics education. 63.16% of year three bioinformatics students showed a positive response regarding the usefulness of the Learning Activity Management System (LAMS) e-learning tool in guiding the learning and discussion process involved in PBL and in enhancing the learning experience by breaking down PBL activities into a sequential workflow. On the other hand, 89.81% of year two bioinformatics students indicated that their revision process was positively impacted with the use of LAMS for guiding the learning process, while 60.19% agreed that the breakdown of activities into a sequential step-by-step workflow by LAMS enhances the learning experience CONCLUSION: We show that e-learning tools are useful for supplementing PBL in bioinformatics education. The results suggest that it is feasible to develop and adopt e-learning tools to supplement a variety of instructional strategies in the future. Shen Jean Lim, Asif M. Khan, Mark De Silva, Kuan Siong Lim, Yongli Hu, Chay Hoon Tan, Tin Wee Tan |
BMC Bioinform. | 7 |
| 2009 | Bioinformatics in Malaysia: Hope, Initiative, Effort, Reality, and Challengesabstract10.1371/journal.pcbi.1000457 Azura M. H. Zeti, Mohd Shahir Shamsir, Khairina Tajul-Arifin, Amir F. Merican, Rahmah Mohamed, Sheila Nathan, Nor Muhammad Mahadi, Suhaimi Napis, Tin Wee Tan |
PLoS Comput. Biol. | 9 |
| 2008 | Automatic synchronization and distribution of biological databases and software over low-bandwidth networks among developing countriesabstractUNLABELLED: Bioinformatics involves the collection, organization and analysis of large amounts of biological data, using networks of computers and databases. Developing countries in the Asia-Pacific region are just moving into this new field of information-based biotechnology. However, the computational infrastructure and network bandwidths available in these countries are still at a basic level compared to that in developed countries. In this study, we assessed the utility of a BitTorrent-based Peer-to-Peer (btP2P) file distribution model for automatic synchronization and distribution of large amounts of biological data among developing countries. The initial country-level nodes in the Asia-Pacific region comprised Thailand, Korea and Singapore. The results showed a significant improvement in download performance using btP2P--three times faster overall download performance than conventional File Transfer Protocol (FTP). This study demonstrated the reliability of btP2P in the dissemination of continuously growing multi-gigabyte biological databases across the three Asia-Pacific countries. The download performance for btP2P can be further improved by including more nodes from other countries into the network. This suggests that the btP2P technology is appropriate for automatic synchronization and distribution of biological databases and software over low-bandwidth networks among developing countries in the Asia-Pacific region. AVAILABILITY: http://everest.bic.nus.edu.sg/p2p/ Unitsa Sangket, Amornrat Phongdara, Wilaiwan Chotigeat, Darran Nathan, Woo-Yeon Kim, Jong Bhak, Chumpol Ngamphiw, Sissades Tongsima, Asif M. Khan, Honghuang Lin, Tin Wee Tan |
Bioinform. | 11 |
| 2008 | Identification of human-to-human transmissibility factors in PB2 proteins of influenza A by large-scale mutual information analysisabstractBACKGROUND: The identification of mutations that confer unique properties to a pathogen, such as host range, is of fundamental importance in the fight against disease. This paper describes a novel method for identifying amino acid sites that distinguish specific sets of protein sequences, by comparative analysis of matched alignments. The use of mutual information to identify distinctive residues responsible for functional variants makes this approach highly suitable for analyzing large sets of sequences. To support mutual information analysis, we developed the AVANA software, which utilizes sequence annotations to select sets for comparison, according to user-specified criteria. The method presented was applied to an analysis of influenza A PB2 protein sequences, with the objective of identifying the components of adaptation to human-to-human transmission, and reconstructing the mutation history of these components. RESULTS: We compared over 3,000 PB2 protein sequences of human-transmissible and avian isolates, to produce a catalogue of sites involved in adaptation to human-to-human transmission. This analysis identified 17 characteristic sites, five of which have been present in human-transmissible strains since the 1918 Spanish flu pandemic. Sixteen of these sites are located in functional domains, suggesting they may play functional roles in host-range specificity. The catalogue of characteristic sites was used to derive sequence signatures from historical isolates. These signatures, arranged in chronological order, reveal an evolutionary timeline for the adaptation of the PB2 protein to human hosts. CONCLUSION: By providing the most complete elucidation to date of the functional components participating in PB2 protein adaptation to humans, this study demonstrates that mutual information is a powerful tool for comparative characterization of sequence sets. In addition to confirming previously reported findings, several novel characteristic sites within PB2 are reported. Sequence signatures generated using the characteristic sites catalogue characterize concisely the adaptation characteristics of individual isolates. Evolutionary timelines derived from signatures of early human influenza isolates suggest that characteristic variants emerged rapidly, and remained remarkably stable through subsequent pandemics. In addition, the signatures of human-infecting H5N1 isolates suggest that this avian subtype has low pandemic potential at present, although it presents more human adaptation components than most avian subtypes. Olivo Miotto, A. T. Heiny, Tin Wee Tan, J. Thomas August, Vladimir Brusic |
BMC Bioinform. | 3 |
| 2008 | Rule-based knowledge aggregation for large-scale protein sequence analysis of influenza A virusesabstractBACKGROUND: The explosive growth of biological data provides opportunities for new statistical and comparative analyses of large information sets, such as alignments comprising tens of thousands of sequences. In such studies, sequence annotations frequently play an essential role, and reliable results depend on metadata quality. However, the semantic heterogeneity and annotation inconsistencies in biological databases greatly increase the complexity of aggregating and cleaning metadata. Manual curation of datasets, traditionally favoured by life scientists, is impractical for studies involving thousands of records. In this study, we investigate quality issues that affect major public databases, and quantify the effectiveness of an automated metadata extraction approach that combines structural and semantic rules. We applied this approach to more than 90,000 influenza A records, to annotate sequences with protein name, virus subtype, isolate, host, geographic origin, and year of isolation. RESULTS: Over 40,000 annotated Influenza A protein sequences were collected by combining information from more than 90,000 documents from NCBI public databases. Metadata values were automatically extracted, aggregated and reconciled from several document fields by applying user-defined structural rules. For each property, values were recovered from >/=88.8% of records, with accuracy exceeding 96% in most cases. Because of semantic heterogeneity, each property required up to six different structural rules to be combined. Significant quality differences between databases were found: GenBank documents yield values more reliably than documents extracted from GenPept. Using a simple set of semantic rules and a reasoner, we reconstructed relationships between sequences from the same isolate, thus identifying 7640 isolates. Validation of isolate metadata against a simple ontology highlighted more than 400 inconsistencies, leading to over 3,000 property value corrections. CONCLUSION: To overcome the quality issues inherent in public databases, automated knowledge aggregation with embedded intelligence is needed for large-scale analyses. Our results show that user-controlled intuitive approaches, based on combination of simple rules, can reliably automate various curation tasks, reducing the need for manual corrections to approximately 5% of the records. Emerging semantic technologies possess desirable features to support today's knowledge aggregation tasks, with a potential to bring immediate benefits to this field. Olivo Miotto, Tin Wee Tan, Vladimir Brusic |
BMC Bioinform. | 2 |
| 2008 | Bioinformatics research in the Asia Pacific: a 2007 updateabstractWe provide a 2007 update on the bioinformatics research in the Asia-Pacific from the Asia Pacific Bioinformatics Network (APBioNet), Asia's oldest bioinformatics organisation set up in 1998. From 2002, APBioNet has organized the first International Conference on Bioinformatics (InCoB) bringing together scientists working in the field of bioinformatics in the region. This year, the InCoB2007 Conference was organized as the 6th annual conference of the Asia-Pacific Bioinformatics Network, on Aug. 27-30, 2007 at Hong Kong, following a series of successful events in Bangkok (Thailand), Penang (Malaysia), Auckland (New Zealand), Busan (South Korea) and New Delhi (India). Besides a scientific meeting at Hong Kong, satellite events organized are a pre-conference training workshop at Hanoi, Vietnam and a post-conference workshop at Nansha, China. This Introduction provides a brief overview of the peer-reviewed manuscripts accepted for publication in this Supplement. We have organized the papers into thematic areas, highlighting the growing contribution of research excellence from this region, to global bioinformatics endeavours. Shoba Ranganathan, Michael Gribskov, Tin Wee Tan |
BMC Bioinform. | 3 |
| 2008 | Emerging strengths in Asia Pacific bioinformaticsabstractThe 2008 annual conference of the Asia Pacific Bioinformatics Network (APBioNet), Asia's oldest bioinformatics organisation set up in 1998, was organized as the 7th International Conference on Bioinformatics (InCoB), jointly with the Bioinformatics and Systems Biology in Taiwan (BIT 2008) Conference, Oct. 20-23, 2008 at Taipei, Taiwan. Besides bringing together scientists from the field of bioinformatics in this region, InCoB is actively involving researchers from the area of systems biology, to facilitate greater synergy between these two groups. Marking the 10th Anniversary of APBioNet, this InCoB 2008 meeting followed on from a series of successful annual events in Bangkok (Thailand), Penang (Malaysia), Auckland (New Zealand), Busan (South Korea), New Delhi (India) and Hong Kong. Additionally, tutorials and the Workshop on Education in Bioinformatics and Computational Biology (WEBCB) immediately prior to the 20th Federation of Asian and Oceanian Biochemists and Molecular Biologists (FAOBMB) Taipei Conference provided ample opportunity for inducting mainstream biochemists and molecular biologists from the region into a greater level of awareness of the importance of bioinformatics in their craft. In this editorial, we provide a brief overview of the peer-reviewed manuscripts accepted for publication herein, grouped into thematic areas. As the regional research expertise in bioinformatics matures, the papers fall into thematic areas, illustrating the specific contributions made by APBioNet to global bioinformatics efforts. Shoba Ranganathan, Wen-Lian Hsu, Ueng-Cheng Yang, Tin Wee Tan |
BMC Bioinform. | 4 |
| 2007 | Methods and protocols for prediction of immunogenic epitopesabstractT-cell recognition of peptide/major histocompatibility complex (MHC) is a prerequisite for cellular immunity. Recently, there has been an influx of bioinformatics tools to facilitate the identification of T-cell epitopes to specific MHC alleles. This article examines existing computational strategies for the study of peptide/MHC interactions. The most important bioinformatics tools and methods with relevance to the study of peptide/MHC interactions have been reviewed. We have also provided guidelines for predicting antigenic peptides based on the availability of existing experimental data. Joo Chuan Tong, Tin Wee Tan, Shoba Ranganathan |
Briefings Bioinform. | 2 |
| 2007 | In silico grouping of peptide/HLA class I complexes using structural interaction characteristicsabstractMOTIVATION: Classification of human leukocyte antigen (HLA) proteins into supertypes underpins the development of epitope-based vaccines with wide population coverage. Current methods for HLA supertype definition, based on common structural features of HLA proteins and/or their functional binding specificities, leave structural interaction characteristics among different HLA supertypes with antigenic peptides unexplored. METHODS: We describe the use of structural interaction descriptors for the analysis of 68 peptide/HLA class I crystallographic structures. Interaction parameters computed include the number of intermolecular hydrogen bonds between each HLA protein and its corresponding bound peptide, solvent accessibility, gap volume and gap index. RESULTS: The structural interactions patterns of peptide/HLA class I complexes investigated herein vary among individual alleles and may be grouped in a supertype dependent manner. Using the proposed methodology, eight HLA class I supertypes were defined based on existing experimental crystallographic structures which largely overlaps (77% consensus) with the definitions by binding motifs. This mode of classification, which considers conformational information of both peptide and HLA proteins, provides an alternative to the characterization of supertypes using either peptide or HLA protein information alone. Joo Chuan Tong, Tin Wee Tan, Shoba Ranganathan |
Bioinform. | 2 |
| 2007 | CASVM: web server for SVM-based prediction of caspase substrates cleavage sitesabstractUNLABELLED: Caspases belong to a unique class of cysteine proteases which function as critical effectors of apoptosis, inflammation and other important cellular processes. Caspases cleave substrates at specific tetrapeptide sites after a highly conserved aspartic acid residue. Prediction of such cleavage sites will complement structural and functional studies on substrates cleavage as well as discovery of new substrates. We have recently developed a support vector machines (SVM) method to address this issue. Our algorithm achieved an accuracy ranging from 81.25 to 97.92%, making it one of the best methods currently available. CASVM is the web server implementation of our SVM algorithms, written in Perl and hosted on a Linux platform. The server can be used for predicting non-canonical caspase substrate cleavage sites. We have also included a relational database containing experimentally verified caspase substrates retrievable using accession IDs, keywords or sequence similarity. AVAILABILITY: http://www.casbase.org/casvm/index.html Lawrence J. K. Wee, Tin Wee Tan, Shoba Ranganathan |
Bioinform. | 2 |
| 2006 | Prediction of HLA-DQ3.2ß Ligands: evidence of multiple registers in class II binding peptidesabstractMOTIVATION: While processing of MHC class II antigens for presentation to helper T-cells is essential for normal immune response, it is also implicated in the pathogenesis of autoimmune disorders and hypersensitivity reactions. Sequence-based computational techniques for predicting HLA-DQ binding peptides have encountered limited success, with few prediction techniques developed using three-dimensional models. METHODS: We describe a structure-based prediction model for modeling peptide-DQ3.2beta complexes. We have developed a rapid and accurate protocol for docking candidate peptides into the DQ3.2beta receptor and a scoring function to discriminate binders from the background. The scoring function was rigorously trained, tested and validated using experimentally verified DQ3.2beta binding and non-binding peptides obtained from biochemical and functional studies. RESULTS: Our model predicts DQ3.2beta binding peptides with high accuracy [area under the receiver operating characteristic (ROC) curve A(ROC) > 0.90], compared with experimental data. We investigated the binding patterns of DQ3.2beta peptides and illustrate that several registers exist within a candidate binding peptide. Further analysis reveals that peptides with multiple registers occur predominantly for high-affinity binders. Joo Chuan Tong, Guanglan Zhang, Tin Wee Tan, J. Thomas August, Vladimir Brusic, Shoba Ranganathan |
Bioinform. | 3 |
| 2006 | Large-scale analysis of antigenic diversity of T-cell epitopes in dengue virusabstractBACKGROUND: Antigenic diversity in dengue virus strains has been studied, but large-scale and detailed systematic analyses have not been reported. In this study, we report a bioinformatics method for analyzing viral antigenic diversity in the context of T-cell mediated immune responses. We applied this method to study the relationship between short-peptide antigenic diversity and protein sequence diversity of dengue virus. We also studied the effects of sequence determinants on viral antigenic diversity. Short peptides, principally 9-mers were studied because they represent the predominant length of binding cores of T-cell epitopes, which are important for formulation of vaccines. RESULTS: Our analysis showed that the number of unique protein sequences required to represent complete antigenic diversity of short peptides in dengue virus is significantly smaller than that required to represent complete protein sequence diversity. Short-peptide antigenic diversity shows an asymptotic relationship to the number of unique protein sequences, indicating that for large sequence sets (approximately 200) the addition of new protein sequences has marginal effect to increasing antigenic diversity. A near-linear relationship was observed between the extent of antigenic diversity and the length of protein sequences, suggesting that, for the practical purpose of vaccine development, antigenic diversity of short peptides from dengue virus can be represented by short regions of sequences (approximately <100 aa) within viral antigens that are specific targets of immune responses (such as T-cell epitopes specific to particular human leukocyte antigen alleles). CONCLUSION: This study provides evidence that there are limited numbers of antigenic combinations in protein sequence variants of a viral species and that short regions of the viral protein are sufficient to capture antigenic diversity of T-cell epitopes. The approach described herein has direct application to the analysis of other viruses, in particular those that show high diversity and/or rapid evolution, such as influenza A virus and human immunodeficiency virus (HIV). Asif M. Khan, A. T. Heiny, Kenneth X. Lee, Kellathur N. Srinivasan, Tin Wee Tan, J. Thomas August, Vladimir Brusic |
BMC Bioinform. | 5 |
| 2006 | Establishing bioinformatics research in the Asia PacificabstractIn 1998, the Asia Pacific Bioinformatics Network (APBioNet), Asia's oldest bioinformatics organisation was set up to champion the advancement of bioinformatics in the Asia Pacific. By 2002, APBioNet was able to gain sufficient critical mass to initiate the first International Conference on Bioinformatics (InCoB) bringing together scientists working in the field of bioinformatics in the region. This year, the InCoB2006 Conference was organized as the 5 th annual conference of the Asia-Pacific Bioinformatics Network, on Dec. 18–20, 2006 in New Delhi, India, following a series of successful events in Bangkok (Thailand), Penang (Malaysia), Auckland (New Zealand) and Busan (South Korea). This Introduction provides a brief overview of the peer-reviewed manuscripts accepted for publication in this Supplement. It exemplifies a typical snapshot of the growing research excellence in bioinformatics of the region as we embark on a trajectory of establishing a solid bioinformatics research culture in the Asia Pacific that is able to contribute fully to the global bioinformatics community. Shoba Ranganathan, Martti Tammi, Michael Gribskov, Tin Wee Tan |
BMC Bioinform. | 4 |
| 2006 | Prediction of desmoglein-3 peptides reveals multiple shared T-cell epitopes in HLA DR4- and DR6- associated Pemphigus vulgarisabstractBACKGROUND: Pemphigus vulgaris (PV) is a severe autoimmune blistering skin disorder that is strongly associated with major histocompatibility complex class II alleles DRB1*0402 and DQB1*0503. The target antigen of PV, desmoglein 3 (Dsg3), is crucial for initiating T-cell response in early disease. Although a number of T-cell specificities within Dsg3 have been reported, the number is limited and the role of T-cells in the pathogenesis of PV remains poorly understood. We report here a structure-based model for the prediction of peptide binding to DRB1*0402 and DQB1*0503. The scoring functions were rigorously trained, tested and validated using experimentally verified peptide sequences. RESULTS: High predictivity is obtained for both DRB1*0402 (r2 = 0.90, s = 1.20 kJ/mol, q2 = 0.82, s(press) = 1.61 kJ/mol) and DQB1*0503 (r2 = 0.95, s = 1.20 kJ/mol, q2 = 0.75, s(press) = 2.15 kJ/mol) models, compared to experimental data. We investigated the binding patterns of Dsg3 peptides and illustrate the existence of multiple immunodominant epitopes that may be responsible for both disease initiation and propagation in PV. Further analysis reveals that DRB1*0402 and DQB1*0503 may share similar specificities by binding peptides at different binding registers, thus providing a molecular mechanism for the dual HLA association observed in PV. CONCLUSION: Collectively, the results of this study provide interesting new insights into the pathology of PV. This is the first report illustrating high-level of cross-reactivity between both PV-implicated alleles, DRB1*0402 and DQB1*0503, as well as the existence of a potentially large number of T-cell epitopes throughout the entire Dsg3 extracellular domain (ECD) and transmembrane region. Our results reveal that DR4 and DR6 PV may initiate in the ECD and transmembrane region respectively, with implications for immunotherapeutic strategies for the treatment of this autoimmune disease. Joo Chuan Tong, Tin Wee Tan, Animesh A. Sinha, Shoba Ranganathan |
BMC Bioinform. | 2 |
| 2006 | SVM-based prediction of caspase substrate cleavage sitesabstractBACKGROUND: Caspases belong to a class of cysteine proteases which function as critical effectors in apoptosis and inflammation by cleaving substrates immediately after unique sites. Prediction of such cleavage sites will complement structural and functional studies on substrates cleavage as well as discovery of new substrates. Recently, different computational methods have been developed to predict the cleavage sites of caspase substrates with varying degrees of success. As the support vector machines (SVM) algorithm has been shown to be useful in several biological classification problems, we have implemented an SVM-based method to investigate its applicability to this domain. RESULTS: A set of unique caspase substrates cleavage sites were obtained from literature and used for evaluating the SVM method. Datasets containing (i) the tetrapeptide cleavage sites, (ii) the tetrapeptide cleavage sites, augmented by two adjacent residues, P1' and P2' amino acids and (iii) the tetrapeptide cleavage sites with ten additional upstream and downstream flanking sequences (where available) were tested. The SVM method achieved an accuracy ranging from 81.25% to 97.92% on independent test sets. The SVM method successfully predicted the cleavage of a novel caspase substrate and its mutants. CONCLUSION: This study presents an SVM approach for predicting caspase substrate cleavage sites based on the cleavage sites and the downstream and upstream flanking sequences. The method shows an improvement over existing methods and may be useful for predicting hitherto undiscovered cleavage sites. Lawrence J. K. Wee, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 2 |
| 2005 | Extraction by Example: Induction of Structural Rules for the Analysis of Molecular Sequence Data from Heterogeneous Sources
Olivo Miotto, Tin Wee Tan, Vladimir Brusic |
IDEAL | 2 |
| 2005 | SPdb - a signal peptide databaseabstractBACKGROUND: The signal peptide plays an important role in protein targeting and protein translocation in both prokaryotic and eukaryotic cells. This transient, short peptide sequence functions like a postal address on an envelope by targeting proteins for secretion or for transfer to specific organelles for further processing. Understanding how signal peptides function is crucial in predicting where proteins are translocated. To support this understanding, we present SPdb signal peptide database http://proline.bic.nus.edu.sg/spdb, a repository of experimentally determined and computationally predicted signal peptides. RESULTS: SPdb integrates information from two sources (a) Swiss-Prot protein sequence database which is now part of UniProt and (b) EMBL nucleotide sequence database. The database update is semi-automated with human checking and verification of the data to ensure the correctness of the data stored. The latest release SPdb release 3.2 contains 18,146 entries of which 2,584 entries are experimentally verified signal sequences; the remaining 15,562 entries are either signal sequences that fail to meet our filtering criteria or entries that contain unverified signal sequences. CONCLUSION: SPdb is a manually curated database constructed to support the understanding and analysis of signal peptides. SPdb tracks the major updates of the two underlying primary databases thereby ensuring that its information remains up-to-date. Khar Heng Choo, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 2 |
| 2004 | Bio-Mirror project for public bio-data distributionabstractTimely worldwide distribution of biosequence and bioinformatics data depends on high performance networking and advances in Internet transport methods. The Bio-Mirror project focuses on providing up-to-date distribution of this rapidly growing and changing data. It offers FTP, Web and Rsync access to many high-volume databanks from several sites around the world. Experiments with data grids and other methods offer future improvements in biology data distribution. Don Gilbert, Yoshihiro Ugawa, Markus Buchhorn, Tin Wee Tan, Akira Mizushima, Hyunchul Kim, Kilnam Chon, Seyeon Weon, Juncai Ma, Yoshihiro Ichiyanagi, Der-Ming Liou, Somnuk Keretho, Suhaimi Napis |
Bioinform. | 4 |
| 2004 | DEDB: a database of Drosophila melanogaster exons in splicing graph formabstractBACKGROUND: A wealth of quality genomic and mRNA/EST sequences in recent years has provided the data required for large-scale genome-wide analysis of alternative splicing. We have capitalized on this by constructing a database that contains alternative splicing information organized as splicing graphs, where all transcripts arising from a single gene are collected, organized and classified. The splicing graph then serves as the basis for the classification of the various types of alternative splicing events. DESCRIPTION: DEDB http://proline.bic.nus.edu.sg/dedb/index.html is a database of Drosophila melanogaster exons obtained from FlyBase arranged in a splicing graph form that permits the creation of simple rules allowing for the classification of alternative splicing events. Pfam domains were also mapped onto the protein sequences allowing users to access the impact of alternative splicing events on domain organization. CONCLUSIONS: DEDB's catalogue of splicing graphs facilitates genome-wide classification of alternative splicing events for genome analysis. The splicing graph viewer brings together genome, transcript, protein and domain information to facilitate biologists in understanding the implications of alternative splicing. Bernett Lee, Tin Wee Tan, Shoba Ranganathan |
BMC Bioinform. | 2 |
| 2003 | MPID: MHC-Peptide Interaction Database for sequence-structure-function information on peptides binding to MHC moleculesabstractSUMMARY: Binding of short antigenic peptides to Major histocompatibility complex (MHC) proteins is the first step in T-cell mediated immune response. To understand the structural principles governing MHC-specific peptide recognition and binding, we have developed the MHC-Peptide Interaction Database (MPID), containing sequence-structure-function information. MPID (version 1.2) contains curated x-ray crystallographic data on 86 MHC peptide complexes, with precomputed interaction parameters (solvent accessibility, hydrogen bonds, gap volume and gap index). A user-friendly web interface and query tools will facilitate the development of predictive algorithms for MHC-peptide binding from a structural viewpoint. AVAILABILITY: Freely accessible from http://surya.bic.nus.edu.sg/mpid. Kunde Ramamoorthy Govindarajan, Pandjassarame Kangueane, Tin Wee Tan, Shoba Ranganathan |
Bioinform. | 3 |
| 2003 | XdomView: protein domain and exon position visualizationabstractSUMMARY: The relationship between intron distribution in the eukaryotic gene and protein structural elements is essential for understanding the origin and evolution of genes. XdomView is a web-based viewer mapping protein structural domains and intron positions in eukaryotic homologues to its tertiary structure. The association of sequence signals to 3D structure in XdomView provides a valuable visualization environment for eukaryotic gene organization, gene evolution, protein folding and protein structure classification. AVAILABILITY: Freely available from http://surya.bic.nus.edu.sg/xdom. Vivek Gopalan, Tin Wee Tan, Shoba Ranganathan |
Bioinform. | 2 |
| 2001 | Generation of a database containing discordant intron positions in eukaryotic genes (MIDB)abstractMOTIVATION: Intron sliding is the relocation of intron-exon boundaries over short distances and is often also referred to as intron slippage or intron migration or intron drift. We have generated a database containing discordant intron positions in homologous genes (MIDB--Mismatched Intron DataBase). Discordant intron positions are those that are either closely located in homologous genes (within a window of 10 nucleotides) or an intron position that is present in one gene but not in any of its homologs. The MIDB database aims at systematically collecting information about mismatched introns in the genes from GenBank and organizing it into a form useful for understanding the genomics and dynamics of introns thereby helping understand the evolution of genes. RESULTS: Intron displacement or sliding is critically important for explaining the present distribution of introns among orthologous and paralogous genes. MIDB allows examining of intron movements and allows mapping of intron positions from homologous proteins onto a single sequence. The database is of potential use for molecular biologists in general and for researchers who are interested in gene evolution and eukaryotic gene structure. Partial analysis of this database allowed us to identify a few putative cases of intron sliding. AVAILABILITY: http://intron.bic.nus.edu.sg/midb/midb.html Meena Kishore Sakharkar, Tin Wee Tan, Sandro J. de Souza |
Bioinform. | 2 |
| 2001 | Cladogramer: incorporating haplotype frequency into cladogram analysisabstractAbstract Summary: We implement a program that incorporates polymorphic sites data, haplotype frequency arrays, and other factors, into cladogram estimation. Availability: It is available at http://sdmc.krdl.org.sg:8080/~lxzhang/cladogramer Contact: [email protected] Louxin Zhang, Chew-Kiat Heng, Tin Wee Tan |
Bioinform. | 3 |
| 2000 | IE-Kb: intron exon knowledge baseabstractAbstract Summary: IE-Kb (Intron Exon-Knowledge base) illustrates the intron–exon dynamics in eukaryotic genes. We have developed three different knowledge sets, namely ‘Non-redundant ExInt’, ‘Non-redundant Pfam-ExInt complement’ and ‘Non-redundant GenBank eukaryotic subdivisional sets’ to understand this phenomenon. Statistical analysis is performed on each knowledge set and the results are made available online. The entries in knowledge sets are ranked based on their intron length, exon length and protein length with relational hyper-links to the corresponding intron phase, intron position, intron sequence, gene definition and parent GenBank entry. Availability: http://intron.bic.nus.edu.sg/iekb/iekb.html Contact: [email protected] * To whom correspondence should be addressed. Meena Kishore Sakharkar, Pandjassarame Kangueane, Tong W. Woon, Tin Wee Tan, Prasanna R. Kolatkar, Manyuan Long, Sandro J. de Souza |
Bioinform. | 4 |
| 1996 | Would Internet Meet Global Expectation?abstractMeeting global expectation of the Internet is like aiming at multiple rapidly moving targets. Today, the Internet has met with international acceptance in terms of global interconnectivity and the conduit for communications and information exchange. It has already become the de facto global information infrastructure (GII). But as it becomes so, expectations rapidly change and some are conflicting expectations. As the Internet evolves and as its connectivity ramifies, new emerging Internet technologies will continue to drive demand and escalate expectation, which in turn, drives more technological progress. Whether the Internet will meet the expectations of the global village may be a moot point, but there is no doubt that these are exciting times for us as we enter the new millennium. Tin Wee Tan |
COMPSAC | 1 |
| 1993 | A hypertext-like approach to navigating through the GCG sequence analysis packageabstractPrograms of the recently released Unix version of the Genetics Computing Group (GCG) sequence analysis package can now be accessed via a user-friendly hypertext-like navigation system, HYGCG. The resultant system organizes the diverse suite of programs into logical groups, and provides a guide and explanation of commands. In addition, for users unfamiliar with the Unix operating system, the program also provides a similar interface to commonly used Unix commands. Options for personal customization and expansion to accommodate GCG extensions and other software are also provided. This system should be useful especially to the inexperienced or infrequent user as context-sensitive on-line help is provided within this simple and consistent approach. Written in the C language and using the curses and termcap libraries, the system is easily portable to most Unix environments and has been made freely available via anonymous file transfer protocol (FTP) through the Internet global computer network. No modification of the GCG package is needed. B. K. Kiong, Tin Wee Tan |
Comput. Appl. Biosci. | 2 |
| 1993 | A general UNIX interface for biocomputing and network information retrieval softwareabstractWe describe a UNIX program, HYBROW, which can integrate without modification a wide range of UNIX biocomputing and network information retrieval software. HYBROW works in conjunction with a separate set of ASCII files containing embedded hypertext-like links. The program operates like a hypertext browser featuring five basic links: file link, execute-only link, execute-display link, directory-browse link and field-filling link. Useful features of the interface may be developed using combinations of these links with simple shell scripts and examples of these are briefly described. The system manager who supports biocomputing users should find the program easy to maintain, and useful in assisting new and infrequent users; it is also simple to incorporate new programs. Moreover, the individual user can customize the interface, create dynamic menus, hypertext a document, invoke shell scripts and new programs simply with a basic understanding of the UNIX operating system and any text editor. This program was written in C language and uses the UNIX curses and termcap libraries. It is freely available as a tar compressed file (by anonymous FTP from nuscc.nus.sg). B. K. Kiong, Tin Wee Tan |
Comput. Appl. Biosci. | 2 |