EDBT 2026 Demo / reviewers in the wild / expert
Weichuan Yu
dblp:34/4460
· DBLP profile ↗
61ranked-venue papers
13as first author
8since 2021 · last 2026
0000-0002-5510-6916ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 12 · 9 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | gSV: a general structural variant detector using the third-generation sequencing dataabstractStructural variants (SVs) are major contributors to genome diversity and disease susceptibility, particularly in cancer. Although third-generation sequencing technologies have substantially improved SV detection sensitivity, accurate detection of complex SVs remains challenging due to fragmented and heterogeneous alignment signals, as well as the dependence of many existing methods on predefined variant models. In this paper, we propose gSV, a general SV detector that integrates alignment-based and assembly-based approaches with the maximum exact match strategy, with particular emphasis on resolving SVs with complex or atypical alignment signatures. Without predefined assumptions about SV types, gSV captures diverse variant signals, enabling the detection of SVs that are usually missed by conventional tools. Benchmarking using both simulated datasets and real long-read sequencing data demonstrates that gSV achieves improved sensitivity and overall detection performance compared with current state-of-the-art SV callers, particularly for simple and complex SV events with complex alignment patterns. Unique SV discoveries in four breast cancer cell lines, particularly in cancer-associated genes, demonstrate the potential biological relevance of gSV-enabled discoveries. Furthermore, analysis of a breast cancer cohort from the Chinese population highlights the utility of gSV for population-scale genomic studies. Collectively, gSV provides a unified framework for comprehensive SV discovery in both research and clinical genomics settings. Jingyu Hao, Jiandong Shi, Sheng Lian, Zhen Zhang 0016, Yongyi Luo, Taobo Hu, Toyotaka Ishibashi, De-Peng Wang, Xiaodan Fan, Weichuan Yu |
Briefings Bioinform. | 11 |
| 2026 | RoBep: a region-oriented deep learning model for B-cell epitope predictionabstractMOTIVATION: Accurate in silico identification of B-cell epitope residues is crucial for antibody design and structure-guided vaccine development. Although recent protein language models and structure-aware methods can capture spatial information of tertiary structure when generating residue embeddings, most existing epitope predictors use these embeddings to perform classification for individual residues one by one, without enforcing spatial continuity for reported epitope residues. Such methods often result in biologically implausible predictions because B-cell epitope residues always cluster together on the antigen surface. RESULTS: We present RoBep, a region-oriented B-cell epitope predictor that explicitly models the spatial clustering of epitope residues. RoBep introduces a novel region constraint mechanism and combines the advanced protein language model ESM-Cambrian with an equivariant graph neural network. Our method outperforms existing structure-based methods on the benchmark dataset, demonstrating improvements of 26%, 45%, 13%, and 43% in F1, Matthews correlation coefficient, area under the precision-recall curve, and AUROC0.1, respectively. In addition to residue-level predictions, RoBep can also provide antibody-antigen binding regions. Importantly, the predicted epitope residues are ensured to be spatially compact, enhancing biological plausibility and practical relevance for immunotherapeutic design. AVAILABILITY AND IMPLEMENTATION: A user-friendly website for using RoBep is provided at https://huggingface.co/spaces/NielTT/RoBep. All datasets, source code used in this work, and implementation instructions of the website are publicly available at https://github.com/YitaoXU/RoBep. Guanyun Wei, Jingying Zhou, Yuanhua Huang, Weichuan Yu, Zhixiang Lin, Xiaodan Fan |
Bioinform. | 5 |
| 2025 | A Comparative Performance Study of Protein Structure Alignment Tools on Homology Detection with Biological RelevanceabstractIn this study, we investigate the performance of nine protein structure alignment tools on three tasks: superposition derivation, structural classification, and function inference. These tools include (1) traditional sequential methods using both 3D and 2D structure representations, (2) non-sequential methods, (3) flexible methods, and (4) deep-learning methods. We also include two canonical sequence alignment methods the Needleman-Wunsch algorithm and BLASTp, as baselines. We have the following results: In deriving superposition, deep learning methods, which focus on library search, produce less accurate patterns than traditional pairwise methods, but are more precise than sequence alignment; In structural classification, structure-based methods are much better than sequence alignment, with DALI performing the best; In function inference, structure-based methods recover more functional homology hits than sequence alignment, with KPAX recovering the most. We identify key factors that affect performance, including scoring metrics, the ability to capture side-chain information, and partial alignment mechanisms. Notably, deep learning methods, such as Foldseek, perform well on large-scale queries, in terms of both speed and accuracy. Our findings highlight the importance of integrating structural and sequence data for performance benefits, providing insight into the development of protein structure alignment tools in the future. Code and data are available at https://github.com/georgedashen/StructAlign Zhuoyang Chen, Xuechen Zhang 0003, Weichuan Yu, Qiong Luo 0001 |
BIBM | 3 |
| 2025 | epLSAP-Align: a non-sequential protein structural alignment solver with entropy-regularized partial linear sum assignment problem formulationabstractMOTIVATION: The three-dimensional protein tertiary structure alignment is a fundamental problem that seeks insights into functions and evolution. Previous structure alignment algorithms have adopted the sequential assumption and used dynamic programming solvers. However, many distantly related structures exhibit non-sequential similarities, and non-sequential alignment tools are less efficient and accurate than sequential ones. In this paper, we formulate the non-sequential alignment as the Entropy-regularized Partial Linear Sum Assignment Problem (epLSAP) and propose a solver based on Sinkhorn algorithms, referred to as epLSAP-Align. RESULTS: Compared with existing non-sequential alignment solvers, our epLSAP-Align can explicitly model the gap penalty, efficiently achieve global optimality and balance coverage and fidelity. We show that epLSAP-Align can be easily integrated into the existing frameworks, such as TM-align and MICAN, resulting in the non-sequential alignment tool epLSAP-TM and epLSAP-MICAN, respectively. Both epLSAP-TM and epLSAP-MICAN achieve better performance than the existing non-sequential alignment tools in terms of biologically meaningful structure overlaps on two sequential alignment test sets MALIDUP and MALISAM, and four non-sequential alignment test sets MALIDUP-ns, MALISAM-ns, 64-difficult-case and RIPC datasets. Also, compared with the most recent non-sequential alignment tool USalign2, our epLSAP-TM is at least 22% faster under the same setting. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/xzhangem/epLSAP-align. Xuechen Zhang 0003, Zhuoyang Chen, Qiong Luo 0001, Longjun Wu, Weichuan Yu |
Bioinform. | 6 |
| 2024 | A Mixed Integer Linear Program for Post-translational Modification CharacterizationabstractCharacterizing post-translational modifications (PTMs) is crucial due to their vital role in regulating cellular activities. Current database search methods for peptide identification and PTM characterization face the challenge of exponentially increasing number of PTM combinations. Consequently, these methods resort to enumerating a small fraction of the PTM combinations. However, based on our analysis, more than 99% of the PTM combinations (up to three PTMs) are infeasible as their total mass does not match the mass shift between the precursor mass and theoretical mass. Enumerating those infeasible combinations is unnecessary and would lead to unsatisfactory performance. To address this issue, we propose a two-step method for characterizing multiple PTMs with a reduced search space of feasible PTM combinations only. The method first uses a mixed integer linear program (MILP) to collect feasible peptide sequences and PTM mass combinations. Then, these feasible peptides are exhaustively examined to find the optimal one. We conducted comprehensive experiments to evaluate the method’s performance. When applied to a soybean data set, our method has better performance than a representative tag-based search engine. The effectiveness of search space reduction and the sensitivity and precision of our method are demonstrated. Shengzhi Lai, Peize Zhao, Weichuan Yu |
BIBM | 4 |
| 2023 | Combining Tags of Various Lengths Benefits Peptide Identification in Bottom-up ProteomicsabstractPeptide identification provides key information for protein inference in bottom-up proteomics. Post-translational modifications (PTMs) are essential to understand cellular activities at the protein level. In current database search methods for peptide identification, precursor mass is a critical parameter to narrow down the search space. However, true peptides may be excluded from the search space if precursor masses are modified by PTMs. Thus, many researchers use peptide sequence segments called tags which are invariant to PTMs in database search. Shorter tags are more sensitive but less accurate, whereas longer tags are more accurate but less frequent. Current methods use tags of fixed lengths, ignoring the effect of different tag lengths. To address this issue, we propose to combine tags of various lengths to improve tag-based peptide identification methods. Using combined tags, true peptides are included in the search space in more cases, resulting in at least 35% and 49% more peptide identifications and PTM results compared to benchmark methods using the same quality control parameters. Shengzhi Lai, Weichuan Yu |
BIBM | 3 |
| 2023 | ECL 3.0: a sensitive peptide identification tool for cross-linking mass spectrometry data analysisabstractBACKGROUND: Cross-linking mass spectrometry (XL-MS) is a powerful technique for detecting protein-protein interactions (PPIs) and modeling protein structures in a high-throughput manner. In XL-MS experiments, proteins are cross-linked by a chemical reagent (namely cross-linker), fragmented, and then fed into a tandem mass spectrum (MS/MS). Cross-linkers are either cleavable or non-cleavable, and each type requires distinct data analysis tools. However, both types of cross-linkers suffer from imbalanced fragmentation efficiency, resulting in a large number of unidentifiable spectra that hinder the discovery of PPIs and protein conformations. To address this challenge, researchers have sought to improve the sensitivity of XL-MS through invention of novel cross-linking reagents, optimization of sample preparation protocols, and development of data analysis algorithms. One promising approach to developing new data analysis methods is to apply a protein feedback mechanism in the analysis. It has significantly improved the sensitivity of analysis methods in the cleavable cross-linking data. The application of the protein feedback mechanism to the analysis of non-cleavable cross-linking data is expected to have an even greater impact because the majority of XL-MS experiments currently employs non-cleavable cross-linkers. RESULTS: In this study, we applied the protein feedback mechanism to the analysis of both non-cleavable and cleavable cross-linking data and observed a substantial improvement in cross-link spectrum matches (CSMs) compared to conventional methods. Furthermore, we developed a new software program, ECL 3.0, that integrates two algorithms and includes a user-friendly graphical interface to facilitate wider applications of this new program. CONCLUSIONS: ECL 3.0 source code is available at https://github.com/yuweichuan/ECL-PF.git . A quick tutorial is available at https://youtu.be/PpZgbi8V2xI . Chen Zhou 0004, Shuaijian Dai, Shengzhi Lai, Yuanqiao Lin, Xuechen Zhang 0003, Weichuan Yu |
BMC Bioinform. | 7 |
| 2021 | Understanding the Limit of Open Search in the Identification of Peptides With Post-translational Modifications - A Simulation-Based StudyabstractPeptide identification from tandem mass spectrometry data is a fundamental task in computational proteomics. Traditional algorithms perform well when facing unmodified peptides. However, when peptides have post-translational modifications (PTMs), these methods cannot provide satisfactory results. Recently, open search methods have been proposed to identify peptides with PTMs. While the performance of these new methods is promising, the identification results vary greatly with respect to the quality of tandem mass spectra and the number of PTMs in peptides. This motivates us to systematically study the relationship between the performance of open search methods and the quality parameters of tandem mass spectrometry data as well as the number of PTMs in peptides. In this paper, we have proposed an analytical model derived from simulated data to describe the relationship between the probability of obtaining correct results and the spectrum quality as well as the number of PTMs. The proposed model is verified using 1,464,146 real experimental spectra. The consistent trend observed in both simulated data and real data reveals the necessary conditions to effectively apply open search methods. Source code of our study is available at http://bioinformatics.ust.hk/PST.html. Jiaan Dai, Fengchao Yu, Chen Zhou 0004, Weichuan Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Data imbalance in CRISPR off-target predictionabstractFor genome-wide CRISPR off-target cleavage sites (OTS) prediction, an important issue is data imbalance-the number of true OTS recognized by whole-genome off-target detection techniques is much smaller than that of all possible nucleotide mismatch loci, making the training of machine learning model very challenging. Therefore, computational models proposed for OTS prediction and scoring should be carefully designed and properly evaluated in order to avoid bias. In our study, two tools are taken as examples to further emphasize the data imbalance issue in CRISPR off-target prediction to achieve better sensitivity and specificity for optimized CRISPR gene editing. We would like to indicate that (1) the benchmark of CRISPR off-target prediction should be properly evaluated and not overestimated by considering data imbalance issue; (2) incorporation of efficient computational techniques (including ensemble learning and data synthesis techniques) can help to address the data imbalance issue and improve the performance of CRISPR off-target prediction. Taking together, we call for more efforts to address the data imbalance issue in CRISPR off-target prediction to facilitate clinical utility of CRISPR-based gene editing techniques. Yuli Gao, Guohui Chuai, Weichuan Yu, Shen Qu, Qi Liu 0019 |
Briefings Bioinform. | 3 |
| 2019 | Xolik: finding cross-linked peptides with maximum paired scores in linear timeabstractMotivation: Cross-linking technique coupled with mass spectrometry (MS) is widely used in the analysis of protein structures and protein-protein interactions. In order to identify cross-linked peptides from MS data, we need to consider all pairwise combinations of peptides, which is computationally prohibitive when the sequence database is large. To alleviate this problem, some heuristic screening strategies are used to reduce the number of peptide pairs during the identification. However, heuristic screening strategies may miss some true cross-linked peptides. Results: We directly tackle the combination challenge without using any screening strategies. With the data structure of double-ended queue, the proposed algorithm reduces the quadratic time complexity of exhaustive searching down to the linear time complexity. We implement the algorithm in a tool named Xolik. The running time of Xolik is validated using databases with different numbers of proteins. Experiments using synthetic and empirical datasets show that Xolik outperforms existing tools in terms of running time and statistical power. Availability and implementation: Source code and binaries of Xolik are freely available at http://bioinformatics.ust.hk/Xolik.html. Supplementary information: Supplementary data are available at Bioinformatics online. Jiaan Dai, Wei Jiang 0019, Fengchao Yu, Weichuan Yu |
Bioinform. | 4 |
| 2018 | A network approach to exploring the functional basis of gene-gene epistatic interactions in disease susceptibilityabstractMotivation: Individual genetic variants explain only a small fraction of heritability in some diseases. Some variants have weak marginal effects on disease risk, but their joint effects are significantly stronger when occurring together. Most studies on such epistatic interactions have focused on methods for identifying the interactions and interpreting individual cases, but few have explored their general functional basis. This was due to the lack of a comprehensive list of epistatic interactions and uncertainties in associating variants to genes. Results: We conducted a large-scale survey of published research articles to compile the first comprehensive list of epistatic interactions in human diseases with detailed annotations. We used various methods to associate these variants to genes to ensure robustness. We found that these genes are significantly more connected in protein interaction networks, are more co-expressed and participate more often in the same pathways. We demonstrate using the list to discover novel disease pathways. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Danny Kit-Sang Yip, Landon L. Chan, Iris K. Pang, Wei Jiang 0019, Nelson L. S. Tang, Weichuan Yu, Kevin Y. Yip |
Bioinform. | 6 |
| 2017 | What is the probability of replicating a statistically significant association in genome-wide association studies?abstractThe goal of genome-wide association studies (GWASs) is to discover genetic variants associated with diseases/traits. Replication is a common validation method in GWASs. We regard an association as true finding when it shows significance in both primary and replication studies. A question worth pondering is what is the probability of a primary association (i.e. a statistically significant association in the primary study) being validated in the replication study? This article systematically reviews the answers to this question from different points of view. As Bayesian methods can help us integrate out the uncertainty about the underlying effect of the primary association, we will mainly focus on the Bayesian view in this article. We refer the Bayesian replication probability as the replication rate (RR). We further describe an estimation method for RR, which makes use of the summary statistics from the primary study. We can use the estimated RR to determine the sample size of the replication study and to check the consistency between the results of the primary study and those of the replication study. We describe an R-package to estimate and apply RR in GWASs. Simulation and real data experiments show that the estimated RR has good prediction and calibration performance. We also use these data to demonstrate the usefulness of RR. The R-package is available at http://bioinformatics.ust.hk/RRate.html. Wei Jiang 0019, Jing-Hao Xue, Weichuan Yu |
Briefings Bioinform. | 3 |
| 2017 | Controlling the joint local false discovery rate is more powerful than meta-analysis methods in joint analysis of summary statistics from multiple genome-wide association studiesabstractMotivation: In genome-wide association studies (GWASs) of common diseases/traits, we often analyze multiple GWASs with the same phenotype together to discover associated genetic variants with higher power. Since it is difficult to access data with detailed individual measurements, summary-statistics-based meta-analysis methods have become popular to jointly analyze datasets from multiple GWASs. Results: In this paper, we propose a novel summary-statistics-based joint analysis method based on controlling the joint local false discovery rate (Jlfdr). We prove that our method is the most powerful summary-statistics-based joint analysis method when controlling the false discovery rate at a certain level. In particular, the Jlfdr-based method achieves higher power than commonly used meta-analysis methods when analyzing heterogeneous datasets from multiple GWASs. Simulation experiments demonstrate the superior power of our method over meta-analysis methods. Also, our method discovers more associations than meta-analysis methods from empirical datasets of four phenotypes. Availability and Implementation: The R-package is available at: http://bioinformatics.ust.hk/Jlfdr.html . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Wei Jiang 0019, Weichuan Yu |
Bioinform. | 2 |
| 2016 | GBOOST 2.0: A GPU-based tool for detecting gene-gene interactions with covariates adjustment in genome-wide association studiesabstractDetecting gene-gene interaction patterns is important to reveal associations between genotype and complex diseases. This task, however, is computationally challenging. For example, in order to exhaustively detect interactions of 1,000,000 single nucleotide polymorphisms (SNPs) genotyped from thousands of individuals, we need to carry out 5×1011statistical tests. To address the computational challenge, Wan et. al. [1] proposed a fast method named BOOST to exhaustively detect interactions of all SNP pairs. BOOST completes pairwise analysis of 360,000 SNPs in 60 hours on a standard desktop PC. As the interaction tests of SNP pairs are highly parallel, Yung et. al. [2] implemented the BOOST method in GPU and named it GBOOST. GBOOST usually takes about one and a half hours to finish genome-wide interaction analysis of a data set containing about 350,000 SNPs and 5,000 samples using Nvidia GeForce GTX 285 dispaly card. Wei Jiang 0019, Ronald C. Ma, Weichuan Yu |
BIBM | 4 |
| 2016 | FPGA implementation of the coupled filtering methodabstractIn ultrasound image analysis, speckle tracking methods are widely applied to study the elasticity of body tissue. However, “feature-motion decorrelation” still remains as a challenge for speckle tracking methods. Recently, a coupled filtering method was proposed to accurately estimate strain values when the tissue deformation is large. The major drawback of the new method is its high computational complexity. Even the GPU-based program requires a few hours to finish the analysis. In this paper, we propose an FPGA-based implementation for further acceleration. The capability of FPGAs on handling different image processing components in this method is discussed. The algorithm is reformulated to build a highly efficient pipeline on FPGA. The final implementation on a Xilinx Virtex-7 FPGA is 15 times faster than the GPU implementation on two NVIDIA graphic cards (GeForce GTX 580). Tianzhu Liang, Philip K. T. Mok, Weichuan Yu |
BIBM | 4 |
| 2016 | ECL: an exhaustive search tool for the identification of cross-linked peptides using whole databaseabstractBACKGROUND: Chemical cross-linking combined with mass spectrometry (CX-MS) is a high-throughput approach to studying protein-protein interactions. The number of peptide-peptide combinations grows quadratically with respect to the number of proteins, resulting in a high computational complexity. Widely used methods including xQuest (Rinner et al., Nat Methods 5(4):315-8, 2008; Walzthoeni et al., Nat Methods 9(9):901-3, 2012), pLink (Yang et al., Nat Methods 9(9):904-6, 2012), ProteinProspector (Chu et al., Mol Cell Proteomics 9:25-31, 2010; Trnka et al., 13(2):420-34, 2014) and Kojak (Hoopmann et al., J Proteome Res 14(5):2190-198, 2015) avoid searching all peptide-peptide combinations by pre-selecting peptides with heuristic approaches. However, pre-selection procedures may cause missing findings. The most intuitive approach is searching all possible candidates. A tool that can exhaustively search a whole database without any heuristic pre-selection procedure is therefore desirable. RESULTS: We have developed a cross-linked peptides identification tool named ECL. It can exhaustively search a whole database in a reasonable period of time without any heuristic pre-selection procedure. Tests showed that searching a database containing 5200 proteins took 7 h. ECL identified more non-redundant cross-linked peptides than xQuest, pLink, and ProteinProspector. Experiments showed that about 30 % of these additional identified peptides were not pre-selected by Kojak. We used protein crystal structures from the protein data bank to check the intra-protein cross-linked peptides. Most of the distances between cross-linking sites were smaller than 30 Å. CONCLUSIONS: To the best of our knowledge, ECL is the first tool that can exhaustively search all candidates in cross-linked peptides identification. The experiments showed that ECL could identify more peptides than xQuest, pLink, and ProteinProspector. A further analysis indicated that some of the additional identified results were thanks to the exhaustive search. Fengchao Yu, Weichuan Yu |
BMC Bioinform. | 3 |
| 2016 | Erratum to "On Feature Motion Decorrelation in Ultrasound Speckle Tracking"abstractIn the above paper (ibid., IEEE Trans. Med. Imag., vol. 32, no. 2, pp. 435-448, Feb. 2013), Section VII, the second sentence in the last paragraph should be corrected as "The program takes about 5 hours for 2-D ultrasound images and is more than 9 times faster than the CPU-based program in a standard PC." Tianzhu Liang, Ling Sing Yung, Weichuan Yu |
IEEE Trans. Medical Imaging | 3 |
| 2015 | PBOOST: a GPU-based tool for parallel permutation tests in genome-wide association studiesabstractMOTIVATION: The importance of testing associations allowing for interactions has been demonstrated by Marchini et al. (2005). A fast method detecting associations allowing for interactions has been proposed by Wan et al. (2010a). The method is based on likelihood ratio test with the assumption that the statistic follows the χ(2) distribution. Many single nucleotide polymorphism (SNP) pairs with significant associations allowing for interactions have been detected using their method. However, the assumption of χ(2) test requires the expected values in each cell of the contingency table to be at least five. This assumption is violated in some identified SNP pairs. In this case, likelihood ratio test may not be applicable any more. Permutation test is an ideal approach to checking the P-values calculated in likelihood ratio test because of its non-parametric nature. The P-values of SNP pairs having significant associations with disease are always extremely small. Thus, we need a huge number of permutations to achieve correspondingly high resolution for the P-values. In order to investigate whether the P-values from likelihood ratio tests are reliable, a fast permutation tool to accomplish large number of permutations is desirable. RESULTS: We developed a permutation tool named PBOOST. It is based on GPU with highly reliable P-value estimation. By using simulation data, we found that the P-values from likelihood ratio tests will have relative error of >100% when 50% cells in the contingency table have expected count less than five or when there is zero expected count in any of the contingency table cells. In terms of speed, PBOOST completed 10(7) permutations for a single SNP pair from the Wellcome Trust Case Control Consortium (WTCCC) genome data (Wellcome Trust Case Control Consortium, 2007) within 1 min on a single Nvidia Tesla M2090 device, while it took 60 min in a single CPU Intel Xeon E5-2650 to finish the same task. More importantly, when simultaneously testing 256 SNP pairs for 10(7) permutations, our tool took only 5 min, while the CPU program took 10 h. By permuting on a GPU cluster consisting of 40 nodes, we completed 10(12) permutations for all 280 SNP pairs reported with P-values smaller than 1.6 × 10⁻¹² in the WTCCC datasets in 1 week. AVAILABILITY AND IMPLEMENTATION: The source code and sample data are available at http://bioinformatics.ust.hk/PBOOST.zip. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Guangyuan Yang, Wei Jiang 0019, Qiang Yang 0001, Weichuan Yu |
Bioinform. | 4 |
| 2014 | Piecewise-constant and low-rank approximation for identification of recurrent copy number variationsabstractMOTIVATION: The post-genome era sees urgent need for more novel approaches to extracting useful information from the huge amount of genetic data. The identification of recurrent copy number variations (CNVs) from array-based comparative genomic hybridization (aCGH) data can help understand complex diseases, such as cancer. Most of the previous computational methods focused on single-sample analysis or statistical testing based on the results of single-sample analysis. Finding recurrent CNVs from multi-sample data remains a challenging topic worth further study. RESULTS: We present a general and robust method to identify recurrent CNVs from multi-sample aCGH profiles. We express the raw dataset as a matrix and demonstrate that recurrent CNVs will form a low-rank matrix. Hence, we formulate the problem as a matrix recovering problem, where we aim to find a piecewise-constant and low-rank approximation (PLA) to the input matrix. We propose a convex formulation for matrix recovery and an efficient algorithm to globally solve the problem. We demonstrate the advantages of PLA compared with alternative methods using synthesized datasets and two breast cancer datasets. The experimental results show that PLA can successfully reconstruct the recurrent CNV patterns from raw data and achieve better performance compared with alternative methods under a wide range of scenarios. AVAILABILITY AND IMPLEMENTATION: The MATLAB code is available at http://bioinformatics.ust.hk/pla.zip. Xiaowei Zhou 0001, Jiming Liu 0001, Weichuan Yu |
Bioinform. | 4 |
| 2013 | Active Contours with Group SimilarityabstractActive contours are widely used in image segmentation. To cope with missing or misleading features in images, researchers have introduced various ways to model the prior of shapes and use the prior to constrain active contours. However, the shape prior is usually learnt from a large set of annotated data, which is not always accessible in practice. Moreover, it is often doubted that the existing shapes in the training set will be sufficient to model the new instance in the testing image. In this paper, we propose to use the group similarity of object shapes in multiple images as a prior to aid segmentation, which can be interpreted as an unsupervised approach of shape prior modeling. We show that the rank of the matrix consisting of multiple shapes is a good measure of the group similarity of the shapes, and the nuclear norm minimization is a simple and effective way to impose the proposed constraint on existing active contour models. Moreover, we develop a fast algorithm to solve the proposed model by using the accelerated proximal method. Experiments using echocardiographic image sequences acquired from acute canine experiments demonstrate that the proposed method can consistently improve the performance of active contour models and increase the robustness against image defects such as missing boundaries. Xiaowei Zhou 0001, James S. Duncan, Weichuan Yu |
CVPR | 4 |
| 2013 | Moving Object Detection by Detecting Contiguous Outliers in the Low-Rank RepresentationabstractObject detection is a fundamental step for automated video analysis in many vision applications. Object detection in a video is usually performed by object detectors or background subtraction techniques. Often, an object detector requires manually labeled examples to train a binary classifier, while background subtraction needs a training sequence that contains no objects to build a background model. To automate the analysis, object detection without a separate training phase becomes a critical task. People have tried to tackle this task by using motion information. But existing motion-based methods are usually limited when coping with complex scenarios such as nonrigid motion and dynamic background. In this paper, we show that the above challenges can be addressed in a unified framework named DEtecting Contiguous Outliers in the LOw-rank Representation (DECOLOR). This formulation integrates object detection and background learning into a single process of optimization, which can be solved by an alternating algorithm efficiently. We explain the relations between DECOLOR and other sparsity-based methods. Experiments on both simulated data and real sequences demonstrate that DECOLOR outperforms the state-of-the-art approaches and it can work effectively on a wide range of complex scenarios. Xiaowei Zhou 0001, Can Yang 0002, Weichuan Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | HapBoost: A Fast Approach to Boosting Haplotype Association Analyses in Genome-Wide Association StudiesabstractGenome-wide association study (GWAS) has been successful in identifying genetic variants that are associated with complex human diseases. In GWAS, multilocus association analyses through linkage disequilibrium (LD), named haplotype-based analyses, may have greater power than single-locus analyses for detecting disease susceptibility loci. However, the large number of SNPs genotyped in GWAS poses great computational challenges in the detection of haplotype associations. We present a fast method named HapBoost for finding haplotype associations, which can be applied to quickly screen the whole genome. The effectiveness of HapBoost is demonstrated by using both synthetic and real data sets. The experimental results show that the proposed approach can achieve comparably accurate results while it performs much faster than existing methods. Can Yang 0002, Qiang Yang 0001, Hongyu Zhao 0003, Weichuan Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2013 | A Combinatorial Perspective of the Protein Inference ProblemabstractIn a shotgun proteomics experiment, proteins are the most biologically meaningful output. The success of proteomics studies depends on the ability to accurately and efficiently identify proteins. Many methods have been proposed to facilitate the identification of proteins from peptide identification results. However, the relationship between protein identification and peptide identification has not been thoroughly explained before. In this paper, we devote ourselves to a combinatorial perspective of the protein inference problem. We employ combinatorial mathematics to calculate the conditional protein probabilities (protein probability means the probability that a protein is correctly identified) under three assumptions, which lead to a lower bound, an upper bound, and an empirical estimation of protein probabilities, respectively. The combinatorial perspective enables us to obtain an analytical expression for protein inference. Our method achieves comparable results with ProteinProphet in a more efficient manner in experiments on two data sets of standard protein mixtures and two data sets of real samples. Based on our model, we study the impact of unique peptides and degenerate peptides (degenerate peptides are peptides shared by at least two proteins) on protein probabilities. Meanwhile, we also study the relationship between our model and ProteinProphet. We name our program ProteinInfer. Its Java source code, our supplementary document and experimental results are available at: >http://bioinformatics.ust.hk/proteininfer. Zengyou He, Weichuan Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2013 | Multisample aCGH Data Analysis via Total Variation and Spectral RegularizationabstractDNA copy number variation (CNV) accounts for a large proportion of genetic variation. One commonly used approach to detecting CNVs is array-based comparative genomic hybridization (aCGH). Although many methods have been proposed to analyze aCGH data, it is not clear how to combine information from multiple samples to improve CNV detection. In this paper, we propose to use a matrix to approximate the multisample aCGH data and minimize the total variation of each sample as well as the nuclear norm of the whole matrix. In this way, we can make use of the smoothness property of each sample and the correlation among multiple samples simultaneously in a convex optimization framework. We also developed an efficient and scalable algorithm to handle large-scale data. Experiments demonstrate that the proposed method outperforms the state-of-the-art techniques under a wide range of scenarios and it is capable of processing large data sets with millions of probes. Xiaowei Zhou 0001, Can Yang 0002, Hongyu Zhao 0003, Weichuan Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2013 | On Feature Motion Decorrelation in Ultrasound Speckle TrackingabstractSpeckle tracking methods refer to motion tracking methods based on speckle patterns in ultrasound images. They are commonly used in ultrasound based elasticity imaging techniques to reveal mechanical properties of tissues for clinical diagnosis. In speckle tracking, feature motion decorrelation exists when speckle patterns are not identical before and after tissue motion and deformation. Feature motion decorrelation violates the underlying assumption of most speckle tracking methods. Consequently, the estimation accuracy of current methods is greatly limited. In this paper, two types of speckle pattern variations, the geometric transformation and the intensity change of speckle patterns, are studied. We show that a coupled filtering method is able to compensate for both types of variations. It provides accurate strain estimations even when tissue deformation or rotation is extremely large. We also show that in most cases, an affine warping method that only compensates for the geometric transformation is able to achieve a similar performance as the coupled filtering method. Feature motion decorrelation in B-mode images is also studied. Finally, we show that in typical elastography studies, speckle tracking methods without modeling local shearing or rotation will fail when tissue deformation is large. Tianzhu Liang, Ling Sing Yung, Weichuan Yu |
IEEE Trans. Medical Imaging | 3 |
| 2012 | Automatic mitral leaflet tracking in echocardiography by outlier detection in the low-rank representationabstractTracking the mitral valve leaflet in an ultrasound sequence is a challenging task because of the poor image quality and fast and irregular leaflet motion. Previous algorithms usually applied standard segmentation methods based on edges, object intensity and anatomical information to segment the mitral leaflet in static frames. However, they are limited in practical applications due to the requirement of manual input for initialization or large annotated datasets for training. In this paper we present a completely automatic and unsupervised algorithm for mitral leaflet detection and tracking. We demonstrate that the image sequence of a cardiac cycle can be well approximated with a low-rank matrix, except for the mitral leaflet region with fast motion and tissue deformation. Based on this difference, we propose to track the mitral leaflet by detecting contiguous outliers in the low-rank representation. With this formulation, the leaflet is tracked using the motion cue, but the complicated motion computation is avoided. To the best of our knowledge, the proposed algorithm is the first unsupervised method for mitral leaflet tracking. The algorithm was tested on both 2D and 3D echocardiography, which achieved accurate segmentation with an average distance of 0.87 ± 0.42mm compared to the manual tracing. Xiaowei Zhou 0001, Can Yang 0002, Weichuan Yu |
CVPR | 3 |
| 2012 | Protein inference: a reviewabstractAssembling peptides identified from tandem mass spectra into a list of proteins, referred to as protein inference, is a critical step in proteomics research. Due to the existence of degenerate peptides and 'one-hit wonders', it is very difficult to determine which proteins are present in the sample. In this paper, we review existing protein inference methods and classify them according to the source of peptide identifications and the principle of algorithms. It is hoped that the readers will gain a good understanding of the current development in this field after reading this review and come up with new protein inference algorithms. Weichuan Yu, Zengyou He |
Briefings Bioinform. | 3 |
| 2012 | Comments on 'An empirical comparison of several recent epistatic interaction detection methods'abstractAbstract Contact: [email protected] Can Yang 0002, Weichuan Yu |
Bioinform. | 3 |
| 2012 | Peptide Reranking with Protein-Peptide Correspondence and Precursor Peak Intensity InformationabstractSearching tandem mass spectra against a protein database has been a mainstream method for peptide identification. Improving peptide identification results by ranking true Peptide-Spectrum Matches (PSMs) over their false counterparts leads to the development of various reranking algorithms. In peptide reranking, discriminative information is essential to distinguish true PSMs from false PSMs. Generally, most peptide reranking methods obtain discriminative information directly from database search scores or by training machine learning models. Information in the protein database and MS1 spectra (i.e., single stage MS spectra) is ignored. In this paper, we propose to use information in the protein database and MS1 spectra to rerank peptide identification results. To quantitatively analyze their effects to peptide reranking results, three peptide reranking methods are proposed: PPMRanker, PPIRanker, and MIRanker. PPMRanker only uses Protein-Peptide Map (PPM) information from the protein database, PPIRanker only uses Precursor Peak Intensity (PPI) information, and MIRanker employs both PPM information and PPI information. According to our experiments on a standard protein mixture data set, a human data set and a mouse data set, PPMRanker and MIRanker achieve better peptide reranking results than PetideProphet, PeptideProphet+NSP (number of sibling peptides) and a score regularization method SRPI. The source codes of PPMRanker, PPIRanker, and MIRanker, and all supplementary documents are available at our website: http://bioinformatics.ust.hk/pepreranking/. Alternatively, these documents can also be downloaded from: http://sourceforge.net/projects/pepreranking/. Zengyou He, Can Yang 0002, Weichuan Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2011 | Identifying disease-associated SNP clusters via contiguous outlier detectionabstractMOTIVATION: Although genome-wide association studies (GWAS) have identified many disease-susceptibility single-nucleotide polymorphisms (SNPs), these findings can only explain a small portion of genetic contributions to complex diseases, which is known as the missing heritability. A possible explanation is that genetic variants with small effects have not been detected. The chance is < 8 that a causal SNP will be directly genotyped. The effects of its neighboring SNPs may be too weak to be detected due to the effect decay caused by imperfect linkage disequilibrium. Moreover, it is still challenging to detect a causal SNP with a small effect even if it has been directly genotyped. RESULTS: In order to increase the statistical power when detecting disease-associated SNPs with relatively small effects, we propose a method using neighborhood information. Since the disease-associated SNPs account for only a small fraction of the entire SNP set, we formulate this problem as Contiguous Outlier DEtection (CODE), which is a discrete optimization problem. In our formulation, we cast the disease-associated SNPs as outliers and further impose a spatial continuity constraint for outlier detection. We show that this optimization can be solved exactly using graph cuts. We also employ the stability selection strategy to control the false positive results caused by imperfect parameter tuning. We demonstrate its advantage in simulations and real experiments. In particular, the newly identified SNP clusters are replicable in two independent datasets. AVAILABILITY: The software is available at: http://bioinformatics.ust.hk/CODE.zip. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Can Yang 0002, Xiaowei Zhou 0001, Qiang Yang 0001, Hong Xue 0001, Weichuan Yu |
Bioinform. | 6 |
| 2011 | GBOOST: a GPU-based tool for detecting gene-gene interactions in genome-wide case control studiesabstractMOTIVATION: Collecting millions of genetic variations is feasible with the advanced genotyping technology. With a huge amount of genetic variations data in hand, developing efficient algorithms to carry out the gene-gene interaction analysis in a timely manner has become one of the key problems in genome-wide association studies (GWAS). Boolean operation-based screening and testing (BOOST), a recent work in GWAS, completes gene-gene interaction analysis in 2.5 days on a desktop computer. Compared with central processing units (CPUs), graphic processing units (GPUs) are highly parallel hardware and provide massive computing resources. We are, therefore, motivated to use GPUs to further speed up the analysis of gene-gene interactions. RESULTS: We implement the BOOST method based on a GPU framework and name it GBOOST. GBOOST achieves a 40-fold speedup compared with BOOST. It completes the analysis of Wellcome Trust Case Control Consortium Type 2 Diabetes (WTCCC T2D) genome data within 1.34 h on a desktop computer equipped with Nvidia GeForce GTX 285 display card. AVAILABILITY: GBOOST code is available at http://bioinformatics.ust.hk/BOOST.html#GBOOST. Ling Sing Yung, Can Yang 0002, Weichuan Yu |
Bioinform. | 4 |
| 2011 | Motif-All: discovering all phosphorylation motifsabstractBACKGROUND: Phosphorylation motifs represent common patterns around the phosphorylation site. The discovery of such kinds of motifs reveals the underlying regulation mechanism and facilitates the prediction of unknown phosphorylation event. To date, people have gathered large amounts of phosphorylation data, making it possible to perform substrate-driven motif discovery using data mining techniques. RESULTS: We describe an algorithm called Motif-All that is able to efficiently identify all statistically significant motifs. The proposed method explores a support constraint to reduce search space and avoid generating random artifacts. As the number of phosphorylated peptides are far less than that of unphosphorylated ones, we divide the mining process into two stages: The first step generates candidates from the set of phosphorylated sequences using only support constraint and the second step tests the statistical significance of each candidate using the odds ratio derived from the whole data set. Experimental results on real data show that Motif-All outperforms current algorithms in terms of both effectiveness and efficiency. CONCLUSIONS: Motif-All is a useful tool for discovering statistically significant phosphorylation motifs. Source codes and data sets are available at: http://bioinformatics.ust.hk/MotifAll.rar. Zengyou He, Can Yang 0002, Weichuan Yu |
BMC Bioinform. | 5 |
| 2011 | Score regularization for peptide identificationabstractPeptide identification from tandem mass spectrometry (MS/MS) data is one of the most important problems in computational proteomics. This technique relies heavily on the accurate assessment of the quality of peptide-spectrum matches (PSMs). However, current MS technology and PSM scoring algorithm are far from perfect, leading to the generation of incorrect peptide-spectrum pairs. Thus, it is critical to develop new post-processing techniques that can distinguish true identifications from false identifications effectively. In this paper, we present a consistency-based PSM re-ranking method to improve the initial identification results. This method uses one additional assumption that two peptides belonging to the same protein should be correlated to each other. We formulate an optimization problem that embraces two objectives through regularization: the smoothing consistency among scores of correlated peptides and the fitting consistency between new scores and initial scores. This optimization problem can be solved analytically. The experimental study on several real MS/MS data sets shows that this re-ranking method improves the identification performance. The score regularization method can be used as a general post-processing step for improving peptide identifications. Source codes and data sets are available at: http://bioinformatics.ust.hk/SRPI.rar . Zengyou He, Hongyu Zhao 0003, Weichuan Yu |
BMC Bioinform. | 3 |
| 2011 | The choice of null distributions for detecting gene-gene interactions in genome-wide association studiesabstractBACKGROUND: In genome-wide association studies (GWAS), the number of single-nucleotide polymorphisms (SNPs) typically ranges between 500,000 and 1,000,000. Accordingly, detecting gene-gene interactions in GWAS is computationally challenging because it involves hundreds of billions of SNP pairs. Stage-wise strategies are often used to overcome the computational difficulty. In the first stage, fast screening methods (e.g. Tuning ReliefF) are applied to reduce the whole SNP set to a small subset. In the second stage, sophisticated modeling methods (e.g., multifactor-dimensionality reduction (MDR)) are applied to the subset of SNPs to identify interesting interaction models and the corresponding interaction patterns. In the third stage, the significance of the identified interaction patterns is evaluated by hypothesis testing. RESULTS: In this paper, we show that this stage-wise strategy could be problematic in controlling the false positive rate if the null distribution is not appropriately chosen. This is because screening and modeling may change the null distribution used in hypothesis testing. In our simulation study, we use some popular screening methods and the popular modeling method MDR as examples to show the effect of the inappropriate choice of null distributions. To choose appropriate null distributions, we suggest to use the permutation test or testing on the independent data set. We demonstrate their performance using synthetic data and a real genome wide data set from an Aged-related Macular Degeneration (AMD) study. CONCLUSIONS: The permutation test or testing on the independent data set can help choosing appropriate null distributions in hypothesis testing, which provides more reliable results in practice. Can Yang 0002, Zengyou He, Qiang Yang 0001, Hong Xue 0001, Weichuan Yu |
BMC Bioinform. | 6 |
| 2011 | A hidden two-locus disease association pattern in genome-wide association studiesabstractBACKGROUND: Recent association analyses in genome-wide association studies (GWAS) mainly focus on single-locus association tests (marginal tests) and two-locus interaction detections. These analysis methods have provided strong evidence of associations between genetics variances and complex diseases. However, there exists a type of association pattern, which often occurs within local regions in the genome and is unlikely to be detected by either marginal tests or interaction tests. This association pattern involves a group of correlated single-nucleotide polymorphisms (SNPs). The correlation among SNPs can lead to weak marginal effects and the interaction does not play a role in this association pattern. This phenomenon is due to the existence of unfaithfulness: the marginal effects of correlated SNPs do not express their significant joint effects faithfully due to the correlation cancelation. RESULTS: In this paper, we develop a computational method to detect this association pattern masked by unfaithfulness. We have applied our method to analyze seven data sets from the Wellcome Trust Case Control Consortium (WTCCC). The analysis for each data set takes about one week to finish the examination of all pairs of SNPs. Based on the empirical result of these real data, we show that this type of association masked by unfaithfulness widely exists in GWAS. CONCLUSIONS: These newly identified associations enrich the discoveries of GWAS, which may provide new insights both in the analysis of tagSNPs and in the experiment design of GWAS. Since these associations may be easily missed by existing analysis tools, we can only connect some of them to publicly available findings from other association studies. As independent data set is limited at this moment, we also have difficulties to replicate these findings. More biological implications need further investigation. AVAILABILITY: The software is freely available at http://bioinformatics.ust.hk/hidden_pattern_finder.zip. Can Yang 0002, Qiang Yang 0001, Hong Xue 0001, Nelson L. S. Tang, Weichuan Yu |
BMC Bioinform. | 6 |
| 2011 | A Partial Set Covering Model for Protein Mixture Identification Using Mass Spectrometry DataabstractProtein identification is a key and essential step in mass spectrometry (MS) based proteome research. To date, there are many protein identification strategies that employ either MS data or MS/MS data for database searching. While MS-based methods provide wider coverage than MS/MS-based methods, their identification accuracy is lower since MS data have less information than MS/MS data. Thus, it is desired to design more sophisticated algorithms that achieve higher identification accuracy using MS data. Peptide Mass Fingerprinting (PMF) has been widely used to identify single purified proteins from MS data for many years. In this paper, we extend this technology to protein mixture identification. First, we formulate the problem of protein mixture identification as a Partial Set Covering (PSC) problem. Then, we present several algorithms that can solve the PSC problem efficiently. Finally, we extend the partial set covering model to both MS/MS data and the combination of MS data and MS/MS data. The experimental results on simulated data and real data demonstrate the advantages of our method: 1) it outperforms previous MS-based approaches significantly; 2) it is useful in the MS/MS-based protein inference; and 3) it combines MS data and MS/MS data in a unified model such that the identification performance is further improved. Zengyou He, Can Yang 0002, Weichuan Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2010 | Predictive rule inference for epistatic interaction detection in genome-wide association studiesabstractMOTIVATION: Under the current era of genome-wide association study (GWAS), finding epistatic interactions in the large volume of SNP data is a challenging and unsolved issue. Few of previous studies could handle genome-wide data due to the difficulties in searching the combinatorially explosive search space and statistically evaluating high-order epistatic interactions given the limited number of samples. In this work, we propose a novel learning approach (SNPRuler) based on the predictive rule inference to find disease-associated epistatic interactions. RESULTS: Our extensive experiments on both simulated data and real genome-wide data from Wellcome Trust Case Control Consortium (WTCCC) show that SNPRuler significantly outperforms its recent competitor. To our knowledge, SNPRuler is the first method that guarantees to find the epistatic interactions without exhaustive search. Our results indicate that finding epistatic interactions in GWAS is computationally attainable in practice. AVAILABILITY: http://bioinformatics.ust.hk/SNPRuler.zip Can Yang 0002, Qiang Yang 0001, Hong Xue 0001, Nelson L. S. Tang, Weichuan Yu |
Bioinform. | 6 |
| 2010 | Detecting two-locus associations allowing for interactions in genome-wide association studiesabstractMOTIVATION: Genome-wide association studies (GWASs) aim to identify genetic susceptibility to complex diseases by assaying and analyzing hundreds of thousands of single nucleotide polymorphisms (SNPs). Although traditional single-locus statistical tests have identified many genetic determinants of susceptibility, those findings cannot completely explain genetic contributions to complex diseases. Marchini and coauthors demonstrated the importance of testing two-locus associations allowing for interactions through a wide range of simulation studies. However, such a test is computationally demanding as we need to test hundreds of billions of SNP pairs in GWAS. Here, we provide a method to address this computational burden for dichotomous phenotypes. RESULTS: We have applied our method on nine datasets from GWAS, including the aged-related macular degeneration (AMD) dataset, the Parkinson's disease dataset and seven datasets from the Wellcome Trust Case Control Consortium (WTCCC). Our method has discovered many associations that were not identified before. The running time for the AMD dataset, the Parkinson's disease dataset and each of seven WTCCC datasets are 2.5, 82 and 90 h on a standard 3.0 GHz desktop with 4 G memory running Windows XP system. Our experiment results demonstrate that our method is feasible for the full-scale analyses of both single- and two-locus associations allowing for interactions in GWAS. AVAILABILITY: http://bioinformatics.ust.hk/SNPAssociation.zip CONTACT: [email protected]; [email protected]; SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Can Yang 0002, Qiang Yang 0001, Hong Xue 0001, Nelson L. S. Tang, Weichuan Yu |
Bioinform. | 6 |
| 2010 | Identifying main effects and epistatic interactions from large-scale SNP data via adaptive group LassoabstractBACKGROUND: Single nucleotide polymorphism (SNP) based association studies aim at identifying SNPs associated with phenotypes, for example, complex diseases. The associated SNPs may influence the disease risk individually (main effects) or behave jointly (epistatic interactions). For the analysis of high throughput data, the main difficulty is that the number of SNPs far exceeds the number of samples. This difficulty is amplified when identifying interactions. RESULTS: In this paper, we propose an Adaptive Group Lasso (AGL) model for large-scale association studies. Our model enables us to analyze SNPs and their interactions simultaneously. We achieve this by introducing a sparsity constraint in our model based on the fact that only a small fraction of SNPs is disease-associated. In order to reduce the number of false positive findings, we develop an adaptive reweighting scheme to enhance sparsity. In addition, our method treats SNPs and their interactions as factors, and identifies them in a grouped manner. Thus, it is flexible to analyze various disease models, especially for interaction detection. However, due to the intensive computation when millions of interaction terms needs to be searched in the model fitting, our method needs to combined with some filtering methods when applied to genome-wide data for detecting interactions. CONCLUSION: By using a wide range of simulated datasets and a real dataset from WTCCC, we demonstrate the advantages of our method. Can Yang 0002, Qiang Yang 0001, Hong Xue 0001, Weichuan Yu |
BMC Bioinform. | 5 |
| 2009 | Optimization-Based Peptide Mass Fingerprinting for Protein Mixture Identification
Zengyou He, Can Yang 0002, Robert Z. Qi, Jason Po-Ming Tam, Weichuan Yu |
RECOMB | 6 |
| 2009 | Improving peptide identification with single-stage mass spectrum peaksabstractMOTIVATION: Database searching is the major peptide identification method in shotgun proteomics. It searches tandem mass spectrometry (MS/MS) spectra against a protein database to identify target peptides. The success of such a database searching method relies on a scoring algorithm that can evaluate the quality of peptide-spectrum matches (PSMs) accurately. However, current scoring algorithms frequently generate inaccurate assignments due to variations and noises in the MS/MS spectra. To address this issue, we like to improve peptide identification by using additional information from other data sources. RESULTS: Single-stage MS data is complementary to MS/MS data in the sense that it provides broader mass coverage but less sequence information. In this article, we show that single-stage MS data can be used to re-rank PSMs. The proposed method explores a linear combination of scores between MS and MS/MS data to perform re-ranking. Experimental results on real data show that such a re-ranking strategy improves the identification performance significantly. AVAILABILITY: http://bioinformatics.ust.hk/ReRankPSMwMS1.rar Zengyou He, Weichuan Yu |
Bioinform. | 2 |
| 2009 | SNPHarvester: a filtering-based approach for detecting epistatic interactions in genome-wide association studiesabstractMOTIVATION: Hundreds of thousands of single nucleotide polymorphisms (SNPs) are available for genome-wide association (GWA) studies nowadays. The epistatic interactions of SNPs are believed to be very important in determining individual susceptibility to complex diseases. However, existing methods for SNP interaction discovery either suffer from high computation complexity or perform poorly when marginal effects of disease loci are weak or absent. Hence, it is desirable to develop an effective method to search epistatic interactions in genome-wide scale. RESULTS: We propose a new method SNPHarvester to detect SNP-SNP interactions in GWA studies. SNPHarvester creates multiple paths in which the visited SNP groups tend to be statistically associated with diseases, and then harvests those significant SNP groups which pass the statistical tests. It greatly reduces the number of SNPs. Consequently, existing tools can be directly used to detect epistatic interactions. By using a wide range of simulated data and a real genome-wide data, we demonstrate that SNPHarvester outperforms its recent competitor significantly and is promising for practical disease prognosis. AVAILABILITY: http://bioinformatics.ust.hk/SNPHarvester.html. Can Yang 0002, Zengyou He, Qiang Yang 0001, Hong Xue 0001, Weichuan Yu |
Bioinform. | 6 |
| 2009 | MegaSNPHunter: a learning approach to detect disease predisposition SNPs and high level interactions in genome wide association studyabstractBACKGROUND: The interactions of multiple single nucleotide polymorphisms (SNPs) are highly hypothesized to affect an individual's susceptibility to complex diseases. Although many works have been done to identify and quantify the importance of multi-SNP interactions, few of them could handle the genome wide data due to the combinatorial explosive search space and the difficulty to statistically evaluate the high-order interactions given limited samples. RESULTS: Three comparative experiments are designed to evaluate the performance of MegaSNPHunter. The first experiment uses synthetic data generated on the basis of epistasis models. The second one uses a genome wide study on Parkinson disease (data acquired by using Illumina HumanHap300 SNP chips). The third one chooses the rheumatoid arthritis study from Wellcome Trust Case Control Consortium (WTCCC) using Affymetrix GeneChip 500K Mapping Array Set. MegaSNPHunter outperforms the best solution in this area and reports many potential interactions for the two real studies. CONCLUSION: The experimental results on both synthetic data and two real data sets demonstrate that our proposed approach outperforms the best solution that is currently available in handling large-scale SNP data both in terms of speed and in terms of detection of potential interactions that were not identified before. To our knowledge, MegaSNPHunter is the first approach that is capable of identifying the disease-associated SNP interactions from WTCCC studies and is promising for practical disease prognosis. Can Yang 0002, Qiang Yang 0001, Hong Xue 0001, Nelson L. S. Tang, Weichuan Yu |
BMC Bioinform. | 6 |
| 2009 | Semi-supervised protein subcellular localizationabstractBACKGROUND: Protein subcellular localization is concerned with predicting the location of a protein within a cell using computational method. The location information can indicate key functionalities of proteins. Accurate predictions of subcellular localizations of protein can aid the prediction of protein function and genome annotation, as well as the identification of drug targets. Computational methods based on machine learning, such as support vector machine approaches, have already been widely used in the prediction of protein subcellular localization. However, a major drawback of these machine learning-based approaches is that a large amount of data should be labeled in order to let the prediction system learn a classifier of good generalization ability. However, in real world cases, it is laborious, expensive and time-consuming to experimentally determine the subcellular localization of a protein and prepare instances of labeled data. RESULTS: In this paper, we present an approach based on a new learning framework, semi-supervised learning, which can use much fewer labeled instances to construct a high quality prediction model. We construct an initial classifier using a small set of labeled examples first, and then use unlabeled instances to refine the classifier for future predictions. CONCLUSION: Experimental results show that our methods can effectively reduce the workload for labeling data using the unlabeled data. Our method is shown to enhance the state-of-the-art prediction results of SVM classifiers by more than 10%. Qian Xu 0005, Derek Hao Hu, Hong Xue 0001, Weichuan Yu, Qiang Yang 0001 |
BMC Bioinform. | 4 |
| 2009 | Comparison of public peak detection algorithms for MALDI mass spectrometry data analysisabstractBACKGROUND: In mass spectrometry (MS) based proteomic data analysis, peak detection is an essential step for subsequent analysis. Recently, there has been significant progress in the development of various peak detection algorithms. However, neither a comprehensive survey nor an experimental comparison of these algorithms is yet available. The main objective of this paper is to provide such a survey and to compare the performance of single spectrum based peak detection methods. RESULTS: In general, we can decompose a peak detection procedure into three consequent parts: smoothing, baseline correction and peak finding. We first categorize existing peak detection algorithms according to the techniques used in different phases. Such a categorization reveals the differences and similarities among existing peak detection algorithms. Then, we choose five typical peak detection algorithms to conduct a comprehensive experimental study using both simulation data and real MALDI MS data. CONCLUSION: The results of comparison show that the continuous wavelet-based algorithm provides the best average performance. Zengyou He, Weichuan Yu |
BMC Bioinform. | 3 |
| 2008 | Peak bagging for peptide mass fingerprintingabstractMOTIVATION: Mass Spectrometry (MS)-based protein identification via peptide mass fingerprinting (PMF) is a key component in high-throughput proteome research. While PMF was the first commonly used protein identification method, provided higher throughput than the tandem MS-based method, its accuracy is lower than that of the tandem MS method. Thus, it is desirable to develop PMF-based algorithm with higher protein identification accuracy to facilitate proteome research. RESULTS: We propose a peak bagging method for single MS-based protein identification. It combines results from multiple PMF algorithms, where each PMF algorithm takes a random peak subset as input. Evaluation with a set of real MALDI-TOF MS spectra shows that the new peak bagging method provides consistent improvements over the single PMF algorithm. Zengyou He, Can Yang 0002, Weichuan Yu |
Bioinform. | 3 |
| 2006 | Towards pointwise motion tracking in echocardiographic image sequences - Comparing the reliability of different features for speckle tracking
Weichuan Yu, Albert J. Sinusas, Karl Thiele, James S. Duncan |
Medical Image Anal. | 1 |
| 2006 | Multiple Peak Alignment in Sequential Data Analysis: A Scale-Space-Based ApproachabstractIn this paper, we address the multiple peak alignment problem in sequential data analysis with an approach based on the Gaussian scale-space theory. We assume that multiple sets of detected peaks are the observed samples of a set of common peaks. We also assume that the locations of the observed peaks follow unimodal distributions (e.g., normal distribution) with their means equal to the corresponding locations of the common peaks and variances reflecting the extension of their variations. Under these assumptions, we convert the problem of estimating locations of the unknown number of common peaks from multiple sets of detected peaks into a much simpler problem of searching for local maxima in the scale-space representation. The optimization of the scale parameter is achieved using an energy minimization approach. We compare our approach with a hierarchical clustering method using both simulated data and real mass spectrometry data. We also demonstrate the merit of extending the binary peak detection method (i.e., a candidate is considered either as a peak or as a nonpeak) with a quantitative scoring measure-based approach (i.e., we assign to each candidate a possibility of being a peak). Weichuan Yu, Xiaoye Li, Baolin Wu, Kenneth R. Williams, Hongyu Zhao 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2005 | Using skew Gabor filter in source signal separation and local spectral orientation analysis
Weichuan Yu, Gerald Sommer, Kostas Daniilidis, James S. Duncan |
Image Vis. Comput. | 1 |
| 2004 | Using Skew Gabor Filter in Source Signal Separation and Local Spectral Multi-Orientation Analysis
Weichuan Yu, Gerald Sommer, Kostas Daniilidis |
CVPR (1) | 1 |
| 2004 | Pointwise Motion Tracking in Echocardiographic Images
Weichuan Yu, Albert J. Sinusas, Karl Thiele, James S. Duncan |
CVPR (1) | 1 |
| 2003 | Multiple motion analysis: in spatial or in spectral domain?
Weichuan Yu, Gerald Sommer, Kostas Daniilidis |
Comput. Vis. Image Underst. | 1 |
| 2003 | Three dimensional orientation signatures with conic kernel filtering for multiple motion analysis
Weichuan Yu, Gerald Sommer, Kostas Daniilidis |
Image Vis. Comput. | 1 |
| 2003 | Combinative multi-scale level set framework for echocardiographic image segmentation
Ning Lin, Weichuan Yu, James S. Duncan |
Medical Image Anal. | 2 |
| 2002 | Combinative Multi-scale Level Set Framework for Echocardiographic Image Segmentation
Ning Lin, Weichuan Yu, James S. Duncan |
MICCAI (1) | 2 |
| 2002 | Oriented Structure of the Occlusion Distortion: Is It Reliable?abstractIn the energy spectrum of an occlusion sequence, the distortion term has the same orientation as the velocity of the occluding signal. Other works claimed that this oriented structure can be used to distinguish the occluding velocity from the occluded one. We argue that the orientation structure of the distortion cannot always work as a reliable feature due to the rapidly decreasing energy contribution. This already weak orientation structure is further blurred by a superposition of distinct distortion components. We also indicate that the superposition principle of Shizawa and Mase (1991) for multiple motion estimation needs to be adjusted. Weichuan Yu, Gerald Sommer, Steven S. Beauchemin, Kostas Daniilidis |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | 3D-Orientation Signatures with Conic Kernel Filtering for Multiple Motion AnalysisabstractIn this paper we propose a new 3D kernel for the recovery of 3D-orientation signatures. The kernel is a Gaussian function defined in local spherical coordinates and its Cartesian support has the shape of a truncated cone with its axis in the radial direction and very small angular support. A set of such kernels is obtained by uniformly sampling the 2D space of polar and azimuth angles. The projection of a local neighborhood on such a kernel set produces a local 3D-orientation signature. In the case of spatiotemporal analysis, such a kernel set can be applied either on the derivative space of a local neighborhood or on the local Fourier transform. The well known planes arising from single or multiple motion produce maxima in the orientation signature. Due to the kernel's local support spatiotemporal signatures possess higher orientation resolution than 3D steerable filters and motion maxima can be detected and localized more accurately. We describe and show in experiments the superiority of the proposed kernels compared to Hough transformation or EM-based multiple motion detection. Weichuan Yu, Gerald Sommer, Kostas Daniilidis |
CVPR (1) | 1 |
| 2001 | Approximate orientation steerability based on angular GaussiansabstractJunctions are significant features in images with intensity variation that exhibits multiple orientations. This makes the detection and characterization of junctions a challenging problem. The characterization of junctions would ideally be given by the response of a filter at every orientation. This can be achieved by the principle of steerability that enables the decomposition of a filter into a linear combination of basis functions. However, current steerability approaches suffer from the consequences of the uncertainty principle: in order to achieve high resolution in orientation they need a large number of basis filters increasing, thus, the computational complexity. Furthermore, these functions have usually a wide support which only accentuates the computational burden. We propose a novel alternative to current steerability approaches. It is based on utilizing a set of polar separable filters with small support to sample orientation information. The orientation signature is then obtained by interpolating orientation samples using Gaussian functions with small support. Compared with current steerability techniques our approach achieves a higher orientation resolution with a lower complexity. In addition, we build a polar pyramid to characterize junctions of arbitrary inherent orientation scales. Weichuan Yu, Kostas Daniilidis, Gerald Sommer |
IEEE Trans. Image Process. | 1 |
| 1999 | Detection and Characterization of Multiple Motion PointsabstractThe computation of optical flow is a well studied topic in biological and computational vision. However, the existence of multiple motions in dynamic imagery due to occlusion or even transparency still raises challenging questions. In this paper, we propose an approach for the detection and characterization of occlusion and transparency. We propose a theoretical framework for both types of multiple motions which explicitly shows the difference between occlusion and transparency in the frequency domain. Then, we employ an EM-algorithm for the computation of one or two image velocities and a simple test for the detection of occlusion. Our approach differs from other EM-approaches which blindly assume the superposition of two models in the spatial domain without providing with a separate formal model for occlusion. We test and compare the characterization performance on synthetic and real data. Weichuan Yu, Gerald Sommer, Steven S. Beauchemin, Kostas Daniilidis |
CVPR | 1 |
| 1998 | Rotated Wedge Averaging Method for Junction ClassificationabstractThe computational cost of conventional filter methods for junction characterization is very high. This burden can be attenuated by using steerable filters. However, in order to achieve a high orientational selectivity to characterize complex junctions a large number of basis filters is necessary. From this results a yet too high computational effort for steerable filters. In this paper we present a new method for characterizing junctions which keeps the high orientational resolution and is computationally efficient. It is based on applying rotated copies of a wedge averaging filter and estimating the derivative with respect to the polar angle. The new method is compared with the steerable wedge filter method in experiments with real images. We show the superiority of our method as well as its adaptability to scale changes and robustness against noise. Weichuan Yu, Kostas Daniilidis, Gerald Sommer |
CVPR | 1 |
| 1998 | Low-Cost Junction Characterization using Polar Averaging Filters
Weichuan Yu, Kostas Daniilidis, Gerald Sommer |
ICIP (3) | 1 |