Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kai Ye 0001

dblp:85/1383-1 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-2851-6741ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
11 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 70% Geometric modeling and processing · 30%
Artificial intelligence
1 paper
3D vision · 77% Graph learning · 23%
Databases, data mining, and information retrieval
2 papers
Indexing and storage engines · 48% Data stream processing · 48% Data mining · 4%

Topics — the 30 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image restoration
image inpainting
1.012026
Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Image and video processing
image restoration
1.012026
Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Geometric modeling and processing
implicit neural representation
1.012026
Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › 3D vision
implicit neural representation
0.912025
STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation · CVPR 2025
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics
0.912025
STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation · CVPR 2025
Bioinformatics and computational biology
cancer genomics
0.732019
PRESM: personalized reference editor for somatic mutation discovery in cancer genomics · Bioinform. 2019
MSIsensor: microsatellite instability detection using paired tumor-normal sequence data · Bioinform. 2014
MEpurity: estimating tumor purity using DNA methylation data · Bioinform. 2019
Indexing and storage engines › membership query
approximate membership query
0.512021
Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership Querying · SIGMOD Conference 2021
Data stream processing
sketch
0.512021
Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership Querying · SIGMOD Conference 2021
Bioinformatics and computational biology
sequence analysis
0.422019
PRESM: personalized reference editor for somatic mutation discovery in cancer genomics · Bioinform. 2019
Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads · Bioinform. 2009
Bioinformatics and computational biology › cancer genomics › somatic mutation analysis
somatic mutation detection
0.412019
PRESM: personalized reference editor for somatic mutation discovery in cancer genomics · Bioinform. 2019
Bioinformatics and computational biology › cancer genomics
tumor purity estimation
0.412019
MEpurity: estimating tumor purity using DNA methylation data · Bioinform. 2019
Image and video processing › image restoration
denoising
0.312026
Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Bioinformatics and computational biology › genomics › structural variation
structural variant analysis
0.312017
BreakPoint Surveyor: a pipeline for structural variant visualization · Bioinform. 2017
Bioinformatics and computational biology › genomics › genome visualization
structural variant visualization
0.312017
BreakPoint Surveyor: a pipeline for structural variant visualization · Bioinform. 2017
Machine learning › Graph learning › graph neural network › graph neural network architecture
spatial graph neural network
0.312025
STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation · CVPR 2025
Bioinformatics and computational biology › sequence analysis
high-throughput sequencing data analysis
0.212016
Detecting dispersed duplications in high-throughput sequencing data using a database-free approach · Bioinform. 2016
Bioinformatics and computational biology › genomics
sequencing
0.212016
Detecting dispersed duplications in high-throughput sequencing data using a database-free approach · Bioinform. 2016
Bioinformatics and computational biology › genomics › structural variation
structural variation detection
0.212016
Detecting dispersed duplications in high-throughput sequencing data using a database-free approach · Bioinform. 2016
Bioinformatics and computational biology
protein sequence analysis
0.222008
Tracing evolutionary pressure · Bioinform. 2008
An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences · Bioinform. 2007
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis
0.112012
PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data · Bioinform. 2012
Bioinformatics and computational biology › transcriptomics › RNA splicing analysis
splice junction detection
0.112012
PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data · Bioinform. 2012
Bioinformatics and computational biology
transcriptomics
0.112012
PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data · Bioinform. 2012
Bioinformatics and computational biology
genomics
0.112009
Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads · Bioinform. 2009
Bioinformatics and computational biology › genomics › structural variation
structural variant detection
0.112009
Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads · Bioinform. 2009
Bioinformatics and computational biology
protein function prediction
0.112008
Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weighting · Bioinform. 2008
Bioinformatics and computational biology
phylogenetics
0.012008
Tracing evolutionary pressure · Bioinform. 2008
Bioinformatics and computational biology › phylogenetics
phylogenetic tree analysis
0.012008
Tracing evolutionary pressure · Bioinform. 2008
Bioinformatics and computational biology
protein structure analysis
0.012008
Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weighting · Bioinform. 2008
Data mining › pattern mining
frequent pattern mining
0.012007
An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences · Bioinform. 2007
Data mining
pattern mining
0.012007
An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences · Bioinform. 2007

Methods — techniques the papers use, named apart from their topics

implicit neural representation · 1.7gradient descent · 1.7tensor decomposition · 1.0self-evolving optimization · 1.0haar wavelet transform · 1.0circular shift and coalesce · 0.5read mapping · 0.4germline substitution integration · 0.4beta mixture model · 0.4paired-end sequencing · 0.3RNA-seq expression integration · 0.3paired-end read alignment · 0.2paired-end read mapping · 0.2length distribution comparison · 0.2sequence pattern mining · 0.1pattern growth · 0.1
YearPublicationVenuePosition
2026 Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters
abstract
Implicit neural representation (INR) has emerged as a powerful paradigm for visual data representation. However, classical INR methods represent data in the original space mixed with different frequency components, and several feature encoding parameters (e.g., the frequency parameter $\omega$ω or the rank $R$R) need manual configurations. In this work, we propose a self-evolving cross-frequency INR using the Haar wavelet transform (termed CF-INR), which decouples data into four frequency components and employs INRs in the wavelet space. CF-INR allows the characterization of different frequency components separately, thus enabling higher accuracy for data representation. To more precisely characterize cross-frequency components, we propose a cross-frequency tensor decomposition paradigm for CF-INR with self-evolving parameters, which automatically updates the rank parameter $R$R and the frequency parameter $\omega$ω for each frequency component through self-evolving optimization. This self-evolution paradigm eliminates the laborious manual tuning of these parameters, and learns a customized cross-frequency feature encoding configuration for each dataset. We evaluate CF-INR on a variety of visual data representation and inverse imaging problems, including image regression, inpainting, denoising, and cloud removal. Extensive experiments demonstrate that CF-INR outperforms state-of-the-art methods in each case.
Yi-Si Luo, Kai Ye 0001, Xi-Le Zhao, Deyu Meng
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation
abstract
Spatial transcriptomics (ST) are emerging technologies that reveal spatial distributions of gene expressions within tissues, serving as important ways to uncover biological insights. However, the irregular spatial profiles and variability of genes make it challenging to integrate spatial information with gene expression under a computational framework. Current algorithms mostly utilize spatial graph neural networks to encode spatial information, which may incur increased computational costs and may not be flexible enough to depict complex spatial configurations. In this study, we introduce a concise yet effective representation framework, STINR, for deciphering ST data. STINR leverages an implicit neural representation (INR) to continuously represent ST data, which efficiently characterizes spatial and slice-wise correlations of ST data by inheriting the implicit smoothness of INR. STINR allows easier integration of multiple slices and multi-omics without any alignment, and serves as a potent tool for various biological tasks including gene imputation, gene denoising, spatial domain detection, and cell-type deconvolution stemed from ST data. In particular, STINR identifies the thinnest cortex layer in the dorsolateral prefrontal cortex which previous methods were unable to achieve, and more accurately identifies tumor regions in the human squamous cell carcinoma, showcasing its practical value for biological discoveries. Code at https://github.com/YisiLuo/STINR.
Yi-Si Luo, Xi-Le Zhao, Kai Ye 0001, Deyu Meng
CVPR3
2025 NeurTV: Total Variation on the Neural Domain
abstract
Abstract. Recently, we have witnessed the success of total variation (TV) for many imaging applications. However, traditional TV is defined on the original pixel domain, which limits its potential. In this work, we suggest a new TV regularization defined on the neural domain. Concretely, the discrete data is implicitly and continuously represented by a deep neural network (DNN), and we use the derivatives of DNN outputs with respect to (w.r.t.) input coordinates to capture local correlations of data. As compared with classical TV on the original domain, the proposed TV on the neural domain (termed NeurTV) enjoys the following advantages. First, NeurTV is free of discretization error induced by the discrete difference operator. Second, NeurTV is not limited to meshgrid but is suitable for both meshgrid and non-meshgrid data. Third, NeurTV can more exactly capture local correlations across data for any direction and any order of derivatives attributed to the implicit and continuous nature of neural domain. We theoretically reinterpret NeurTV under the variational approximation framework, which allows us to build the connection between NeurTV and classical TV and inspires us to develop variants (e.g., space-variant NeurTV). Extensive numerical experiments with meshgrid data (e.g., color and hyperspectral images) and non-meshgrid data (e.g., point clouds and spatial transcriptomics) showcase the effectiveness of the proposed methods.
Yi-Si Luo, Xi-Le Zhao, Kai Ye 0001, Deyu Meng
SIAM J. Imaging Sci.3
2023 Comparison and benchmark of structural variants detected from long read and long-read assembly
abstract
Structural variant (SV) detection is essential for genomic studies, and long-read sequencing technologies have advanced our capacity to detect SVs directly from read or de novo assembly, also known as read-based and assembly-based strategy. However, to date, no independent studies have compared and benchmarked the two strategies. Here, on the basis of SVs detected by 20 read-based and eight assembly-based detection pipelines from six datasets of HG002 genome, we investigated the factors that influence the two strategies and assessed their performance with well-curated SVs. We found that up to 80% of the SVs could be detected by both strategies among different long-read datasets, whereas variant type, size, and breakpoint detected by read-based strategy were greatly affected by aligners. For the high-confident insertions and deletions at non-tandem repeat regions, a remarkable subset of them (82% in assembly-based calls and 93% in read-based calls), accounting for around 4000 SVs, could be captured by both reads and assemblies. However, discordance between two strategies was largely caused by complex SVs and inversions, which resulted from inconsistent alignment of reads and assemblies at these loci. Finally, benchmarking with SVs at medically relevant genes, the recall of read-based strategy reached 77% on 5X coverage data, whereas assembly-based strategy required 20X coverage data to achieve similar performance. Therefore, integrating SVs from read and assembly is suggested for general-purpose detection because of inconsistently detected complex SVs and inversions, whereas assembly-based strategy is optional for applications with limited resources.
Jiadong Lin, Peng Jia 0004, Songbo Wang, Walter A. Kosters, Kai Ye 0001
Briefings Bioinform.5
2022 Homotopic Convex Transformation: A New Landscape Smoothing Method for the Traveling Salesman Problem
abstract
This article proposes a novel landscape smoothing method for the symmetric traveling salesman problem (TSP). We first define the homotopic convex (HC) transformation of a TSP as a convex combination of a well-constructed simple TSP and the original TSP. The simple TSP, called the convex-hull TSP, is constructed by transforming a known local or global optimum. We observe that controlled by the coefficient of the convex combination, with local or global optimum: 1) the landscape of the HC transformed TSP is smoothed in terms that its number of local optima is reduced compared to the original TSP and 2) the fitness distance correlation of the HC transformed TSP is increased. Furthermore, we observe that the smoothing effect of the HC transformation depends highly on the quality of the used optimum. A high-quality optimum leads to a better smoothing effect than a low-quality optimum. We then propose an iterative algorithmic framework in which the proposed HC transformation is combined within a heuristic TSP solver. It works as an escaping scheme from local optima aiming to improve the global searchability of the combined heuristic. Case studies using the 3-Opt and the Lin-Kernighan local search as the heuristic solver show that the resultant algorithms significantly outperform their counterparts and two other smoothing-based TSP heuristic solvers on most of the test instances with up to 20 000 cities.
Jialong Shi, Jianyong Sun, Qingfu Zhang 0001, Kai Ye 0001
IEEE Trans. Cybern.4
2021 Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership Querying
abstract
Given a set S, Membership Querying (MQ) answers whether a query element $q\in S$. It is a fundamental task in areas like database systems and computer networks. In this paper, we consider a more general problem, Multi-Set Multi-Membership Querying (MS-MMQ). Given n sets $S_0,łdots,S_n-1 $, MS-MMQ answers which sets contain element q. A direct way to address MS-MMQ is to build an MQ structure (e.g., Bloom Filter) for each set. However, the query and space complexities grow linearly with n and become prohibitive for a large n. To address this challenge, we propose a novel Circular Shift and Coalesce (CSC) framework to efficiently achieve approximate MS-MMQ. Instead of building an MQ data structure for each set, the CSC index encodes all n sets into a compact sketch and retrieves only a few bytes in the sketch for a query, which achieves high memory-efficiency and boosts the query speed by several times. CSC is compatible with mainstream data structures for Approximate MQ. We conduct experiments on real-world datasets and results demonstrate that our framework is up to 91.2 times faster and up to 48.9 times more accurate than state-of-the-art methods.
Rundong Li 0002, Pinghui Wang, Jiongli Zhu, Junzhou Zhao, Jia Di, Xiaofei Yang 0003, Kai Ye 0001
SIGMOD Conference7
2020 From Innovations to Prospects: What Is Hidden Behind Cryptocurrencies?
abstract
The great influence of Bitcoin has promoted the rapid development of blockchain-based digital currencies, especially the altcoins, since 2013. However, most altcoins share similar source codes, resulting in concerns about code innovations. In this paper, an empirical study on existing altcoins is carried out to offer a thorough understanding of various aspects associated with altcoin innovations. Firstly, we construct the dataset of altcoins, including source code repositories, GitHub fork relations, and market capitalizations (cap). Then, we analyze the altcoin innovations from the perspective of source code similarities. The results demonstrate that more than 85% of altcoin repositories present high code similarities. Next, a temporal clustering algorithm is proposed to mine the inheritance relationship among various altcoins. The family pedigrees of altcoin are constructed, in which the altcoin presents similar evolution features as biology, such as power-law in family size, variety in family evolution, etc. Finally, we investigate the correlation between code innovations and market capitalization. Although we fail to predict the price of altcoins based on their code similarities, the results show that altcoins with higher innovations reflect better market prospects.
Ang Jia, Ming Fan 0002, Wenying Wei, Zijiang Yang 0006, Kai Ye 0001, Ting Liu 0002
MSR7
2019 PRESM: personalized reference editor for somatic mutation discovery in cancer genomics
abstract
MOTIVATION: Accurate detection of somatic mutations is a crucial step toward understanding cancer. Various tools have been developed to detect somatic mutations from cancer genome sequencing data by mapping reads to a universal reference genome and inferring likelihoods from complex statistical models. However, read mapping is frequently obstructed by mismatches between germline and somatic mutations on a read and the reference genome. Previous attempts to develop personalized genome tools are not compatible with downstream statistical models for somatic mutation detection. RESULTS: We present PRESM, a tool that builds personalized reference genomes by integrating germline mutations into the reference genome. The aforementioned obstacle is circumvented by using a two-step germline substitution procedure, maintaining positional fidelity using an innovative workaround. Reads derived from tumor tissue can be positioned more accurately along a personalized reference than a universal reference due to the reduced genetic distance between the subject (tumor genome) and the target (the personalized genome). Application of PRESM's personalized genome reduced false-positive (FP) somatic mutation calls by as much as 55.5%, and facilitated the discovery of a novel somatic point mutation on a germline insertion in PDE1A, a phosphodiesterase associated with melanoma. Moreover, all improvements in calling accuracy were achieved without parameter optimization, as PRESM itself is parameter-free. Hence, similar increases in read mapping and decreases in the FP rate will persist when PRESM-built genomes are applied to any user-provided dataset. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/precisionomics/PRESM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chen Cao 0002, Lauren Mak, Guangxu Jin, Paul Gordon, Kai Ye 0001
Bioinform.5
2019 MEpurity: estimating tumor purity using DNA methylation data
abstract
MOTIVATION: Tumor purity is a fundamental property of each cancer sample and affects downstream investigations. Current tumor purity estimation methods either require matched normal sample or report moderately high tumor purity even on normal samples. It is critical to develop a novel computational approach to estimate tumor purity with sufficient precision based on tumor-only sample. RESULTS: In this study, we developed MEpurity, a beta mixture model-based algorithm, to estimate the tumor purity based on tumor-only Illumina Infinium 450k methylation microarray data. We applied MEpurity to both The Cancer Genome Atlas (TCGA) cancer data and cancer cell line data, demonstrating that MEpurity reports low tumor purity on normal samples and comparable results on tumor samples with other state-of-art methods. AVAILABILITY AND IMPLEMENTATION: MEpurity is a C++ program which is available at https://github.com/xjtu-omics/MEpurity. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaofei Yang 0003, Tingjie Wang, Jiadong Lin, Yongyong Kang, Peng Jia 0004, Kai Ye 0001
Bioinform.7
2019 Learning From a Stream of Nonstationary and Dependent Data in Multiobjective Evolutionary Optimization
abstract
Combining machine learning techniques has shown great potentials in evolutionary optimization since the domain knowledge of an optimization problem, if well learned, can be a great help for creating high-quality solutions. However, existing learning-based multiobjective evolutionary algorithms (MOEAs) spend too much computational overhead on learning. To address this problem, we propose a learning-based MOEA where an online learning algorithm is embedded within the evolutionary search procedure. The online learning algorithm takes the stream of sequentially generated solutions along the evolution as its training data. It is noted that the stream of solutions are temporal, dependent, nonstationary, and nonstatic. These data characteristics make existing online learning algorithm not suitable for the evolution data. We hence modify an existing online agglomerative clustering algorithm to accommodate these characteristics. The modified online clustering algorithm is applied to adaptively discover the structure of the Pareto optimal set; and the learned structure is used to guide new solution creation. Experimental results have shown significant improvement over four state-of-the-art MOEAs on a variety of benchmark problems.
Jianyong Sun, Hu Zhang 0002, Aimin Zhou, Qingfu Zhang 0001, Ke Zhang 0020, Zhenbiao Tu, Kai Ye 0001
IEEE Trans. Evol. Comput.7
2017 BreakPoint Surveyor: a pipeline for structural variant visualization
abstract
SUMMARY: BreakPoint Surveyor (BPS) is a computational pipeline for the discovery, characterization, and visualization of complex genomic rearrangements, such as viral genome integration, in paired-end sequence data. BPS facilitates interpretation of structural variants by merging structural variant breakpoint predictions, gene exon structure, read depth, and RNA-sequencing expression into a single comprehensive figure. AVAILABILITY AND IMPLEMENTATION: Source code and sample data freely available for download at https://github.com/ding-lab/BreakPointSurveyor, distributed under the GNU GPLv3 license, implemented in R, Python and BASH scripts, and supported on Unix/Linux/OS X operating systems. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Matthew A. Wyczalkowski, Kristine M. Wylie, Michael D. McLellan, Jennifer Flynn, Mo Huang, Kai Ye 0001, Xian Fan, Ken Chen 0001, Michael C. Wendl
Bioinform.7
2016 Detecting dispersed duplications in high-throughput sequencing data using a database-free approach
abstract
MOTIVATION: Dispersed duplications (DDs) such as transposon element insertions and copy number variations are ubiquitous in the human genome. They have attracted the interest of biologists as well as medical researchers due to their role in both evolution and disease. The efforts of discovering DDs in high-throughput sequencing data are currently dominated by database-oriented approaches that require pre-existing knowledge of the DD elements to be detected. RESULTS: We present DD_DETECTION, a database-free approach to finding DD events in high-throughput sequencing data. DD_DETECTION is able to detect DDs purely from paired-end read alignments. We show in a comparative study that this method is able to compete with database-oriented approaches in recovering validated transposon insertion events. We also experimentally validate the predictions of DD_DETECTION on a human DNA sample, showing that it can find not only duplicated elements present in common databases but also DDs of novel type. AVAILABILITY AND IMPLEMENTATION: The software presented in this article is open source and available from https://bitbucket.org/mkroon/dd_detection.
M. Kroon, Eric-Wubbo Lameijer, N. Lakenberg, Jayne Y. Hehir-Kwa, D. T. Thung, P. Eline Slagboom, Joost N. Kok, Kai Ye 0001
Bioinform.8
2014 MSIsensor: microsatellite instability detection using paired tumor-normal sequence data
abstract
MOTIVATION: Microsatellite instability (MSI) is an important indicator of larger genome instability and has been linked to many genetic diseases, including Lynch syndrome. MSI status is also an independent prognostic factor for favorable survival in multiple cancer types, such as colorectal and endometrial. It also informs the choice of chemotherapeutic agents. However, the current PCR-electrophoresis-based detection procedure is laborious and time-consuming, often requiring visual inspection to categorize samples. We developed MSIsensor, a C++ program for automatically detecting somatic microsatellite changes. It computes length distributions of microsatellites per site in paired tumor and normal sequence data, subsequently using these to statistically compare observed distributions in both samples. Comprehensive testing indicates MSIsensor is an efficient and effective tool for deriving MSI status from standard tumor-normal paired sequence data. AVAILABILITY AND IMPLEMENTATION: https://github.com/ding-lab/msisensor
Beifang Niu, Kai Ye 0001, Qunyuan Zhang, Charles Lu 0002, Mingchao Xie, Michael D. McLellan, Michael C. Wendl
Bioinform.2
2012 PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data
abstract
MOTIVATION: RNA-seq is a powerful technology for the study of transcriptome profiles that uses deep-sequencing technologies. Moreover, it may be used for cellular phenotyping and help establishing the etiology of diseases characterized by abnormal splicing patterns. In RNA-Seq, the exact nature of splicing events is buried in the reads that span exon-exon boundaries. The accurate and efficient mapping of these reads to the reference genome is a major challenge. RESULTS: We developed PASSion, a pattern growth algorithm-based pipeline for splice site detection in paired-end RNA-Seq reads. Comparing the performance of PASSion to three existing RNA-Seq analysis pipelines, TopHat, MapSplice and HMMSplicer, revealed that PASSion is competitive with these packages. Moreover, the performance of PASSion is not affected by read length and coverage. It performs better than the other three approaches when detecting junctions in highly abundant transcripts. PASSion has the ability to detect junctions that do not have known splicing motifs, which cannot be found by the other tools. Of the two public RNA-Seq datasets, PASSion predicted ≈ 137,000 and 173,000 splicing events, of which on average 82 are known junctions annotated in the Ensembl transcript database and 18% are novel. In addition, our package can discover differential and shared splicing patterns among multiple samples. AVAILABILITY: The code and utilities can be freely downloaded from https://trac.nbic.nl/passion and ftp://ftp.sanger.ac.uk/pub/zn1/passion.
Yanju Zhang, Eric-Wubbo Lameijer, Peter A. C. 't Hoen, Zemin Ning, P. Eline Slagboom, Kai Ye 0001
Bioinform.6
2009 Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads
abstract
MOTIVATION: There is a strong demand in the genomic community to develop effective algorithms to reliably identify genomic variants. Indel detection using next-gen data is difficult and identification of long structural variations is extremely challenging. RESULTS: We present Pindel, a pattern growth approach, to detect breakpoints of large deletions and medium-sized insertions from paired-end short reads. We use both simulated reads and real data to demonstrate the efficiency of the computer program and accuracy of the results. AVAILABILITY: The binary code and a short user manual can be freely downloaded from http://www.ebi.ac.uk/ approximately kye/pindel/. CONTACT: [email protected]; [email protected].
Kai Ye 0001, Marcel H. Schulz, Rolf Apweiler, Zemin Ning
Bioinform.1
2008 Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weighting
abstract
MOTIVATION: Identification of residues that account for protein function specificity is crucial, not only for understanding the nature of functional specificity, but also for protein engineering experiments aimed at switching the specificity of an enzyme, regulator or transporter. Available algorithms generally use multiple sequence alignments to identify residue positions conserved within subfamilies but divergent in between. However, many biological examples show a much subtler picture than simple intra-group conservation versus inter-group divergence. RESULTS: We present multi-RELIEF, a novel approach for identifying specificity residues that is based on RELIEF, a state-of-the-art Machine-Learning technique for feature weighting. It estimates the expected 'local' functional specificity of residues from an alignment divided in multiple classes. Optionally, 3D structure information is exploited by increasing the weight of residues that have high-weight neighbors. Using ROC curves over a large body of experimental reference data, we show that (a) multi-RELIEF identifies specificity residues for the seven test sets used, (b) incorporating structural information improves prediction for specificity of interaction with small molecules and (c) comparison of multi-RELIEF with four other state-of-the-art algorithms indicates its robustness and best overall performance. AVAILABILITY: A web-server implementation of multi-RELIEF is available at www.ibi.vu.nl/programs/multirelief. Matlab source code of the algorithm and data sets are available on request for academic use.
Kai Ye 0001, K. Anton Feenstra, Jaap Heringa, Adriaan P. IJzerman, Elena Marchiori
Bioinform.1
2008 Tracing evolutionary pressure
abstract
MOTIVATION: Recent advances in sequencing techniques have yielded enormous amounts of protein sequence data from various species. This large dataset allows sequence comparison between paralogous and orthologous proteins to identify motifs or functional positions that account for the differences of functional subgroups ('specificity' positions). Algorithms such as SDPpred and the two-entropies analysis (TEA) have been developed to detect such specificity positions from a multiple sequence alignment (MSA) grouped into classes according to certain biological functions. Other algorithms such as TreeDet compute a classification and then predict specificity positions associated with it. However, there are still many unresolved questions: Was the optimal subdivision of a protein family achieved? Do the definitions at different levels of the phylogenetic tree affect the prediction of specificity positions? Can the whole phylogenetic tree be used instead of only one level in it to predict specificity positions? RESULTS: Here we present a novel method, TEA-O (Two-entropies analysis-Objective), to trace the evolutionary pressure from the root to the branches of the phylogenetic tree. At each level of the tree, a TEA plot is produced to capture the signal of the evolutionary pressure. A consensus TEA-O plot is composed from the whole series of plots to provide a condensed representation. Positions related to functions that evolved early (conserved) or later (specificity) are close to the lower-left or upper-left corner of the TEA-O plot, respectively. This novel approach allows an unbiased, user-independent, analysis of residue relevance in a protein family. We compared our TEA-O method with various algorithms using both synthetic and real protein sequences. The results show that our method is robust, sensitive to subtle differences in evolutionary pressure during evolution and comprehensive because all positions in the MSA are presented in the consensus plot. AVAILABILITY: All computer programs and datasets used in this work are available at http://nava.liacs.nl/kye/TEA-O/ for academic use.
Kai Ye 0001, Gert Vriend, Adriaan P. IJzerman
Bioinform.1
2007 An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences
abstract
MOTIVATION: Pattern discovery in protein sequences is often based on multiple sequence alignments (MSA). The procedure can be computationally intensive and often requires manual adjustment, which may be particularly difficult for a set of deviating sequences. In contrast, two algorithms, PRATT2 (http//www.ebi.ac.uk/pratt/) and TEIRESIAS (http://cbcsrv.watson.ibm.com/) are used to directly identify frequent patterns from unaligned biological sequences without an attempt to align them. Here we propose a new algorithm with more efficiency and more functionality than both PRATT2 and TEIRESIAS, and discuss some of its applications to G protein-coupled receptors, a protein family of important drug targets. RESULTS: In this study, we designed and implemented six algorithms to mine three different pattern types from either one or two datasets using a pattern growth approach. We compared our approach to PRATT2 and TEIRESIAS in efficiency, completeness and the diversity of pattern types. Compared to PRATT2, our approach is faster, capable of processing large datasets and able to identify the so-called type III patterns. Our approach is comparable to TEIRESIAS in the discovery of the so-called type I patterns but has additional functionality such as mining the so-called type II and type III patterns and finding discriminating patterns between two datasets. AVAILABILITY: The source code for pattern growth algorithms and their pseudo-code are available at http://www.liacs.nl/home/kosters/pg/.
Kai Ye 0001, Walter A. Kosters, Adriaan P. IJzerman
Bioinform.1