VLDB 2026 Research / reviewers in the wild / expert
Kai Ye 0001
dblp:85/1383-1
· DBLP profile ↗
18ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-2851-6741ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 100% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 70% Geometric modeling and processing · 30% | |
| Artificial intelligence
1 paper |
3D vision · 77% Graph learning · 23% | |
| Databases, data mining, and information retrieval
2 papers |
Indexing and storage engines · 48% Data stream processing · 48% Data mining · 4% |
Topics — the 30 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › image restoration
image inpainting |
1.0 | 1 | 2026 | Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Image and video processing
image restoration |
1.0 | 1 | 2026 | Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Geometric modeling and processing
implicit neural representation |
1.0 | 1 | 2026 | Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › 3D vision
implicit neural representation |
0.9 | 1 | 2025 | STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation · CVPR 2025 |
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics |
0.9 | 1 | 2025 | STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation · CVPR 2025 |
Bioinformatics and computational biology
cancer genomics |
0.7 | 3 | 2019 | PRESM: personalized reference editor for somatic mutation discovery in cancer genomics · Bioinform. 2019 MSIsensor: microsatellite instability detection using paired tumor-normal sequence data · Bioinform. 2014 MEpurity: estimating tumor purity using DNA methylation data · Bioinform. 2019 |
Indexing and storage engines › membership query
approximate membership query |
0.5 | 1 | 2021 | Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership Querying · SIGMOD Conference 2021 |
Data stream processing
sketch |
0.5 | 1 | 2021 | Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership Querying · SIGMOD Conference 2021 |
Bioinformatics and computational biology
sequence analysis |
0.4 | 2 | 2019 | PRESM: personalized reference editor for somatic mutation discovery in cancer genomics · Bioinform. 2019 Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads · Bioinform. 2009 |
Bioinformatics and computational biology › cancer genomics › somatic mutation analysis
somatic mutation detection |
0.4 | 1 | 2019 | PRESM: personalized reference editor for somatic mutation discovery in cancer genomics · Bioinform. 2019 |
Bioinformatics and computational biology › cancer genomics
tumor purity estimation |
0.4 | 1 | 2019 | MEpurity: estimating tumor purity using DNA methylation data · Bioinform. 2019 |
Image and video processing › image restoration
denoising |
0.3 | 1 | 2026 | Cross-Frequency Implicit Neural Representation With Self-Evolving Parameters · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Bioinformatics and computational biology › genomics › structural variation
structural variant analysis |
0.3 | 1 | 2017 | BreakPoint Surveyor: a pipeline for structural variant visualization · Bioinform. 2017 |
Bioinformatics and computational biology › genomics › genome visualization
structural variant visualization |
0.3 | 1 | 2017 | BreakPoint Surveyor: a pipeline for structural variant visualization · Bioinform. 2017 |
Machine learning › Graph learning › graph neural network › graph neural network architecture
spatial graph neural network |
0.3 | 1 | 2025 | STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation · CVPR 2025 |
Bioinformatics and computational biology › sequence analysis
high-throughput sequencing data analysis |
0.2 | 1 | 2016 | Detecting dispersed duplications in high-throughput sequencing data using a database-free approach · Bioinform. 2016 |
Bioinformatics and computational biology › genomics
sequencing |
0.2 | 1 | 2016 | Detecting dispersed duplications in high-throughput sequencing data using a database-free approach · Bioinform. 2016 |
Bioinformatics and computational biology › genomics › structural variation
structural variation detection |
0.2 | 1 | 2016 | Detecting dispersed duplications in high-throughput sequencing data using a database-free approach · Bioinform. 2016 |
Bioinformatics and computational biology
protein sequence analysis |
0.2 | 2 | 2008 | Tracing evolutionary pressure · Bioinform. 2008 An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences · Bioinform. 2007 |
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis |
0.1 | 1 | 2012 | PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data · Bioinform. 2012 |
Bioinformatics and computational biology › transcriptomics › RNA splicing analysis
splice junction detection |
0.1 | 1 | 2012 | PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data · Bioinform. 2012 |
Bioinformatics and computational biology
transcriptomics |
0.1 | 1 | 2012 | PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq data · Bioinform. 2012 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2009 | Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads · Bioinform. 2009 |
Bioinformatics and computational biology › genomics › structural variation
structural variant detection |
0.1 | 1 | 2009 | Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads · Bioinform. 2009 |
Bioinformatics and computational biology
protein function prediction |
0.1 | 1 | 2008 | Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weighting · Bioinform. 2008 |
Bioinformatics and computational biology
phylogenetics |
0.0 | 1 | 2008 | Tracing evolutionary pressure · Bioinform. 2008 |
Bioinformatics and computational biology › phylogenetics
phylogenetic tree analysis |
0.0 | 1 | 2008 | Tracing evolutionary pressure · Bioinform. 2008 |
Bioinformatics and computational biology
protein structure analysis |
0.0 | 1 | 2008 | Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weighting · Bioinform. 2008 |
Data mining › pattern mining
frequent pattern mining |
0.0 | 1 | 2007 | An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences · Bioinform. 2007 |
Data mining
pattern mining |
0.0 | 1 | 2007 | An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences · Bioinform. 2007 |
Methods — techniques the papers use, named apart from their topics
implicit neural representation · 1.7gradient descent · 1.7tensor decomposition · 1.0self-evolving optimization · 1.0haar wavelet transform · 1.0circular shift and coalesce · 0.5read mapping · 0.4germline substitution integration · 0.4beta mixture model · 0.4paired-end sequencing · 0.3RNA-seq expression integration · 0.3paired-end read alignment · 0.2paired-end read mapping · 0.2length distribution comparison · 0.2sequence pattern mining · 0.1pattern growth · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Frequency Implicit Neural Representation With Self-Evolving ParametersabstractImplicit neural representation (INR) has emerged as a powerful paradigm for visual data representation. However, classical INR methods represent data in the original space mixed with different frequency components, and several feature encoding parameters (e.g., the frequency parameter $\omega$ω or the rank $R$R) need manual configurations. In this work, we propose a self-evolving cross-frequency INR using the Haar wavelet transform (termed CF-INR), which decouples data into four frequency components and employs INRs in the wavelet space. CF-INR allows the characterization of different frequency components separately, thus enabling higher accuracy for data representation. To more precisely characterize cross-frequency components, we propose a cross-frequency tensor decomposition paradigm for CF-INR with self-evolving parameters, which automatically updates the rank parameter $R$R and the frequency parameter $\omega$ω for each frequency component through self-evolving optimization. This self-evolution paradigm eliminates the laborious manual tuning of these parameters, and learns a customized cross-frequency feature encoding configuration for each dataset. We evaluate CF-INR on a variety of visual data representation and inverse imaging problems, including image regression, inpainting, denoising, and cloud removal. Extensive experiments demonstrate that CF-INR outperforms state-of-the-art methods in each case. Yi-Si Luo, Kai Ye 0001, Xi-Le Zhao, Deyu Meng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | STINR: Deciphering Spatial Transcriptomics via Implicit Neural RepresentationabstractSpatial transcriptomics (ST) are emerging technologies that reveal spatial distributions of gene expressions within tissues, serving as important ways to uncover biological insights. However, the irregular spatial profiles and variability of genes make it challenging to integrate spatial information with gene expression under a computational framework. Current algorithms mostly utilize spatial graph neural networks to encode spatial information, which may incur increased computational costs and may not be flexible enough to depict complex spatial configurations. In this study, we introduce a concise yet effective representation framework, STINR, for deciphering ST data. STINR leverages an implicit neural representation (INR) to continuously represent ST data, which efficiently characterizes spatial and slice-wise correlations of ST data by inheriting the implicit smoothness of INR. STINR allows easier integration of multiple slices and multi-omics without any alignment, and serves as a potent tool for various biological tasks including gene imputation, gene denoising, spatial domain detection, and cell-type deconvolution stemed from ST data. In particular, STINR identifies the thinnest cortex layer in the dorsolateral prefrontal cortex which previous methods were unable to achieve, and more accurately identifies tumor regions in the human squamous cell carcinoma, showcasing its practical value for biological discoveries. Code at https://github.com/YisiLuo/STINR. Yi-Si Luo, Xi-Le Zhao, Kai Ye 0001, Deyu Meng |
CVPR | 3 |
| 2025 | NeurTV: Total Variation on the Neural DomainabstractAbstract. Recently, we have witnessed the success of total variation (TV) for many imaging applications. However, traditional TV is defined on the original pixel domain, which limits its potential. In this work, we suggest a new TV regularization defined on the neural domain. Concretely, the discrete data is implicitly and continuously represented by a deep neural network (DNN), and we use the derivatives of DNN outputs with respect to (w.r.t.) input coordinates to capture local correlations of data. As compared with classical TV on the original domain, the proposed TV on the neural domain (termed NeurTV) enjoys the following advantages. First, NeurTV is free of discretization error induced by the discrete difference operator. Second, NeurTV is not limited to meshgrid but is suitable for both meshgrid and non-meshgrid data. Third, NeurTV can more exactly capture local correlations across data for any direction and any order of derivatives attributed to the implicit and continuous nature of neural domain. We theoretically reinterpret NeurTV under the variational approximation framework, which allows us to build the connection between NeurTV and classical TV and inspires us to develop variants (e.g., space-variant NeurTV). Extensive numerical experiments with meshgrid data (e.g., color and hyperspectral images) and non-meshgrid data (e.g., point clouds and spatial transcriptomics) showcase the effectiveness of the proposed methods. Yi-Si Luo, Xi-Le Zhao, Kai Ye 0001, Deyu Meng |
SIAM J. Imaging Sci. | 3 |
| 2023 | Comparison and benchmark of structural variants detected from long read and long-read assemblyabstractStructural variant (SV) detection is essential for genomic studies, and long-read sequencing technologies have advanced our capacity to detect SVs directly from read or de novo assembly, also known as read-based and assembly-based strategy. However, to date, no independent studies have compared and benchmarked the two strategies. Here, on the basis of SVs detected by 20 read-based and eight assembly-based detection pipelines from six datasets of HG002 genome, we investigated the factors that influence the two strategies and assessed their performance with well-curated SVs. We found that up to 80% of the SVs could be detected by both strategies among different long-read datasets, whereas variant type, size, and breakpoint detected by read-based strategy were greatly affected by aligners. For the high-confident insertions and deletions at non-tandem repeat regions, a remarkable subset of them (82% in assembly-based calls and 93% in read-based calls), accounting for around 4000 SVs, could be captured by both reads and assemblies. However, discordance between two strategies was largely caused by complex SVs and inversions, which resulted from inconsistent alignment of reads and assemblies at these loci. Finally, benchmarking with SVs at medically relevant genes, the recall of read-based strategy reached 77% on 5X coverage data, whereas assembly-based strategy required 20X coverage data to achieve similar performance. Therefore, integrating SVs from read and assembly is suggested for general-purpose detection because of inconsistently detected complex SVs and inversions, whereas assembly-based strategy is optional for applications with limited resources. Jiadong Lin, Peng Jia 0004, Songbo Wang, Walter A. Kosters, Kai Ye 0001 |
Briefings Bioinform. | 5 |
| 2022 | Homotopic Convex Transformation: A New Landscape Smoothing Method for the Traveling Salesman ProblemabstractThis article proposes a novel landscape smoothing method for the symmetric traveling salesman problem (TSP). We first define the homotopic convex (HC) transformation of a TSP as a convex combination of a well-constructed simple TSP and the original TSP. The simple TSP, called the convex-hull TSP, is constructed by transforming a known local or global optimum. We observe that controlled by the coefficient of the convex combination, with local or global optimum: 1) the landscape of the HC transformed TSP is smoothed in terms that its number of local optima is reduced compared to the original TSP and 2) the fitness distance correlation of the HC transformed TSP is increased. Furthermore, we observe that the smoothing effect of the HC transformation depends highly on the quality of the used optimum. A high-quality optimum leads to a better smoothing effect than a low-quality optimum. We then propose an iterative algorithmic framework in which the proposed HC transformation is combined within a heuristic TSP solver. It works as an escaping scheme from local optima aiming to improve the global searchability of the combined heuristic. Case studies using the 3-Opt and the Lin-Kernighan local search as the heuristic solver show that the resultant algorithms significantly outperform their counterparts and two other smoothing-based TSP heuristic solvers on most of the test instances with up to 20 000 cities. Jialong Shi, Jianyong Sun, Qingfu Zhang 0001, Kai Ye 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership QueryingabstractGiven a set S, Membership Querying (MQ) answers whether a query element $q\in S$. It is a fundamental task in areas like database systems and computer networks. In this paper, we consider a more general problem, Multi-Set Multi-Membership Querying (MS-MMQ). Given n sets $S_0,łdots,S_n-1 $, MS-MMQ answers which sets contain element q. A direct way to address MS-MMQ is to build an MQ structure (e.g., Bloom Filter) for each set. However, the query and space complexities grow linearly with n and become prohibitive for a large n. To address this challenge, we propose a novel Circular Shift and Coalesce (CSC) framework to efficiently achieve approximate MS-MMQ. Instead of building an MQ data structure for each set, the CSC index encodes all n sets into a compact sketch and retrieves only a few bytes in the sketch for a query, which achieves high memory-efficiency and boosts the query speed by several times. CSC is compatible with mainstream data structures for Approximate MQ. We conduct experiments on real-world datasets and results demonstrate that our framework is up to 91.2 times faster and up to 48.9 times more accurate than state-of-the-art methods. Rundong Li 0002, Pinghui Wang, Jiongli Zhu, Junzhou Zhao, Jia Di, Xiaofei Yang 0003, Kai Ye 0001 |
SIGMOD Conference | 7 |
| 2020 | From Innovations to Prospects: What Is Hidden Behind Cryptocurrencies?abstractThe great influence of Bitcoin has promoted the rapid development of blockchain-based digital currencies, especially the altcoins, since 2013. However, most altcoins share similar source codes, resulting in concerns about code innovations. In this paper, an empirical study on existing altcoins is carried out to offer a thorough understanding of various aspects associated with altcoin innovations. Firstly, we construct the dataset of altcoins, including source code repositories, GitHub fork relations, and market capitalizations (cap). Then, we analyze the altcoin innovations from the perspective of source code similarities. The results demonstrate that more than 85% of altcoin repositories present high code similarities. Next, a temporal clustering algorithm is proposed to mine the inheritance relationship among various altcoins. The family pedigrees of altcoin are constructed, in which the altcoin presents similar evolution features as biology, such as power-law in family size, variety in family evolution, etc. Finally, we investigate the correlation between code innovations and market capitalization. Although we fail to predict the price of altcoins based on their code similarities, the results show that altcoins with higher innovations reflect better market prospects. Ang Jia, Ming Fan 0002, Wenying Wei, Zijiang Yang 0006, Kai Ye 0001, Ting Liu 0002 |
MSR | 7 |
| 2019 | PRESM: personalized reference editor for somatic mutation discovery in cancer genomicsabstractMOTIVATION: Accurate detection of somatic mutations is a crucial step toward understanding cancer. Various tools have been developed to detect somatic mutations from cancer genome sequencing data by mapping reads to a universal reference genome and inferring likelihoods from complex statistical models. However, read mapping is frequently obstructed by mismatches between germline and somatic mutations on a read and the reference genome. Previous attempts to develop personalized genome tools are not compatible with downstream statistical models for somatic mutation detection. RESULTS: We present PRESM, a tool that builds personalized reference genomes by integrating germline mutations into the reference genome. The aforementioned obstacle is circumvented by using a two-step germline substitution procedure, maintaining positional fidelity using an innovative workaround. Reads derived from tumor tissue can be positioned more accurately along a personalized reference than a universal reference due to the reduced genetic distance between the subject (tumor genome) and the target (the personalized genome). Application of PRESM's personalized genome reduced false-positive (FP) somatic mutation calls by as much as 55.5%, and facilitated the discovery of a novel somatic point mutation on a germline insertion in PDE1A, a phosphodiesterase associated with melanoma. Moreover, all improvements in calling accuracy were achieved without parameter optimization, as PRESM itself is parameter-free. Hence, similar increases in read mapping and decreases in the FP rate will persist when PRESM-built genomes are applied to any user-provided dataset. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/precisionomics/PRESM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chen Cao 0002, Lauren Mak, Guangxu Jin, Paul Gordon, Kai Ye 0001 |
Bioinform. | 5 |
| 2019 | MEpurity: estimating tumor purity using DNA methylation dataabstractMOTIVATION: Tumor purity is a fundamental property of each cancer sample and affects downstream investigations. Current tumor purity estimation methods either require matched normal sample or report moderately high tumor purity even on normal samples. It is critical to develop a novel computational approach to estimate tumor purity with sufficient precision based on tumor-only sample. RESULTS: In this study, we developed MEpurity, a beta mixture model-based algorithm, to estimate the tumor purity based on tumor-only Illumina Infinium 450k methylation microarray data. We applied MEpurity to both The Cancer Genome Atlas (TCGA) cancer data and cancer cell line data, demonstrating that MEpurity reports low tumor purity on normal samples and comparable results on tumor samples with other state-of-art methods. AVAILABILITY AND IMPLEMENTATION: MEpurity is a C++ program which is available at https://github.com/xjtu-omics/MEpurity. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaofei Yang 0003, Tingjie Wang, Jiadong Lin, Yongyong Kang, Peng Jia 0004, Kai Ye 0001 |
Bioinform. | 7 |
| 2019 | Learning From a Stream of Nonstationary and Dependent Data in Multiobjective Evolutionary OptimizationabstractCombining machine learning techniques has shown great potentials in evolutionary optimization since the domain knowledge of an optimization problem, if well learned, can be a great help for creating high-quality solutions. However, existing learning-based multiobjective evolutionary algorithms (MOEAs) spend too much computational overhead on learning. To address this problem, we propose a learning-based MOEA where an online learning algorithm is embedded within the evolutionary search procedure. The online learning algorithm takes the stream of sequentially generated solutions along the evolution as its training data. It is noted that the stream of solutions are temporal, dependent, nonstationary, and nonstatic. These data characteristics make existing online learning algorithm not suitable for the evolution data. We hence modify an existing online agglomerative clustering algorithm to accommodate these characteristics. The modified online clustering algorithm is applied to adaptively discover the structure of the Pareto optimal set; and the learned structure is used to guide new solution creation. Experimental results have shown significant improvement over four state-of-the-art MOEAs on a variety of benchmark problems. Jianyong Sun, Hu Zhang 0002, Aimin Zhou, Qingfu Zhang 0001, Ke Zhang 0020, Zhenbiao Tu, Kai Ye 0001 |
IEEE Trans. Evol. Comput. | 7 |
| 2017 | BreakPoint Surveyor: a pipeline for structural variant visualizationabstractSUMMARY: BreakPoint Surveyor (BPS) is a computational pipeline for the discovery, characterization, and visualization of complex genomic rearrangements, such as viral genome integration, in paired-end sequence data. BPS facilitates interpretation of structural variants by merging structural variant breakpoint predictions, gene exon structure, read depth, and RNA-sequencing expression into a single comprehensive figure. AVAILABILITY AND IMPLEMENTATION: Source code and sample data freely available for download at https://github.com/ding-lab/BreakPointSurveyor, distributed under the GNU GPLv3 license, implemented in R, Python and BASH scripts, and supported on Unix/Linux/OS X operating systems. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthew A. Wyczalkowski, Kristine M. Wylie, Michael D. McLellan, Jennifer Flynn, Mo Huang, Kai Ye 0001, Xian Fan, Ken Chen 0001, Michael C. Wendl |
Bioinform. | 7 |
| 2016 | Detecting dispersed duplications in high-throughput sequencing data using a database-free approachabstractMOTIVATION: Dispersed duplications (DDs) such as transposon element insertions and copy number variations are ubiquitous in the human genome. They have attracted the interest of biologists as well as medical researchers due to their role in both evolution and disease. The efforts of discovering DDs in high-throughput sequencing data are currently dominated by database-oriented approaches that require pre-existing knowledge of the DD elements to be detected. RESULTS: We present DD_DETECTION, a database-free approach to finding DD events in high-throughput sequencing data. DD_DETECTION is able to detect DDs purely from paired-end read alignments. We show in a comparative study that this method is able to compete with database-oriented approaches in recovering validated transposon insertion events. We also experimentally validate the predictions of DD_DETECTION on a human DNA sample, showing that it can find not only duplicated elements present in common databases but also DDs of novel type. AVAILABILITY AND IMPLEMENTATION: The software presented in this article is open source and available from https://bitbucket.org/mkroon/dd_detection. M. Kroon, Eric-Wubbo Lameijer, N. Lakenberg, Jayne Y. Hehir-Kwa, D. T. Thung, P. Eline Slagboom, Joost N. Kok, Kai Ye 0001 |
Bioinform. | 8 |
| 2014 | MSIsensor: microsatellite instability detection using paired tumor-normal sequence dataabstractMOTIVATION: Microsatellite instability (MSI) is an important indicator of larger genome instability and has been linked to many genetic diseases, including Lynch syndrome. MSI status is also an independent prognostic factor for favorable survival in multiple cancer types, such as colorectal and endometrial. It also informs the choice of chemotherapeutic agents. However, the current PCR-electrophoresis-based detection procedure is laborious and time-consuming, often requiring visual inspection to categorize samples. We developed MSIsensor, a C++ program for automatically detecting somatic microsatellite changes. It computes length distributions of microsatellites per site in paired tumor and normal sequence data, subsequently using these to statistically compare observed distributions in both samples. Comprehensive testing indicates MSIsensor is an efficient and effective tool for deriving MSI status from standard tumor-normal paired sequence data. AVAILABILITY AND IMPLEMENTATION: https://github.com/ding-lab/msisensor Beifang Niu, Kai Ye 0001, Qunyuan Zhang, Charles Lu 0002, Mingchao Xie, Michael D. McLellan, Michael C. Wendl |
Bioinform. | 2 |
| 2012 | PASSion: a pattern growth algorithm-based pipeline for splice junction detection in paired-end RNA-Seq dataabstractMOTIVATION: RNA-seq is a powerful technology for the study of transcriptome profiles that uses deep-sequencing technologies. Moreover, it may be used for cellular phenotyping and help establishing the etiology of diseases characterized by abnormal splicing patterns. In RNA-Seq, the exact nature of splicing events is buried in the reads that span exon-exon boundaries. The accurate and efficient mapping of these reads to the reference genome is a major challenge. RESULTS: We developed PASSion, a pattern growth algorithm-based pipeline for splice site detection in paired-end RNA-Seq reads. Comparing the performance of PASSion to three existing RNA-Seq analysis pipelines, TopHat, MapSplice and HMMSplicer, revealed that PASSion is competitive with these packages. Moreover, the performance of PASSion is not affected by read length and coverage. It performs better than the other three approaches when detecting junctions in highly abundant transcripts. PASSion has the ability to detect junctions that do not have known splicing motifs, which cannot be found by the other tools. Of the two public RNA-Seq datasets, PASSion predicted ≈ 137,000 and 173,000 splicing events, of which on average 82 are known junctions annotated in the Ensembl transcript database and 18% are novel. In addition, our package can discover differential and shared splicing patterns among multiple samples. AVAILABILITY: The code and utilities can be freely downloaded from https://trac.nbic.nl/passion and ftp://ftp.sanger.ac.uk/pub/zn1/passion. Yanju Zhang, Eric-Wubbo Lameijer, Peter A. C. 't Hoen, Zemin Ning, P. Eline Slagboom, Kai Ye 0001 |
Bioinform. | 6 |
| 2009 | Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short readsabstractMOTIVATION: There is a strong demand in the genomic community to develop effective algorithms to reliably identify genomic variants. Indel detection using next-gen data is difficult and identification of long structural variations is extremely challenging. RESULTS: We present Pindel, a pattern growth approach, to detect breakpoints of large deletions and medium-sized insertions from paired-end short reads. We use both simulated reads and real data to demonstrate the efficiency of the computer program and accuracy of the results. AVAILABILITY: The binary code and a short user manual can be freely downloaded from http://www.ebi.ac.uk/ approximately kye/pindel/. CONTACT: [email protected]; [email protected]. Kai Ye 0001, Marcel H. Schulz, Rolf Apweiler, Zemin Ning |
Bioinform. | 1 |
| 2008 | Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weightingabstractMOTIVATION: Identification of residues that account for protein function specificity is crucial, not only for understanding the nature of functional specificity, but also for protein engineering experiments aimed at switching the specificity of an enzyme, regulator or transporter. Available algorithms generally use multiple sequence alignments to identify residue positions conserved within subfamilies but divergent in between. However, many biological examples show a much subtler picture than simple intra-group conservation versus inter-group divergence. RESULTS: We present multi-RELIEF, a novel approach for identifying specificity residues that is based on RELIEF, a state-of-the-art Machine-Learning technique for feature weighting. It estimates the expected 'local' functional specificity of residues from an alignment divided in multiple classes. Optionally, 3D structure information is exploited by increasing the weight of residues that have high-weight neighbors. Using ROC curves over a large body of experimental reference data, we show that (a) multi-RELIEF identifies specificity residues for the seven test sets used, (b) incorporating structural information improves prediction for specificity of interaction with small molecules and (c) comparison of multi-RELIEF with four other state-of-the-art algorithms indicates its robustness and best overall performance. AVAILABILITY: A web-server implementation of multi-RELIEF is available at www.ibi.vu.nl/programs/multirelief. Matlab source code of the algorithm and data sets are available on request for academic use. Kai Ye 0001, K. Anton Feenstra, Jaap Heringa, Adriaan P. IJzerman, Elena Marchiori |
Bioinform. | 1 |
| 2008 | Tracing evolutionary pressureabstractMOTIVATION: Recent advances in sequencing techniques have yielded enormous amounts of protein sequence data from various species. This large dataset allows sequence comparison between paralogous and orthologous proteins to identify motifs or functional positions that account for the differences of functional subgroups ('specificity' positions). Algorithms such as SDPpred and the two-entropies analysis (TEA) have been developed to detect such specificity positions from a multiple sequence alignment (MSA) grouped into classes according to certain biological functions. Other algorithms such as TreeDet compute a classification and then predict specificity positions associated with it. However, there are still many unresolved questions: Was the optimal subdivision of a protein family achieved? Do the definitions at different levels of the phylogenetic tree affect the prediction of specificity positions? Can the whole phylogenetic tree be used instead of only one level in it to predict specificity positions? RESULTS: Here we present a novel method, TEA-O (Two-entropies analysis-Objective), to trace the evolutionary pressure from the root to the branches of the phylogenetic tree. At each level of the tree, a TEA plot is produced to capture the signal of the evolutionary pressure. A consensus TEA-O plot is composed from the whole series of plots to provide a condensed representation. Positions related to functions that evolved early (conserved) or later (specificity) are close to the lower-left or upper-left corner of the TEA-O plot, respectively. This novel approach allows an unbiased, user-independent, analysis of residue relevance in a protein family. We compared our TEA-O method with various algorithms using both synthetic and real protein sequences. The results show that our method is robust, sensitive to subtle differences in evolutionary pressure during evolution and comprehensive because all positions in the MSA are presented in the consensus plot. AVAILABILITY: All computer programs and datasets used in this work are available at http://nava.liacs.nl/kye/TEA-O/ for academic use. Kai Ye 0001, Gert Vriend, Adriaan P. IJzerman |
Bioinform. | 1 |
| 2007 | An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequencesabstractMOTIVATION: Pattern discovery in protein sequences is often based on multiple sequence alignments (MSA). The procedure can be computationally intensive and often requires manual adjustment, which may be particularly difficult for a set of deviating sequences. In contrast, two algorithms, PRATT2 (http//www.ebi.ac.uk/pratt/) and TEIRESIAS (http://cbcsrv.watson.ibm.com/) are used to directly identify frequent patterns from unaligned biological sequences without an attempt to align them. Here we propose a new algorithm with more efficiency and more functionality than both PRATT2 and TEIRESIAS, and discuss some of its applications to G protein-coupled receptors, a protein family of important drug targets. RESULTS: In this study, we designed and implemented six algorithms to mine three different pattern types from either one or two datasets using a pattern growth approach. We compared our approach to PRATT2 and TEIRESIAS in efficiency, completeness and the diversity of pattern types. Compared to PRATT2, our approach is faster, capable of processing large datasets and able to identify the so-called type III patterns. Our approach is comparable to TEIRESIAS in the discovery of the so-called type I patterns but has additional functionality such as mining the so-called type II and type III patterns and finding discriminating patterns between two datasets. AVAILABILITY: The source code for pattern growth algorithms and their pseudo-code are available at http://www.liacs.nl/home/kosters/pg/. Kai Ye 0001, Walter A. Kosters, Adriaan P. IJzerman |
Bioinform. | 1 |