VLDB 2026 Research / reviewers in the wild / expert
Hiroto Saigo
dblp:12/898
· DBLP profile ↗
25ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0001-5314-5367ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 3 first-authorDatabases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAPSE-Synergy: Balancing Formal Soundness and Empirical Coverage in Neural Theorem ProvingabstractAbstract Large language models (LLMs) can propose candidate lemmas, yet translating natural-language reasoning into verified Coq artifacts remains difficult due to semantic drift, missing abstractions, and weak type discipline. Most existing systems treat the prover as a black-box filter: as long as some candidates verify, little is known about which transformations are semantically safe, or how much coverage must be sacrificed to obtain formal guarantees. We present Semantic Alignment for Proof Strategy Extraction (SAPSE), a retrieval-first framework that links natural-language proof steps with formal lemmas through dual-domain embeddings and type-aware abstraction. At its core, SAPSE introduces an abstract syntax tree (AST) sanitizer whose verified fragment, mechanized in Rocq 9.1.0 for a minimal calculus, consists of two concretely implemented transformations, require injection and equality canonicalization, together with a parametric binder-normalization schema that is currently instantiated as the identity on terms. For all admissible inputs in this calculus, these verified passes are proved to preserve both typing and logical equivalence under import-only context extension. A production sanitizer extends this core with unverified but practical passes: scope resolution, list parameterization, and formatting for complete Coq syntax, while reusing the same verified interface. On top of this verified core, we develop an adaptive Synergy pipeline that first applies unverified heuristic repairs to maximize empirical coverage and then invokes the verified core under admissibility guards to enforce soundness. Rather than optimizing raw verification accuracy, Synergy exposes a reproducible safety-coverage frontier: on a 2,000-lemma real Coq benchmark, a retrieval-only baseline verifies 37.6% of generated candidates, while the Synergy configuration verifies 32.8% with zero unsafe rewrites among guarded repairs and comparable runtime. Fragment-coverage analysis shows that 98.1% of benchmark lemmas lie in the mechanized fragment and that every Synergy success falls inside this fragment, so the AST-level soundness theorem applies directly. A differential analysis of the 96 lemmas that the retrieval-only baseline verifies but Synergy fails on decomposes these “lost successes” by semantic category, structural complexity, and failure mode, turning the observed performance gap into a diagnostic tool for future verified transformations. Overall, SAPSE-Synergy provides a partially mechanized but practically effective bridge between probabilistic lemma generation and formally verified transformation. It quantifies, for the first time, how much empirical coverage must be traded for a small, compositional verified core that safely mediates heuristic repairs in neural theorem proving. (Source code and full experimental artifacts are publicly available at: https://github.com/leochenminrui/SAPSE-Synergy ) Minrui Chen 0001, Huidong Jiang, Hiroto Saigo |
FM (1) | 3 |
| 2024 | A Branch-and-Bound Approach to Efficient Classification and Retrieval of Documents
Kotaro Ii, Hiroto Saigo, Yasuo Tabei |
ICPRAM | 2 |
| 2024 | Benchmarking a Wide Range of Unsupervised Learning Methods for Detecting Anomaly in Blast FurnaceabstractSteel plays important roles in our daily lives, as it surrounds us in the form of various products. Blast furnace, one of the main facility in steel production process, is traditionally monitored by skilled workers to prevent incidents. However, there is a growing demand to automate the monitoring process by leveraging machine learning. This paper focuses on investigating the suitability of unsupervised learning methods for detecting anomalies in blast furnaces. Extensive benchmarking is conducted using a dataset collected from blast furnaces, encompassing a wide range of unsupervised learning methods, including both traditional approaches and recent deep learning-based techniques. The computational experiments yield results that suggest the effectiveness of traditional methods over deep learning-based methods. To validate this observation, additional experiments are performed on publicly available non time series datasets and complex time series datasets. These experiments serve to confirm the superiority of traditional methods in handling non time series datasets, while deep learning methods exhibit better performance in dealing with complex time series datasets. We have also discovered that dimensionality reduction before anomaly detection is beneficial in eliminating outliers and effectively modeling the normal data points in the blast furnace dataset. Kendai Itakura, Dukka B. KC, Hiroto Saigo |
ICPRAM | 3 |
| 2023 | pLMSNOSite: an ensemble-based approach for predicting protein S-nitrosylation sites by integrating supervised word embedding and embedding from pre-trained protein language modelabstractBACKGROUND: Protein S-nitrosylation (SNO) plays a key role in transferring nitric oxide-mediated signals in both animals and plants and has emerged as an important mechanism for regulating protein functions and cell signaling of all main classes of protein. It is involved in several biological processes including immune response, protein stability, transcription regulation, post translational regulation, DNA damage repair, redox regulation, and is an emerging paradigm of redox signaling for protection against oxidative stress. The development of robust computational tools to predict protein SNO sites would contribute to further interpretation of the pathological and physiological mechanisms of SNO. RESULTS: Using an intermediate fusion-based stacked generalization approach, we integrated embeddings from supervised embedding layer and contextualized protein language model (ProtT5) and developed a tool called pLMSNOSite (protein language model-based SNO site predictor). On an independent test set of experimentally identified SNO sites, pLMSNOSite achieved values of 0.340, 0.735 and 0.773 for MCC, sensitivity and specificity respectively. These results show that pLMSNOSite performs better than the compared approaches for the prediction of S-nitrosylation sites. CONCLUSION: Together, the experimental results suggest that pLMSNOSite achieves significant improvement in the prediction performance of S-nitrosylation sites and represents a robust computational approach for predicting protein S-nitrosylation sites. pLMSNOSite could be a useful resource for further elucidation of SNO and is publicly available at https://github.com/KCLabMTU/pLMSNOSite . Pawel Pratyush, Suresh Pokharel, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 3 |
| 2022 | Correction: DeepSuccinylSite: a deep learning based approach for protein succinylation site predictionabstractResults: Using an independent test set of experimentally identified succinylation sites, our method achieved efficiency scores of 79%, 68.7% and 0.27 for sensitivity, specificity and MCC respectively, with an area under the receiver operator characteristic (ROC) curve of 0.8.In side-by-side comparisons with previously described succinylation site predictors, DeepSuccinylSite produces similar or better results compared to the other state-of-the-art predictors.On page 7, Last paragraph on right should be changed from Consequently, DeepSuccinylSite achieved a significantly higher performance as measured by MCC.Indeed, DeepSuccinylSite exhibited an ~ 62% increase in MCC when compared to the next highest method, GPSuc.to: Consequently, DeepSuccinylSite achieved an MCC score (at decision boundary of 0.5) on par with the top performingmethod, GPSuc.On page 2, in Table 1, the negative data of Independent Test should be 2977 rather than 254.On page 8, in Table 6, the MCC data of DeepSuccinylSite should be 0.27 rather than 0.48. Niraj Thapa, Meenal Chaudhari, Sean McManus, Kaushik Roy 0003, Robert H. Newman, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 6 |
| 2021 | Topic modeling for sequential documents based on hybrid inter-document topic dependency
Wenbo Li 0011, Hiroto Saigo, Bin Tong, Einoshin Suzuki |
J. Intell. Inf. Syst. | 2 |
| 2020 | Automatically mining Relevant Variable Interactions Via Sparse Bayesian Learning
Ryoichiro Yafune, Daisuke Sakuma, Mirai Takayanagi, Yasuo Tabei, Noritaka Saito, Hiroto Saigo |
ICPR | 6 |
| 2020 | Context-Aware Latent Dirichlet Allocation for Topic Segmentation
Wenbo Li 0011, Tetsu Matsukawa, Hiroto Saigo, Einoshin Suzuki |
PAKDD (1) | 3 |
| 2020 | DeepSuccinylSite: a deep learning based approach for protein succinylation site predictionabstractBACKGROUND: Protein succinylation has recently emerged as an important and common post-translation modification (PTM) that occurs on lysine residues. Succinylation is notable both in its size (e.g., at 100 Da, it is one of the larger chemical PTMs) and in its ability to modify the net charge of the modified lysine residue from + 1 to - 1 at physiological pH. The gross local changes that occur in proteins upon succinylation have been shown to correspond with changes in gene activity and to be perturbed by defects in the citric acid cycle. These observations, together with the fact that succinate is generated as a metabolic intermediate during cellular respiration, have led to suggestions that protein succinylation may play a role in the interaction between cellular metabolism and important cellular functions. For instance, succinylation likely represents an important aspect of genomic regulation and repair and may have important consequences in the etiology of a number of disease states. In this study, we developed DeepSuccinylSite, a novel prediction tool that uses deep learning methodology along with embedding to identify succinylation sites in proteins based on their primary structure. RESULTS: Using an independent test set of experimentally identified succinylation sites, our method achieved efficiency scores of 79%, 68.7% and 0.48 for sensitivity, specificity and MCC respectively, with an area under the receiver operator characteristic (ROC) curve of 0.8. In side-by-side comparisons with previously described succinylation predictors, DeepSuccinylSite represents a significant improvement in overall accuracy for prediction of succinylation sites. CONCLUSION: Together, these results suggest that our method represents a robust and complementary technique for advanced exploration of protein succinylation. Niraj Thapa, Meenal Chaudhari, Sean McManus, Kaushik Roy 0003, Robert H. Newman, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 6 |
| 2018 | Entire Regularization Path for Sparse Nonnegative Interaction ModelabstractBuilding sparse combinatorial model with non-negative constraint is essential in solving real-world problems such as in biology, in which the target response is often formulated by additive linear combination of features variables. This paper presents a solution to this problem by combining itemset mining with non-negative least squares. However, once incorporation of modern regularization is considered, then a naive solution requires to solve expensive enumeration problem many times for every regularization parameter. In this paper, we devise a regularization path tracking algorithm such that combinatorial feature is searched and included one by one to the solution set. Our contribution is a proposal of novel bounds specifically designed for the feature search problem. In synthetic dataset, the proposed method is demonstrated to run orders of magnitudes faster than a naive counterpart which does not employ tree pruning. We also empirically show that non-negativity constraints can reduce the number of active features much less than that of LASSO, leading to significant speed-ups in pattern search. In experiments using HIV-1 drug resistance dataset, the proposed method could successfully model the rapidly increasing drug resistance triggered by accumulation of mutations in HIV-1 genetic sequences. We also demonstrate the effectiveness of non-negativity constraints in suppressing false positive features, resulting in a model with smaller number of features and thereby improved interpretability. Mirai Takayanagi, Yasuo Tabei, Hiroto Saigo |
ICDM | 3 |
| 2018 | RF-NR: Random Forest Based Approach for Improved Classification of Nuclear ReceptorsabstractThe Nuclear Receptor (NR) superfamily plays an important role in key biological, developmental, and physiological processes. Developing a method for the classification of NR proteins is an important step towards understanding the structure and functions of the newly discovered NR protein. The recent studies on NR classification are either unable to achieve optimum accuracy or are not designed for all the known NR subfamilies. In this study, we developed RF-NR, which is a Random Forest based approach for improved classification of nuclear receptors. The RF-NR can predict whether a query protein sequence belongs to one of the eight NR subfamilies or it is a non-NR sequence. The RF-NR uses spectrum-like features namely: Amino Acid Composition, Di-peptide Composition, and Tripeptide Composition. Benchmarking on two independent datasets with varying sequence redundancy reduction criteria, the RF-NR achieves better (or comparable) accuracy than other existing methods. The added advantage of our approach is that we can also obtain biological insights about the important features that are required to classify NR subfamilies. RF-NR is freely available at http://bcb.ncat.edu/RF_NR. Hamid D. Ismail, Hiroto Saigo, Dukka B. KC |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Structural Class Classification of 3D Protein Structure Based on Multi-View 2D ImagesabstractComputing similarity or dissimilarity between protein structures is an important task in structural biology. A conventional method to compute protein structure dissimilarity requires structural alignment of the proteins. However, defining one best alignment is difficult, especially when the structures are very different. In this paper, we propose a new similarity measure for protein structure comparisons using a set of multi-view 2D images of 3D protein structures. In this approach, each protein structure is represented by a subspace from the image set. The similarity between two protein structures is then characterized by the canonical angles between the two subspaces. The primary advantage of our method is that precise alignment is not needed. We employed Grassmann Discriminant Analysis (GDA) as the subspace-based learning in the classification framework. We applied our method for the classification problem of seven SCOP structural classes of protein 3D structures. The proposed method outperformed the k-nearest neighbor method (k-NN) based on conventional alignment-based methods CE, FATCAT, and TM-align. Our method was also applied to the classification of SCOP folds of membrane proteins, where the proposed method could recognize the fold HEM-binding four-helical bundle (f.21) much better than TM-Align. Chendra Hadi Suryanto, Hiroto Saigo, Kazuhiro Fukui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | CNN-BLPred: a Convolutional neural network based predictor for β-Lactamases (BL) and their classesabstractBACKGROUND: The β-Lactamase (BL) enzyme family is an important class of enzymes that plays a key role in bacterial resistance to antibiotics. As the newly identified number of BL enzymes is increasing daily, it is imperative to develop a computational tool to classify the newly identified BL enzymes into one of its classes. There are two types of classification of BL enzymes: Molecular Classification and Functional Classification. Existing computational methods only address Molecular Classification and the performance of these existing methods is unsatisfactory. RESULTS: We addressed the unsatisfactory performance of the existing methods by implementing a Deep Learning approach called Convolutional Neural Network (CNN). We developed CNN-BLPred, an approach for the classification of BL proteins. The CNN-BLPred uses Gradient Boosted Feature Selection (GBFS) in order to select the ideal feature set for each BL classification. Based on the rigorous benchmarking of CCN-BLPred using both leave-one-out cross-validation and independent test sets, CCN-BLPred performed better than the other existing algorithms. Compared with other architectures of CNN, Recurrent Neural Network, and Random Forest, the simple CNN architecture with only one convolutional layer performs the best. After feature extraction, we were able to remove ~95% of the 10,912 features using Gradient Boosted Trees. During 10-fold cross validation, we increased the accuracy of the classic BL predictions by 7%. We also increased the accuracy of Class A, Class B, Class C, and Class D performance by an average of 25.64%. The independent test results followed a similar trend. CONCLUSIONS: We implemented a deep learning algorithm known as Convolutional Neural Network (CNN) to develop a classifier for BL classification. Combined with feature selection on an exhaustive feature set and using balancing method such as Random Oversampling (ROS), Random Undersampling (RUS) and Synthetic Minority Oversampling Technique (SMOTE), CNN-BLPred performs significantly better than existing algorithms for BL classification. Clarence White, Hamid D. Ismail, Hiroto Saigo, Dukka B. KC |
BMC Bioinform. | 3 |
| 2016 | Scalable Partial Least Squares Regression on Grammar-Compressed Data MatricesabstractWith massive high-dimensional data now commonplace in research and industry, there is a strong and growing demand for more scalable computational techniques for data analysis and knowledge discovery. Key to turning these data into knowledge is the ability to learn statistical models with high interpretability. Current methods for learning statistical models either produce models that are not interpretable or have prohibitive computational costs when applied to massive data. In this paper we address this need by presenting a scalable algorithm for partial least squares regression (PLS), which we call compression-based PLS (cPLS), to learn predictive linear models with a high interpretability from massive high-dimensional data. We propose a novel grammar-compressed representation of data matrices that supports fast row and column access while the data matrix is in a compressed form. The original data matrix is grammar-compressed and then the linear model in PLS is learned on the compressed data matrix, which results in a significant reduction in working space, greatly improving scalability. We experimentally test cPLS on its ability to learn linear models for classification, regression and feature extraction with various massive high-dimensional data, and show that cPLS performs superiorly in terms of prediction accuracy, computational efficiency, and interpretability. Yasuo Tabei, Hiroto Saigo, Yoshihiro Yamanishi, Simon J. Puglisi |
KDD | 2 |
| 2010 | Reaction graph kernels predict EC numbers of unknown enzymatic reactions in plant secondary metabolismabstractBACKGROUND: Understanding of secondary metabolic pathway in plant is essential for finding druggable candidate enzymes. However, there are many enzymes whose functions are not yet discovered in organism-specific metabolic pathways. Towards identifying the functions of those enzymes, assignment of EC numbers to the enzymatic reactions they catalyze plays a key role, since EC numbers represent the categorization of enzymes on one hand, and the categorization of enzymatic reactions on the other hand. RESULTS: We propose reaction graph kernels for automatically assigning EC numbers to unknown enzymatic reactions in a metabolic network. Reaction graph kernels compute similarity between two chemical reactions considering the similarity of chemical compounds in reaction and their relationships. In computational experiments based on the KEGG/REACTION database, our method successfully predicted the first three digits of the EC number with 83% accuracy. We also exhaustively predicted missing EC numbers in plant's secondary metabolism pathway. The prediction results of reaction graph kernels on 36 unknown enzymatic reactions are compared with an expert's knowledge. Using the same data for evaluation, we compared our method with E-zyme, and showed its ability to assign more number of accurate EC numbers. CONCLUSION: Reaction graph kernels are a new metric for comparing enzymatic reactions. Hiroto Saigo, Masahiro Hattori, Hisashi Kashima, Koji Tsuda |
BMC Bioinform. | 1 |
| 2009 | A Bayesian Approach to Graphy Regression with Relevant Subgraph SelectionabstractMany real-world applications with graph data require the solution of a given regression task as well as the identification of the subgraphs which are relevant for the task. In these cases graphs are commonly represented as high dimensional binary vectors of indicators of subgraphs. However, since the dimensionality of such indicator vectors can be high even for small datasets, traditional regression algorithms become intractable and past approaches used to preselect a feasible subset of subgraphs. A different approach was recently proposed by a Lasso-type method where the objective function optimization with a large number of variables is reformulated as a dual mathematical programming problem with a small number of variables but a large number of constraints. The dual problem is then solved by column generation, where the subgraphs corresponding to the most violated constraints are found by weighted subgraph mining. This paper proposes an extension of this method to a Bayesian approach in which the regression parameters are considered as random variables and integrated out from the model likelihood, thus providing a posterior distribution on the target variable as opposed to a point estimate. We focus on a linear regression model with a Gaussian prior distribution on the parameters. We evaluate our approach on several molecular graph datasets and analyze whether the uncertainty in the target estimate given by the target posterior distribution variance can be used to improve model performance and therefore provides useful additional information. Silvia Chiappa, Hiroto Saigo, Koji Tsuda |
SDM | 2 |
| 2009 | gBoost: a mathematical programming approach to graph classification and regressionabstractGraph mining methods enumerate frequently appearing subgraph patterns, which can be used as features for subsequent classification or regression. However, frequent patterns are not necessarily informative for the given learning problem. We propose a mathematical programming boosting method (gBoost) that progressively collects informative patterns. Compared to AdaBoost, gBoost can build the prediction rule with fewer iterations. To apply the boosting method to graph data, a branch-and-bound pattern search algorithm is developed based on the DFS code tree. The constructed search space is reused in later iterations to minimize the computation time. Our method can learn more efficiently than the simpler method based on frequent substructure mining, because the output labels are used as an extra information source for pruning the search space. Furthermore, by engineering the mathematical program, a wide range of machine learning problems can be solved without modifying the pattern search algorithm. Hiroto Saigo, Sebastian Nowozin, Tadashi Kadowaki, Taku Kudo, Koji Tsuda |
Mach. Learn. | 1 |
| 2008 | Iterative Subgraph Mining for Principal Component AnalysisabstractGraph mining methods enumerate frequent subgraphs efficiently, but they are not necessarily good features for machine learning due to high correlation among features. Thus it makes sense to perform principal component analysis to reduce the dimensionality and create decorrelated features. We present a novel iterative mining algorithm that captures informative patterns corresponding to major entries of top principal components. It repeatedly calls weighted substructure mining where example weights are updated in each iteration. The Lanczos algorithm, a standard algorithm of eigen decomposition, is employed to update the weights. In experiments, our patterns are shown to approximate the principal components obtained by frequent mining. Hiroto Saigo, Koji Tsuda |
ICDM | 1 |
| 2008 | Regression with interval output valuesabstractWe consider a regression problem where target values are given as intervals, and propose a statistical approach to it. Although it is hard to solve the optimization problem directly, we propose an approximation method based on the EM algorithm. Experiments using the benchmark datasets show effectiveness of our approach. Hisashi Kashima, Kazutaka Yamasaki, Akihiro Inokuchi, Hiroto Saigo |
ICPR | 4 |
| 2008 | Partial least squares regression for graph miningabstractAttributed graphs are increasingly more common in many application domains such as chemistry, biology and text processing. A central issue in graph mining is how to collect informative subgraph patterns for a given learning task. We propose an iterative mining method based on partial least squares regression (PLS). To apply PLS to graph data, a sparse version of PLS is developed first and then it is combined with a weighted pattern mining algorithm. The mining algorithm is iteratively called with different weight vectors, creating one latent component per one mining call. Our method, graph PLS, is efficient and easy to implement, because the weight vector is updated with elementary matrix calculations. In experiments, our graph PLS algorithm showed competitive prediction accuracies in many chemical datasets and its efficiency was significantly superior to graph boosting (gBoost) and the naive method based on frequent graph mining. Hiroto Saigo, Nicole Krämer 0002, Koji Tsuda |
KDD | 1 |
| 2007 | Mining complex genotypic features for predicting HIV-1 drug resistanceabstractMOTIVATION: Human immunodeficiency virus type 1 (HIV-1) evolves in human body, and its exposure to a drug often causes mutations that enhance the resistance against the drug. To design an effective pharmacotherapy for an individual patient, it is important to accurately predict the drug resistance based on genotype data. Notably, the resistance is not just the simple sum of the effects of all mutations. Structural biological studies suggest that the association of mutations is crucial: even if mutations A or B alone do not affect the resistance, a significant change might happen when the two mutations occur together. Linear regression methods cannot take the associations into account, while decision tree methods can reveal only limited associations. Kernel methods and neural networks implicitly use all possible associations for prediction, but cannot select salient associations explicitly. RESULTS: Our method, itemset boosting, performs linear regression in the complete space of power sets of mutations. It implements a forward feature selection procedure where, in each iteration, one mutation combination is found by an efficient branch-and-bound search. This method uses all possible combinations, and salient associations are explicitly shown. In experiments, our method worked particularly well for predicting the resistance of nucleotide reverse transcriptase inhibitors (NRTIs). Furthermore, it successfully recovered many mutation associations known in biological literature. AVAILABILITY: http://www.kyb.mpg.de/bs/people/hiroto/iboost/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hiroto Saigo, Takeaki Uno, Koji Tsuda |
Bioinform. | 1 |
| 2006 | Optimizing amino acid substitution matrices with a local alignment kernelabstractBACKGROUND: Detecting remote homologies by direct comparison of protein sequences remains a challenging task. We had previously developed a similarity score between sequences, called a local alignment kernel, that exhibits good performance for this task in combination with a support vector machine. The local alignment kernel depends on an amino acid substitution matrix. Since commonly used BLOSUM or PAM matrices for scoring amino acid matches have been optimized to be used in combination with the Smith-Waterman algorithm, the matrices optimal for the local alignment kernel can be different. RESULTS: Contrary to the local alignment score computed by the Smith-Waterman algorithm, the local alignment kernel is differentiable with respect to the amino acid substitution and its derivative can be computed efficiently by dynamic programming. We optimized the substitution matrix by classical gradient descent by setting an objective function that measures how well the local alignment kernel discriminates homologs from non-homologs in the COG database. The local alignment kernel exhibits better performance when it uses the matrices and gap parameters optimized by this procedure than when it uses the matrices optimized for the Smith-Waterman algorithm. Furthermore, the matrices and gap parameters optimized for the local alignment kernel can also be used successfully by the Smith-Waterman algorithm. CONCLUSION: This optimization procedure leads to useful substitution matrices, both for the local alignment kernel and the Smith-Waterman algorithm. The best performance for homology detection is obtained by the local alignment kernel. Hiroto Saigo, Jean-Philippe Vert, Tatsuya Akutsu |
BMC Bioinform. | 1 |
| 2006 | Functional Census of Mutation Sequence Spaces: The Example of p53 Cancer Rescue MutantsabstractMany biomedical problems relate to mutant functional properties across a sequence space of interest, e.g., flu, cancer, and HIV. Detailed knowledge of mutant properties and function improves medical treatment and prevention. A functional census of p53 cancer rescue mutants would aid the search for cancer treatments from p53 mutant rescue. We devised a general methodology for conducting a functional census of a mutation sequence space by choosing informative mutants early. The methodology was tested in a double-blind predictive test on the functional rescue property of 71 novel putative p53 cancer rescue mutants iteratively predicted in sets of three (24 iterations). The first double-blind 15-point moving accuracy was 47 percent and the last was 86 percent; r = 0.01 before an epiphanic 16th iteration and r = 0.92 afterward. Useful mutants were chosen early (overall r = 0.80). Code and data are freely available (http://www.igb.uci.edu/research/research.html, corresponding authors: R.H.L. for computation and R.K.B. for biology). Samuel A. Danziger, Sanjay Joshua Swamidass, Jue Zeng, Lawrence R. Dearth, Jonathan H. Chen, Jianlin Cheng, Vinh P. Hoang, Hiroto Saigo, Ray Luo 0001, Pierre Baldi, Rainer K. Brachmann, Richard H. Lathrop |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2005 | Graph kernels for chemical informatics
Liva Ralaivola, Sanjay Joshua Swamidass, Hiroto Saigo, Pierre Baldi |
Neural Networks | 3 |
| 2004 | Protein homology detection using string alignment kernelsabstractMOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. Discriminative methods involving support vector machines (SVMs) are currently the most effective methods for the problem of superfamily recognition in the Structural Classification Of Proteins (SCOP) database. The performance of SVMs depends critically on the kernel function used to quantify the similarity between sequences. RESULTS: We propose new kernels for strings adapted to biological sequences, which we call local alignment kernels. These kernels measure the similarity between two sequences by summing up scores obtained from local alignments with gaps of the sequences. When tested in combination with SVM on their ability to recognize SCOP superfamilies on a benchmark dataset, the new kernels outperform state-of-the-art methods for remote homology detection. AVAILABILITY: Software and data available upon request. Hiroto Saigo, Jean-Philippe Vert, Nobuhisa Ueda, Tatsuya Akutsu |
Bioinform. | 1 |