VLDB 2026 Research / reviewers in the wild / expert
Tsuyoshi Kato
dblp:58/1278
· DBLP profile ↗
21ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0001-5311-4966ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
2 papers |
Mathematical optimization · 100% | |
| Artificial intelligence
2 papers |
Representation and self-supervised learning · 78% Learning paradigms · 22% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.2 | 1 | 2016 | Stochastic Dykstra Algorithms for Metric Learning with Positive Definite Covariance Descriptors · ECCV (6) 2016 |
Mathematical optimization › continuous optimization › convex optimization
conic optimization |
0.2 | 2 | 2010 | Conic Programming for Multitask Learning · IEEE Trans. Knowl. Data Eng. 2010 Multi-Task Learning via Conic Programming · NIPS 2007 |
Mathematical optimization › continuous optimization › convex optimization › conic optimization
second-order cone programming |
0.2 | 2 | 2010 | Conic Programming for Multitask Learning · IEEE Trans. Knowl. Data Eng. 2010 Multi-Task Learning via Conic Programming · NIPS 2007 |
Bioinformatics and computational biology
protein structure analysis |
0.1 | 1 | 2010 | Metric learning for enzyme active-site search · Bioinform. 2010 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological network inference |
0.1 | 1 | 2009 | Simultaneous inference of biological networks of multiple species from genome-wide data and evolutionary information: a semi-supervised approach · Bioinform. 2009 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2007 | Multi-Task Learning via Conic Programming · NIPS 2007 |
Bioinformatics and computational biology › biological network › network biology
network inference |
0.1 | 1 | 2005 | Selective integration of multiple biological data for supervised network inference · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.3stochastic dykstra algorithms · 0.2task network · 0.1template matching · 0.1metric learning · 0.1convex optimization · 0.1semi-supervised learning · 0.1link propagation · 0.1multiple kernel learning · 0.1kernel matrix completion · 0.1expectation-maximization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Linearly Convergent Mixup LearningabstractLearning in the reproducing kernel Hilbert space (RKHS) such as the support vector machine has been recognized as a promising technique. It continues to be highly effective and competitive in numerous prediction tasks, particularly in settings where there is a shortage of training data or computational limitations exist. Mixup data augmentation has widely been used to address the issue to limited training data. However, this augmentation technique remains challenging when applied to learning in RKHS, due to the generation of intermediate class labels. Although gradient descent methods handle these labels effectively, dual optimization approaches are typically not directly applicable. In this study, we present two novel algorithms applicable to a broad range of binary classification models. Unlike gradient-based approaches, our algorithms do not require hyperparameters like learning rates, simplifying their implementation and optimization. Both the number of iterations to converge and the computational cost per iteration scale linearly with respect to the dataset size. The numerical experiments demonstrate that our algorithms achieve faster convergence to the optimal solution compared to gradient descent approaches, and that mixup data augmentation consistently improves the predictive performance across various loss functions. Gakuto Obi, Ayato Saito, Yuto Sasaki, Tsuyoshi Kato |
ISIT | 4 |
| 2021 | Adaptive Signal Variances: CNN Initialization Through Modern ArchitecturesabstractDeep convolutional neural networks (CNNs), renowned for their consistent performance, are widely understood by practitioners that the stability of learning depends on the initialization of the model parameters in each layer. Kaiming initialization, the de facto standard, is derived from a much simpler CNN model which consists of only the convolution and fully connected layers. Compared to the current CNN models, the basis CNN model for the Kaiming initialization does not include the max pooling or global average pooling layers. In this study, we derive an new initialization scheme formulated from modern CNN architectures, and empirically investigate the performance of the new initialization methods compared to the standard initialization methods widely used today. Takahiko Henmi, Esmeraldo Ronnie Rey Zara, Yoshihiro Hirohashi, Tsuyoshi Kato |
ICIP | 4 |
| 2021 | Frank-Wolfe algorithm for learning SVM-type multi-category classifiersabstractThe multi-category support vector machine (MC-SVM) is one of the most popular machine learning algorithms. There are numerous MC-SVM variants, although different optimization algorithms were developed for diverse learning machines. In this study, we developed a new optimization algorithm that can be applied to several MC-SVM variants. The algorithm is based on the Frank-Wolfe framework that requires two subproblems, direction-finding and line search, in each iteration. The contribution of this study is the discovery that both subproblems have a closed form solution if the Frank-Wolfe framework is applied to the dual problem. Additionally, the closed form solutions on both the direction-finding and line search exist even for the Moreau envelopes of the loss functions. We used several large datasets to demonstrate that the proposed optimization algorithm rapidly converges and thereby improves the pattern recognition performance. Kenya Tajima, Yoshihiro Hirohashi, Esmeraldo Ronnie Rey Zara, Tsuyoshi Kato |
SDM | 4 |
| 2020 | Learning Sign-Constrained Support Vector MachinesabstractDomain knowledge is useful to improve the generalization performance of learning machines. Sign constraints are a handy representation to combine domain knowledge with learning machine. In this paper, we consider constraining the signs of the weight coefficients in learning the linear support vector machine, and develop two optimization algorithms for minimizing the empirical risk under the sign constraints. One of the two algorithms is based on the projected gradient method, in which each iteration of the projected gradient method takes O(nd) computational cost and the sublinear convergence of the objective error is guaranteed. The second algorithm is based on the Frank-Wolfe method that also converges sublinearly and possesses a clear termination criterion. We show that each iteration of the Frank-Wolfe also requires O(nd) cost. Furthermore, we derive the explicit expression for the minimal iteration number to ensure an ε-accurate solution by analyzing the curvature of the objective function. Finally, we empirically demonstrate that the sign constraints are a promising technique when similarities to the training examples compose the feature vector. Kenya Tajima, Kohei Tsuchida, Esmeraldo Ronnie Rey Zara, Naoya Ohta, Tsuyoshi Kato |
ICPR | 5 |
| 2019 | Learning Weighted Top-k Support Vector MachineabstractNowadays, the top-$k$ accuracy is a major performance criterion when benchmarking multi-class classifier using datasets with a large number of categories. Top-$k$ multiclass SVM has been designed with the aim to minimize the empirical risk based on the top-$k$ accuracy. There already exist two SDCA-based algorithms to learn the top-$k$ SVM, enjoying several preferable properties for optimization, although both the algorithms suffer from two disadvantages. A weak point is that, since the design of the algorithms are specialized only to the top-$k$ hinge, their applicability to other variants is limited. The other disadvantage is that both the two algorithms cannot attain the optimal solution in most cases due to their theoritical imperfections. In this study, a weighted extension of top-$k$ SVM is considered, and novel learning algorithms based on the Frank-Wolfe algorithm is devised. The new learning algorithms possess all the favorable properties of SDCA as well as the applicability not only to the original top-$k$ SVM but also to the weighted extension. Geometrical convergence is achieved by smoothing the loss functions. Numerical simulations demonstrate that only the proposed Frank-Wolfe algorithms can converge to the optimum, in contrast with the failure of the two existing SDCA-based algorithms. Finally, our analytical results for these two studies are presented to shed light on the meaning of the solutions produced from their algorithms. Tsuyoshi Kato, Yoshihiro Hirohashi |
ACML | 1 |
| 2016 | Stochastic Dykstra Algorithms for Metric Learning with Positive Definite Covariance Descriptors
Tomoki Matsuzawa, Raissa Relator, Jun Sese, Tsuyoshi Kato |
ECCV (6) | 4 |
| 2015 | Segmental HOG: new descriptor for glomerulus detection in kidney microscopy imageabstractBACKGROUND: The detection of the glomeruli is a key step in the histopathological evaluation of microscopic images of the kidneys. However, the task of automatic detection of the glomeruli poses challenges owing to the differences in their sizes and shapes in renal sections as well as the extensive variations in their intensities due to heterogeneity in immunohistochemistry staining. Although the rectangular histogram of oriented gradients (Rectangular HOG) is a widely recognized powerful descriptor for general object detection, it shows many false positives owing to the aforementioned difficulties in the context of glomeruli detection. RESULTS: A new descriptor referred to as Segmental HOG was developed to perform a comprehensive detection of hundreds of glomeruli in images of whole kidney sections. The new descriptor possesses flexible blocks that can be adaptively fitted to input images in order to acquire robustness for the detection of the glomeruli. Moreover, the novel segmentation technique employed herewith generates high-quality segmentation outputs, and the algorithm is assured to converge to an optimal solution. Consequently, experiments using real-world image data revealed that Segmental HOG achieved significant improvements in detection performance compared to Rectangular HOG. CONCLUSION: The proposed descriptor for glomeruli detection presents promising results, and it is expected to be useful in pathological evaluation. Tsuyoshi Kato, Raissa Relator, Hayliang Ngouv, Yoshihiro Hirohashi, Osamu Takaki, Tetsuhiro Kakimoto, Kinya Okada |
BMC Bioinform. | 1 |
| 2013 | Impulse noise removal by using one-dimensional switching median filter applied along space-filling curve reflecting structural context of imageabstractA switching median filter (SMF) is effective for impulse noise removal while still preserving edges in an input image. This filter firstly detects pixels corrupted by impulse noise and then filters only the noise-corrupted pixels. However the noise detection process does not always work perfectly. Particularly, pixels constituting thin lines in an input image tend to be incorrectly detected as noise-corrupted pixels, and such pixels are filtered despite the needlessness of the filtering. As the result of the filtering, the image might be over-smoothed and be deteriorated throughout the entire image. To cope with this problem, we propose a new impulse noise removal method based on a one-dimensional SMF and a space-filling curve which reflects structural contexts of an input image. The effectiveness of the proposed method is verified by some experiments. Takanori Koga, Noriaki Suetake, Tsuyoshi Kato, Eiji Uchino |
IECON | 3 |
| 2011 | Discriminative structural approaches for enzyme active-site predictionabstractBACKGROUND: Predicting enzyme active-sites in proteins is an important issue not only for protein sciences but also for a variety of practical applications such as drug design. Because enzyme reaction mechanisms are based on the local structures of enzyme active-sites, various template-based methods that compare local structures in proteins have been developed to date. In comparing such local sites, a simple measurement, RMSD, has been used so far. RESULTS: This paper introduces new machine learning algorithms that refine the similarity/deviation for comparison of local structures. The similarity/deviation is applied to two types of applications, single template analysis and multiple template analysis. In the single template analysis, a single template is used as a query to search proteins for active sites, whereas a protein structure is examined as a query to discover the possible active-sites using a set of templates in the multiple template analysis. CONCLUSIONS: This paper experimentally illustrates that the machine learning algorithms effectively improve the similarity/deviation measurements for both the analyses. Tsuyoshi Kato, Nozomi Nagano |
BMC Bioinform. | 1 |
| 2010 | An Accurate Prediction Method for Protein Structural Class from Signal Patterns of NMR Spectra in the Absence of Chemical Shift AssignmentsabstractThe structural class information about a protein is important to understand its biological properties. NMR is one of the most powerful tools to obtain structural information of proteins in atomic resolution. However, an analysis of protein three-dimensional structure from NMR spectra usually requires laborious chemical shift assignment. We developed a new method for predicting the protein structural class directly from the NMR spectra without any chemical shift assignment. The results show that our method outperforms the methods using current secondary structure prediction. Hiromi Arai, Naoya Tochio, Tsuyoshi Kato, Takanori Kigawa, Masayuki Yamamura |
BIBE | 3 |
| 2010 | Metric learning for enzyme active-site searchabstractMOTIVATION: Finding functionally analogous enzymes based on the local structures of active sites is an important problem. Conventional methods use templates of local structures to search for analogous sites, but their performance depends on the selection of atoms for inclusion in the templates. RESULTS: The automatic selection of atoms so that site matches can be discriminated from mismatches. The algorithm provides not only good predictions, but also some insights into which atoms are important for the prediction. Our experimental results suggest that the metric learning automatically provides more effective templates than those whose atoms are selected manually. AVAILABILITY: Online software is available at http://www.net-machine.net/∼kato/lpmetric1/ Tsuyoshi Kato, Nozomi Nagano |
Bioinform. | 1 |
| 2010 | Conic Programming for Multitask LearningabstractWhen we have several related tasks, solving them simultaneously has been shown to be more effective than solving them individually. This approach is called multitask learning (MTL). In this paper, we propose a novel MTL algorithm. Our method controls the relatedness among the tasks locally, so all pairs of related tasks are guaranteed to have similar solutions. We apply the above idea to support vector machines and show that the optimization problem can be cast as a second-order cone program, which is convex and can be solved efficiently. The usefulness of our approach is demonstrated in ordinal regression, link prediction, and collaborative filtering, each of which can be formulated as a structured multitask problem. Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama, Kiyoshi Asai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2009 | Link Propagation: A Fast Semi-supervised Learning Algorithm for Link PredictionabstractWe propose Link Propagation as a new semi-supervised learning method for link prediction problems, where the task is to predict unknown parts of the network structure by using auxiliary information such as node similarities. Since the proposed method can fill in missing parts of tensors, it is applicable to multi-relational domains, allowing us to handle multiple types of links simultaneously. We also give a novel efficient algorithm for Link Propagation based on an accelerated conjugate gradient method. Hisashi Kashima, Tsuyoshi Kato, Yoshihiro Yamanishi, Masashi Sugiyama, Koji Tsuda |
SDM | 2 |
| 2009 | Simultaneous inference of biological networks of multiple species from genome-wide data and evolutionary information: a semi-supervised approachabstractMOTIVATION: The existing supervised methods for biological network inference work on each of the networks individually based only on intra-species information such as gene expression data. We believe that it will be more effective to use genomic data and cross-species evolutionary information from different species simultaneously, rather than to use the genomic data alone. RESULTS: We created a new semi-supervised learning method called Link Propagation for inferring biological networks of multiple species based on genome-wide data and evolutionary information. The new method was applied to simultaneous reconstruction of three metabolic networks of Caenorhabditis elegans, Helicobacter pylori and Saccharomyces cerevisiae, based on gene expression similarities and amino acid sequence similarities. The experimental results proved that the new simultaneous network inference method consistently improves the predictive performance over the individual network inferences, and it also outperforms in accuracy and speed other established methods such as the pairwise support vector machine. AVAILABILITY: The software and data are available at http://cbio.ensmp.fr/~yyamanishi/LinkPropagation/. Hisashi Kashima, Yoshihiro Yamanishi, Tsuyoshi Kato, Masashi Sugiyama, Koji Tsuda |
Bioinform. | 3 |
| 2009 | Robust Label Propagation on Multiple NetworksabstractTransductive inference on graphs such as label propagation algorithms is receiving a lot of attention. In this paper, we address a label propagation problem on multiple networks and present a new algorithm that automatically integrates structure information brought in by multiple networks. The proposed method is robust in that irrelevant networks are automatically deemphasized, which is an advantage over Tsuda's approach (2005). We also show that the proposed algorithm can be interpreted as an expectation-maximization (EM) algorithm with a student-t prior. Finally, we demonstrate the usefulness of our method in protein function prediction and digit classification, and show analytically and experimentally that our algorithm is much more efficient than existing algorithms. Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama |
IEEE Trans. Neural Networks | 1 |
| 2008 | Integration of Multiple Networks for Robust Label PropagationabstractTransductive inference on graphs such as label propagation algorithms is receiving a lot of attention. In this paper, we address a label propagation problem on multiple networks and present a new algorithm that automatically integrates structure information brought in by multiple networks. The proposed method is robust in that irrelevant networks are automatically deemphasized, which is an advantage over Tsuda et al.'s approach [14]. We also show that the proposed algorithm can be interpreted as an EM algorithm with a Student-t prior. Finally, we demonstrate the usefulness of our method in protein function prediction. Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama |
SDM | 1 |
| 2007 | Multi-Task Learning via Conic ProgrammingabstractWhen we have several related tasks, solving them simultaneously is shown to be more effective than solving them individually. This approach is called multi-task learning (MTL) and has been studied extensively. Existing approaches to MTL often treat all the tasks as \emph{uniformly related to each other and the relatedness of the tasks is controlled globally. For this reason, the existing methods can lead to undesired solutions when some tasks are not highly related to each other, and some pairs of related tasks can have significantly different solutions. In this paper, we propose a novel MTL algorithm that can overcome these problems. Our method makes use of a task network, which describes the relation structure among tasks. This allows us to deal with intricate relation structures in a systematic way. Furthermore, we control the relatedness of the tasks locally, so all pairs of related tasks are guaranteed to have similar solutions. We apply the above idea to support vector machines (SVMs) and show that the optimization problem can be cast as a second order cone program, which is convex and can be solved efficiently. The usefulness of our approach is demonstrated through simulations with protein super-family classification and ordinal regression problems. Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama, Kiyoshi Asai |
NIPS | 1 |
| 2007 | Classification of heterogeneous microarray data by maximum entropy kernelabstractBACKGROUND: There is a large amount of microarray data accumulating in public databases, providing various data waiting to be analyzed jointly. Powerful kernel-based methods are commonly used in microarray analyses with support vector machines (SVMs) to approach a wide range of classification problems. However, the standard vectorial data kernel family (linear, RBF, etc.) that takes vectorial data as input, often fails in prediction if the data come from different platforms or laboratories, due to the low gene overlaps or consistencies between the different datasets. RESULTS: We introduce a new type of kernel called maximum entropy (ME) kernel, which has no pre-defined function but is generated by kernel entropy maximization with sample distance matrices as constraints, into the field of SVM classification of microarray data. We assessed the performance of the ME kernel with three different data: heterogeneous kidney carcinoma, noise-introduced leukemia, and heterogeneous oral cavity carcinoma metastasis data. The results clearly show that the ME kernel is very robust for heterogeneous data containing missing values and high-noise, and gives higher prediction accuracies than the standard kernels, namely, linear, polynomial and RBF. CONCLUSION: The results demonstrate its utility in effectively analyzing promiscuous microarray data of rare specimens, e.g., minor diseases or species, that present difficulty in compiling homogeneous data in a single laboratory. Wataru Fujibuchi, Tsuyoshi Kato |
BMC Bioinform. | 2 |
| 2006 | Network-based de-noising improves prediction from microarray dataabstractBACKGROUND: Prediction of human cell response to anti-cancer drugs (compounds) from microarray data is a challenging problem, due to the noise properties of microarrays as well as the high variance of living cell responses to drugs. Hence there is a strong need for more practical and robust methods than standard methods for real-value prediction. RESULTS: We devised an extended version of the off-subspace noise-reduction (de-noising) method to incorporate heterogeneous network data such as sequence similarity or protein-protein interactions into a single framework. Using that method, we first de-noise the gene expression data for training and test data and also the drug-response data for training data. Then we predict the unknown responses of each drug from the de-noised input data. For ascertaining whether de-noising improves prediction or not, we carry out 12-fold cross-validation for assessment of the prediction performance. We use the Pearson's correlation coefficient between the true and predicted response values as the prediction performance. De-noising improves the prediction performance for 65% of drugs. Furthermore, we found that this noise reduction method is robust and effective even when a large amount of artificial noise is added to the input data. CONCLUSION: We found that our extended off-subspace noise-reduction method combining heterogeneous biological data is successful and quite useful to improve prediction of human cell cancer drug responses from microarray data. Tsuyoshi Kato, Yukio Murata, Koh Miura, Kiyoshi Asai, Paul Horton, Koji Tsuda, Wataru Fujibuchi |
BMC Bioinform. | 1 |
| 2005 | Selective integration of multiple biological data for supervised network inferenceabstractMOTIVATION: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected. RESULTS: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network. AVAILABILITY: Software is available on request. Tsuyoshi Kato, Koji Tsuda, Kiyoshi Asai |
Bioinform. | 1 |
| 2000 | Precise Hand-printed Character Recognition Using Elastic Models via Nonlinear TransformationabstractDistorted character recognition is a difficult but in-evitable problem in hand-printed character recognition. In this paper, we propose a character recognition method us-ing elastic models for recognizing cursive characters with intricate structure. The models are fitted to unknown in-put patterns by applying the EM algorithm to minimize a measure of fittness. To avoid falling into local minima, mul-tiresolutional approach is introduced. Moreover, nonlinear transformation is adopted to realize more flexible matching. Experiments performed on Japanese characters show effec-tiveness of the proposed method. 1. Tsuyoshi Kato, Shinichiro Omachi, Hirotomo Aso |
ICPR | 1 |