Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tsuyoshi Kato

dblp:58/1278 · DBLP profile ↗
← Back
21ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0001-5311-4966ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Mathematical optimization · 100%
Artificial intelligence
2 papers
Representation and self-supervised learning · 78% Learning paradigms · 22%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.212016
Stochastic Dykstra Algorithms for Metric Learning with Positive Definite Covariance Descriptors · ECCV (6) 2016
Mathematical optimization › continuous optimization › convex optimization
conic optimization
0.222010
Conic Programming for Multitask Learning · IEEE Trans. Knowl. Data Eng. 2010
Multi-Task Learning via Conic Programming · NIPS 2007
Mathematical optimization › continuous optimization › convex optimization › conic optimization
second-order cone programming
0.222010
Conic Programming for Multitask Learning · IEEE Trans. Knowl. Data Eng. 2010
Multi-Task Learning via Conic Programming · NIPS 2007
Bioinformatics and computational biology
protein structure analysis
0.112010
Metric learning for enzyme active-site search · Bioinform. 2010
Bioinformatics and computational biology › network bioinformatics › biological network analysis
biological network inference
0.112009
Simultaneous inference of biological networks of multiple species from genome-wide data and evolutionary information: a semi-supervised approach · Bioinform. 2009
Machine learning › Learning paradigms
multi-task learning
0.112007
Multi-Task Learning via Conic Programming · NIPS 2007
Bioinformatics and computational biology › biological network › network biology
network inference
0.112005
Selective integration of multiple biological data for supervised network inference · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.3stochastic dykstra algorithms · 0.2task network · 0.1template matching · 0.1metric learning · 0.1convex optimization · 0.1semi-supervised learning · 0.1link propagation · 0.1multiple kernel learning · 0.1kernel matrix completion · 0.1expectation-maximization · 0.1
YearPublicationVenuePosition
2025 Linearly Convergent Mixup Learning
abstract
Learning in the reproducing kernel Hilbert space (RKHS) such as the support vector machine has been recognized as a promising technique. It continues to be highly effective and competitive in numerous prediction tasks, particularly in settings where there is a shortage of training data or computational limitations exist. Mixup data augmentation has widely been used to address the issue to limited training data. However, this augmentation technique remains challenging when applied to learning in RKHS, due to the generation of intermediate class labels. Although gradient descent methods handle these labels effectively, dual optimization approaches are typically not directly applicable. In this study, we present two novel algorithms applicable to a broad range of binary classification models. Unlike gradient-based approaches, our algorithms do not require hyperparameters like learning rates, simplifying their implementation and optimization. Both the number of iterations to converge and the computational cost per iteration scale linearly with respect to the dataset size. The numerical experiments demonstrate that our algorithms achieve faster convergence to the optimal solution compared to gradient descent approaches, and that mixup data augmentation consistently improves the predictive performance across various loss functions.
Gakuto Obi, Ayato Saito, Yuto Sasaki, Tsuyoshi Kato
ISIT4
2021 Adaptive Signal Variances: CNN Initialization Through Modern Architectures
abstract
Deep convolutional neural networks (CNNs), renowned for their consistent performance, are widely understood by practitioners that the stability of learning depends on the initialization of the model parameters in each layer. Kaiming initialization, the de facto standard, is derived from a much simpler CNN model which consists of only the convolution and fully connected layers. Compared to the current CNN models, the basis CNN model for the Kaiming initialization does not include the max pooling or global average pooling layers. In this study, we derive an new initialization scheme formulated from modern CNN architectures, and empirically investigate the performance of the new initialization methods compared to the standard initialization methods widely used today.
Takahiko Henmi, Esmeraldo Ronnie Rey Zara, Yoshihiro Hirohashi, Tsuyoshi Kato
ICIP4
2021 Frank-Wolfe algorithm for learning SVM-type multi-category classifiers
abstract
The multi-category support vector machine (MC-SVM) is one of the most popular machine learning algorithms. There are numerous MC-SVM variants, although different optimization algorithms were developed for diverse learning machines. In this study, we developed a new optimization algorithm that can be applied to several MC-SVM variants. The algorithm is based on the Frank-Wolfe framework that requires two subproblems, direction-finding and line search, in each iteration. The contribution of this study is the discovery that both subproblems have a closed form solution if the Frank-Wolfe framework is applied to the dual problem. Additionally, the closed form solutions on both the direction-finding and line search exist even for the Moreau envelopes of the loss functions. We used several large datasets to demonstrate that the proposed optimization algorithm rapidly converges and thereby improves the pattern recognition performance.
Kenya Tajima, Yoshihiro Hirohashi, Esmeraldo Ronnie Rey Zara, Tsuyoshi Kato
SDM4
2020 Learning Sign-Constrained Support Vector Machines
abstract
Domain knowledge is useful to improve the generalization performance of learning machines. Sign constraints are a handy representation to combine domain knowledge with learning machine. In this paper, we consider constraining the signs of the weight coefficients in learning the linear support vector machine, and develop two optimization algorithms for minimizing the empirical risk under the sign constraints. One of the two algorithms is based on the projected gradient method, in which each iteration of the projected gradient method takes O(nd) computational cost and the sublinear convergence of the objective error is guaranteed. The second algorithm is based on the Frank-Wolfe method that also converges sublinearly and possesses a clear termination criterion. We show that each iteration of the Frank-Wolfe also requires O(nd) cost. Furthermore, we derive the explicit expression for the minimal iteration number to ensure an ε-accurate solution by analyzing the curvature of the objective function. Finally, we empirically demonstrate that the sign constraints are a promising technique when similarities to the training examples compose the feature vector.
Kenya Tajima, Kohei Tsuchida, Esmeraldo Ronnie Rey Zara, Naoya Ohta, Tsuyoshi Kato
ICPR5
2019 Learning Weighted Top-k Support Vector Machine
abstract
Nowadays, the top-$k$ accuracy is a major performance criterion when benchmarking multi-class classifier using datasets with a large number of categories. Top-$k$ multiclass SVM has been designed with the aim to minimize the empirical risk based on the top-$k$ accuracy. There already exist two SDCA-based algorithms to learn the top-$k$ SVM, enjoying several preferable properties for optimization, although both the algorithms suffer from two disadvantages. A weak point is that, since the design of the algorithms are specialized only to the top-$k$ hinge, their applicability to other variants is limited. The other disadvantage is that both the two algorithms cannot attain the optimal solution in most cases due to their theoritical imperfections. In this study, a weighted extension of top-$k$ SVM is considered, and novel learning algorithms based on the Frank-Wolfe algorithm is devised. The new learning algorithms possess all the favorable properties of SDCA as well as the applicability not only to the original top-$k$ SVM but also to the weighted extension. Geometrical convergence is achieved by smoothing the loss functions. Numerical simulations demonstrate that only the proposed Frank-Wolfe algorithms can converge to the optimum, in contrast with the failure of the two existing SDCA-based algorithms. Finally, our analytical results for these two studies are presented to shed light on the meaning of the solutions produced from their algorithms.
Tsuyoshi Kato, Yoshihiro Hirohashi
ACML1
2016 Stochastic Dykstra Algorithms for Metric Learning with Positive Definite Covariance Descriptors
Tomoki Matsuzawa, Raissa Relator, Jun Sese, Tsuyoshi Kato
ECCV (6)4
2015 Segmental HOG: new descriptor for glomerulus detection in kidney microscopy image
abstract
BACKGROUND: The detection of the glomeruli is a key step in the histopathological evaluation of microscopic images of the kidneys. However, the task of automatic detection of the glomeruli poses challenges owing to the differences in their sizes and shapes in renal sections as well as the extensive variations in their intensities due to heterogeneity in immunohistochemistry staining. Although the rectangular histogram of oriented gradients (Rectangular HOG) is a widely recognized powerful descriptor for general object detection, it shows many false positives owing to the aforementioned difficulties in the context of glomeruli detection. RESULTS: A new descriptor referred to as Segmental HOG was developed to perform a comprehensive detection of hundreds of glomeruli in images of whole kidney sections. The new descriptor possesses flexible blocks that can be adaptively fitted to input images in order to acquire robustness for the detection of the glomeruli. Moreover, the novel segmentation technique employed herewith generates high-quality segmentation outputs, and the algorithm is assured to converge to an optimal solution. Consequently, experiments using real-world image data revealed that Segmental HOG achieved significant improvements in detection performance compared to Rectangular HOG. CONCLUSION: The proposed descriptor for glomeruli detection presents promising results, and it is expected to be useful in pathological evaluation.
Tsuyoshi Kato, Raissa Relator, Hayliang Ngouv, Yoshihiro Hirohashi, Osamu Takaki, Tetsuhiro Kakimoto, Kinya Okada
BMC Bioinform.1
2013 Impulse noise removal by using one-dimensional switching median filter applied along space-filling curve reflecting structural context of image
abstract
A switching median filter (SMF) is effective for impulse noise removal while still preserving edges in an input image. This filter firstly detects pixels corrupted by impulse noise and then filters only the noise-corrupted pixels. However the noise detection process does not always work perfectly. Particularly, pixels constituting thin lines in an input image tend to be incorrectly detected as noise-corrupted pixels, and such pixels are filtered despite the needlessness of the filtering. As the result of the filtering, the image might be over-smoothed and be deteriorated throughout the entire image. To cope with this problem, we propose a new impulse noise removal method based on a one-dimensional SMF and a space-filling curve which reflects structural contexts of an input image. The effectiveness of the proposed method is verified by some experiments.
Takanori Koga, Noriaki Suetake, Tsuyoshi Kato, Eiji Uchino
IECON3
2011 Discriminative structural approaches for enzyme active-site prediction
abstract
BACKGROUND: Predicting enzyme active-sites in proteins is an important issue not only for protein sciences but also for a variety of practical applications such as drug design. Because enzyme reaction mechanisms are based on the local structures of enzyme active-sites, various template-based methods that compare local structures in proteins have been developed to date. In comparing such local sites, a simple measurement, RMSD, has been used so far. RESULTS: This paper introduces new machine learning algorithms that refine the similarity/deviation for comparison of local structures. The similarity/deviation is applied to two types of applications, single template analysis and multiple template analysis. In the single template analysis, a single template is used as a query to search proteins for active sites, whereas a protein structure is examined as a query to discover the possible active-sites using a set of templates in the multiple template analysis. CONCLUSIONS: This paper experimentally illustrates that the machine learning algorithms effectively improve the similarity/deviation measurements for both the analyses.
Tsuyoshi Kato, Nozomi Nagano
BMC Bioinform.1
2010 An Accurate Prediction Method for Protein Structural Class from Signal Patterns of NMR Spectra in the Absence of Chemical Shift Assignments
abstract
The structural class information about a protein is important to understand its biological properties. NMR is one of the most powerful tools to obtain structural information of proteins in atomic resolution. However, an analysis of protein three-dimensional structure from NMR spectra usually requires laborious chemical shift assignment. We developed a new method for predicting the protein structural class directly from the NMR spectra without any chemical shift assignment. The results show that our method outperforms the methods using current secondary structure prediction.
Hiromi Arai, Naoya Tochio, Tsuyoshi Kato, Takanori Kigawa, Masayuki Yamamura
BIBE3
2010 Metric learning for enzyme active-site search
abstract
MOTIVATION: Finding functionally analogous enzymes based on the local structures of active sites is an important problem. Conventional methods use templates of local structures to search for analogous sites, but their performance depends on the selection of atoms for inclusion in the templates. RESULTS: The automatic selection of atoms so that site matches can be discriminated from mismatches. The algorithm provides not only good predictions, but also some insights into which atoms are important for the prediction. Our experimental results suggest that the metric learning automatically provides more effective templates than those whose atoms are selected manually. AVAILABILITY: Online software is available at http://www.net-machine.net/∼kato/lpmetric1/
Tsuyoshi Kato, Nozomi Nagano
Bioinform.1
2010 Conic Programming for Multitask Learning
abstract
When we have several related tasks, solving them simultaneously has been shown to be more effective than solving them individually. This approach is called multitask learning (MTL). In this paper, we propose a novel MTL algorithm. Our method controls the relatedness among the tasks locally, so all pairs of related tasks are guaranteed to have similar solutions. We apply the above idea to support vector machines and show that the optimization problem can be cast as a second-order cone program, which is convex and can be solved efficiently. The usefulness of our approach is demonstrated in ordinal regression, link prediction, and collaborative filtering, each of which can be formulated as a structured multitask problem.
Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama, Kiyoshi Asai
IEEE Trans. Knowl. Data Eng.1
2009 Link Propagation: A Fast Semi-supervised Learning Algorithm for Link Prediction
abstract
We propose Link Propagation as a new semi-supervised learning method for link prediction problems, where the task is to predict unknown parts of the network structure by using auxiliary information such as node similarities. Since the proposed method can fill in missing parts of tensors, it is applicable to multi-relational domains, allowing us to handle multiple types of links simultaneously. We also give a novel efficient algorithm for Link Propagation based on an accelerated conjugate gradient method.
Hisashi Kashima, Tsuyoshi Kato, Yoshihiro Yamanishi, Masashi Sugiyama, Koji Tsuda
SDM2
2009 Simultaneous inference of biological networks of multiple species from genome-wide data and evolutionary information: a semi-supervised approach
abstract
MOTIVATION: The existing supervised methods for biological network inference work on each of the networks individually based only on intra-species information such as gene expression data. We believe that it will be more effective to use genomic data and cross-species evolutionary information from different species simultaneously, rather than to use the genomic data alone. RESULTS: We created a new semi-supervised learning method called Link Propagation for inferring biological networks of multiple species based on genome-wide data and evolutionary information. The new method was applied to simultaneous reconstruction of three metabolic networks of Caenorhabditis elegans, Helicobacter pylori and Saccharomyces cerevisiae, based on gene expression similarities and amino acid sequence similarities. The experimental results proved that the new simultaneous network inference method consistently improves the predictive performance over the individual network inferences, and it also outperforms in accuracy and speed other established methods such as the pairwise support vector machine. AVAILABILITY: The software and data are available at http://cbio.ensmp.fr/~yyamanishi/LinkPropagation/.
Hisashi Kashima, Yoshihiro Yamanishi, Tsuyoshi Kato, Masashi Sugiyama, Koji Tsuda
Bioinform.3
2009 Robust Label Propagation on Multiple Networks
abstract
Transductive inference on graphs such as label propagation algorithms is receiving a lot of attention. In this paper, we address a label propagation problem on multiple networks and present a new algorithm that automatically integrates structure information brought in by multiple networks. The proposed method is robust in that irrelevant networks are automatically deemphasized, which is an advantage over Tsuda's approach (2005). We also show that the proposed algorithm can be interpreted as an expectation-maximization (EM) algorithm with a student-t prior. Finally, we demonstrate the usefulness of our method in protein function prediction and digit classification, and show analytically and experimentally that our algorithm is much more efficient than existing algorithms.
Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama
IEEE Trans. Neural Networks1
2008 Integration of Multiple Networks for Robust Label Propagation
abstract
Transductive inference on graphs such as label propagation algorithms is receiving a lot of attention. In this paper, we address a label propagation problem on multiple networks and present a new algorithm that automatically integrates structure information brought in by multiple networks. The proposed method is robust in that irrelevant networks are automatically deemphasized, which is an advantage over Tsuda et al.'s approach [14]. We also show that the proposed algorithm can be interpreted as an EM algorithm with a Student-t prior. Finally, we demonstrate the usefulness of our method in protein function prediction.
Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama
SDM1
2007 Multi-Task Learning via Conic Programming
abstract
When we have several related tasks, solving them simultaneously is shown to be more effective than solving them individually. This approach is called multi-task learning (MTL) and has been studied extensively. Existing approaches to MTL often treat all the tasks as \emph{uniformly related to each other and the relatedness of the tasks is controlled globally. For this reason, the existing methods can lead to undesired solutions when some tasks are not highly related to each other, and some pairs of related tasks can have significantly different solutions. In this paper, we propose a novel MTL algorithm that can overcome these problems. Our method makes use of a task network, which describes the relation structure among tasks. This allows us to deal with intricate relation structures in a systematic way. Furthermore, we control the relatedness of the tasks locally, so all pairs of related tasks are guaranteed to have similar solutions. We apply the above idea to support vector machines (SVMs) and show that the optimization problem can be cast as a second order cone program, which is convex and can be solved efficiently. The usefulness of our approach is demonstrated through simulations with protein super-family classification and ordinal regression problems.
Tsuyoshi Kato, Hisashi Kashima, Masashi Sugiyama, Kiyoshi Asai
NIPS1
2007 Classification of heterogeneous microarray data by maximum entropy kernel
abstract
BACKGROUND: There is a large amount of microarray data accumulating in public databases, providing various data waiting to be analyzed jointly. Powerful kernel-based methods are commonly used in microarray analyses with support vector machines (SVMs) to approach a wide range of classification problems. However, the standard vectorial data kernel family (linear, RBF, etc.) that takes vectorial data as input, often fails in prediction if the data come from different platforms or laboratories, due to the low gene overlaps or consistencies between the different datasets. RESULTS: We introduce a new type of kernel called maximum entropy (ME) kernel, which has no pre-defined function but is generated by kernel entropy maximization with sample distance matrices as constraints, into the field of SVM classification of microarray data. We assessed the performance of the ME kernel with three different data: heterogeneous kidney carcinoma, noise-introduced leukemia, and heterogeneous oral cavity carcinoma metastasis data. The results clearly show that the ME kernel is very robust for heterogeneous data containing missing values and high-noise, and gives higher prediction accuracies than the standard kernels, namely, linear, polynomial and RBF. CONCLUSION: The results demonstrate its utility in effectively analyzing promiscuous microarray data of rare specimens, e.g., minor diseases or species, that present difficulty in compiling homogeneous data in a single laboratory.
Wataru Fujibuchi, Tsuyoshi Kato
BMC Bioinform.2
2006 Network-based de-noising improves prediction from microarray data
abstract
BACKGROUND: Prediction of human cell response to anti-cancer drugs (compounds) from microarray data is a challenging problem, due to the noise properties of microarrays as well as the high variance of living cell responses to drugs. Hence there is a strong need for more practical and robust methods than standard methods for real-value prediction. RESULTS: We devised an extended version of the off-subspace noise-reduction (de-noising) method to incorporate heterogeneous network data such as sequence similarity or protein-protein interactions into a single framework. Using that method, we first de-noise the gene expression data for training and test data and also the drug-response data for training data. Then we predict the unknown responses of each drug from the de-noised input data. For ascertaining whether de-noising improves prediction or not, we carry out 12-fold cross-validation for assessment of the prediction performance. We use the Pearson's correlation coefficient between the true and predicted response values as the prediction performance. De-noising improves the prediction performance for 65% of drugs. Furthermore, we found that this noise reduction method is robust and effective even when a large amount of artificial noise is added to the input data. CONCLUSION: We found that our extended off-subspace noise-reduction method combining heterogeneous biological data is successful and quite useful to improve prediction of human cell cancer drug responses from microarray data.
Tsuyoshi Kato, Yukio Murata, Koh Miura, Kiyoshi Asai, Paul Horton, Koji Tsuda, Wataru Fujibuchi
BMC Bioinform.1
2005 Selective integration of multiple biological data for supervised network inference
abstract
MOTIVATION: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected. RESULTS: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network. AVAILABILITY: Software is available on request.
Tsuyoshi Kato, Koji Tsuda, Kiyoshi Asai
Bioinform.1
2000 Precise Hand-printed Character Recognition Using Elastic Models via Nonlinear Transformation
abstract
Distorted character recognition is a difficult but in-evitable problem in hand-printed character recognition. In this paper, we propose a character recognition method us-ing elastic models for recognizing cursive characters with intricate structure. The models are fitted to unknown in-put patterns by applying the EM algorithm to minimize a measure of fittness. To avoid falling into local minima, mul-tiresolutional approach is introduced. Moreover, nonlinear transformation is adopted to realize more flexible matching. Experiments performed on Japanese characters show effec-tiveness of the proposed method. 1.
Tsuyoshi Kato, Shinichiro Omachi, Hirotomo Aso
ICPR1