VLDB 2026 Research / reviewers in the wild / expert
Ryosuke Kojima
dblp:161/9289
· DBLP profile ↗
19ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 8 first-author · 5 since 2021Systems, architecture and hardware · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Theory of computation · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Deep Dissipative DynamicsabstractThis study challenges strictly guaranteeing ``dissipativity'' of a dynamical system represented by neural networks learned from given time-series data. Dissipativity is a crucial indicator for dynamical systems that generalizes stability and input-output stability, known to be valid across various systems including robotics, biological systems, and molecular dynamics. By analytically proving the general solution to the nonlinear Kalman–Yakubovich–Popov (KYP) lemma, which is the necessary and sufficient condition for dissipativity, we propose a differentiable projection that transforms any dynamics represented by neural networks into dissipative ones and a learning method for the transformed dynamics. Utilizing the generality of dissipativity, our method strictly guarantee stability, input-output stability, and energy conservation of trained dynamical systems. Finally, we demonstrate the robustness of our method against out-of-domain input through applications to robotic arms and fluid dynamics. Yuji Okamoto, Ryosuke Kojima |
AAAI | 2 |
| 2024 | Reconsidering Stochastic Policy Gradient Methods for Traffic Signal Control
Masahiro Kato, Ryosuke Kojima |
IEA/AIE | 2 |
| 2023 | Network-based prediction approach for cancer-specific driver missense mutations using a graph neural networkabstractBACKGROUND: In cancer genomic medicine, finding driver mutations involved in cancer development and tumor growth is crucial. Machine-learning methods to predict driver missense mutations have been developed because variants are frequently detected by genomic sequencing. However, even though the abnormalities in molecular networks are associated with cancer, many of these methods focus on individual variants and do not consider molecular networks. Here we propose a new network-based method, Net-DMPred, to predict driver missense mutations considering molecular networks. Net-DMPred consists of the graph part and the prediction part. In the graph part, molecular networks are learned by a graph neural network (GNN). The prediction part learns whether variants are driver variants using features of individual variants combined with the graph features learned in the graph part. RESULTS: Net-DMPred, which considers molecular networks, performed better than conventional methods. Furthermore, the prediction performance differed by the molecular network structure used in learning, suggesting that it is important to consider not only the local network related to cancer but also the large-scale network in living organisms. CONCLUSIONS: We propose a network-based machine learning method, Net-DMPred, for predicting cancer driver missense mutations. Our method enables us to consider the entire graph architecture representing the molecular network because it uses GNN. Net-DMPred is expected to detect driver mutations from a lot of missense mutations that are not known to be associated with cancer. Narumi Hatano, Mayumi Kamada, Ryosuke Kojima, Yasushi Okuno |
BMC Bioinform. | 3 |
| 2023 | Individual health-disease phase diagrams for disease prevention based on machine learningabstractEarly disease detection and prevention methods based on effective interventions are gaining attention worldwide. Progress in precision medicine has revealed that substantial heterogeneity exists in health data at the individual level and that complex health factors are involved in chronic disease development. Machine-learning techniques have enabled precise personal-level disease prediction by capturing individual differences in multivariate data. However, it is challenging to identify what aspects should be improved for disease prevention based on future disease-onset prediction because of the complex relationships among multiple biomarkers. Here, we present a health-disease phase diagram (HDPD) that represents an individual's health state by visualizing the future-onset boundary values of multiple biomarkers that fluctuate early in the disease progression process. In HDPDs, future-onset predictions are represented by perturbing multiple biomarker values while accounting for dependencies among variables. We constructed HDPDs for 11 diseases using longitudinal health checkup cohort data of 3,238 individuals, comprising 3,215 measurement items and genetic data. The improvement of biomarker values to the non-onset region in HDPD remarkably prevented future disease onset in 7 out of 11 diseases. HDPDs can represent individual physiological states in the onset process and be used as intervention goals for disease prevention. Eiichiro Uchino, Noriaki Sato, Ayano Araki, Kei Terayama, Ryosuke Kojima, Koichi Murashita, Ken Itoh, Tatsuya Mikami, Yoshinori Tamada, Yasushi Okuno |
J. Biomed. Informatics | 6 |
| 2022 | Learning Deep Input-Output Stable DynamicsabstractLearning stable dynamics from observed time-series data is an essential problem in robotics, physical modeling, and systems biology. Many of these dynamics are represented as an inputs-output system to communicate with the external environment. In this study, we focus on input-output stable systems, exhibiting robustness against unexpected stimuli and noise. We propose a method to learn nonlinear systems guaranteeing the input-output stability. Our proposed method utilizes the differentiable projection onto the space satisfying the Hamilton-Jacobi inequality to realize the input-output stability. The problem of finding this projection can be formulated as a quadratic constraint quadratic programming problem, and we derive the particular solution analytically. Also, we apply our method to a toy bistable model and the task of training a benchmark generated from a glucose-insulin simulator. The results show that the nonlinear system with neural networks by our method achieves the input-output stability, unlike naive neural networks. Our code is available at https://github.com/clinfo/DeepIOStability . Ryosuke Kojima, Yuji Okamoto |
NeurIPS | 1 |
| 2021 | Assessment of von Mises-Bernoulli Deep Neural Network in Sound Source Localization
Katsutoshi Itoyama, Yoshiya Morimoto, Shungo Masaki, Ryosuke Kojima, Kenji Nishida, Kazuhiro Nakadai |
Interspeech | 4 |
| 2021 | Boolean Network Learning in Vector Spaces for Genome-wide Network AnalysisabstractBoolean networks (BNs) are one of the standard tools for modeling gene regulatory networks in biology but their learning has been limited to small networks due to computational difficulty. Aiming at unprecedented scalability, we focus on a subclass of BNs called AND/OR Boolean networks where Boolean formulas are restricted to a conjunction or a disjunction of literals. We represent an AND/OR BN with N nodes by an N x 2N binary matrix Q paired with an N dimensional integer vector theta called a threshold vector, a state of the BN by an N dimensional binary state vector s and a state transition by matrix operations on Q, theta and s. Given a list of state transitions S = s_0...s_L, we learn Q and theta in a continuous space by minimizing a cost function J(Q*,theta,S) w.r.t. a real number matrix Q* and theta while thresholding Q* into a binary matrix Q using theta so that Q represents an AND/OR BN realizing the target state transitions S. We conducted experiments with artificial and real data sets to check scalability and accuracy of our learning algorithm. First we randomly generated AND/OR BNs up to N=5,000 nodes and empirically confirmed O(N^2) learning time behavior using them. We also observed 99.8% bit-by-bit prediction accuracy (prediction accuracy = 1 - test error) with state transition data generated by AND/OR BNs. For real data, we learned genome-wide AND/OR BNs with 10,928 nodes for budding yeast from transcription profiling data sets, each containing 10,928 mRNAs and 40 transitions and achieved for instance 84.3% prediction accuracy and successfully extracted more than 6,000 small AND/ORs whose average prediction accuracy reaches much higher 94.9%. Taisuke Sato, Ryosuke Kojima |
KR | 2 |
| 2019 | Batched Sparse Matrix Multiplication for Accelerating Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) are recently getting much attention in bioinformatics and chemoinformatics as a state-of-the-art machine learning approach with high accuracy. GCNs process convolutional operations along with graph structures, and GPUs are used to process enormous operations including sparse-dense matrix multiplication (SpMM) when the graph structure is expressed as an adjacency matrix with sparse matrix format. However, the SpMM operation on small graph, where the number of nodes is tens or hundreds, hardly exploits high parallelism or compute power of GPU. Therefore, SpMM becomes a bottleneck of training and inference in GCNs applications. In order to improve the performance of GCNs applications, we propose new SpMM algorithm especially for small sparse matrix and Batched SpMM, which exploits high parallelism of GPU by processing multiple SpMM operations with single CUDA kernel. To the best of our knowledge, this is the first work of batched approach for SpMM. We evaluated the performance of the GCNs application on TSUBAME3.0 implementing NVIDIA Tesla P100 GPU, and our batched approach shows significant speedups of up to 1.59x and 1.37x in training and inference, respectively. Yusuke Nagasaka, Akira Nukada, Ryosuke Kojima, Satoshi Matsuoka |
CCGRID | 3 |
| 2018 | HARK-Bird-Box: A Portable Real-time Bird Song Scene Analysis SystemabstractThis paper addresses real-time bird song scene analysis. Observation of animal behavior such as communication of wild birds would be aided by a portable device implementing a real-time system that can localize sound sources, measure their timing, classify their sources, and visualize these factors of sources. The difficulty of such a system is an integration of these functions considering the real-time requirement. To realize such a system, we propose a cascaded approach, cascading sound source detection, localization, separation, feature extraction, classification, and visualization for bird song analysis. Our system is constructed by combining an open source software for robot audition called HARK and a deep learning library to implement a bird song classifier based on a convolutional neural network (CNN). Considering portability, we implemented this system on a single-board computer, Jetson TX2, with a microphone array and developed a prototype device for bird song scene analysis. A preliminary experiment confirms a computational time for the whole system to realize a real-time system. Also, an additional experiment with a bird song dataset revealed a trade-off relationship between classification accuracy and time consuming and the effectiveness of our classifier. Ryosuke Kojima, Osamu Sugiyama, Kotaro Hoshiba, Reiji Suzuki, Kazuhiro Nakadai |
IROS | 1 |
| 2018 | Learning to rank in PRISM
Ryosuke Kojima, Taisuke Sato |
Int. J. Approx. Reason. | 1 |
| 2017 | A Spatial-Cue-Based Probabilistic Model for Bird Song Scene AnalysisabstractThis paper addresses bird song scene analysis focusing on location of birds and acoustic features of bird songs. Such a research area usually requires manual annotation related to positions and/or vocalization types of the target animals for a large amount of observed data. However, this manual annotation has two problems. One is that it is tough to annotate data observed in real environments because environmental noise exist and sound is reflected by trees and the ground, and also several birds at different locations may sing at the same time. The other is that it is inevitable that manual annotation produces inaccurate and inconsistent labels due to human errors and annotators' individual differences. For the first problem, we propose a Spatial-Cue-Based Probabilistic Model (SCBPM), which is a probabilistic model to estimate the maximum likelihood result for a bird song scene analysis by integrating sound source detection, localization, separation and identification based on spatial information of sound sources. For the second problem, we employ a semiautomatic annotation approach, in which a semi-supervised training method is deduced for SCBPM. This method decreases the amount of manual annotation. Preliminary experiments using recorded bird song data from the wild revealed that our system outperformed a conventional bird song scene analysis system by simply connecting sound source detection, localization, separation and identification in a cascade way in terms of identification accuracy. Ryosuke Kojima, Osamu Sugiyama, Kotaro Hoshiba, Reiji Suzuki, Kazuhiro Nakadai |
DSAA | 1 |
| 2017 | Development of microphone-array-embedded UAV for search and rescue taskabstractThis paper addresses online outdoor sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). In addition to sound source localization, sound source enhancement and robust communication method are also described. This system is one instance of deployment of our continuously developing open source software for robot audition called HARK (Honda Research Institute Japan Audition for Robots with Kyoto University). To improve the robustness against outdoor acoustic noise, we propose to combine two sound source localization methods based on MUSIC (multiple signal classification) to cope with trade-off between latency and noise robustness. The standard Eigenvalue decomposition based MUSIC (SEVD-MUSIC) has smaller latency but less noise robustness, whereas the incremental generalized singular value decomposition based MUSIC (iGSVD-MUSIC) has higher noise robustness but larger latency. A UAV operator can use an appropriate method according to the situation. A sound enhancement method called online robust principal component analysis (ORPCA) enables the operator to detect a target sound source more easily. To improve the stability of wireless communication, and robustness of the UAV system against weather changes, we developed data compression based on free lossless audio codec (FLAC) extended to support a 16 ch audio data stream via UDP, and developed a water-resistant microphone array. The resulting system successfully worked in an outdoor search and rescue task in ImPACT Tough Robotics Challenge in November 2016. Kazuhiro Nakadai, Makoto Kumon, Hiroshi G. Okuno, Kotaro Hoshiba, Mizuho Wakabayashi, Kai Washizaki, Takahiro Ishiki, Daniel Gabriel, Yoshiaki Bando, Takayuki Morito, Ryosuke Kojima, Osamu Sugiyama |
IROS | 11 |
| 2016 | Reduction of Computational Cost Using Two-Stage Deep Neural Network for Training for Denoising and Sound Source Identification
Takayuki Morito, Osamu Sugiyama, Satoshi Uemura, Ryosuke Kojima, Kazuhiro Nakadai |
IEA/AIE | 4 |
| 2016 | Semi-automatic bird song analysis by spatial-cue-based integration of sound source detection, localization, separation, and identificationabstractThis paper addresses bird song analysis based on semi-automatic annotation. Research in animal behavior, especially with birds, would be aided by automated (or semiautomated) systems that can localize sounds, measure their timing, and identify their source. This is difficult to achieve in real environments where several birds may be singing from different locations and at the same time. Analysis of recordings from the wild has in the past typically required manual annotation. Such annotation is not always accurate or even consistent, as it may vary both within or between observers. Here we propose a system that uses automated methods from robot audition, including sound source detection, localization, separation and identification. In robot audition these technologies have typically been studied separately; combining them often leads to poor performance in real-time application from the wild. We suggest that integration is aided by placing a primary focus on spatial cues, then combining other features within a Bayesian framework. A second problem has been that supervised machine learning methods typically requires a pre-trained model that may require a large training set of annotated labels. We have employed a semi-automatic annotation approach that requires much less pre-annotation. Preliminary experiments with recordings of bird songs from the wild revealed that for identification accuracy our system outperformed a method based on conventional robot audition. Ryosuke Kojima, Osamu Sugiyama, Reiji Suzuki, Kazuhiro Nakadai, Charles E. Taylor |
IROS | 1 |
| 2016 | Partially Shared Deep Neural Network in sound source separation and identification using a UAV-embedded microphone arrayabstractThis paper addresses sound source separation and identification for noise-contaminated acoustic signals recorded with a microphone array embedded in an Unmanned Aerial Vehicle (UAV), aiming at people's voice detection quickly and widely in a disaster situation. The key approach to achieve this is Deep Neural Network (DNN), but it is well known that training a DNN needs a huge dataset to improve its performance. In a practical application, building such a dataset is not often realistic owing to the cost of manual data annotation. Therefore, we propose a Partially-Shared Deep Neural Network (PS-DNN) which can learn multiple tasks at the same time with a small amount of annotated data. Preliminary results show that the PS-DNN outperforms conventional DNN-based approaches which require fully-annotated data in training in terms of identification accuracy. In addition, it maintains performance even when noise-suppressed signals are used for sound source separation training, and partially annotated data is used for sound source identification training. Takayuki Morito, Osamu Sugiyama, Ryosuke Kojima, Kazuhiro Nakadai |
IROS | 3 |
| 2015 | Scene Understanding Based on Sound and Text Information for a Cooking Support Robot
Ryosuke Kojima, Osamu Sugiyama, Kazuhiro Nakadai |
IEA/AIE | 1 |
| 2015 | Interactive Interface to Optimize Sound Source Localization with HARK
Osamu Sugiyama, Ryosuke Kojima, Kazuhiro Nakadai |
IEA/AIE | 2 |
| 2015 | Audio-visual scene understanding utilizing text information for a cooking support robotabstractThis paper addresses multimodal “scene understanding” for a robot using audio-visual and text information. Scene understanding is defined by extracting six-W information such as What, When, Where, Who, Why, and hoW on the surrounding environment. Although scene understanding for a robot has been studied in the fields of robot vision and audition, only the first four Ws except for why and how information were considered. We, thus, focus on extracting how information, in particular, on cooking scenes. In cooking scenes, we define how information as a cooking procedure, and it is useful that a robot gives appropriate advice for cooking. To realize such cooking support, we propose a multi-modal cooking procedure recognition framework consisting of Convolutional Neural Network (CNN), and Hierarchical Hidden Markov Model (HHMM). CNN is knows as one of the most advanced classifiers, and it is applied to recognize a cooking events from audio and visual information. HHMM models a cooking procedure represented by a sequence of cooking events, which is defined as a relationship between cooking events using text data obtained from web, and the cooking events classified with CNN. Therefore, our proposed framework integrates these three types of modalities. We constructed an interactive cooking support system based on the proposed framework, which advice a next step in the current cooking procedure through human-robot communication. Preliminary results with simulated and real recorded multi-modal scenes showed the robustness of the proposed framework in a noisy and/or occluded situation. Ryosuke Kojima, Osamu Sugiyama, Kazuhiro Nakadai |
IROS | 1 |
| 2014 | Goal and Plan Recognition via Parse Trees Using Prefix and Infix Probability Computation
Ryosuke Kojima, Taisuke Sato |
ILP | 1 |