Ryo Yoshida

dblp:44/2367 · DBLP profile ↗
← Back
37ranked-venue papers
11as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 4 first-author · 11 since 2021Systems, architecture and hardware · 4 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Language Acquisition Device in Large Language Models
abstract
Large Language Models (LLMs) remain substantially less data-efficient than humans.Prepretraining (PPT) on synthetic languages has been proposed to close this gap, with prior work emphasizing highly expressive formal languages such as k-Shuffle Dyck.Inspired by the Language Acquisition Device (LAD) hypothesis, which posits that innate constraints preemptively restrict the learner's hypothesis space to natural-language-like structure, we propose LAD-inspired PPT: pre-pretraining on MP-STRUCT, a formal language whose strings encode hierarchical composition, feature-based dependencies, and long-distance displacement via MERGE, AGREE, and MOVE.A brief 500step PPT with MP-STRUCT matches strong formal-language baselines in token efficiency while additionally imparting a human-like resistance to structurally implausible languages.Analyzing simplified variants, we find that MP-STRUCT CORE outperforms k-Shuffle Dyck despite not being definable in C-RASP (a formal bound on transformer expressivity), challenging the prior hypothesis that effective PPT languages must be both hierarchically expressive and circuit-theoretically learnable.We show that functional landmarks, which reduce dependency resolution ambiguity, are a key driver, suggesting that effective PPT design depends not only on expressivity but also on the accessibility of dependency resolution.
Masato Mita, Taiga Someya, Ryo Yoshida, Yohei Oseki
ACL (1)3
2026 An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal
abstract
Surprisal theory hypothesizes that the difficulty of human sentence processing increases linearly with surprisal, the negative logprobability of a word given its context.Computational psycholinguistics has tested this hypothesis using language models (LMs) as proxies for human prediction.While surprisal derived from recent neural LMs generally captures human processing difficulty on naturalistic corpora that predominantly consist of simple sentences, it severely underestimates processing difficulty on sentences that require syntactic disambiguation (garden-path effects).This leads to the claim that the processing difficulty of such sentences cannot be reduced to surprisal, although it remains possible that neural LMs simply differ from humans in nextword prediction.In this paper, we investigate whether it is truly impossible to construct a neural LM that can explain garden-path effects via surprisal.Specifically, instead of evaluating offthe-shelf neural LMs, we fine-tune these LMs on garden-path sentences so as to better align surprisal-based reading-time estimates with actual human reading times.Our results show that fine-tuned LMs do not overfit and successfully capture human reading slowdowns on held-out garden-path items; they even improve predictive power for human reading times on naturalistic corpora and preserve their general LM capabilities.These results provide an existence proof for a neural LM that can explain both garden-path effects and naturalistic reading times via surprisal, but also raise a theoretical question: what kind of evidence can truly falsify surprisal theory?
Ryo Yoshida, Shinnosuke Isono, Taiga Someya, Yohei Oseki, Tatsuki Kuribayashi
ACL (1)1
2025 Developmentally-plausible Working Memory Shapes a Critical Period for Language Acquisition
abstract
Large language models possess general linguistic abilities but acquire language less efficiently than humans. This study proposes a method for integrating the developmental characteristics of working memory during the critical period, a stage when human language acquisition is particularly efficient, into the training process of language models. The proposed method introduces a mechanism that initially constrains working memory during the early stages of training and gradually relaxes this constraint in an exponential manner as learning progresses. Targeted syntactic evaluation shows that the proposed method outperforms conventional methods without memory constraints or with static memory constraints. These findings not only provide new directions for designing data-efficient language models but also offer indirect evidence supporting the role of the developmental characteristics of working memory as the underlying mechanism of the critical period in language acquisition.
Masato Mita, Ryo Yoshida, Yohei Oseki
ACL (1)2
2025 If Attention Serves as a Cognitive Model of Human Memory Retrieval, What is the Plausible Memory Representation?
abstract
Recent work in computational psycholinguistics has revealed intriguing parallels between attention mechanisms and human memory retrieval, focusing primarily on vanilla Transformers that operate on token-level representations. However, computational psycholinguistic research has also established that syntactic structures provide compelling explanations for human sentence processing that token-level factors cannot fully account for. In this paper, we investigate whether the attention mechanism of Transformer Grammar (TG), which uniquely operates on syntactic structures as representational units, can serve as a cognitive model of human memory retrieval, using Normalized Attention Entropy (NAE) as a linking hypothesis between models and humans. Our experiments demonstrate that TG’s attention achieves superior predictive power for self-paced reading times compared to vanilla Transformer’s, with further analyses revealing independent contributions from both models. These findings suggest that human sentence processing involves dual memory representations—one based on syntactic structures and another on token sequences—with attention serving as the general memory retrieval algorithm, while highlighting the importance of incorporating syntactic structures as representational units.
Ryo Yoshida, Shinnosuke Isono, Kohei Kajikawa, Taiga Someya, Yushi Sugimoto, Yohei Oseki
ACL (1)1
2025 On the Realism of LiDAR Spoofing Attacks against Autonomous Driving Vehicle at High Speed and Long Distance
Takami Sato, Yuki Hayakawa, Kazuma Ikeda, Ozora Sako, Rokuto Nagata, Ryo Yoshida, Qi Alfred Chen, Kentaro Yoshioka
NDSS7
2024 Emergent Word Order Universals from Cognitively-Motivated Language Models
abstract
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Tatsuki Kuribayashi, Ryo Ueda, Ryo Yoshida, Yohei Oseki, Ted Briscoe, Timothy Baldwin
ACL (1)3
2024 Dissociating Syntactic Operations via Composition Count
Kohei Kajikawa, Ryo Yoshida, Yohei Oseki
CogSci2
2024 Targeted Syntactic Evaluation on the Chomsky Hierarchy
abstract
In this paper, we propose a novel evaluation paradigm for Targeted Syntactic Evaluations, where we assess how well language models can recognize linguistic phenomena situated at different levels of the Chomsky hierarchy. Specifically, we create formal languages that abstract four syntactic phenomena in natural languages, each identified at a different level of the Chomsky hierarchy, and use these to evaluate the capabilities of language models: (1) (Adj)ˆn NP type, (2) NPˆn VPˆn type, (3) Nested Dependency type, and (4) Cross Serial Dependency type. We first train three different language models (LSTM, Transformer LM, and Stack-RNN) on language modeling tasks and then evaluate them using pairs of a positive and a negative sentence by investigating whether they can assign a higher probability to the positive sentence than the negative one. Our result demonstrated that all language models have the ability to capture the structural patterns of the (Adj)ˆn NP type formal language. However, LSTM and Transformer LM failed to capture NPˆn VPˆn type language and no architectures can recognize nested dependency and Cross Serial dependency correctly. Neural language models, especially Transformer LMs, have exhibited high performance across a multitude of downstream tasks, leading to the perception that they possess an understanding of natural languages. However, our findings suggest that these models may not necessarily comprehend the syntactic structures that underlie natural language phenomena such as dependency. Rather, it appears that they may extend grammatical rules equivalent to regular grammars to approximate the rules governing dependencies.
Taiga Someya, Ryo Yoshida, Yohei Oseki
LREC/COLING2
2024 Ensemble dynamics and information flow deduction from whole-brain imaging data
abstract
The recent advancements in large-scale activity imaging of neuronal ensembles offer valuable opportunities to comprehend the process involved in generating brain activity patterns and understanding how information is transmitted between neurons or neuronal ensembles. However, existing methodologies for extracting the underlying properties that generate overall dynamics are still limited. In this study, we applied previously unexplored methodologies to analyze time-lapse 3D imaging (4D imaging) data of head neurons of the nematode Caenorhabditis elegans. By combining time-delay embedding with the independent component analysis, we successfully decomposed whole-brain activities into a small number of component dynamics. Through the integration of results from multiple samples, we extracted common dynamics from neuronal activities that exhibit apparent divergence across different animals. Notably, while several components show common cooperativity across samples, some component pairs exhibited distinct relationships between individual samples. We further developed time series prediction models of synaptic communications. By combining dimension reduction using the general framework, gradient kernel dimension reduction, and probabilistic modeling, the overall relationships of neural activities were incorporated. By this approach, the stochastic but coordinated dynamics were reproduced in the simulated whole-brain neural network. We found that noise in the nervous system is crucial for generating realistic whole-brain dynamics. Furthermore, by evaluating synaptic interaction properties in the models, strong interactions within the core neural circuit, variable sensory transmission and importance of gap junctions were inferred. Virtual optogenetics can be also performed using the model. These analyses provide a solid foundation for understanding information flow in real neural networks.
Yu Toyoshima, Hirofumi Sato, Daiki Nagata, Manami Kanamori, Moon Sun Jang, Koyo Kuze, Suzu Oe, Takayuki Teramoto, Yuishi Iwasaki, Ryo Yoshida, Takeshi Ishihara, Yuichi Iino
PLoS Comput. Biol.10
2023 Transfer Learning with Affine Model Transformation
abstract
Supervised transfer learning has received considerable attention due to its potential to boost the predictive power of machine learning in scenarios where data are scarce. Generally, a given set of source models and a dataset from a target domain are used to adapt the pre-trained models to a target domain by statistically learning domain shift and domain-specific factors. While such procedurally and intuitively plausible methods have achieved great success in a wide range of real-world applications, the lack of a theoretical basis hinders further methodological development. This paper presents a general class of transfer learning regression called affine model transfer, following the principle of expected-square loss minimization. It is shown that the affine model transfer broadly encompasses various existing methods, including the most common procedure based on neural feature extractors. Furthermore, the current paper clarifies theoretical properties of the affine model transfer such as generalization error and excess risk. Through several case studies, we demonstrate the practical benefits of modeling and estimating inter-domain commonality and domain-specific factors separately with the affine-type transfer models.
Shunya Minami, Kenji Fukumizu, Yoshihiro Hayashi, Ryo Yoshida
NeurIPS4
2021 A General Class of Transfer Learning Regression without Implementation Cost
abstract
We propose a novel framework that unifies and extends existing methods of transfer learning (TL) for regression. To bridge a pretrained source model to the model on a target task, we introduce a density-ratio reweighting function, which is estimated through the Bayesian framework with a specific prior distribution. By changing two intrinsic hyperparameters and the choice of the density-ratio model, the proposed method can integrate three popular methods of TL: TL based on cross-domain similarity regularization, a probabilistic TL using the density-ratio estimation, and fine-tuning of pretrained neural networks. Moreover, the proposed method can benefit from its simple implementation without any additional cost; the regression model can be fully trained using off-the-shelf libraries for supervised learning in which the original output variable is simply transformed to a new output variable. We demonstrate its simplicity, generality, and applicability using various real data applications.
Shunya Minami, Song Liu 0002, Stephen Wu 0001, Kenji Fukumizu, Ryo Yoshida
AAAI5
2021 Lower Perplexity is Not Always Human-Like
abstract
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tatsuki Kuribayashi, Yohei Oseki, Takumi Ito, Ryo Yoshida, Masayuki Asahara, Kentaro Inui
ACL/IJCNLP (1)4
2021 Modeling Human Sentence Processing with Left-Corner Recurrent Neural Network Grammars
abstract
In computational linguistics, it has been shown that hierarchical structures make language models (LMs) more human-like.However, the previous literature has been agnostic about a parsing strategy of the hierarchical models.In this paper, we investigated whether hierarchical structures make LMs more human-like, and if so, which parsing strategy is most cognitively plausible.In order to address this question, we evaluated three LMs against human reading times in Japanese with head-final leftbranching structures: Long Short-Term Memory (LSTM) as a sequential model and Recurrent Neural Network Grammars (RNNGs) with top-down and left-corner parsing strategies as hierarchical models.Our computational modeling demonstrated that left-corner RNNGs outperformed top-down RNNGs and LSTM, suggesting that hierarchical and leftcorner architectures are more cognitively plausible than top-down or sequential architectures.In addition, the relationships between the cognitive plausibility and (i) perplexity, (ii) parsing, and (iii) beam size will also be discussed.1
Ryo Yoshida, Hiroshi Noji, Yohei Oseki
EMNLP (1)1
2019 Statistical inference of the rate of RNA polymerase II elongation by total RNA sequencing
abstract
MOTIVATION: Sequencing total RNA without poly-A selection enables us to obtain a transcriptomic profile of nascent RNAs undergoing transcription with co-transcriptional splicing. In general, the RNA-seq reads exhibit a sawtooth pattern in a gene, which is characterized by a monotonically decreasing gradient across introns in the 5'-3' direction, and by substantially higher levels of RNA-seq reads present in exonic regions. Such patterns result from the process of underlying transcription elongation by RNA polymerase II, which traverses the DNA strand in a 5'-3' direction as it performs a complex series of mRNA synthesis and processing. Therefore, data of sequenced total RNAs could be utilized to infer the rate of transcription elongation by solving the inverse problem. RESULTS: Though solving the inverse problem in total RNA-seq has the great potential, statistical methods have not yet been fully developed. We demonstrate what extent the newly developed method can be useful. The objective is to reconstruct the spatial distribution of transcription elongation rates in a gene from a given noisy, sawtooth-like profile. It is necessary to recover the signal source of the elongation rates separately from several types of nuisance factors, such as unobserved modes of co-transcriptionally occurring mRNA splicing, which exert significant influences on the sawtooth shape. The present method was tested using published total RNA-seq data derived from mouse embryonic stem cells. We investigated the spatial characteristics of the estimated elongation rates, focusing especially on the relation to promoter-proximal pausing of RNA polymerase II, nucleosome occupancy and histone modification patterns. AVAILABILITY AND IMPLEMENTATION: A C implementation of PolSter and sample data are available at https://github.com/yoshida-lab/PolSter. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yumi Kawamura, Shinsuke Koyama, Ryo Yoshida
Bioinform.3
2018 SPF-CellTracker: Tracking Multiple Cells with Strongly-Correlated Moves Using a Spatial Particle Filter
abstract
Tracking many cells in time-lapse 3D image sequences is an important challenging task of bioimage informatics. Motivated by a study of brain-wide 4D imaging of neural activity in C. elegans, we present a new method of multi-cell tracking. Data types to which the method is applicable are characterized as follows: (i) cells are imaged as globular-like objects, (ii) it is difficult to distinguish cells on the basis of shape and size only, (iii) the number of imaged cells in the several-hundred range, (iv) movements of nearly-located cells are strongly correlated, and (v) cells do not divide. We developed a tracking software suite that we call SPF-CellTracker. Incorporating dependency on the cells' movements into the prediction model is the key for reducing the tracking errors: the cell switching and the coalescence of the tracked positions. We model the target cells' correlated movements as a Markov random field and we also derive a fast computation algorithm, which we call spatial particle filter. With the live-imaging data of the nuclei of C. elegans neurons in which approximately 120 nuclei of neurons were imaged, the proposed method demonstrated improved accuracy compared to the standard particle filter and the method developed by Tokunaga et al. (2014).
Osamu Hirose, Shotaro Kawaguchi, Terumasa Tokunaga, Yu Toyoshima, Takayuki Teramoto, Sayuri Kuge, Takeshi Ishihara, Yuichi Iino, Ryo Yoshida
IEEE ACM Trans. Comput. Biol. Bioinform.9
2016 Accurate Automatic Detection of Densely Distributed Cell Nuclei in 3D Space
abstract
To measure the activity of neurons using whole-brain activity imaging, precise detection of each neuron or its nucleus is required. In the head region of the nematode C. elegans, the neuronal cell bodies are distributed densely in three-dimensional (3D) space. However, no existing computational methods of image analysis can separate them with sufficient accuracy. Here we propose a highly accurate segmentation method based on the curvatures of the iso-intensity surfaces. To obtain accurate positions of nuclei, we also developed a new procedure for least squares fitting with a Gaussian mixture model. Combining these methods enables accurate detection of densely distributed cell nuclei in a 3D space. The proposed method was implemented as a graphical user interface program that allows visualization and correction of the results of automatic detection. Additionally, the proposed method was applied to time-lapse 3D calcium imaging data, and most of the nuclei in the images were successfully tracked and measured.
Yu Toyoshima, Terumasa Tokunaga, Osamu Hirose, Manami Kanamori, Takayuki Teramoto, Moon Sun Jang, Sayuri Kuge, Takeshi Ishihara, Ryo Yoshida, Yuichi Iino
PLoS Comput. Biol.9
2015 Repulsive parallel MCMC algorithm for discovering diverse motifs from large sequence sets
abstract
Abstract Motivation The motif discovery problem consists of finding recurring patterns of short strings in a set of nucleotide sequences. This classical problem is receiving renewed attention as most early motif discovery methods lack the ability to handle large data of recent genome-wide ChIP studies. New ChIP-tailored methods focus on reducing computation time and pay little regard to the accuracy of motif detection. Unlike such methods, our method focuses on increasing the detection accuracy while maintaining the computation efficiency at an acceptable level. The major advantage of our method is that it can mine diverse multiple motifs undetectable by current methods. Results The repulsive parallel Markov chain Monte Carlo (RPMCMC) algorithm that we propose is a parallel version of the widely used Gibbs motif sampler. RPMCMC is run on parallel interacting motif samplers. A repulsive force is generated when different motifs produced by different samplers near each other. Thus, different samplers explore different motifs. In this way, we can detect much more diverse motifs than conventional methods can. Through application to 228 transcription factor ChIP-seq datasets of the ENCODE project, we show that the RPMCMC algorithm can find many reliable cofactor interacting motifs that existing methods are unable to discover. Availability and implementation A C++ implementation of RPMCMC and discovered cofactor motifs for the 228 ENCODE ChIP-seq datasets are available from http://daweb.ism.ac.jp/yoshidalab/motif. Supplementary information Supplementary data are available from Bioinformatics online.
Hisaki Ikebata, Ryo Yoshida
Bioinform.2
2014 Automated detection and tracking of many cells by using 4D live-cell imaging data
abstract
MOTIVATION: Automated fluorescence microscopes produce massive amounts of images observing cells, often in four dimensions of space and time. This study addresses two tasks of time-lapse imaging analyses; detection and tracking of the many imaged cells, and it is especially intended for 4D live-cell imaging of neuronal nuclei of Caenorhabditis elegans. The cells of interest appear as slightly deformed ellipsoidal forms. They are densely distributed, and move rapidly in a series of 3D images. Thus, existing tracking methods often fail because more than one tracker will follow the same target or a tracker transits from one to other of different targets during rapid moves. RESULTS: The present method begins by performing the kernel density estimation in order to convert each 3D image into a smooth, continuous function. The cell bodies in the image are assumed to lie in the regions near the multiple local maxima of the density function. The tasks of detecting and tracking the cells are then addressed with two hill-climbing algorithms. The positions of the trackers are initialized by applying the cell-detection method to an image in the first frame. The tracking method keeps attacking them to near the local maxima in each subsequent image. To prevent the tracker from following multiple cells, we use a Markov random field (MRF) to model the spatial and temporal covariation of the cells and to maximize the image forces and the MRF-induced constraint on the trackers. The tracking procedure is demonstrated with dynamic 3D images that each contain >100 neurons of C.elegans. AVAILABILITY: http://daweb.ism.ac.jp/yoshidalab/crest/ismb2014 SUPPLEMENTARY INFORMATION: Supplementary data are available at http://daweb.ism.ac.jp/yoshidalab/crest/ismb2014
Terumasa Tokunaga, Osamu Hirose, Shotaro Kawaguchi, Yu Toyoshima, Takayuki Teramoto, Hisaki Ikebata, Sayuri Kuge, Takeshi Ishihara, Yuichi Iino, Ryo Yoshida
Bioinform.10
2012 Identifying Gene Pathways Associated with Cancer Characteristics via Sparse Statistical Methods
abstract
We propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the Sparse Probabilistic Principal Component Analysis (SPPCA). A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data.
Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.7
2011 SiGN-SSM: open source parallel software for estimating gene networks with state space models
abstract
UNLABELLED: SiGN-SSM is an open-source gene network estimation software able to run in parallel on PCs and massively parallel supercomputers. The software estimates a state space model (SSM), that is a statistical dynamic model suitable for analyzing short time and/or replicated time series gene expression profiles. SiGN-SSM implements a novel parameter constraint effective to stabilize the estimated models. Also, by using a supercomputer, it is able to determine the gene network structure by a statistical permutation test in a practical time. SiGN-SSM is applicable not only to analyzing temporal regulatory dependencies between genes, but also to extracting the differentially regulated genes from time series expression profiles. AVAILABILITY: SiGN-SSM is distributed under GNU Affero General Public Licence (GNU AGPL) version 3 and can be downloaded at http://sign.hgc.jp/signssm/. The pre-compiled binaries for some architectures are available in addition to the source code. The pre-installed binaries are also available on the Human Genome Center supercomputer system. The online manual and the supplementary information of SiGN-SSM is available on our web site. CONTACT: [email protected].
Yoshinori Tamada, Rui Yamaguchi, Seiya Imoto, Osamu Hirose, Ryo Yoshida, Masao Nagasaki, Satoru Miyano
Bioinform.5
2010 Discovering functional gene pathways associated with cancer heterogeneity via sparse supervised learning
abstract
We propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the sparse probabilistic principal component analysis. A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data.
Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano
BIBM7
2010 Implementation of Sequential Importance Sampling in GPGPU
Keisuke Hayashi, Masaya M. Saito, Ryo Yoshida, Tomoyuki Higuchi
FUSION3
2010 Bayesian experts in exploring reaction kinetics of transcription circuits
abstract
MOTIVATION: Biochemical reactions in cells are made of several types of biological circuits. In current systems biology, making differential equation (DE) models simulatable in silico has been an appealing, general approach to uncover a complex world of biochemical reaction dynamics. Despite of a need for simulation-aided studies, our research field has yet provided no clear answers: how to specify kinetic values in models that are difficult to measure from experimental/theoretical analyses on biochemical kinetics. RESULTS: We present a novel non-parametric Bayesian approach to this problem. The key idea lies in the development of a Dirichlet process (DP) prior distribution, called Bayesian experts, which reflects substantive knowledge on reaction mechanisms inherent in given models and experimentally observable kinetic evidences to the subsequent parameter search. The DP prior identifies significant local regions of unknown parameter space before proceeding to the posterior analyses. This article reports that a Bayesian expert-inducing stochastic search can effectively explore unknown parameters of in silico transcription circuits such that solutions of DEs reproduce transcriptomic time course profiles. AVAILABILITY: A sample source code is available at the URL http://daweb.ism.ac.jp/~yoshidar/lisdas/.
Ryo Yoshida, Masaya M. Saito, Hiromichi Nagao, Tomoyuki Higuchi
Bioinform.1
2010 Bayesian Learning in Sparse Graphical Factor Models via Variational Mean-Field Annealing
Ryo Yoshida, Mike West
J. Mach. Learn. Res.1
2009 Development of novel self-oscillating molecular robot fueled by organic acid
abstract
In our previous study, we first succeeded in construction of a novel-types self-oscillting molecular robot. However, the driving environment of the molecular robot was firmly restricted in strong acid conditions. This is because the molecular robots drive induced by the Belousov-Zhabotinsky (BZ) reaction, which is well known for exhibiting temporal and spatiotemporal oscillating phenomena. The overall process of the BZ reaction is the oxidation of an organic substrate, such as malonic acid (MA) or citric acid, by an oxidizing agent (bromate ion) in the presence of a strong acid and a metal catalyst. In this study, in order to drive the novel molecular robot under the physiological condition, we conducted the modification of the molecular structure of the self-oscillating polymer chain. In order to cause the self-oscillation under the biological condition, we synthesized a built-in system where the BZ substrates other than organic acid were incorporated into the molecular robot itself. As a result, the novel molecular robot drives under the biological condition. We believe that the development of the novel molecular robot lead to construction of the novel biomimetic soft robots and actuators, and may inspire novel nonlinear experimental and theoretical considerations.
Yusuke Hara, Shingo Maeda, Ryo Yoshida, Shuji Hashimoto
IROS3
2009 Chemical robot-design of peristaltic polymer gel actuator-
abstract
In this proceeding, we introduce peristaltic gel actuators as a chemical robot. The polymer gels prepared here have a cyclic reaction network like metabolic process in itself. With a cyclic reaction, the polymer gel swells-shrinks autonomously. The periodic self-oscillating motion of the gel is produced by the dissipating chemical energy of the oscillatory Belouzov-Zhabotinsky (BZ) reaction. We have succeeded in convey the object automatically by utilizing the synthetic polymer gel. This experimental fact represents the great possibility of the chemical robot.
Shingo Maeda, Yusuke Hara, Ryo Yoshida, Shuji Hashimoto
IROS3
2008 A New Cell Phone Remote Control for People with Visual Impairment
Ryo Yoshida, Michiaki Yasumura
ICCHP1
2008 Statistical inference of transcriptional module-based gene networks from time course gene expression profiles by using state space models
abstract
MOTIVATION: Statistical inference of gene networks by using time-course microarray gene expression profiles is an essential step towards understanding the temporal structure of gene regulatory mechanisms. Unfortunately, most of the current studies have been limited to analysing a small number of genes because the length of time-course gene expression profiles is fairly short. One promising approach to overcome such a limitation is to infer gene networks by exploring the potential transcriptional modules which are sets of genes sharing a common function or involved in the same pathway. RESULTS: In this article, we present a novel approach based on the state space model to identify the transcriptional modules and module-based gene networks simultaneously. The state space model has the potential to infer large-scale gene networks, e.g. of order 10(3), from time-course gene expression profiles. Particularly, we succeeded in the identification of a cell cycle system by using the gene expression profiles of Saccharomyces cerevisiae in which the length of the time-course and number of genes were 24 and 4382, respectively. However, when analysing shorter time-course data, e.g. of length 10 or less, the parameter estimations of the state space model often fail due to overfitting. To extend the applicability of the state space model, we provide an approach to use the technical replicates of gene expression profiles, which are often measured in duplicate or triplicate. The use of technical replicates is important for achieving highly-efficient inferences of gene networks with short time-course data. The potential of the proposed method has been demonstrated through the time-course analysis of the gene expression profiles of human umbilical vein endothelial cells (HUVECs) undergoing growth factor deprivation-induced apoptosis. AVAILABILITY: Supplementary Information and the software (TRANS-MNET) are available at http://daweb.ism.ac.jp/~yoshidar/software/ssm/.
Osamu Hirose, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Tomoyuki Higuchi, Stephen D. Charnock-Jones, Cristin G. Print, Satoru Miyano
Bioinform.2
2008 Bayesian learning of biological pathways on genomic data assimilation
abstract
MOTIVATION: Mathematical modeling and simulation, based on biochemical rate equations, provide us a rigorous tool for unraveling complex mechanisms of biological pathways. To proceed to simulation experiments, it is an essential first step to find effective values of model parameters, which are difficult to measure from in vivo and in vitro experiments. Furthermore, once a set of hypothetical models has been created, any statistical criterion is needed to test the ability of the constructed models and to proceed to model revision. RESULTS: The aim of our research is to present a new statistical technology towards data-driven construction of in silico biological pathways. The method starts with a knowledge-based modeling with hybrid functional Petri net. It then proceeds to the Bayesian learning of model parameters for which experimental data are available. This process exploits quantitative measurements of evolving biochemical reactions, e.g. gene expression data. Another important issue that we consider is statistical evaluation and comparison of the constructed hypothetical pathways. For this purpose, we have developed a new Bayesian information-theoretic measure that assesses the predictability and the biological robustness of in silico pathways. AVAILABILITY: The FORTRAN source codes are available at the URL http://daweb.ism.ac.jpyoshidar/GDA/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ryo Yoshida, Masao Nagasaki, Rui Yamaguchi, Seiya Imoto, Satoru Miyano, Tomoyuki Higuchi
Bioinform.1
2008 ExonMiner: Web service for analysis of GeneChip Exon array data
abstract
BACKGROUND: Some splicing isoform-specific transcriptional regulations are related to disease. Therefore, detection of disease specific splice variations is the first step for finding disease specific transcriptional regulations. Affymetrix Human Exon 1.0 ST Array can measure exon-level expression profiles that are suitable to find differentially expressed exons in genome-wide scale. However, exon array produces massive datasets that are more than we can handle and analyze on personal computer. RESULTS: We have developed ExonMiner that is the first all-in-one web service for analysis of exon array data to detect transcripts that have significantly different splicing patterns in two cells, e.g. normal and cancer cells. ExonMiner can perform the following analyses: (1) data normalization, (2) statistical analysis based on two-way ANOVA, (3) finding transcripts with significantly different splice patterns, (4) efficient visualization based on heatmaps and barplots, and (5) meta-analysis to detect exon level biomarkers. We implemented ExonMiner on a supercomputer system in order to perform genome-wide analysis for more than 300,000 transcripts in exon array data, which has the potential to reveal the aberrant splice variations in cancer cells as exon level biomarkers. CONCLUSION: ExonMiner is well suited for analysis of exon array data and does not require any installation of software except for internet browsers. What all users need to do is to access the ExonMiner URL http://ae.hgc.jp/exonminer. Users can analyze full dataset of exon array data within hours by high-level statistical analysis with sound theoretical basis that finds aberrant splice variants as biomarkers.
Kazuyuki Numata, Ryo Yoshida, Masao Nagasaki, Ayumu Saito, Seiya Imoto, Satoru Miyano
BMC Bioinform.2
2007 Computational Genome-Wide Discovery of Aberrant Splice Variations with Exon Expression Profiles
abstract
Alternative splicing plays a prominent role in eukaryotic gene regulations that allow a single gene to generate the multiple mRNA products. The recent advent of GeneChipregHuman Exon 1.0 ST Array enables us to measure the exon expression profiles of human cells on a genome-wide scale. With this advent, analysis of functional gene regulation could be extended to detect not only differentially expressed genes, but also specific splicing events that occur in target cells, but not in normal controls. We address some statistical issues for the identification of biomarker splice variations with exon expression data. The proposed method involves the following steps: (1) Whole transcript analysis with the nonparametric analysis of variance (ANOVA) to identify potential biomarkers that present specific splice variations. (2) Meta-analysis for discriminating non-specific splice variations that are caused by clinical heterogeneity in the collected samples. In the analysis of human cells, controlling non-specific splicing factors is essential for success in the detection of biomarker splice variations because splice patterns are possibly affected by inter-individual differences in the collected samples. We demonstrate its utility and perform a whole transcript analysis of exon expression profiles of colorectal carcinoma.
Ryo Yoshida, Kazuyuki Numata, Seiya Imoto, Masao Nagasaki, Atsushi Doi, Kazuko Ueno, Satoru Miyano
BIBE1
2007 Chemical robot - Design of self-walking gel
abstract
Stimuli-responsive polymers and gels have been applied to biomimetic actuators or artificial muscles. Biomimetic robots using such materials can move like a living creature. Especially, electroactive polymers are said to be promising materials as advanced soft actuators. Its mechanical motion is controlled by external stimuli such as a change in solvent composition, pH, temperature and electric field etc. On the other hand, many living organisms generate an autonomous motion without external driving stimuli. In this paper, we report a novel biomimetic gel actuator that can bend and stretch with worm-like motion spontaneously without external driving stimuli. The gel actuator converts the chemical energy of oscillating reaction, i.e., the Belouzov-Zhabotinsky (BZ) reaction into the kinetic energy inside the gel. Although the gel is completely composed of synthetic polymer, it shows autonomous motion as if it is alive. To cause anisotropic contraction with curvature changes, the gel strip with gradient structure was prepared. Furthermore, the gel can walk with repeated bending and stretching motion by itself like a looper by coupling with ratchet mechanism. The "self-walking" gel actuator will create a new framework of biomimetic chemical robot.
Shingo Maeda, Yusuke Hara, Ryo Yoshida, Shuji Hashimoto
IROS3
2007 Statistical Absolute Evaluation of Gene Ontology Terms with Gene Expression Data
Pramod K. Gupta, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Satoru Miyano
ISBRA2
2006 Prototyping and Evaluation of New Remote Controls for People with Visual Impairment
Michiaki Yasumura, Ryo Yoshida, Megumi Yoshida
ICCHP2
2006 ArrayCluster: an analytic tool for clustering, data visualization and module finder on gene expression profiles
abstract
SUMMARY: One of the significant challenges in gene expression analysis is to find unknown subtypes of several diseases at the molecular levels. This task can be addressed by grouping gene expression patterns of the collected samples on the basis of a large number of genes. Application of commonly used clustering methods to such a dataset however are likely to fail owing to over-learning, because the number of samples to be grouped is much smaller than the data dimension which is equal to the number of genes involved in the dataset. To overcome such difficulty, we developed a novel model-based clustering method, referred to as the mixed factors analysis. The ArrayCluster is a freely available software to perform the mixed factors analysis. It provides us some analytic tools for clustering DNA microarray experiments, data visualization and an automatic detector for module transcriptional of genes that are relevant to the calibrated molecular subtypes and so on.
Ryo Yoshida, Tomoyuki Higuchi, Seiya Imoto, Satoru Miyano
Bioinform.1
2005 A Penalized Likelihood Estimation on Transcriptional Module-Based Clustering
Ryo Yoshida, Seiya Imoto, Tomoyuki Higuchi
ICCSA (3)1
2000 3D web environment for knowledge management
Ryo Yoshida, Takaaki Murao, Tatsuo Miyazawa
Future Gener. Comput. Syst.1