EDBT 2026 Demo / reviewers in the wild / expert
Takashi Matsubara 0001
dblp:70/6748-1
· DBLP profile ↗
44ranked-venue papers
19as first author
17since 2021 · last 2026
0000-0003-0642-4800ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 18 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Object-Centric World Models for Causality-Aware Reinforcement LearningabstractWorld models have been developed to support sample-efficient deep reinforcement learning agents. However, it remains challenging for world models to accurately replicate environments that are high-dimensional, non-stationary, and composed of multiple objects with rich interactions since most world models learn holistic representations of all environmental components. By contrast, humans perceive the environment by decomposing it into discrete objects, facilitating efficient decision-making. Motivated by this insight, we propose Slot Transformer Imagination with CAusality-aware reinforcement learning (STICA), a unified framework in which object-centric Transformers serve as the world model and causality-aware policy and value networks. STICA represents each observation as a set of object-centric tokens, together with tokens for the agent action and the resulting reward, enabling the world model to predict token-level dynamics and interactions. The policy and value networks then estimate token-level cause--effect relations and use them in the attention layers, yielding causality-guided decision-making. Experiments on object-rich benchmarks demonstrate that STICA consistently outperforms state-of-the-art agents in both sample efficiency and final performance. Yosuke Nishimoto, Takashi Matsubara 0001 |
AAAI | 2 |
| 2025 | Number Theoretic Accelerated Learning of Physics-Informed Neural NetworksabstractPhysics-informed neural networks solve partial differential equations by training neural networks. Since this method approximates infinite-dimensional PDE solutions with finite collocation points, minimizing discretization errors by selecting suitable points is essential for accelerating the learning process. Inspired by number theoretic methods for numerical analysis, we introduce good lattice training and periodization tricks, which ensure the conditions required by the theory. Our experiments demonstrate that GLT requires 2-7 times fewer collocation points, resulting in lower computational cost, while achieving competitive performance compared to typical sampling methods. Takashi Matsubara 0001, Takaharu Yaguchi |
AAAI | 1 |
| 2025 | Poisson-Dirac Neural Networks for Modeling Coupled Dynamical Systems across DomainsabstractDeep learning has achieved great success in modeling dynamical systems, providing data-driven simulators to predict complex phenomena, even without known governing equations. However, existing models have two major limitations: their narrow focus on mechanical systems and their tendency to treat systems as monolithic. These limitations reduce their applicability to dynamical systems in other domains, such as electrical and hydraulic systems, and to coupled systems. To address these limitations, we propose Poisson-Dirac Neural Networks (PoDiNNs), a novel framework based on the Dirac structure that unifies the port-Hamiltonian and Poisson formulations from geometric mechanics. This framework enables a unified representation of various dynamical systems across multiple domains as well as their interactions and degeneracies arising from couplings. Our experiments demonstrate that PoDiNNs offer improved accuracy and interpretability in modeling unknown coupled dynamical systems from data. Razmik Arman Khosrovian, Takaharu Yaguchi, Hiroaki Yoshimura, Takashi Matsubara 0001 |
ICLR | 4 |
| 2025 | Deep Energy-Based Discrete-Time Physical Model for Reproducing Energetic BehaviorabstractModeling and simulating physical phenomena, especially those governed by partial differential equations (PDEs), pose significant challenges in computational physics and scientific machine learning. While neural network approaches have made strides in learning continuous-time dynamics, they have struggled with discrete-time scenarios and often fail to adhere to fundamental laws of physics, such as the conservation of energy and mass. This study addresses this gap by introducing a novel deep energy-based discrete-time model. In the real world, energy-based modeling theories like Hamiltonian mechanics and the Landau theory are pivotal, as they support various laws of physics. By integrating differential geometric structures into neural networks as coefficient matrices, our model successfully simulates the conservation and dissipation laws of energy and mass. Furthermore, we propose an automatic discrete differentiation algorithm, which enables neural networks to utilize the discrete gradient method, ensuring adherence to these laws in discrete-time settings. This capability also facilitates the identification of such laws directly from data by learning matrices that represent geometric structures. These advantages are verified using simulation results of physical phenomena, namely the 1- and 2-D Korteweg-de Vries (KdV) equation and the Cahn-Hilliard equation. Takashi Matsubara 0001, Takehiro Aoshima, Ai Ishikawa, Takaharu Yaguchi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Predicated Diffusion: Predicate Logic-Based Attention Guidance for Text-to-Image Diffusion ModelsabstractDiffusion models have achieved remarkable success in generating high-quality, diverse, and creative images. However, in text-based image generation, they often struggle to accurately capture the intended meaning of the text. For instance, a specified object might not be generated, or an adjective might incorrectly alter unintended objects. Moreover, we found that relationships indicating possession between objects are frequently overlooked. Despite the diversity of users' intentions in text, existing methods often focus on only some aspects of these intentions. In this paper, we propose Predicated Diffusion, a unified framework designed to more effectively express users' intentions. It represents the intended meaning as propositions using predicate logic and treats the pixels in attention maps as fuzzy predicates. This approach provides a differentiable loss function that offers guidance for the image generation process to better fulfill the propositions. Comparative evaluations with existing methods demonstrated that Predicated Diffusion excels in generating images faithful to various text prompts, while maintaining high image quality, as validated by human evaluators and pretrained image-text models. Kota Sueyoshi, Takashi Matsubara 0001 |
CVPR | 2 |
| 2024 | The Symplectic Adjoint Method: Memory-Efficient Backpropagation of Neural-Network-Based Differential EquationsabstractThe combination of neural networks and numerical integration can provide highly accurate models of continuous-time dynamical systems and probabilistic distributions. However, if a neural network is used n times during numerical integration, the whole computation graph can be considered as a network n times deeper than the original. The backpropagation algorithm consumes memory in proportion to the number of uses times of the network size, causing practical difficulties. This is true even if a checkpointing scheme divides the computation graph into subgraphs. Alternatively, the adjoint method obtains a gradient by a numerical integration backward in time; although this method consumes memory only for single-network use, the computational cost of suppressing numerical errors is high. The symplectic adjoint method proposed in this study, an adjoint method solved by a symplectic integrator, obtains the exact gradient (up to rounding error) with memory proportional to the number of uses plus the network size. The theoretical analysis shows that it consumes much less memory than the naive backpropagation algorithm and checkpointing schemes. The experiments verify the theory, and they also demonstrate that the symplectic adjoint method is faster than the adjoint method and is more robust to rounding errors. Takashi Matsubara 0001, Yuto Miyatake, Takaharu Yaguchi |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Deep Curvilinear Editing: Commutative and Nonlinear Image Manipulation for Pretrained Deep Generative ModelabstractSemantic editing of images is the fundamental goal of computer vision. Although deep learning methods, such as generative adversarial networks (GANs), are capable of producing high-quality images, they often do not have an inherent way of editing generated images semantically. Recent studies have investigated a way of manipulating the latent variable to determine the images to be generated. However, methods that assume linear semantic arithmetic have certain limitations in terms of the quality of image editing, whereas methods that discover nonlinear semantic pathways provide non-commutative editing, which is inconsistent when applied in different orders. This study proposes a novel method called deep curvilinear editing (DeCurvEd) to determine semantic commuting vector fields on the latent space. We theoretically demonstrate that owing to commutativity, the editing of multiple attributes depends only on the quantities and not on the order. Furthermore, we experimentally demonstrate that compared to previous methods, the nonlinear and commutative nature of DeCurvEdfacilitates the disentanglement of image attributes and provides higher-quality editing. Takehiro Aoshima, Takashi Matsubara 0001 |
CVPR | 2 |
| 2023 | FINDE: Neural Differential Equations for Finding and Preserving Invariant Quantities
Takashi Matsubara 0001, Takaharu Yaguchi |
ICLR | 1 |
| 2023 | A Two-View EEG Representation for Brain Cognition by Composite Temporal-Spatial Contrastive LearningabstractElectroencephalography (EEG) is a major tool for studying neurophysiological processes. Investigating reliable representations from highly noisy measurements is a pending challenge, however, the medically treasured and insufficient labeled data have driven this process away from a supervised learning manner. Recent works have turned their attention to self-supervised learning (SSL), putting the contrastive strategy on capturing the spatio-temporal characteristics of the neuronal events of interest. We argue that the temporal-spatial view is not the best choice for the SSL contrastive objective because there is a missing piece of the EEG representation that is usually ignored: dynamic fluctuations in brain neurons and the statistical learning of analog/artificial neural networks cannot handle the dynamic characteristics well. This paper proposes a novel two-view contrastive learning framework to refine EEG features from local-global and past-future views. An array of spiking neural networks is embedded to project spatio-temporal features onto the spike sequences to represent the dynamic fluctuation information of EEG. Experimenting with sleep stage classification and prediction of lethal epileptic seizures, we verify the proposal competes favorably against the state-of-the-art methods and offers high-quality features, that is, supervised learning on top of them observes a significant improvement in classification after only one training iteration. Zheng Chen 0012, Lingwei Zhu, Haohui Jia, Takashi Matsubara 0001 |
SDM | 4 |
| 2022 | KAM Theory Meets Statistical Learning Theory: Hamiltonian Neural Networks with Non-zero Training LossabstractMany physical phenomena are described by Hamiltonian mechanics using an energy function (Hamiltonian). Recently, the Hamiltonian neural network, which approximates the Hamiltonian by a neural network, and its extensions have attracted much attention. This is a very powerful method, but theoretical studies are limited. In this study, by combining the statistical learning theory and KAM theory, we provide a theoretical analysis of the behavior of Hamiltonian neural networks when the learning error is not completely zero. A Hamiltonian neural network with non-zero errors can be considered as a perturbation from the true dynamics, and the perturbation theory of the Hamilton equation is widely known as KAM theory. To apply KAM theory, we provide a generalization error bound for Hamiltonian neural networks by deriving an estimate of the covering number of the gradient of the multi-layer perceptron, which is the key ingredient of the model. This error bound gives a sup-norm bound on the Hamiltonian that is required in the application of KAM theory. Takashi Matsubara 0001, Takaharu Yaguchi |
AAAI | 2 |
| 2022 | Automated Cancer Subtyping via Vector Quantization Mutual Information Maximization
Zheng Chen 0012, Lingwei Zhu, Ziwei Yang 0002, Takashi Matsubara 0001 |
ECML/PKDD (1) | 4 |
| 2022 | Topology-Aware Flow-Based Point Cloud GenerationabstractPoint clouds have attracted attention as a representation of an object’s surface. Deep generative models have typically used a continuous map from a dense set in a latent space to express their variations. However, a continuous map cannot adequately express the varying numbers of holes. That is, previous approaches disregarded the topological structure of point clouds. Furthermore, a point cloud comprises several subparts, making it difficult to express it using a continuous map. This paper proposes ChartPointFlow, a flow-based deep generative model that forms a map conditioned on a label. Similar to a manifold chart, a map conditioned on a label is assigned to a continuous subset of a point cloud. Thus, ChartPointFlow is able to maintain the topological structure with clear boundaries and holes, whereas previous approaches generated blurry point clouds with fuzzy holes. The experimental results show that ChartPointFlow achieves state-of-the-art performance in various tasks, including generation, reconstruction, upsampling, and segmentation. Takumi Kimura, Takashi Matsubara 0001, Kuniaki Uehara |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Generative Model Using Unregularized Score for Anomaly Detection With Heterogeneous ComplexityabstractAccurate and automated detection of anomalous samples in an image dataset can be accomplished with a probabilistic model. Such images have heterogeneous complexity, however, and a probabilistic model tends to overlook simply shaped objects with small anomalies. The reason is that a probabilistic model assigns undesirable lower likelihoods to complexly shaped objects, which are nevertheless consistent with the current set standards. This difficulty is critical, especially for a defect detection task, where the anomaly can be a small scratch or grime. To overcome this difficulty, we propose an unregularized score for deep generative models (DGMs). We found that the regularization terms of the DGMs considerably influence the anomaly score depending on the complexity of the samples. By removing these terms, we obtain an unregularized score, which we evaluated on toy datasets, two in-house manufacturing datasets, and on open manufacturing and medical datasets. The empirical results demonstrate that the unregularized score is robust to the apparent complexity of given samples and detects anomalies selectively. Takashi Matsubara 0001, Kazuki Sato, Kenta Hama, Ryosuke Tachibana, Kuniaki Uehara |
IEEE Trans. Cybern. | 1 |
| 2021 | ChartPointFlow for Topology-Aware 3D Point Cloud GenerationabstractA point cloud serves as a representation of the surface of a three-dimensional (3D) shape. Deep generative models have been adapted to model their variations typically using a map from a ball-like set of latent variables. However, previous approaches did not pay much attention to the topological structure of a point cloud, despite that a continuous map cannot express the varying numbers of holes and intersections. Moreover, a point cloud is often composed of multiple subparts, and it is also difficult to express. In this study, we propose ChartPointFlow, a flow-based generative model with multiple latent labels for 3D point clouds. Each label is assigned to points in an unsupervised manner. Then, a map conditioned on a label is assigned to a continuous subset of a point cloud, similar to a chart of a manifold. This enables our proposed model to preserve the topological structure with clear boundaries, whereas previous approaches tend to generate blurry point clouds and fail to generate holes. The experimental results demonstrate that ChartPointFlow achieves state-of-the-art performance in terms of generation and reconstruction compared with other point cloud generators. Moreover, ChartPointFlow divides an object into semantic subparts using charts, and it demonstrates superior performance in case of unsupervised segmentation. Takumi Kimura, Takashi Matsubara 0001, Kuniaki Uehara |
ACM Multimedia | 2 |
| 2021 | Neural Symplectic Form: Learning Hamiltonian Equations on General Coordinate SystemsabstractIn recent years, substantial research on the methods for learning Hamiltonian equations has been conducted. Although these approaches are very promising, the commonly used representation of the Hamilton equation uses the generalized momenta, which are generally unknown. Therefore, the training data must be represented in this unknown coordinate system, and this causes difficulty in applying the model to real data. Meanwhile, Hamiltonian equations also have a coordinate-free expression that is expressed by using the symplectic 2-form. In this study, we propose a model that learns the symplectic form from data using neural networks, thereby providing a method for learning Hamiltonian equations from data represented in general coordinate systems, which are not limited to the generalized coordinates and the generalized momenta. Consequently, the proposed method is capable not only of modeling target equations of both Hamiltonian and Lagrangian formalisms but also of extracting unknown Hamiltonian structures hidden in the data. For example, many polynomial ordinary differential equations such as the Lotka-Volterra equation are known to admit non-trivial Hamiltonian structures, and our numerical experiments show that such structures can be certainly learned from data. Technically, each symplectic 2-form is associated with a skew-symmetric matrix, but not all skew-symmetric matrices define the symplectic 2-form. In the proposed method, using the fact that symplectic 2-forms are derived as the exterior derivative of certain differential 1-forms, we model the differential 1-form by neural networks, thereby improving the efficiency of learning. Takashi Matsubara 0001, Takaharu Yaguchi |
NeurIPS | 2 |
| 2021 | Symplectic Adjoint Method for Exact Gradient of Neural ODE with Minimal MemoryabstractA neural network model of a differential equation, namely neural ODE, has enabled the learning of continuous-time dynamical systems and probabilistic distributions with high accuracy. The neural ODE uses the same network repeatedly during a numerical integration. The memory consumption of the backpropagation algorithm is proportional to the number of uses times the network size. This is true even if a checkpointing scheme divides the computation graph into sub-graphs. Otherwise, the adjoint method obtains a gradient by a numerical integration backward in time. Although this method consumes memory only for a single network use, it requires high computational cost to suppress numerical errors. This study proposes the symplectic adjoint method, which is an adjoint method solved by a symplectic integrator. The symplectic adjoint method obtains the exact gradient (up to rounding error) with memory proportional to the number of uses plus the network size. The experimental results demonstrate that the symplectic adjoint method consumes much less memory than the naive backpropagation algorithm and checkpointing schemes, performs faster than the adjoint method, and is more robust to rounding errors. Takashi Matsubara 0001, Yuto Miyatake, Takaharu Yaguchi |
NeurIPS | 1 |
| 2021 | Exploring Uncertainty Measures for Image-caption Embedding-and-retrieval TaskabstractWith the significant development of black-box machine learning algorithms, particularly deep neural networks, the practical demand for reliability assessment is rapidly increasing. On the basis of the concept that “Bayesian deep learning knows what it does not know,” the uncertainty of deep neural network outputs has been investigated as a reliability measure for classification and regression tasks. By considering an embedding task as a regression task, several existing studies have quantified the uncertainty of embedded features and improved the retrieval performance of cutting-edge models by model averaging. However, in image-caption embedding-and-retrieval tasks, well-known samples are not always easy to retrieve. This study shows that the existing method has poor performance in reliability assessment and investigates another aspect of image-caption embedding-and-retrieval tasks. We propose posterior uncertainty by considering the retrieval task as a classification task, which can accurately assess the reliability of retrieval results. The consistent performance of the two uncertainty measures is observed with different datasets (MS-COCO and Flickr30k), different deep-learning architectures (dropout and batch normalization), and different similarity functions. To the best of our knowledge, this is the first study to perform a reliability assessment on image-caption embedding-and-retrieval tasks. Kenta Hama, Takashi Matsubara 0001, Kuniaki Uehara, Jianfei Cai 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Att-DARTS: Differentiable Neural Architecture Search for AttentionabstractNeural architecture search (NAS) is a promising method to automatically identify neural network architectures. Differentiable architecture search (DARTS) is method that significantly reduces search time and finds architectures that can achieve state-of-the-art performance. For computer vision tasks, DARTS searches convolutional neural networks (CNNs) via stacking convolution layers and pooling operations. Recent studies on neural architectures indicate that attention modules can improve the performances of CNNs by discarding information of no interest, while existing NAS methods have put little focus on it. In this study, we propose Att-DARTS, which searches attention modules as well as convolution and pooling operations simultaneously. In our experiments on CIFAR-10 and CIFAR-100 datasets, we demonstrate that Att-DARTS can find architectures that achieve lower classification error rates and require fewer parameters compared to those found by DARTS. Kohei Nakai, Takashi Matsubara 0001, Kuniaki Uehara |
IJCNN | 2 |
| 2020 | Deep Energy-based Modeling of Discrete-Time PhysicsabstractPhysical phenomena in the real world are often described by energy-based modeling theories, such as Hamiltonian mechanics or the Landau theory, which yield various physical laws. Recent developments in neural networks have enabled the mimicking of the energy conservation law by learning the underlying continuous-time differential equations. However, this may not be possible in discrete time, which is often the case in practical learning and computation. Moreover, other physical laws have been overlooked in the previous neural network models. In this study, we propose a deep energy-based physical model that admits a specific differential geometric structure. From this structure, the conservation or dissipation law of energy and the mass conservation law follow naturally. To ensure the energetic behavior in discrete time, we also propose an automatic discrete differentiation algorithm that enables neural networks to employ the discrete gradient method. Takashi Matsubara 0001, Ai Ishikawa, Takaharu Yaguchi |
NeurIPS | 1 |
| 2020 | Data Augmentation Using Random Image Cropping and Patching for Deep CNNsabstractDeep convolutional neural networks (CNNs) have achieved remarkable results in image processing tasks. However, their high expression ability risks overfitting. Consequently, data augmentation techniques have been proposed to prevent overfitting while enriching datasets. Recent CNN architectures with more parameters are rendering traditional data augmentation techniques insufficient. In this study, we propose a new data augmentation technique called random image cropping and patching (RICAP) which randomly crops four images and patches them to create a new training image. Moreover, RICAP mixes the class labels of the four images, resulting in an advantage of the soft labels. We evaluated RICAP with current state-of-the-art CNNs (e.g., the shake-shake regularization model) by comparison with competitive data augmentation techniques such as cutout and mixup. RICAP achieves a new state-of-the-art test error of 2.19% on CIFAR-10. We also confirmed that deep CNNs with RICAP achieve better results on classification tasks using CIFAR-100 and ImageNet, an image-caption retrieval task using Microsoft COCO, and other computer vision tasks. Ryo Takahashi 0003, Takashi Matsubara 0001, Kuniaki Uehara |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | A Human-Like Agent Based on a Hybrid of Reinforcement and Imitation LearningabstractReinforcement learning (RL) builds an effective agent that handles tasks in complex and uncertain environments by maximizing future reward. However, the efficiency is insufficient for practical use such as game AI and autonomous driving. An effective but selfish agent conflicts with other humans, and hence the demand of a human-like behavior arises. Imitation learning (IL) has been employed to train an agent to mimic the actions of expert behaviors provided as training data. However, IL tends to build an agent limited in performance by the expert skill, and even worse, the agent exhibits an inconsistent behavior since IL is not goal-oriented. In this paper, we propose a training scheme by mixing RL and IL for both discrete and continuous action space problems. The proposed scheme builds an agent that achieves a performance higher than an agent trained by only IL and exhibits a more human-like behavior than agents trained by RL or IL, validated by human sensitivity. Rousslan Fernand Julien Dossa, Xinyu Lian, Hirokazu Nomoto, Takashi Matsubara 0001, Kuniaki Uehara |
IJCNN | 4 |
| 2019 | Deep Generative State-Space Modeling of FMRI Images for Psychiatric Disorder DiagnosisabstractAn early and accurate diagnosis of psychiatric disorders is critical for patients' quality of life and deep understanding of the disorders. For this reason, many studies have proposed machine learning-based diagnostic procedures for functional magnetic resonance imaging (fMRI) data. Especially, these procedures often employed temporal models due to the time-varying nature of the brain activities and probabilistic generative models for understanding the underlying mechanism of the disorders. For leveraging the recent advantage of deep learning, we proposed a state-space model of fMRI images based on deep learning. The proposed deep state-space model is more flexible than conventional models and less likely to suffer from overfitting than a straightforward deep learning-based classifier. The proposed model estimates the subjects' conditions more accurately than existing diagnostic procedures. Also, the proposed model potentially identifies brain regions related to the disorders. Koki Kusano, Tetsuo Tashiro, Takashi Matsubara 0001, Kuniaki Uehara |
IJCNN | 3 |
| 2019 | Predictable Uncertainty-Aware Unsupervised Deep Anomaly SegmentationabstractImage-based anomaly segmentation is a fundamental topic for image analysis. For medical use, it supports treatments via refined diagnosis and growth rate evaluation of tumors and lesions. Especially, an unsupervised training is expected to generalize to unknown anomalies. Probabilistic models have been used for this purpose, whereby these models are trained to maximize the likelihood of known samples and detect anomalous samples by assigning low likelihoods. Recent studies have proposed a probabilistic model based on deep neural networks (DNNs) called AEs and they achieved significant performance thanks to their flexibility. However, AEs are sensitive to complex structure (e.g., ridges and grooves of a brain) rather than semantic anomalies (e.g., tumors and lesions). We decomposed the approximated log-likelihood into two terms; predictable uncertainty and normalized error. We found that the former represents the complexity of structure. Hence, we propose the normalized error as a novel uncertainty-sensitive score by removing the predictable uncertainty. We evaluated our score by experiments with head magnetic resonance imaging (MRI) datasets and demonstrate the robustness of the proposed normalized error to data complexity. Kazuki Sato, Kenta Hama, Takashi Matsubara 0001, Kuniaki Uehara |
IJCNN | 3 |
| 2019 | A Novel Weight-Shared Multi-Stage CNN for Scale RobustnessabstractConvolutional neural networks (CNNs) have demonstrated remarkable results in image classification for benchmark tasks and practical applications. The CNNs with deeper architectures have achieved even higher performance recently thanks to their robustness to the parallel shift of objects in images and their numerous parameters and the resulting high expression ability. However, CNNs have a limited robustness to other geometric transformations such as scaling and rotation. This limits the performance improvement of the deep CNNs, but there is no established solution. This paper focuses on scale transformation and proposes a network architecture called the weight-shared multi-stage network (WSMS-Net), which consists of multiple stages of CNNs. The proposed WSMS-Net is easily combined with existing deep CNNs such as residual network and densely connected convolutional network and enables them to acquire robustness to object scaling. Experimental results on the CIFAR-10, CIFAR-100, and ImageNet datasets demonstrate that existing deep CNNs combined with the proposed WSMS-Net achieve higher accuracies for image classification tasks with only a minor increase in the number of parameters and computation time. Ryo Takahashi 0003, Takashi Matsubara 0001, Kuniaki Uehara |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | RICAP: Random Image Cropping and Patching Data Augmentation for Deep CNNsabstractDeep convolutional neural networks (CNNs) have demonstrated remarkable results in image recognition owing to their rich expression ability and numerous parameters. However, an excessive expression ability compared to the variety of training images often has a risk of overfitting. Data augmentation techniques have been proposed to address this problem as they enrich datasets by flipping, cropping, resizing, and color-translating images. They enable deep CNNs to achieve an impressive performance. In this study, we propose a new data augmentation technique called \emph{random image cropping and patching} (\emph{RICAP}), which randomly crops four images and patches them to construct a new training image. Hence, RICAP randomly picks up subsets of original features among the four images and discard others, enriching the variety of training images. Also, RICAP mixes the class labels of the four images and enjoys a benefit similar to label smoothing. We evaluated RICAP with current state-of-the-art CNNs (e.g., shake-shake regularization model) and achieved a new state-of-the-art test error of \textcolor{red}{$2.23%$} on CIFAR-10 among competitive data augmentation techniques such as cutout and mixup. We also confirmed that deep CNNs with RICAP achieved better results on CIFAR-100 and ImageNet than those results obtained by other techniques. Ryo Takahashi 0003, Takashi Matsubara 0001, Kuniaki Uehara |
ACML | 2 |
| 2018 | Hypernetwork-based Implicit Posterior Estimation and Model Averaging of CNNabstractDeep neural networks have a rich ability to learn complex representations and achieved remarkable results in various tasks. However, they are prone to overfitting due to the limited number of training samples; regularizing the learning process of neural networks is critical. In this paper, we propose a novel regularization method, which estimates parameters of a large convolutional neural network as implicit probabilistic distributions generated by a hypernetwork. Also, we can perform model averaging to improve the network performance. Experimental results demonstrate our regularization method outperformed the commonly-used maximum a posterior (MAP) estimation. Kenya Ukai, Takashi Matsubara 0001, Kuniaki Uehara |
ACML | 2 |
| 2018 | Anomaly Machine Component Detection by Deep Generative Model with Unregularized ScoreabstractOne of the most common needs in manufacturing plants is rejecting products not coincident with the standards as anomalies. Accurate and automatic anomaly detection improves product reliability and reduces inspection cost. Probabilistic models have been employed to detect test samples with lower likelihoods as anomalies in unsupervised manner. Recently, a probabilistic model called deep generative model (DGM) has been proposed for end-to-end modeling of natural images and already achieved a certain success. However, anomaly detection of machine components with complicated structures is still challenging because they produce a wide variety of normal image patches with low likelihoods. For overcoming this difficulty, we propose unregularized score for the DGM. As its name implies, the unregularized score is the anomaly score of the DGM without the regularization terms. The unregularized score is robust to the inherent complexity of a sample and has a smaller risk of rejecting a sample appearing less frequently but being coincident with the standards. Takashi Matsubara 0001, Ryosuke Tachibana, Kuniaki Uehara |
IJCNN | 1 |
| 2018 | Structured Deep Generative Model of fMRI Signals for Mental Disorder Diagnosis
Takashi Matsubara 0001, Tetsuo Tashiro, Kuniaki Uehara |
MICCAI (3) | 1 |
| 2017 | Scale-Invariant Recognition by Weight-Shared CNNs in ParallelabstractDeep convolutional neural networks (CNNs) have become one of the most successful methods for image processing tasks in past few years. Recent studies on modern residual architectures, enabling CNNs to be much deeper, have achieved much better results thanks to their high expressive ability by numerous parameters. In general, CNNs are known to have the robustness to the small parallel shift of objects in images by their local receptive fields, weight parameters shared by each unit, and pooling layers sandwiching them. However, CNNs have a limited robustness to the other geometric transformations such as scaling and rotation, and this lack becomes an obstacle to performance improvement even now. This paper proposes a novel network architecture, the \emphweight-shared multi-stage network (WSMS-Net), and focuses on acquiring the scale invariance by constructing of multiple stages of CNNs. The WSMS-Net is easily combined with existing deep CNNs, enables existing deep CNNs to acquire a robustness to the scaling, and therefore, achieves higher classification accuracy on CIFAR-10, CIFAR-100 and ImageNet datasets. Ryo Takahashi 0003, Takashi Matsubara 0001, Kuniaki Uehara |
ACML | 2 |
| 2017 | Spike timing-dependent conduction delay learning model classifying spatio-temporal spike patternsabstractPrecise spike timing is considered to play a fundamental role in communication and signal processing in biological neural networks. Understanding such mechanism contributes to both deep understanding of biological system and development of engineering applications such as efficient computational architectures. However, the biological mechanism which adjusts and maintains the spike timing still remains unclear. Previous studies have proposed algorithms adjusting synaptic efficacy and axonal conduction delay so that the spike timings get close to the desired spike timings in supervised manner. Supervised learning always requires desired spike timings as teacher signal, and thus it should not be dominant in biological system, which is considered to adapt to environment without teacher. This study proposes a spike timing-dependent learning model adjusting synaptic efficacy and axonal conduction delay in both unsupervised and supervised manners. The proposed learning algorithm approximates Expectation-Maximization algorithm and can classify the input data coded into spatio-temporal spike patterns. Furthermore, the proposed learning algorithm agrees with various results of existing biological experiments such as spike timing-dependent plasticity, and therefore it could be a good candidate of a model of biological delay learning. Takashi Matsubara 0001 |
IJCNN | 1 |
| 2017 | Automatic manga colorization with color style by generative adversarial netsabstractMany comic books are now published as digital books, which easily provide color contents compared to the physical books. The motivation of automatic colorization of comic books now arises. Previous studies colorize sketches without other clues or with spatial color annotations. They are expected to reduce workloads of comic artists but still require spatial color annotations for desirable colorizations. This study introduces a color style information and combines it with conditional adversarially learned inference. The experimental results demonstrate that the objects are painted with colors depending on the color style information and that the color style information extracted from a color image supports to painting an object with a desirable color. Yuusuke Kataoka, Takashi Matsubara 0001, Kuniaki Uehara |
SNPD | 2 |
| 2017 | Developing game AI agent behaving like human by mixing reinforcement learning and supervised learningabstractArtificial intelligence (AI) agent created with Deep Q-Networks (DQN) can defeat human agents in video games. Despite its high performance, DQN often exhibits odd behaviors, which could be immersion-breaking against the purpose of creating game AI. Moreover, DQN is capable of reacting to the game environment much faster than humans, making itself invincible (thus not fun to play with) in certain types of games. On the other hand, supervised learning framework trains an AI agent using historical play data of human agents as training data. Supervised learning agent exhibits a more human-like behavior than reinforcement learning agents because of imitating training data. However, its performance is often no better than human agents. The ultimate purpose of AI agents is to entertain human players. A good performance and a humanlike behavior are important factors of the AI agents, and both of them should be achieved simultaneously. This study proposes frameworks combining reinforcement learning and supervised learning and we call then separated network model and shared network model. We evaluated their performances by the game scores and behaviors by Turing test. The experimental results demonstrate that the proposed frameworks develop an AI agent of better performance than human agent and natural behavior than reinforcement learning agents. Shohei Miyashita, Xinyu Lian, Takashi Matsubara 0001, Kuniaki Uehara |
SNPD | 4 |
| 2016 | Deep learning for stock prediction using numerical and textual informationabstractThis paper proposes a novel application of deep learning models, Paragraph Vector, and Long Short-Term Memory (LSTM), to financial time series forecasting. Investors make decisions according to various factors, including consumer price index, price-earnings ratio, and miscellaneous events reported in newspapers. In order to assist their decisions in a timely manner, many automatic ways to analyze those information have been proposed in the last decade. However, many of them used either numerical or textual information, but not both for a single company. In this paper, we propose an approach that converts newspaper articles into their distributed representations via Paragraph Vector and models the temporal effects of past events on opening prices about multiple companies with LSTM. The performance of the proposed approach is demonstrated on real-world data of fifty companies listed on Tokyo Stock Exchange. Ryo Akita, Akira Yoshihara, Takashi Matsubara 0001, Kuniaki Uehara |
ICIS | 3 |
| 2016 | Image generation using generative adversarial networks and attention mechanismabstractFor image generation, deep neural networks are trained to extract high-level features on natural images and to reconstruct the images from the features. However it is difficult to learn to generate images containing enormous contents. To overcome this difficulty, a network with an attention mechanism has been proposed. It is trained to attend to parts of the image and to generate images step by step. This enables the network to deal with the details of a part of the image and the rough structure of the entire image. The attention mechanism is implemented by recurrent neural networks. Additionally, the Generative Adversarial Networks (GANs) approach has been proposed to generate more realistic images. In this study, we present image generation where leverages effectiveness of attention mechanism and the GANs approach. We show our method enables the iterative construction of images and more realistic image generation than standard GANs and the attention mechanism of DRAW. Yuusuke Kataoka, Takashi Matsubara 0001, Kuniaki Uehara |
ICIS | 2 |
| 2016 | Semi-Supervised learning using adversarial networksabstractSemi-supervised learning is a topic of practical importance because of the difficulty of obtaining numerous labeled data. In this paper, we apply an extension of adversarial autoencoder to semi-supervised learning tasks. In attempt to separate style and content, we divide the latent representation of the autoencoder into two parts. We regularize the autoencoder by imposing a prior distribution on both parts to make them independent. As a result, one of the latent representations is associated with content, which is useful to classify the images. We demonstrate that our method disentangles style and content of the input images and achieves less test error rate than vanilla autoencoder on MNIST semi-supervised classification tasks. Ryosuke Tachibana, Takashi Matsubara 0001, Kuniaki Uehara |
ICIS | 2 |
| 2016 | A novel homeostatic plasticity model realized by random fluctuations in excitatory synapsesabstractHomeostatic plasticity in mammalian central nervous system is considered to maintain activity in neuronal circuits within a functional range. In the absence of homeostatic plasticity neuronal activity is prone to be destabilized because correlation-based synaptic modification, Hebbian plasticity, induces positive feedback change. Several studies on homeostatic plasticity assumed the existence of a process for monitoring neuronal activity and adjusting synaptic efficacy on a time scale of hours, but its biological mechanism still remains unclear. Excitatory synaptic efficacy is associated with the size of a post-synaptic element, dendritic spine, and the size of the dendritic spine fluctuates even after neuronal activity is silenced. These fluctuations could be a non-Hebbian form of synaptic plasticity that serves such a homeostatic function. This study proposed and analyzed a synaptic plasticity model incorporating random fluctuations and Hebbian plasticity at excitatory synapses, and found that it prevents excessive changes in neuronal activity by adjusting synaptic efficacy. Random fluctuations do not monitor neuronal activity, but their relative influence depends on neuronal activity. The proposed synaptic plasticity model acts as a form of homeostatic plasticity, regardless of neuronal activity monitoring. Thus, random fluctuations play an important role in homeostatic plasticity and contribute to development and functions of neural networks. Takashi Matsubara 0001, Kuniaki Uehara |
IJCNN | 1 |
| 2016 | An Asynchronous Recurrent Network of Cellular Automaton-Based Neurons and Its Reproduction of Spiking Neural Network ActivitiesabstractModeling and implementation approaches for the reproduction of input-output relationships in biological nervous tissues contribute to the development of engineering and clinical applications. However, because of high nonlinearity, the traditional modeling and implementation approaches encounter difficulties in terms of generalization ability (i.e., performance when reproducing an unknown data set) and computational resources (i.e., computation time and circuit elements). To overcome these difficulties, asynchronous cellular automaton-based neuron (ACAN) models, which are described as special kinds of cellular automata that can be implemented as small asynchronous sequential logic circuits have been proposed. This paper presents a novel type of such ACAN and a theoretical analysis of its excitability. This paper also presents a novel network of such neurons, which can mimic input-output relationships of biological and nonlinear ordinary differential equation model neural networks. Numerical analyses confirm that the presented network has a higher generalization ability than other major modeling and implementation approaches. In addition, Field-Programmable Gate Array-implementations confirm that the presented network requires lower computational resources. Takashi Matsubara 0001, Hiroyuki Torikai |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | A nonlinear model of fMRI BOLD signal including the trend componentabstractThis paper presents a nonlinear model of the human brain activity response to visual stimuli according to Blood-Oxygen-Level-Dependent (BOLD) signals scanned by functional Magnetic Resonance Imaging (fMRI). A BOLD signal usually contains a low frequency signal component (trend), which is often ignored by the existing models or removed by approximation methods. However, such detrending could also destroy the dynamics of the BOLD signal and miss an important response. This paper shows a model that, in the absence of detrending, can predict the BOLD signal with smaller errors than existing models. For detrending, the presented model has also a lower Schwarz information criterion than existing models, which implies that the presented model will be less likely to overfit the experimental data. Takashi Matsubara 0001, Hiroyuki Torikai, Tetsuya Shimokawa, Kenji Leibnitz, Ferdinand Peper |
IJCNN | 1 |
| 2013 | A novel reservoir network of asynchronous cellular automaton based neurons for MIMO neural system reproductionabstractModeling and implementation of input-output relationships in biological nervous tissues contribute to the development of engineering and clinical applications. However, because of the high nonlinearity, the traditional modeling and implementation approaches have difficulties in terms of generalization ability (i.e., performance on reproducing an unknown data) and computational resources. To overcome these difficulties, asynchronous cellular automaton based neuron models has been presented, which are neuron models described as special kinds of cellular automata and can be implemented as small asynchronous sequential logic circuits. This paper presents a novel network of such models, which can mimic input-output relationships of biological and nonlinear ODE model neural networks. Computer simulations confirm that the presented network has a higher generalization ability than another modeling and implementation approach. In addition, brief comparisons of the computational resources for execution and learning shows that the presented network requires less computational resources. Takashi Matsubara 0001, Hiroyuki Torikai |
IJCNN | 1 |
| 2013 | Asynchronous Cellular Automaton-Based Neuron: Theoretical Analysis and On-FPGA LearningabstractA generalized asynchronous cellular automaton-based neuron model is a special kind of cellular automaton that is designed to mimic the nonlinear dynamics of neurons. The model can be implemented as an asynchronous sequential logic circuit and its control parameter is the pattern of wires among the circuit elements that is adjustable after implementation in a field-programmable gate array (FPGA) device. In this paper, a novel theoretical analysis method for the model is presented. Using this method, stabilities of neuron-like orbits and occurrence mechanisms of neuron-like bifurcations of the model are clarified theoretically. Also, a novel learning algorithm for the model is presented. An equivalent experiment shows that an FPGA-implemented learning algorithm enables an FPGA-implemented model to automatically reproduce typical nonlinear responses and occurrence mechanisms observed in biological and model neurons. Takashi Matsubara 0001, Hiroyuki Torikai |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | A Novel Bifurcation-Based Synthesis of Asynchronous Cellular Automaton Based Neuron
Takashi Matsubara 0001, Hiroyuki Torikai |
ICANN (1) | 1 |
| 2012 | A generalized asynchronous digital spiking neuron: Theoretical analysis and compartmental modelabstractThe most generalized version of asynchronous sequential logic circuit based neuron models is introduced, where the dynamics of the model is modeled by an asynchronous cellular automaton. In this paper, a new theoretical analysis method is presented, and stabilities of neuron-like orbits and occurrence mechanisms of relational neuron-like bifurcations are clarified theoretically. A synapse unit and a simple compartmental model are also presented, and their functions are confirmed numerically. Takashi Matsubara 0001, Hiroyuki Torikai |
IJCNN | 1 |
| 2011 | Dynamic Response Behaviors of a Generalized Asynchronous Digital Spiking Neuron Model
Takashi Matsubara 0001, Hiroyuki Torikai |
ICONIP (3) | 1 |
| 2011 | A novel asynchronous digital spiking neuron model and its various neuron-like bifurcations and responsesabstractA novel spiking neuron model whose nonlinear dynamics is described by an asynchronous cellular automaton is presented. The model can be implemented by a simple digital sequential logic circuit but can exhibit various neuron-like bifurcations and responses. Using the Poincaré mapping technique, it is clarified that the model can reproduce major bifurcation mechanisms of excitabilities and spikings of biological and model neurons. It is also clarified that the model can reproduce major excitatory responses of the neurons. Takashi Matsubara 0001, Hiroyuki Torikai |
IJCNN | 1 |