Akisato Kimura

dblp:55/1636 · DBLP profile ↗
← Back
79ranked-venue papers
14as first author
34since 2021 · last 2026
0009-0007-3042-6810ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 34 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 16 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Theory of computation · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Accelerating Graph Construction for MIPS without Search Accuracy Loss
Yasuhiro Fujiwara, Ángel López García-Arias, Yu Mitsuzumi, Yasutoshi Ida, Atsutoshi Kumagai, Masahiro Nakano, Makoto Nakatsuji, Akisato Kimura
EDBT8
2026 Fast Vector Quantization Algorithm for ScaNN
abstract
Maximum Inner Product Search (MIPS) is a popular task to find the vector with the highest inner product for a given query. ScaNN is a score-aware quantization approach for MIPS that effectively transforms vectors with higher inner products into short sequences of codewords within codebooks. When quantizing vectors, it iteratively updates codebooks by assigning vectors to codewords and computing inverse matrices obtained from the assigned vectors. ScaNN, however, incurs a high computation cost when quantizing large-scale data. This is because (1) it computes quantization losses for all pairs of vectors and codewords, and (2) the size of the inverse matrices is quadratic in the number of dimensions. Our proposal, F-ScaNN, increases the efficiency of ScaNN through two techniques: (1) it computes the upper and lower bounds of the losses to assign vectors, and (2) it employs the conjugate gradient method to avoid computing the inverse matrix. Theoretically, we can obtain the same quantization results as ScaNN. Furthermore, we can improve search accuracy by using scaled codewords. Experiments show that our approach is significantly faster than previous approaches.
Yasuhiro Fujiwara, Ángel López García-Arias, Yasutoshi Ida, Atsutoshi Kumagai, Masahiro Nakano, Makoto Nakatsuji, Akisato Kimura
KDD (1)7
2026 Model-free Domain Adaptation for Concealed Multimodal Large-Language Models
abstract
Multimodal large-language models (MLLMs) exhibit remarkable capability for various vision tasks but still struggle with the domain-shift problem, in which their performance degrades for data from unfamiliar domains. Since the latest MLLMs often conceal their model resources (i.e., data, parameters, and outputs) from training purposes, current domain adaptation methods cannot satisfactorily address this problem due to their dependence on those resources. To this end, we introduce a novel domain adaptation setup, "model-free domain adaptation (MFDA)" of MLLMs, to investigate whether we can address domain adaptation problems without using any resources of the concealed models. As a proof of concept for MFDA, we built a method named model-transferable domain-adaptable visual prompting (MTDA-VP). In the training, this method executes cross-model visual prompting on surrogate models with a domain adaptation objective so that the visual prompts simultaneously acquire model transferability and domain adaptability. In the testing, we can adapt the concealed MLLMs to the target domain by just inputting the test images with the trained prompt into the models. Besides, we developed two techniques, cross-model pseudo labeling (CMPL) and cross-model gradient alignment (CMGA), to further enhance model transferability and domain adaptability of the visual prompts. We empirically confirmed that MFDA-VP improved the performance of several MLLMs with large margins on two datasets.
Yu Mitsuzumi, Akisato Kimura, Hisashi Kashima
WACV2
2026 IPCD: Intrinsic Point-Cloud Decomposition
abstract
Point clouds are widely used in various fields, including augmented reality (AR) and robotics, where relighting and texture editing are crucial for realistic visualization. Achieving these tasks requires accurately separating albedo from shade. However, performing this separation on point clouds presents two key challenges: (1) the non-grid structure of point clouds makes conventional image-based decomposition models ineffective, and (2) point-cloud models designed for other tasks do not explicitly consider global-light direction, resulting in inaccurate shade. In this paper, we introduce Intrinsic Point-Cloud Decomposition (IPCD), which extends image decomposition to the direct decomposition of colored point clouds into albedo and shade. To overcome challenge (1), we propose IPCD-Net that extends image-based model with point-wise feature aggregation for non-grid data processing. For challenge (2), we introduce Projection-based Luminance Distribution (PLD) with a hierarchical feature refinement, capturing global-light ques via multi-view projection. For comprehensive evaluation, we create a synthetic outdoor-scene dataset. Experimental results demonstrate that IPCD-Net reduces cast shadows in albedo and enhances color accuracy in shade. Furthermore, we showcase its applications in texture editing, relighting, and point-cloud registration under varying illumination. Finally, we verify the real-world applicability of IPCD-Net.
Shogo Sato, Takuhiro Kaneko, Shoichiro Takeda, Tomoyasu Shimada, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida, Akisato Kimura
WACV8
2025 Multi-Task Learning for Ultrasonic Echo-based Depth Estimation with Audible Frequency Recovery
abstract
While depth maps of indoor scenes are often essential for a variety of applications, measuring depth maps usually requires dedicated depth sensors, which are not always available. Echo-based depth estimation has been explored as a promising alternative solution. However, most existing methods assume the use of audible echoes, with the major problem that prevents their use in quiet spaces or in situations where the generation of audible sound is prohibited. In this paper, we explore depth estimation based on ultrasonic echoes, which has scarcely been explored so far. The key idea of our method is to learn a depth estimation model that can exploit useful, but missing information in the audible frequency band. To this end, we perform multi-task learning that requires estimation of depth maps from ultrasound echoes while simultaneously restoring the audible frequency range. Furthermore, to evaluate the performance with real echo data, we develop a data collection device and collect a real sound dataset. Experimental results on this real echo dataset and public simulation benchmark dataset demonstrate that our method outperforms existing methods. Our real echo dataset and the code will be publicly available if the paper is accepted.
Junpei Honma, Akisato Kimura, Go Irie
ICASSP2
2025 Flexible Source-free Domain Generalization via Domain Prompt-Discriminator Collaborative Learning
abstract
Source-free domain generalization (SFDG) is an emerging paradigm of domain generalization problems that aims to train a model generalizable to unseen target domains without any actual source data. Existing methods have pragmatically tackled this challenging problem by using large-scale vision-language pre-trained models (VLMs). However, they cannot fully exploit the potential of VLMs. Although the favorable configurations should inherently differ for each input domain, those methods employ a domain-shared configuration (e.g., classifier, representation) for data inputs from any domain, which undermines the flexibility for various domains. In this paper, we propose a novel SFDG method called Domain Prompt-Discriminator Collaborative Learning. Our method endows the model with domain flexibility by jointly training the following two modules: (1) various domain-specific prompts to enhance generalizability to unseen domains and (2) a domain discriminator to choose favorable domains for the input image. Notably, this training can be carried out in the form of simple domain classification learning. Our method design also opens the door to a simple yet effective test-time adaptation technique that further boosts recognition accuracy. We empirically demonstrate that our method consistently outperforms state-of-the-art SFDG methods on four domain generalization benchmarks.
Yu Mitsuzumi, Akisato Kimura, Hisashi Kashima
IJCNN2
2025 Unsupervised Single-Image Intrinsic Image Decomposition with LiDAR Intensity Enhanced Training
abstract
Unsupervised intrinsic image decomposition (IID) is the task of separating a natural image into albedo and shade without ground truth during training. Although a recent model employing light detection and ranging (LiDAR) intensity demonstrated impressive performance, the necessity of LiDAR intensity during inference restricts its practicality. To expand the usage scenario while maintaining the IID quality achieved by using both an image and its corresponding LiDAR intensity, we propose a novel approach that utilizes an image without LiDAR intensity during inference while utilizing both an image and LiDAR intensity during training. Specifically, our proposed model processes an image and LiDAR intensity individually using distinct encoder paths during training, but utilizes only an imageencoder path during inference. Additionally, we introduce an albedo-alignment loss aligning the gray-scale albedo from an image to that from its corresponding LiDAR intensity. LiDAR intensity is not affected by illumination effects including cast shadows, thus albedo-alignment loss transfers the illumination-invariant property of LiDAR intensity to the image-encoder path. Furthermore, we also propose image-LiDAR conversion (ILC) paths that mutually translates the style of an image and LiDAR intensity. IID models translate an image into albedo and shade styles while keeping the image contents, thus it is important to separate the image into contents and style. Trained with pairs of an image and its corresponding LiDAR intensity which share contents but differ in style, the mutual translation in ILC paths improve the accuracy of the separation. Consequently, our model achieves comparable IID quality to the existing model with LiDAR intensity, while utilizing only an image without LiDAR intensity during inference.
Shogo Sato, Takuhiro Kaneko, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida, Akisato Kimura
WACV6
2024 Acoustic-based 3D Human Pose Estimation Robust to Human Position
Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa
BMVC4
2024 Understanding and Improving Source-Free Domain Adaptation from a Theoretical Perspective
abstract
Source-free Domain Adaptation (SFDA) is an emerging and challenging research area that addresses the problem of unsupervised domain adaptation (UDA) without source data. Though numerous successful methods have been proposed for SFDA, a theoretical understanding of why these methods work well is still absent. In this paper, we shed light on the theoretical perspective of existing SFDA methods. Specifically, we find that SFDA loss functions comprising discriminability and diversity losses work in the same way as the training objective in the theory of self-training based on the expansion assumption, which shows the existence of the target error bound. This finding brings two novel insights that enable us to build an improved SFDA method comprising 1) Model Training with Auto-Adjusting Diversity Constraint and 2) Augmentation Training with Teacher-Student Framework, yielding a better recognition performance. Extensive experiments on three benchmark datasets demonstrate the validity of the theoretical analysis and our method.
Yu Mitsuzumi, Akisato Kimura, Hisashi Kashima
CVPR2
2024 Estimating Indoor Scene Depth Maps From Ultrasonic Echoes
abstract
Measuring 3D geometric structures of indoor scenes requires dedicated depth sensors, which are not always available. Echo-based depth estimation has recently been studied as a promising alternative solution. All previous studies have assumed the use of echoes in the audible range. However, one major problem is that audible echoes cannot be used in quiet spaces or other situations where producing audible sounds is prohibited. In this paper, we consider echo-based depth estimation using inaudible ultrasonic echoes. While ultrasonic waves provide high measurement accuracy in theory, the actual depth estimation accuracy when ultrasonic echoes are used has remained unclear, due to its disadvantage of being sensitive to noise and susceptible to attenuation. We first investigate the depth estimation accuracy when the frequency of the sound source is restricted to the high-frequency band, and found that the accuracy decreased when the frequency was limited to ultrasonic ranges. Based on this observation, we propose a novel deep learning method to improve the accuracy of ultrasonic echo-based depth estimation by using audible echoes as auxiliary data only during training. Experimental results with a public dataset demonstrate that our method improves the estimation accuracy.
Junpei Honma, Akisato Kimura, Go Irie
ICIP2
2024 Cross-Action Cross-Subject Skeleton Action Recognition Via Simultaneous Action-Subject Learning With Two-Step Feature Removal
abstract
In this paper, we tackle a novel skeleton-based action recognition problem named Cross-Action Cross-Subject (CACS) Skeleton Action Recognition, where we can access the data of only a part of the target action classes for each training subject. Existing skeleton-based action recognition methods suffer from solving this problem because there are scarce clues to resolve the cross-entanglement of action and subject information, and the trained model will confuse those two features. To solve this challenging problem, we propose a method that consists of simultaneous action-subject learning with feature removal. In our method, 1) we use two data augmentation techniques, Bone Randomization and Phase Randomization, to roughly remove unnecessary features for respective recognitions, and then, 2) we introduce a debiased learning approach to remove the confusing features by minimizing mutual information with an action-subject-shared discriminator network. Extensive experiments on three datasets demonstrate that our method is consistently effective for several CACS problems.
Yu Mitsuzumi, Akisato Kimura, Go Irie, Atsushi Nakazawa
ICIP2
2024 Efficient Algorithm for K-Multiple-Means
abstract
K-Multiple-Means is an extension of K-means for the clustering of multiple means used in many applications, such as image segmentation, load balancing, and blind-source separation. Since K-means uses only one mean to represent each cluster, it fails to capture non-spherical cluster structures of data points. However, since K-Multiple-Means represents the cluster by computing multiple means and grouping them into specified c clusters, it can effectively capture the non-spherical clusters of the data points. To obtain the clusters, K-Multiple-Means updates a similarity matrix of a bipartite graph between the data points and the multiple means by iteratively computing the leading c singular vectors of the matrix. K-Multiple-Means, however, incurs a high computation cost for large-scale data due to the iterative SVD computations. Our proposal, F-KMM, increases the efficiency of K-Multiple-Means by computing the singular vectors from a smaller similarity matrix between the multiple means obtained from the similarity matrix of the bipartite graph. To compute the similarity matrix of the bipartite graph efficiently, we skip unnecessary distance computations and estimate lower bounding distances between the data points and the multiple means. Theoretically, the proposed approach guarantees the same clustering results as K-Multiple-Means since it can exactly compute the singular vectors from the similarity matrix between the multiple means. Experiments show that our approach is several orders of magnitude faster than previous clustering approaches that use multiple means.
Yasuhiro Fujiwara, Atsutoshi Kumagai, Yasutoshi Ida, Masahiro Nakano, Makoto Nakatsuji, Akisato Kimura
Proc. ACM Manag. Data6
2024 Phase Randomization: A data augmentation for domain adaptation in human action recognition
Yu Mitsuzumi, Go Irie, Akisato Kimura, Atsushi Nakazawa
Pattern Recognit.3
2023 Selective Scene Text Removal
Hayato Mitani, Akisato Kimura, Seiichi Uchida
BMVC2
2023 Listening Human Behavior: 3D Human Pose Estimation with Acoustic Signals
abstract
Given only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actions, we explore how low-level acoustic signals can provide enough clues to estimate 3D human poses by active acoustic sensing with a single pair of microphones and loudspeakers (see Fig. 1). This is a challenging task since sound is much more diffractive than other signals and therefore covers up the shape of objects in a scene. Accordingly, we introduce a framework that encodes multichannel audio features into 3D human poses. Aiming to capture subtle sound changes to reveal detailed pose information, we explicitly extract phase features from the acoustic signals together with typical spectrum features and feed them into our human pose estimation network. Also, we show that reflected or diffracted sounds are easily influenced by subjects' physique differences e.g., height and muscularity, which deteriorates prediction accuracy. We reduce these gaps by using a subject discriminator to improve accuracy. Our experiments suggest that with the use of only low-dimensional acoustic information, our method outperforms baseline methods. The datasets and codes used in this project will be publicly available.
Yuto Shibata, Yutaka Kawashima, Mariko Isogawa, Go Irie, Akisato Kimura, Yoshimitsu Aoki
CVPR5
2023 Efficient Network Representation Learning via Cluster Similarity
Yasuhiro Fujiwara, Yasutoshi Ida, Atsutoshi Kumagai, Masahiro Nakano, Akisato Kimura, Naonori Ueda
DASFAA (3)5
2023 Deep Quantigraphic Image Enhancement via Comparametric Equations
abstract
Most recent methods of deep image enhancement can be generally classified into two types: decompose-and-enhance and illumination estimation-centric. The former is usually less efficient, and the latter is constrained by a strong assumption regarding image reflectance as the desired enhancement result. To alleviate this constraint while retaining high efficiency, we propose a novel trainable module that diversifies the conversion from the low-light image and illumination map to the enhanced image. It formulates image enhancement as a comparametric equation parameterized by a camera response function and an exposure compensation ratio. By incorporating this module in an illumination estimation-centric DNN, our method improves the flexibility of deep image enhancement, limits the computational burden to illumination estimation, and allows for fully unsupervised learning adaptable to the diverse demands of different tasks.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura
ICASSP3
2023 Efficient Network Representation Learning via Cluster Similarity
abstract
Abstract Network representation learning is a de facto tool for graph analytics. The mainstream of the previous approaches is to factorize the proximity matrix between nodes. However, if n is the number of nodes, since the size of the proximity matrix is $$n \times n$$ n × n , it needs $$O(n^3)$$ O ( n 3 ) time and $$O(n^2)$$ O ( n 2 ) space to perform network representation learning; they are significantly high for large-scale graphs. This paper introduces the novel idea of using similarities between clusters instead of proximities between nodes; the proposed approach computes the representations of the clusters from similarities between clusters and computes the representations of nodes by referring to them. If l is the number of clusters, since $$l \ll n$$ l ≪ n , we can efficiently obtain the representations of clusters from a small $$l \times l$$ l × l similarity matrix. Furthermore, since nodes in each cluster share similar structural properties, we can effectively compute the representation vectors of nodes. Experiments show that our approach can perform network representation learning more efficiently and effectively than existing approaches.
Yasuhiro Fujiwara, Yasutoshi Ida, Atsutoshi Kumagai, Masahiro Nakano, Akisato Kimura, Naonori Ueda
Data Sci. Eng.5
2023 Deep attentive time warping
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.4
2022 Nonparametric Relational Models with Superrectangulation
abstract
This paper addresses the question, ”What is the smallest object that contains all rectangular partitions with n or fewer blocks?” and shows its application to relational data analysis using a new strategy we call super Bayes as an alternative to Bayesian nonparametric (BNP) methods. Conventionally, standard BNP methods have combined the Aldous-Hoover-Kallenberg representation with parsimonious stochastic processes on rectangular partitioning to construct BNP relational models. As a result, conventional methods face the great difficulty of searching for a parsimonious random rectangular partition that fits the observed data well in Bayesian inference. As a way to essentially avoid such a problem, we propose a strategy to combine an extremely redundant rectangular partition as a deterministic (non-probabilistic) object. Specifically, we introduce a special kind of rectangular partitioning, which we call superrectangulation, that contains all possible rectangular partitions. Delightfully, this strategy completely eliminates the difficult task of searching around for random rectangular partitions, since the superrectangulation is deterministically fixed in inference. Experiments on predictive performance in relational data analysis show that the super Bayesian model provides a more stable analysis than the existing BNP models, which are less likely to be trapped in bad local optima.
Masahiro Nakano, Ryo Nishikimi, Yasuhiro Fujiwara, Akisato Kimura, Takeshi Yamada, Naonori Ueda
AISTATS4
2022 Fast Binary Network Hashing via Graph Clustering
abstract
Network hashing converts each node of a graph into a compact binary code, and it is a useful graph analytics tool since it can reduce memory cost. INH-MF is a network hashing approach to factorize the high-order proximity matrix representing similarities between nodes. However, since it cuts small nonzero elements from the proximity matrix, it fails to effectively extract insights from the graph. Moreover, it incurs high memory and computational costs since the proximity matrix is large and dense. We propose Graph Clustering-based Network Hashing, a novel network hashing approach. To compute the proximities effectively, it uses the structural relationships between nodes and clusters obtained from a graph clustering approach. Moreover, it can efficiently compute hash codes from eigenvectors of the matrix corresponding to the graph Laplacian by using its low-rank property. Experiments show that it can more efficiently and effectively compute hash codes than previous approaches.
Yasuhiro Fujiwara, Masahiro Nakano, Atsutoshi Kumagai, Yasutoshi Ida, Akisato Kimura, Naonori Ueda
IEEE Big Data5
2022 Font Shape-to-Impression Translation
Masaya Ueda, Akisato Kimura, Seiichi Uchida
DAS2
2022 Co-Attention-Guided Bilinear Model for Echo-Based Depth Estimation
abstract
Echoes reflect a geometric structure of a scene surrounding a sound source. In this paper, we address the problem of estimating depth maps of indoor scenes based on echoes. First, we experimentally show that fusing multiple acoustic features, especially spectrogram and angular spectrum, can improve estimation accuracy. We then propose a novel bilinear model that incorporates dense co-attention for effective feature fusion. Our model is able to obtain a compact fused feature while capturing the second-order correlations of intra-and inter-features. Thorough evaluations on two datasets demonstrate the superiority of the proposed method over the state-of-the-art echo-based depth estimation and feature fusion methods.
Go Irie, Takashi Shibata 0001, Akisato Kimura
ICASSP3
2022 Font Generation with Missing Impression Labels
abstract
Our goal is to generate fonts with specific impressions, by training a generative adversarial network with a font dataset with impression labels. The main difficulty is that font impression is ambiguous and the absence of an impression label does not always mean that the font does not have the impression. This paper proposes a font generation model that is robust against missing impression labels. The key ideas of the proposed method are (1) a co-occurrence-based missing label estimator and (2) an impression label space compressor. The first is to interpolate missing impression labels based on the co-occurrence of labels in the dataset and use them for training the model as completed label conditions. The second is an encoder-decoder module to compress the high-dimensional impression space into low-dimensional. We proved that the proposed model generates high-quality font images using multi-label data with missing labels through qualitative and quantitative evaluations. Our code is available at https://github.com/SeiyaMatsuda/Font-Generation-with-Missing-Impression-Labels.
Seiya Matsuda, Akisato Kimura, Seiichi Uchida
ICPR2
2022 ConceptBeam: Concept Driven Target Speech Extraction
abstract
We propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speaker in a mixture. Typical approaches have been exploiting properties of audio signals, such as harmonic structure and direction of arrival. In contrast, ConceptBeam tackles the problem with semantic clues. Specifically, we extract the speech of speakers speaking about a concept, i.e., a topic of interest, using a concept specifier such as an image or speech. Solving this novel problem would open the door to innovative applications such as listening systems that focus on a particular topic discussed in a conversation. Unlike keywords, concepts are abstract notions, making it challenging to directly represent a target concept. In our scheme, a concept is encoded as a semantic embedding by mapping the concept specifier to a shared embedding space. This modality-independent space can be built by means of deep metric learning using paired data consisting of images and their spoken captions. We use it to bridge modality-dependent information, i.e., the speech segments in the mixture, and the specified, modality-independent concept. As a proof of our scheme, we performed experiments using a set of images associated with spoken captions. That is, we generated speech mixtures from these spoken captions and used the images or speech signals as the concept specifiers. We then extracted the target speech using the acoustic characteristics of the identified segments. We compare ConceptBeam with two methods: one based on keywords obtained from recognition systems and another based on sound source separation. We show that ConceptBeam clearly outperforms the baseline methods and effectively extracts speech based on the semantic representation.
Yasunori Ohishi, Marc Delcroix, Tsubasa Ochiai, Shoko Araki, Daiki Takeuchi, Daisuke Niizumi, Akisato Kimura, Noboru Harada, Kunio Kashino
ACM Multimedia7
2022 Shared Latent Space of Font Shapes and Their Noisy Impressions
Daichi Haraguchi, Seiya Matsuda, Akisato Kimura, Seiichi Uchida
MMM (2)4
2022 Contrast enhancement based on reflectance-oriented probabilistic equalization
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
Signal Process.3
2021 Bayesian nonparametric model for arbitrary cubic partitioning
abstract
In this paper, we propose a continuous-time Markov process for cubic partitioning models of three-dimensional (3D) arrays and its application to Bayesian nonparametric relational data analysis of 3D array data. Relational data analysis is a topic that has been actively studied in the field of Bayesian nonparametrics, and in particular, models for analyzing 3D arrays have attracted much attention in recent years. In particular, the cubic partitioning model is very popular due to its practical usefulness, and various models such as the infinite relational model and the Mondrian process have been proposed. However, these conventional models have the disadvantage that they are limited to a certain class of cubic partitions, and there is a need for a model that can represent a broader class of arbitrary cubic partitions, which has long been an open issue in this field. In this study, we propose a stochastic process that can represent arbitrary cubic partitions of 3D arrays as a continuous-time Markov process. Furthermore, by combining it with the Aldous-Hoover-Kallenberg representation theorem, we construct an infinitely exchangeable 3D relational model and apply it to real data to show its application to relational data analysis. Experiments show that the proposed model improves the prediction performance by expanding the class of representable cubic partitioning.
Masahiro Nakano, Yasuhiro Fujiwara, Akisato Kimura, Takeshi Yamada, Naonori Ueda
ACML3
2021 Reflectance-Oriented Probabilistic Equalization for Image Enhancement
abstract
Despite recent advances in image enhancement, it remains difficult for existing approaches to adaptively improve the brightness and contrast for both low-light and normal-light images. To solve this problem, we propose a novel 2D histogram equalization approach. It assumes intensity occurrence and co-occurrence to be dependent on each other and derives the distribution of intensity occurrence (1D histogram) by marginalizing over the distribution of intensity co-occurrence (2D histogram). This scheme improves global contrast more effectively and reduces noise amplification. The 2D histogram is defined by incorporating the local pixel value differences in image reflectance into the density estimation to alleviate the adverse effects of dark lighting conditions. Over 500 images were used for evaluation, demonstrating the superiority of our approach over existing studies. It can sufficiently improve the brightness of low-light images while avoiding over-enhancement in normal-light images.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
ICASSP3
2021 Impressions2Font: Generating Fonts by Specifying Impressions
Seiya Matsuda, Akisato Kimura, Seiichi Uchida
ICDAR (3)2
2021 Attention to Warp: Deep Metric Learning for Multivariate Time Series
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
ICDAR (3)4
2021 Which Parts Determine the Impression of the Font?
Masaya Ueda, Akisato Kimura, Seiichi Uchida
ICDAR (3)2
2021 Deep Reinforcement Image Matching with Self-Termination
abstract
Deep reinforcement learning-based image matching sequentially searches only the promising regions in the reference image that match the query, leading to a significantly small number of steps compared to traditional methods. Since existing methods do not have any function to judge whether the target region has been successfully identified or not, they continue to search until the preset maximum number of search steps is reached. In this paper, we propose a deep image matching network that can terminate the matching process by itself. Our network is designed to have a halting module that identifies whether the current reference region matches the query based on the image features and the search history. The entire network is effectively trained end-to-end in a framework of deep reinforcement learning that incorporates a new loss function to evaluate the accuracy of the termination decision. Experimental results demonstrate that our method can achieve highly competitive or better matching accuracy with fewer search steps than the existing methods.
Onkar Krishna, Go Irie, Xiaomeng Wu, Akisato Kimura, Kunio Kashino
ICIP4
2021 Permuton-induced Chinese Restaurant Process
abstract
This paper proposes the permuton-induced Chinese restaurant process (PCRP), a stochastic process on rectangular partitioning of a matrix. This distribution is suitable for use as a prior distribution in Bayesian nonparametric relational model to find hidden clusters in matrices and network data. Our main contribution is to introduce the notion of permutons into the well-known Chinese restaurant process (CRP) for sequence partitioning: a permuton is a probability measure on $[0,1]\times [0,1]$ and can be regarded as a geometric interpretation of the scaling limit of permutations. Specifically, we extend the model that the table order of CRPs has a random geometric arrangement on $[0,1]\times [0,1]$ drawn from the permuton. By analogy with the relationship between the stick-breaking process (SBP) and CRP for the infinite mixture model of a sequence, this model can be regarded as a multi-dimensional extension of CRP paired with the block-breaking process (BBP), which has been recently proposed as a multi-dimensional extension of SBP. While BBP always has an infinite number of redundant intermediate variables, PCRP can be composed of varying size intermediate variables in a data-driven manner depending on the size and quality of the observation data. Experiments show that PCRP can improve the prediction performance in relational data analysis by reducing the local optima and slow mixing problems compared with the conventional BNP models because the local transitions of PCRP in Markov chain Monte Carlo inference are more flexible than the previous models.
Masahiro Nakano, Yasuhiro Fujiwara, Akisato Kimura, Takeshi Yamada, Naonori Ueda
NeurIPS3
2020 Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms
abstract
We propose a trilingual semantic embedding model that associates visual objects in images with segments of speech signals corresponding to spoken words in an unsupervised manner. Unlike the existing models, our model incorporates three different languages, namely, English, Hindi, and Japanese. To build the model, we used the existing English and Hindi datasets and collected a new corpus of Japanese speech captions. These spoken captions are spontaneous descriptions by individual speakers, rather than readings based on prepared transcripts. Therefore, we introduce a self-attention mechanism into the model to better map the spoken captions associated with the same image into the embedding space. We hope that the self-attention mechanism efficiently captures relationships between widely separated word-like segments. Experimental results show that the introduction of a third language improves the average performance in terms of cross-modal and cross-lingual retrieval accuracy, and that the self-attention mechanism added to the model works effectively.
Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David F. Harwath, James R. Glass
ICASSP2
2020 A Generative Self-Ensemble Approach To Simulated+Unsupervised Learning
abstract
In this paper, we consider Simulated and Unsupervised (S+U) learning which is a problem of learning from labeled synthetic and unlabeled real images. After translating the synthetic images to real ones, existing S+U learning methods use only the labeled synthetic images for training a predictor (e.g., a regression function) and ignore the target real images, which may result in unsatisfactory prediction performance. Our approach utilizes both synthetic and real images to train the predictor. The main idea of ours is to involve a self-ensemble learning framework into S+U learning. More specifically, we require the prediction results for an unlabeled real image to be consistent between “teacher” and “student” predictors, even after some perturbations are added to the image. Furthermore, aiming at generating diverse perturbations along the underlying data manifold, we introduce one-to-many image translation between synthetic and real images. Evaluation experiments on an appearance-based gaze estimation task demonstrate that the proposed ideas can improve the prediction accuracy and our full method can outperform existing S+U learning methods.
Yu Mitsuzumi, Go Irie, Akisato Kimura, Atsushi Nakazawa
ICIP3
2020 Total Whitening for Online Signature Verification Based on Deep Representation
abstract
In deep metric learning targeted at time series, the correlation between feature activations may be easily enlarged through highly nonlinear neural networks, leading to suboptimal embedding effectiveness. An effective solution to this problem is whitening. For example, in online signature verification, whitening can be derived for three individual Gaussian distributions, namely the distributions of local features at all temporal positions 1) for all signatures of all subjects, 2) for all signatures of each particular subject, and 3) for each particular signature of each particular subject. This study proposes a unified method called total whitening that integrates these individual Gaussians. Total whitening rectifies the layout of multiple individual Gaussians to resemble a standard normal distribution, improving the balance between intraclass invariance and interclass discriminative power. Experimental results demonstrate that total whitening achieves state-of-the-art accuracy when tested on online signature verification benchmarks.
Xiaomeng Wu, Akisato Kimura, Kunio Kashino, Seiichi Uchida
ICPR2
2020 Pair Expansion for Learning Multilingual Semantic Embeddings Using Disjoint Visually-Grounded Speech Audio Datasets
Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David F. Harwath, James R. Glass
INTERSPEECH2
2020 Baxter Permutation Process
abstract
In this paper, a Bayesian nonparametric (BNP) model for Baxter permutations (BPs), termed BP process (BPP) is proposed and applied to relational data analysis. The BPs are a well-studied class of permutations, and it has been demonstrated that there is one-to-one correspondence between BPs and several interesting objects including floorplan partitioning (FP), which constitutes a subset of rectangular partitioning (RP). Accordingly, the BPP can be used as an FP model. We combine the BPP with a multi-dimensional extension of the stick-breaking process called the {\it block-breaking process} to fill the gap between FP and RP, and obtain a stochastic process on arbitrary RPs. Compared with conventional BNP models for arbitrary RPs, the proposed model is simpler and has a high affinity with Bayesian inference.
Masahiro Nakano, Akisato Kimura, Takeshi Yamada, Naonori Ueda
NeurIPS2
2019 Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals
abstract
Sounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of semantic segmentation with given multichannel audio signals. Our approach uses a convolutional neural network that is designed to directly output semantic segmentation results by taking audio features as its inputs. A bilinear feature fusion scheme is incorporated that efficiently models underlying higher-order interactions between audio and visual sources. Experimental evaluations with both synthetic and real sound datasets show that our approach can recover the desired segmented images reasonably well.
Go Irie, Mirela Ostrek, Hirokazu Kameoka, Akisato Kimura, Takahito Kawanishi, Kunio Kashino
ICASSP5
2019 Prewarping Siamese Network: Learning Local Representations for Online Signature Verification
abstract
We propose a neural network-based framework for learning local representations of multivariate time series, and demonstrate its effectiveness for online signature verification. In contrast to related works that optimize a global distance objective, we incorporate a Siamese network into dynamic time warping (DTW), leading to a novel prewarping Siamese network (PSN) optimized with a local embedding loss. PSN learns a feature space that preserves the temporal location-wise distances of local structures. Local embedding, along with the alignment conditions of DTW, imposes a temporal consistency constraint on the sequence-level distance measure while achieving invariance as regards non-linear distortions. Validation on online signature verification datasets demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Seiichi Uchida, Kunio Kashino
ICASSP2
2019 Deep Dynamic Time Warping: End-to-End Local Representation Learning for Online Signature Verification
abstract
Siamese networks have been shown to be successful in learning deep representations for multivariate time series verification. However, most related studies optimize a global distance objective and suffer from a low discriminative power due to the loss of temporal information. To address this issue, we propose an end-to-end, neural network-based framework for learning local representations of time series, and demonstrate its effectiveness for online signature verification. This framework optimizes a Siamese network with a local embedding loss, and learns a feature space that preserves the temporal location-wise distances between time series. To achieve invariance to non-linear temporal distortion, we propose building a dynamic time warping block on top of the Siamese network, which will greatly improve the accuracy for local correspondences across intra-personal variability. Validation with respect to online signature verification demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Brian Kenji Iwana, Seiichi Uchida, Kunio Kashino
ICDAR2
2018 Weakly Supervised Collective Feature Learning From Curated Media
Yusuke Mukuta, Akisato Kimura, David B. Adrian, Zoubin Ghahramani
AAAI2
2018 Few-shot learning of neural networks from scratch by pseudo example optimization
Akisato Kimura, Zoubin Ghahramani, Koh Takeuchi 0001, Tomoharu Iwata, Naonori Ueda
BMVC1
2018 Introducing Local Distance-Based Features to Temporal Convolutional Neural Networks
abstract
In this paper, we propose the use of local distance-based features determined by Dynamic Time Warping (DTW) for temporal Convolutional Neural Networks (CNN). Traditionally, DTW is used as a robust distance metric for time series patterns. However, this traditional use of DTW only utilizes the scalar distance metric and discards the local distances between the dynamically matched sequence elements. This paper proposes recovering these local distances, or DTW features, and utilizing them for the input of a CNN. We demonstrate that these features can provide additional information for the classification of isolated handwritten digits and characters. Furthermore, we demonstrate that the DTW features can be combined with the spatial coordinate features in multi-modal fusion networks to achieve state-of-the-art accuracy on the Unipen online handwritten character datasets.
Brian Kenji Iwana, Minoru Mori, Akisato Kimura, Seiichi Uchida
ICFHR3
2016 Infinite Plaid Models for Infinite Bi-Clustering
abstract
We propose a probabilistic model for non-exhaustive and overlapping (NEO) bi-clustering. Our goal is to extract a few sub-matrices from the given data matrix, where entries of a sub-matrix are characterized by a specific distribution or parameters. Existing NEO biclustering methods typically require the number of sub-matrices to be extracted, which is essentially difficult to fix a priori. In this paper, we extend the plaid model, known as one of the best NEO bi-clustering algorithms, to allow infinite bi-clustering; NEO bi-clustering without specifying the number of sub-matrices. Our model can represent infinite sub-matrices formally. We develop a MCMC inference without the finite truncation, which potentially addresses all possible numbers of sub-matrices. Experiments quantitatively and qualitatively verify the usefulness of the proposed model. The results reveal that our model can offer more precise and in-depth analysis of sub-matrices.
Katsuhiko Ishiguro, Issei Sato, Masahiro Nakano, Akisato Kimura, Naonori Ueda
AAAI4
2015 Identifying Attractive News Headlines for Social Media
abstract
In the past, leading newspaper companies and broadcasters were the sole distributors of news articles, and thus news consumers simply received news articles from those outlets at regular intervals. However, the growth of social media and smart devices led to a considerable change in this traditional relationship between news providers and consumers. Hundreds of thousands of news articles are now distributed on social media, and consumers can access those articles at any time via smart devices. This has meant that news providers are under pressure to find ways of engaging the attention of consumers. This paper provides a novel solution to this problem by identifying attractive headlines as a gateway to news articles. We first perform one of the first investigations of news headlines on a major viral medium. Using our investigation as a basis, we also propose a learning-to-rank method that suggests promising news headlines. Our experiments with 2,000 news articles demonstrate that our proposed method can accurately identify attractive news headlines from the candidates and reveals several promising factors of making news articles go viral.
Sawa Kourogi, Hiroyuki Fujishiro, Akisato Kimura, Hitoshi Nishikawa
CIKM3
2015 Visual Attention Driven by Auditory Cues - Selecting Visual Features in Synchronization with Attracting Auditory Events
Jiro Nakajima, Akisato Kimura, Akihiro Sugimoto, Kunio Kashino
MMM (2)2
2014 Rectangular Tiling Process
abstract
This paper proposes a novel stochastic process that represents the arbitrary rectangular partitioning of an infinite-dimensional matrix as the conditional projective limit. Rectangular partitioning is used in relational data analysis, and is classified into three types: regular grid, hierarchical, and arbitrary. Conventionally, a variety of probabilistic models have been advanced for the first two, including the product of Chinese restaurant processes and the Mondrian process. However, existing models for arbitrary partitioning are too complicated to permit the analysis of the statistical behaviors of models, which places very severe capability limits on relational data analysis. In this paper, we propose a new probabilistic model of arbitrary partitioning called the rectangular tiling process (RTP). Our model has a sound mathematical base in projective systems and infinite extension of conditional probabilities, and is capable of representing partitions of infinite elements as found in ordinary Bayesian nonparametric models.
Masahiro Nakano, Katsuhiko Ishiguro, Akisato Kimura, Takeshi Yamada, Naonori Ueda
ICML3
2013 Clustering-based anomaly detection in multi-view data
abstract
This paper proposes a simple yet effective anomaly detection method for multi-view data. The proposed approach detects anomalies by comparing the neighborhoods in different views. Specifically, clustering is performed separately in the different views and affinity vectors are derived for each object from the clustering results. Then, the anomalies are detected by comparing affinity vectors in the multiple views. An advantage of the proposed method over existing methods is that the tuning parameters can be determined effectively from the given data. Through experiments on synthetic and benchmark datasets, we show that the proposed method outperforms existing methods.
Alejandro Marcos Alvarez, Makoto Yamada, Akisato Kimura, Tomoharu Iwata
CIKM3
2013 Non-negative Multiple Tensor Factorization
abstract
Non-negative Tensor Factorization (NTF) is a widely used technique for decomposing a non-negative value tensor into sparse and reasonably interpretable factors. However, NTF performs poorly when the tensor is extremely sparse, which is often the case with real-world data and higher-order tensors. In this paper, we propose Non-negative Multiple Tensor Factorization (NMTF), which factorizes the target tensor and auxiliary tensors simultaneously. Auxiliary data tensors compensate for the sparseness of the target data tensor. The factors of the auxiliary tensors also allow us to examine the target data from several different aspects. We experimentally confirm that NMTF performs better than NTF in terms of reconstructing the given data. Furthermore, we demonstrate that the proposed NMTF can successfully extract spatio-temporal patterns of people's daily life such as leisure, drinking, and shopping activity by analyzing several tensors extracted from online review data sets.
Koh Takeuchi 0001, Ryota Tomioka, Katsuhiko Ishiguro, Akisato Kimura, Hiroshi Sawada
ICDM4
2013 Non-Negative Multiple Matrix Factorization
Koh Takeuchi 0001, Katsuhiko Ishiguro, Akisato Kimura, Hiroshi Sawada
IJCAI3
2013 Change-Point Detection with Feature Selection in High-Dimensional Time-Series Data
Makoto Yamada, Akisato Kimura, Futoshi Naya, Hiroshi Sawada
IJCAI2
2013 Image context discovery from socially curated contents
abstract
This paper proposes a novel method of discovering a set of image contents sharing a specific context (attributes or implicit meaning) with the help of image collections obtained from social curation platforms. Socially curated contents are promising to analyze various kinds of multimedia information, since they are manually filtered and organized based on specific individual preferences, interests or perspectives. Our proposed method fully exploits the process of social curation: (1) How image contents are manually grouped together by users, and (2) how image contents are distributed in the platform. Our method reveals the fact that image contents with a specific context are naturally grouped together and every image content includes really various contexts that cannot necessarily be verbalized by texts.% A preliminary experiment with a small collection of a million of images yields a promising result.
Akisato Kimura, Katsuhiko Ishiguro, Makoto Yamada, Alejandro Marcos Alvarez, Kaori Kataoka, Kazuhiko Murasaki
ACM Multimedia1
2012 Single Image Segmentation with Estimated Depth
abstract
Object segmentation is a fundamental problem in computer vision. Although many segmentation methods have been proposed, most of them still rely on the appearances of images (i.e., colors or textures) [1, 2, 3, 4, 6, 8]. Consequently, they have a difficulty in distinguishing an object from the background with a similar appearance to the object. To overcome this difficulty, we employ a depth map of an input image as an additional cue to the object segmentation. The main contribution of this work is to introduce a novel segmentation framework that utilizes the depth map combined with a color image to describe the features of objects and backgrounds, where the depth map is estimated from the color image. While a depth map has great potential for use in segmentation, finding a way of integrating two completely different physical quantities, namely the color and depth, has remained unclear. We introduce an integration of the color and depth likelihood on objectness and backgroundness, which simply and effectively extends a traditional segmentation framework based on the Markov random fields (MRF) [2]. By refining the likelihood with the depth information, our proposed method can suppress the incorrect detection of misleading backgrounds. A single image is expressed by K, where K includes color information C = {Cx ∈R}x∈Ω, and in our case, depth informationZ = {Zx ∈R}x∈Ω (x is a position in the image domain Ω ⊂ N2). Object segmentation is the problem of assigning the label A = {Ax}x∈Ω, which gives a label Ax = {0,1} to each pixel, where the labels 1 and 0 at x respectively correspond to the object and background. The statistical relationship between K and A can be described by an MRF, and the appropriate configuration of the labels can be derived by minimizing the following energy function E:
Ryo Yonetani, Akisato Kimura, Hitoshi Sakano, Ken Fukuchi
BMVC2
2012 Towards Automatic Image Understanding and Mining via Social Curation
abstract
The amount and variety of multimedia data such as images, movies and music available on over social networks are increasing rapidly. However, the ability to analyze and exploit these unorganized multimedia data remains inadequate, even with state-of-the-art media processing techniques. Our finding in this paper is that the emerging social curation service is a promising information source for the automatic understanding and mining of images distributed and exchanged via social media. One remarkable virtue of social curation service datasets is that they are weakly supervised: the content in the service is manually collected, selected and maintained by users. This is very different from other social information sources, and we can utilize this characteristics for media content mining without expensive media processing techniques. In this paper we present a machine learning system for predicting view counts of images in social curation data as the first step to automatic image content evaluation. Our experiments confirm that the simple features extracted from a social curation corpus are much superior in terms of count prediction than the gold-standard image features of computer vision research.
Katsuhiko Ishiguro, Akisato Kimura, Koh Takeuchi 0001
ICDM2
2012 Designing various component analysis at will
Akisato Kimura, Hitoshi Sakano, Hirokazu Kameoka, Masashi Sugiyama
ICPR1
2012 Creating Stories: Social Curation of Twitter Messages
Kevin Duh, Tsutomu Hirao, Akisato Kimura, Katsuhiko Ishiguro, Tomoharu Iwata, Ching-man Au Yeung
ICWSM3
2012 Fully Automatic Extraction of Salient Objects from Videos in Near Real Time
abstract
Automatic video segmentation plays an important role in a wide range of computer vision and image processing applications. Recently, various methods have been proposed for this purpose. The problem is that most of these methods are far from real-time processing even for low-resolution videos due to the complex procedures. To this end, we propose a new and quite fast method for automatic video segmentation with the help of (1) efficient optimization of Markov random fields with polynomial time of the number of pixels by introducing graph cuts, (2) automatic, computationally efficient but stable derivation of segmentation priors using visual saliency and sequential update mechanism and (3) an implementation strategy in the principle of stream processing with graphics processor units. Test results indicate that our method extracts appropriate regions from videos as precisely as and much faster than previous semi-automatic methods even though no supervisions have been incorporated.
Akamine Kazuma, Ken Fukuchi, Akisato Kimura, Shigeru Takagi
Comput. J.3
2011 Automatic video annotation via Hierarchical Topic Trajectory Model considering cross-modal correlations
abstract
We propose a new statistical model, named Hierarchical Topic Trajectory Model (HTTM), for acquiring a dynamically changing topic model that represents the relationship between video frames and associated text labels. Model parameter estimation, annotation and retrieval can be executed within a unified framework with a few computation. It is also easy to add new modals such as audio signal and geotags. Preliminary experiments on video annotation task with manually annotated video dataset indicate that our proposed method can improve the annotation accuracy.
Takuho Nakano, Akisato Kimura, Hirokazu Kameoka, Shigeki Miyabe, Shigeki Sagayama, Nobutaka Ono, Kunio Kashino, Takuya Nishimoto
ICASSP2
2011 Automatic audio tag classification via semi-supervised canonical density estimation
abstract
We propose a novel semi-supervised method for building a statistical model that represents the relationship between sounds and text labels ("tags"). The proposed method, named semi-supervised canonical density estimation, makes use of unlabeled sound data in two ways: 1) a low-dimensional latent space representing topics of sounds is extracted by a semi-supervised variant of canonical correlation analysis, and 2) topic models are learned by multi-class extension of semi-supervised kernel density estimation in the topic space. Real-world audio tagging experiments indicate that our pro posed method improves the accuracy even when only a small number of labeled sounds are available.
Jun Takagi, Yasunori Ohishi, Akisato Kimura, Masashi Sugiyama, Makoto Yamada, Hirokazu Kameoka
ICASSP3
2010 SemiCCA: Efficient Semi-supervised Learning of Canonical Correlations
abstract
Canonical correlation analysis (CCA) is a powerful tool for analyzing multi-dimensional paired data. However, CCA tends to perform poorly when the number of paired samples is limited, which is often the case in practice. To cope with this problem, we propose a semi-supervised variant of CCA named "Semi CCA" that allows us to incorporate additional unpaired samples for mitigating overfitting. The proposed method smoothly bridges the eigenvalue problems of CCA and principal component analysis (PCA), and thus its solution can be computed efficiently just by solving a single (generalized) eigenvalue problem as the original CCA. Preliminary experiments with artificially generated samples and PASCAL VOC data sets demonstrate the effectiveness of the proposed method.
Akisato Kimura, Hirokazu Kameoka, Masashi Sugiyama, Takuho Nakano, Eisaku Maeda, Hitoshi Sakano, Katsuhiko Ishiguro
ICPR1
2010 Universal source coding for multiple decoders with side information
abstract
A multiterminal lossy source coding problem, which includes various problems such as the Wyner-Ziv problem and the complementary delivery problem as special cases, is considered. It is shown that any point in the achievable rate-distortion region can be attained even if the source statistics are not known.
Shigeaki Kuzuoka, Akisato Kimura, Tomohiko Uyematsu
ISIT2
2009 Saliency-based video segmentation with graph cuts and sequentially updated priors
abstract
This paper proposes a new method for achieving precise video segmentation without any supervision or interaction. The main contributions of this report include 1) the introduction of fully automatic segmentation based on the maximum a posteriori (MAP) estimation of the Markov random field (MRF) with graph cuts and saliency-driven priors and 2) the updating of priors and feature likelihoods by integrating the previous segmentation results and the currently estimated saliency-based visual attention. Test results indicate that our new method precisely extracts probable regions from videos without any supervised interactions.
Ken Fukuchi, Kouji Miyazato, Akisato Kimura, Shigeru Takagi, Junji Yamato
ICME3
2009 Real-time estimation of human visual attention with dynamic Bayesian network and MCMC-based particle filter
abstract
Recent studies in signal detection theory suggest that the human responses to the stimuli on a visual display are nondeterministic. People may attend to different locations on the same visual input at the same time. Constructing a stochastic model of human visual attention would be promising to tackle the above problem. This paper proposes a new method to achieve a quick and precise estimation of human visual attention based on our previous stochastic model with a dynamic Bayesian network. A particle filter with Markov chain Monte-Carlo (MCMC) sampling make it possible to achieve a quick and precise estimation through stream processing. Experimental results indicate that the proposed method can estimate human visual attention in real time and more precisely than previous methods.
Kouji Miyazato, Akisato Kimura, Shigeru Takagi, Junji Yamato
ICME2
2009 Universal Source Coding OverGeneralized Complementary Delivery Networks
abstract
This paper deals with a universal coding problem for a certain kind of multiterminal source coding network called a generalized complementary delivery network. In this network, messages from multiple correlated sources are jointly encoded, and each decoder has access to some of the messages to enable it to reproduce the other messages. Both fixed-to-fixed length and fixed-to-variable length lossless coding schemes are considered. Explicit constructions of universal codes and the bounds of the error probabilities are clarified by using methods of types and graph-theoretical analysis.
Akisato Kimura, Tomohiko Uyematsu, Shigeaki Kuzuoka, Shun Watanabe
IEEE Trans. Inf. Theory1
2008 A stochastic model of selective visual attention with a dynamic Bayesian network
abstract
Recent studies in signal detection theory suggest that the human responses to the stimuli on a visual display are nondeterministic. People may attend to different locations on the same visual input at the same time. To predict the likelihood of where humans typically focus on a video scene, we propose a new stochastic model of visual attention by introducing a dynamic Bayesian network. Our model simulates and combines the visual saliency response and the cognitive state of a person to estimate the most probable attended regions. Experimental results have demonstrated that our model performs significantly better in predicting human visual attention compared to the previous deterministic model.
Derek Pang, Akisato Kimura, Tatsuto Takeuchi, Junji Yamato, Kunio Kashino
ICME2
2008 Dynamic Markov random fields for stochastic modeling of visual attention
abstract
This report proposes a new stochastic model of visual attention to predict the likelihood of where humans typically focus on a video scene. The proposed model is composed of a dynamic Bayesian network that simulates and combines a person’s visual saliency response and eye movement patterns to estimate the most probable regions of attention. Dynamic Markov random field (MRF) models are newly introduced to include spatiotemporal relationships of visual saliency responses. Experimental results have revealed that the propose model outperforms the previous deterministic model and the stochastic model without dynamic MRF in predicting human visual attention.
Akisato Kimura, Derek Pang, Tatsuto Takeuchi, Junji Yamato, Kunio Kashino
ICPR1
2008 Universal coding for lossy complementary delivery problem
abstract
This paper deals with a universal lossy coding problem for a certain kind of multiterminal source coding network called a complementary delivery system. A universal coding scheme based on Wyner-Ziv codes is proposed. While the proposed scheme cannot attain the optimal rate-distortion trade off in general, the rate-loss is upper bounded by a universal constant under some mild conditions. Moreover, the proposed scheme allows us to apply (non-universal) Wyner-Ziv codes to construct a universal lossy complementary delivery code.
Shigeaki Kuzuoka, Akisato Kimura, Tomohiko Uyematsu
ISIT2
2007 Robust Search Methods for Music Signals Based on Simple Representation
abstract
Signal similarity search is an important technique for music information retrieval. A basic task is finding identical signal segments on unlabeled music-signal archives, given a short music signal fragment as a query. In such a task, the search must be fast and sufficiently robust against possible signal fluctuations due to noise and distortions. In this special session paper, we describe a search method designed to cope with additive interfering sounds by spectral partitioning. Then, we introduce another method designed to be robust under multiplicative noise or distortion based on binary area representation.
Kunio Kashino, Akisato Kimura, Hidehisa Nagano, Takayuki Kurozumi
ICASSP (4)2
2007 A Computational Model of Saliency Depletion/Recovery Phenomena for the Salient Region Extraction of Videos
abstract
This report proposes a new algorithm for extracting salient regions of videos by introducing two important properties of the early human visual system: (1) Instantaneous saliency depletion with gradual recovery, whereby saliency is instantaneously suppressed and gradually recovered in previously attended regions. (2) Gradual saliency depletion with instantaneous recovery, whereby saliency is gradually decreased over time in non-surprising regions and at the same time recovered in surprising locations. With the introduction of these properties, redundant information in videos can be suppressed and important information is eventually enhanced.
Clement Leung, Akisato Kimura, Tatsuto Takeuchi, Kunio Kashino
ICME2
2007 Universal coding for correlated sources with complementary delivery
abstract
This report deals with a universal coding problem for a certain kind of multiterminal source coding system that we call the complementary delivery coding system. Both fixed-to- fixed length and fixed-to-variable length lossless coding schemes are considered. Explicit constructions of universal codes and the bounds of the error probabilities are clarified via type-theoretical and graph-theoretical analyses.
Akisato Kimura, Tomohiko Uyematsu, Shigeaki Kuzuoka
ISIT1
2004 Similarity-based partial image retrieval guaranteeing same accuracy as exhaustive matching
abstract
We propose a new framework for quick and accurate partial image retrieval from a huge number of images based on a predefined distance measure. Finding partial similarities generally requires a huge amount of storage space for indexes due to the large number of portions of images. The proposed method extracts portions from each database image at a constant spacing, while it extracts all possible portions from a query image. In this way, the proposed method can greatly reduce the size of indexes while theoretically guaranteeing the same accuracy as exhaustive matching.
Akisato Kimura, Takahito Kawanishi, Kunio Kashino
ICME1
2004 Weak variable-length Slepian-Wolf coding with linked encoders for mixed sources
abstract
Coding problems for correlated information sources were first investigated by Slepian and Wolf. They considered the data compression system, called the SW system, where two sequences emitted from correlated sources are separately encoded to codewords, and sent to a single decoder which has to output the original sequence pairs with a small probability or error. In this paper, we investigate the coding problem of a modified SW system allowing two encoders to communicate with zero rate. First, we consider the fixed-length coding and clarify that the admissible rate region for general sources is equal to that of the original SW system. Next, we investigate the variable-length coding having the asymptotically vanishing probability of error. We clarify the admissible rate region for mixed sources characterized by two ergodic sources and show that this region is strictly wider than that for fixed-length codes. Further, we investigate the universal coding problem for memoryless sources in the system and show that the SW system with linked encoders has much more flexibility than the original SW system.
Akisato Kimura, Tomohiko Uyematsu
IEEE Trans. Inf. Theory1
2003 Dynamic-segmentation-based feature dimension reduction for quick audio/video searching
abstract
We propose a new feature dimension reduction method for multimedia search. The main technique in the method is dynamic segmentation that partitions sequential feature trajectories dynamically. While dynamic segmentation reduces the average dimensionality and accelerates the search, it requires huge amount of calculation. Thus, our method quickly executes suboptimal partitioning of the trajectories by using the discreteness of dimension changes. This guarantees the optimal amount of calculation to derive the suboptimal partitioning under the condition that the dimension monotonously increases as the segment length increases. The experiment shows that our method is over 10 times faster than a straightforward dynamic segmentation method.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP (3)1
2003 Dynamic-segmentation-based feature dimension reduction for quick audio/video searching
abstract
We propose a new feature dimension reduction method for multimedia search. The main technique in the method is dynamic segmentation that partitions sequential feature trajectories dynamically. While dynamic segmentation reduces the average dimensionality and accelerates the search, it requires huge amount of calculation. Thus, our method quickly executes suboptimal partitioning of the trajectories by using the discreteness of dimension changes. This guarantees the optimal amount of calculation to derive the suboptimal partitioning under the condition that the dimension monotonously increases as the segment length increases. The experiment shows that our method is over 10 times faster than a straightforward dynamic segmentation method.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICME1
2002 A quick search method for multimedia signals using feature compression based on piecewise linear maps
abstract
We propose a quick algorithm for multimedia signal search. The algorithm comprises two techniques: feature compression based on piecewise linear maps and distance bounding to efficiently limit the search space. When compared with existing multimedia search techniques, they greatly reduce the computational cost required in searching. Although feature compression is employed in our method, our bounding technique mathematically guarantees the same recall rate as the search based on the original features; no segment to be detected is missed. Experiments indicate that the proposed algorithm is approximately 10 times faster than and as accurate as an existing fast method maitaining the same search accuracy.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP1
2001 Very quick audio searching: introducing global pruning to the Time-Series Active Search
abstract
Previously, we proposed a histogram-based quick signal search method called Time-Series Active Search (TAS). TAS is a method of searching through long audio or video recordings for a specified segment, based on signal similarity. TAS is fast; it can search through a 24-hour recording in 1 second after a query-independent preprocessing. However, an even faster method is required when we consider a huge amount of audio archives, for example a month's worth of recordings. Thus, we propose a preprocessing method that significantly accelerates TAS. The core part of this method comprises a global histogram clustering of long signals and a pruning scheme using those clusters. Tests using broadcast recording indicate that the proposed algorithm achieves a search speed approximately 3 to 30 times faster than TAS. In these tests, the search results are exactly the same as with TAS.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP1
2001 Weak variable-length Slepian-Wolf coding with linked encoders for mixed sources
abstract
Slepian and Wolf (see IEEE Trans. Inform. Theory, vol.19, p.471-80, July 1973) first considered the data compression of correlated sources called the SW system, where two sequences emitted from correlated sources are separately encoded to codewords, and sent to a single decoder which has to output original sequence pairs. Recently, Oohama (see IEEE Trans. Inform. Theory, vol.42, p.837-47, May 1996) has extended the SW system and investigated a more general case where there are some mutual linkages between two encoders of the SW system. In this paper, we investigate variable-length coding which allows asymptotically vanishing probability of error for the system considered by Oohama. We clarify the admissible rate region for mixed sources characterized by two ergodic sources, and show that this region is strictly wider than that for fixed-length codes.
Akisato Kimura, Tomohiko Uyematsu
ITW1