Kunio Kashino

dblp:68/651 · DBLP profile ↗
← Back
111ranked-venue papers
11as first author
17since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 97 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 43 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Hyperbolic PHATE: Visualizing Continuous Hierarchy of Latent Differentiation Structures
abstract
This paper proposes a method for embedding diffusion potentials into a hyperbolic space in order to visualize the differentiation structure consisting of diffusion and branching inherent in high-dimensional data. In recent years, the rapid development of single-cell sequencing in the field of biological information processing has made it possible to observe the evolution of gene expression levels in a snapshot-like manner as cells grow from birth to each organ or tissue. Visualization of such high-dimensional (gene pattern dimension) data is expected to provide important insights into the mechanisms of cell differentiation. Therefore, in the visualization of such data, there is a need for a system that emphasizes the "diffusion" structure that gradually shifts with time and the "branching" structure that broadly branches off into individual organs and tissues. Conventionally, the diffusion map and its extension PHATE have been developed as visualization methods specializing in diffusion structures, and hyperbolic embedding has been used as a method specializing in branching structures. However, methods that attempt to explicitly capture diffusion and branching structures simultaneously have not yet received much attention. In this paper, we focus on diffusion mapping (and its extension, PHATE), which specializes in diffusion structures, and hyperbolic embedding, which specializes in branching structures, and propose a visualization method that combines the advantages of both in order to better capture differentiaion structures consisting of diffusion and branching. As a symbolic example, we demonstrate our method using gene-cell expression data in the context of single cell analysis.
Masahiro Nakano, Hiroki Sakuma, Ryo Nishikimi, Kenji Komiya, Tomoharu Iwata, Kunio Kashino
ICASSP6
2024 Warped Diffusion for Latent Differentiation Inference
abstract
This paper proposes a Bayesian nonparametric diffusion model with a black-box warping function represented by a Gaussian process to infer potential diffusion structures latent in observed data, such as differentiation mechanisms of living cells and phylogenetic evolution processes of media information. In general, the task of inferring latent differentiation structures is very difficult to handle due to two interrelated settings. One is that the conversion mechanism between hidden structure and often higher dimensional observations is unknown (and is a complex mechanism). The other is that the topology of the hidden diffuse structure itself is unknown. Therefore, in this paper, we propose a BNP-based strategy as a natural way to deal with these two challenging settings simultaneously. Specifically, as an extension of the Gaussian process latent variable model, we propose a model in which the black box transformation from latent variable space to observed data space is represented by a Gaussian process, and introduce a BNP diffusion model for the latent variable space. We show its application to the visualization of the diffusion structure of media information and to the task of inferring cell differentiation structure from single-cell gene expression levels.
Masahiro Nakano, Hiroki Sakuma, Ryo Nishikimi, Ryohei Shibue, Tomoharu Iwata, Kunio Kashino
AISTATS7
2024 Sunflower Strategy for Bayesian Relational Data Analysis
abstract
This paper proposes a new inference strategy for Bayesian relational data analysis, inspired by the sunflower lemma in extremal combinatorics. Relational data analysis using rectangular partitioning is a particularly popular analysis tool due to its interpretability and explainability, and has been used in a wide range of signal processing and machine learning applications such as network community detection and cluster structure extraction from cell gene expression data. On the other hand, the combinatorial space of rectangular partitions makes the derivation of its inference algorithms difficult. As a result, many inference algorithms, such as Markov chain Monte Carlo (MCMC) methods, have suffered from local mode and slow mixing problems. Therefore, we propose a method of inference to avoid the difficult combinatorial space by using tractable auxiliary spaces with relaxed constraints, utilizing the knowledge of the sunflower lemma in extremal combinatorics. Specifically, we construct a MCMC method to induce a target combinatorial space of rectangular partitions as the intersection (as the core of a sunflower) of sufficiently and redundantly many relaxed and tractable spaces (as petals of the sunflower). Through experiments, we will show that our method has better prediction performance than previous state-of-the-art methods for network community detection and cluster structure extraction of gene-cell expression data tasks.
Masahiro Nakano, Ryohei Shibue, Kunio Kashino
ICASSP3
2024 Masked Modeling Duo: Towards a Universal Audio Pre-Training Framework
abstract
Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals. Unlike conventional methods, M2D obtains a training signal by encoding only the masked part, encouraging the two networks in M2D to model the input. While M2D improves general-purpose audio representations, a specialized representation is essential for real-world applications, such as in industrial and medical domains. The often confidential and proprietary data in such domains is typically limited in size and has a different distribution from that in pre-training datasets. Therefore, we propose M2D for X (M2D-X), which extends M2D to enable the pre-training of specialized representations for an application X. M2D-X learns from M2D and an additional task and inputs background noise. We make the additional task configurable to serve diverse applications, while the background noise helps learn on small data and forms a denoising task that makes representation robust. With these design choices, M2D-X should learn a representation specialized to serve various application needs. Our experiments confirmed that the representations for general-purpose audio, specialized for the highly competitive AudioSet and speech domain, and a small-data medical task achieve top-level performance, demonstrating the potential of using our models as a universal audio pre-training framework. Our code is available online for future studies.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
abstract
Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods learn representations directly by predicting representations of masked patches; however, we think using all patches to encode training signal representations is suboptimal. We propose a new method, Masked Modeling Duo (M2D), that learns representations directly while obtaining training signals using only masked patches. In the M2D, the online network encodes visible patches and predicts masked patch representations, and the target network, a momentum encoder, encodes masked patches. To better predict target representations, the online network should model the input well, while the target network should also model it well to agree with online predictions. Then the learned representations should better model the input. We validated the M2D by learning general-purpose audio representations, and M2D set new state-of-the-art performance on tasks such as UrbanSound8K, VoxCeleb1, AudioSet20K, GTZAN, and SpeechCommandsV2.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
ICASSP5
2023 Masked Modeling Duo for Speech: Specializing General-Purpose Audio Representation to Speech using Denoising Distillation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
INTERSPEECH5
2023 Deep attentive time warping
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
Pattern Recognit.5
2023 BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations
abstract
Pre-trained models are essential as feature extractors in modern machine learning systems in various domains. In this study, we hypothesize that representations effective for general audio tasks should provide multiple aspects of robust features of the input sound. For recognizing sounds regardless of perturbations such as varying pitch or timbre, features should be robust to these perturbations. For serving the diverse needs of tasks such as recognition of emotions or music genres, representations should provide multiple aspects of information, such as local and global features. To implement our principle, we propose a self-supervised learning method: Bootstrap Your Own Latent (BYOL) for Audio (BYOL-A, pronounced “viola”). BYOL-A pre-trains representations of the input sound invariant to audio data augmentations, which makes the learned representations robust to the perturbations of sounds. Whereas the BYOL-A encoder combines local and global features and calculates their statistics to make the representation provide multi-aspect information. As a result, the learned representations should provide robust and multi-aspect information to serve various needs of diverse tasks. We evaluated the general audio task performance of BYOL-A compared to previous state-of-the-art methods, and BYOL-A demonstrated generalizability with the best average result of 72.4% and the best VoxCeleb1 result of 57.6%. Extensive ablation experiments revealed that the BYOL-A encoder architecture contributes to most performance, and the final critical portion resorts to the BYOL framework and BYOL-A augmentations. Our code is available online for future studies.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Introducing Auxiliary Text Query-modifier to Content-based Audio Retrieval
Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi, Noboru Harada, Kunio Kashino
INTERSPEECH5
2022 ConceptBeam: Concept Driven Target Speech Extraction
abstract
We propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speaker in a mixture. Typical approaches have been exploiting properties of audio signals, such as harmonic structure and direction of arrival. In contrast, ConceptBeam tackles the problem with semantic clues. Specifically, we extract the speech of speakers speaking about a concept, i.e., a topic of interest, using a concept specifier such as an image or speech. Solving this novel problem would open the door to innovative applications such as listening systems that focus on a particular topic discussed in a conversation. Unlike keywords, concepts are abstract notions, making it challenging to directly represent a target concept. In our scheme, a concept is encoded as a semantic embedding by mapping the concept specifier to a shared embedding space. This modality-independent space can be built by means of deep metric learning using paired data consisting of images and their spoken captions. We use it to bridge modality-dependent information, i.e., the speech segments in the mixture, and the specified, modality-independent concept. As a proof of our scheme, we performed experiments using a set of images associated with spoken captions. That is, we generated speech mixtures from these spoken captions and used the images or speech signals as the concept specifiers. We then extracted the target speech using the acoustic characteristics of the identified segments. We compare ConceptBeam with two methods: one based on keywords obtained from recognition systems and another based on sound source separation. We show that ConceptBeam clearly outperforms the baseline methods and effectively extracts speech based on the semantic representation.
Yasunori Ohishi, Marc Delcroix, Tsubasa Ochiai, Shoko Araki, Daiki Takeuchi, Daisuke Niizumi, Akisato Kimura, Noboru Harada, Kunio Kashino
ACM Multimedia9
2022 Contrast enhancement based on reflectance-oriented probabilistic equalization
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
Signal Process.4
2021 Reflectance-Oriented Probabilistic Equalization for Image Enhancement
abstract
Despite recent advances in image enhancement, it remains difficult for existing approaches to adaptively improve the brightness and contrast for both low-light and normal-light images. To solve this problem, we propose a novel 2D histogram equalization approach. It assumes intensity occurrence and co-occurrence to be dependent on each other and derives the distribution of intensity occurrence (1D histogram) by marginalizing over the distribution of intensity co-occurrence (2D histogram). This scheme improves global contrast more effectively and reduces noise amplification. The 2D histogram is defined by incorporating the local pixel value differences in image reflectance into the density estimation to alleviate the adverse effects of dark lighting conditions. Over 500 images were used for evaluation, demonstrating the superiority of our approach over existing studies. It can sufficiently improve the brightness of low-light images while avoiding over-enhancement in normal-light images.
Xiaomeng Wu, Yongqing Sun, Akisato Kimura, Kunio Kashino
ICASSP4
2021 Attention to Warp: Deep Metric Learning for Multivariate Time Series
Shinnosuke Matsuo, Xiaomeng Wu, Gantugs Atarsaikhan, Akisato Kimura, Kunio Kashino, Brian Kenji Iwana, Seiichi Uchida
ICDAR (3)5
2021 Deep Reinforcement Image Matching with Self-Termination
abstract
Deep reinforcement learning-based image matching sequentially searches only the promising regions in the reference image that match the query, leading to a significantly small number of steps compared to traditional methods. Since existing methods do not have any function to judge whether the target region has been successfully identified or not, they continue to search until the preset maximum number of search steps is reached. In this paper, we propose a deep image matching network that can terminate the matching process by itself. Our network is designed to have a halting module that identifies whether the current reference region matches the query based on the image features and the search history. The entire network is effectively trained end-to-end in a framework of deep reinforcement learning that incorporates a new loss function to evaluate the accuracy of the termination decision. Experimental results demonstrate that our method can achieve highly competitive or better matching accuracy with fewer search steps than the existing methods.
Onkar Krishna, Go Irie, Xiaomeng Wu, Akisato Kimura, Kunio Kashino
ICIP5
2021 BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
abstract
Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio representation learning approach. We propose learning general-purpose audio representation from a single audio segment without expecting relationships between different time segments of audio samples. To implement this principle, we introduce Bootstrap Your Own Latent (BYOL) for Audio (BYOL-A, pronounced “viola”), an audio self-supervised learning method based on BYOL for learning general-purpose audio representation. Unlike most previous audio self-supervised learning methods that rely on agreement of vicinity audio segments or disagreement of remote ones, BYOL-A creates contrasts in an augmented audio segment pair derived from a single audio segment. With a combination of normalization and augmentation techniques, BYOL-A achieves state-of-the-art results in various downstream tasks. Extensive ablation studies also clarified the contribution of each component and their combinations.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
IJCNN5
2021 Contrast enhancement based on discriminative co-occurrence statistics
Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
Multim. Tools Appl.4
2021 Reflectance-Guided Histogram Equalization and Comparametric Approximation
abstract
Existing image enhancement methods fall short of expectations because with them it is difficult to improve global and local image contrast simultaneously. To address this issue, we propose a histogram equalization-based method called RG-CACHE. It adapts to the data-dependent requirements of brightness enhancement and improves the visibility of details without losing the global contrast. RG-CACHE incorporates the spatial information provided by image context into density estimation for discriminative histogram equalization. To minimize the adverse effect of nonuniform illumination, we propose defining spatial information on the basis of image reflectance estimated with edge-preserving smoothing. RG-CACHE works particularly well for determining how the background brightness should be adaptively adjusted and for revealing useful image details hidden in the dark. To handle the loss of details due to the monotonicity of the intensity mapping function, we further propose a post-processing method to approximate RG-CACHE with a brightness transformation function corresponding to a parameterized camera response function. This method is called comparametric approximation. It takes into account a regression problem, in which the parameters of the camera response function are chosen so that the converted intensities are optimally matched to the image enhanced by RG-CACHE. Comparametric approximation is especially suitable for recovering useful image details that tend to be suppressed due to insufficient reflectance contrast.
Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
IEEE Trans. Circuits Syst. Video Technol.3
2020 Cascaded Transposed Long-Range Convolutions for Monocular Depth Estimation
Go Irie, Daiki Ikami, Takahito Kawanishi, Kunio Kashino
ACCV (3)4
2020 Adaptive Spotting: Deep Reinforcement Object Search in 3D Point Clouds
Onkar Krishna, Go Irie, Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
ACCV (3)5
2020 Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms
abstract
We propose a trilingual semantic embedding model that associates visual objects in images with segments of speech signals corresponding to spoken words in an unsupervised manner. Unlike the existing models, our model incorporates three different languages, namely, English, Hindi, and Japanese. To build the model, we used the existing English and Hindi datasets and collected a new corpus of Japanese speech captions. These spoken captions are spontaneous descriptions by individual speakers, rather than readings based on prepared transcripts. Therefore, we introduce a self-attention mechanism into the model to better map the spoken captions associated with the same image into the embedding space. We hope that the self-attention mechanism efficiently captures relationships between widely separated word-like segments. Experimental results show that the introduction of a third language improves the average performance in terms of cross-modal and cross-lingual retrieval accuracy, and that the self-attention mechanism added to the model works effectively.
Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David F. Harwath, James R. Glass
ICASSP4
2020 Reflectance-Guided, Contrast-Accumulated Histogram Equalization
abstract
Existing image enhancement methods fall short of expectations because with them it is difficult to improve global and local image contrast simultaneously. To address this problem, we propose a histogram equalization-based method that adapts to the data-dependent requirements of brightness enhancement and improves the visibility of details without losing the global contrast. This method incorporates the spatial information provided by image context in density estimation for discriminative histogram equalization. To minimize the adverse effect of non-uniform illumination, we propose defining spatial information on the basis of image reflectance estimated with edge preserving smoothing. Our method works particularly well for determining how the background brightness should be adaptively adjusted and for revealing useful image details hidden in the dark.
Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
ICASSP3
2020 Translating Adult's Focus of Attention to Elderly's
abstract
Predicting which part of a scene elderly people would pay attention to could be useful in assisting their daily activities, such as driving, walking, and searching. Many computational models for predicting focus of attention (FoA) have been developed. However, most of them focus on mimicking adult FoA and do not work well for predicting elderly's, due to age-related changes in human vision. Is it possible to leverage the prediction results made by an FoA model of general adults to accurately predict elderly's FoA, rather than training a new network from scratch? In this paper, we consider a novel problem of translating adult's FoA to elderly's and propose an approach based on deep image-to-image translation. Our model is trained by minimizing both Kullback-Leibler divergence and adversarial loss to approximate the joint probability distribution of adult and elderly FoA. Experiments on two datasets demonstrate that our model gives remarkable prediction accuracy.
Onkar Krishna, Go Irie, Takahito Kawanishi, Kunio Kashino, Kiyoharu Aizawa
ICPR4
2020 Unsupervised Co-Segmentation for Athlete Movements and Live Commentaries Using Crossmodal Temporal Proximity
abstract
Audio-visual co-segmentation is a task to extract segments and regions corresponding to specific events on unlabeled audio and video signals. It is particularly important to accomplish it in an unsupervised way, since it is generally very difficult to manually label all the objects and events appearing in audio-visual signals for supervised learning. Here, we propose to take advantage of the temporal proximity of corresponding audio and video entities included in the signals. For this purpose, we newly employ a guided attention scheme to this task to efficiently detect and utilize temporal co-occurrences of audio and video information. Experiments using a real TV broadcasts of sumo wrestling, a sport event, with live commentaries show that our model can automatically extract specific athlete movements and its spoken descriptions in an unsupervised manner.
Yasunori Ohishi, Kunio Kashino
ICPR3
2020 Total Whitening for Online Signature Verification Based on Deep Representation
abstract
In deep metric learning targeted at time series, the correlation between feature activations may be easily enlarged through highly nonlinear neural networks, leading to suboptimal embedding effectiveness. An effective solution to this problem is whitening. For example, in online signature verification, whitening can be derived for three individual Gaussian distributions, namely the distributions of local features at all temporal positions 1) for all signatures of all subjects, 2) for all signatures of each particular subject, and 3) for each particular signature of each particular subject. This study proposes a unified method called total whitening that integrates these individual Gaussians. Total whitening rectifies the layout of multiple individual Gaussians to resemble a standard normal distribution, improving the balance between intraclass invariance and interclass discriminative power. Experimental results demonstrate that total whitening achieves state-of-the-art accuracy when tested on online signature verification benchmarks.
Xiaomeng Wu, Akisato Kimura, Kunio Kashino, Seiichi Uchida
ICPR3
2020 Pair Expansion for Learning Multilingual Semantic Embeddings Using Disjoint Visually-Grounded Speech Audio Datasets
Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David F. Harwath, James R. Glass
INTERSPEECH4
2020 Harmonic Lowering for Accelerating Harmonic Convolution for Audio Signals
Hirotoshi Takeuchi, Kunio Kashino, Yasunori Ohishi, Hiroshi Saruwatari
INTERSPEECH2
2019 Delving Deep into Least Square Regression Model for Subspace Clustering
Masataka Yamaguchi, Go Irie, Takahito Kawanishi, Kunio Kashino
BMVC4
2019 Seeing through Sounds: Predicting Visual Semantic Segmentation Results from Multichannel Audio Signals
abstract
Sounds provide us with vast amounts of information about surrounding objects and can even remind us visual images of them. Is it possible to implement this noteworthy human ability on machines? In this paper, we study a new task that consists of predicting image recognition results in the form of semantic segmentation with given multichannel audio signals. Our approach uses a convolutional neural network that is designed to directly output semantic segmentation results by taking audio features as its inputs. A bilinear feature fusion scheme is incorporated that efficiently models underlying higher-order interactions between audio and visual sources. Experimental evaluations with both synthetic and real sound datasets show that our approach can recover the desired segmented images reasonably well.
Go Irie, Mirela Ostrek, Hirokazu Kameoka, Akisato Kimura, Takahito Kawanishi, Kunio Kashino
ICASSP7
2019 Learning Search Path for Region-level Image Matching
abstract
Finding a region of an image which matches to a query from a large number of candidates is a fundamental problem in image processing. The exhaustive nature of the sliding window approach has encouraged works that can reduce the run time by skipping unnecessary windows or pixels that do not play a substantial role in search results. However, such a pruning-based approach still needs to evaluate the non-ignorable number of candidates, which leads to a limited efficiency improvement. We propose an approach to learn efficient search paths from data. Our model is based on a CNN-LSTM architecture which is designed to sequentially determine a prospective location to be searched next based on the history of the locations attended. We propose a reinforcement learning algorithm to train the model in an end-to-end manner, which allows to jointly learn the search paths and deep image features for matching. These properties together significantly reduce the number of windows to be evaluated and makes it robust to background clutters. Our model gives remarkable matching accuracy with the reduced number of windows and run time on MNIST and FlickrLogos-32 datasets.
Onkar Krishna, Go Irie, Xiaomeng Wu, Takahito Kawanishi, Kunio Kashino
ICASSP5
2019 Prewarping Siamese Network: Learning Local Representations for Online Signature Verification
abstract
We propose a neural network-based framework for learning local representations of multivariate time series, and demonstrate its effectiveness for online signature verification. In contrast to related works that optimize a global distance objective, we incorporate a Siamese network into dynamic time warping (DTW), leading to a novel prewarping Siamese network (PSN) optimized with a local embedding loss. PSN learns a feature space that preserves the temporal location-wise distances of local structures. Local embedding, along with the alignment conditions of DTW, imposes a temporal consistency constraint on the sequence-level distance measure while achieving invariance as regards non-linear distortions. Validation on online signature verification datasets demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Seiichi Uchida, Kunio Kashino
ICASSP4
2019 Subspace Structure-Aware Spectral Clustering for Robust Subspace Clustering
abstract
Subspace clustering is the problem of partitioning data drawn from a union of multiple subspaces. The most popular subspace clustering framework in recent years is the graph clustering-based approach, which performs subspace clustering in two steps: graph construction and graph clustering. Although both steps are equally important for accurate clustering, the vast majority of work has focused on improving the graph construction step rather than the graph clustering step. In this paper, we propose a novel graph clustering framework for robust subspace clustering. By incorporating a geometry-aware term with the spectral clustering objective, we encourage our framework to be robust to noise and outliers in given affinity matrices. We also develop an efficient expectation-maximization-based algorithm for optimization. Through extensive experiments on four real-world datasets, we demonstrate that the proposed method outperforms existing methods.
Masataka Yamaguchi, Go Irie, Takahito Kawanishi, Kunio Kashino
ICCV4
2019 Deep Dynamic Time Warping: End-to-End Local Representation Learning for Online Signature Verification
abstract
Siamese networks have been shown to be successful in learning deep representations for multivariate time series verification. However, most related studies optimize a global distance objective and suffer from a low discriminative power due to the loss of temporal information. To address this issue, we propose an end-to-end, neural network-based framework for learning local representations of time series, and demonstrate its effectiveness for online signature verification. This framework optimizes a Siamese network with a local embedding loss, and learns a feature space that preserves the temporal location-wise distances between time series. To achieve invariance to non-linear temporal distortion, we propose building a dynamic time warping block on top of the Siamese network, which will greatly improve the accuracy for local correspondences across intra-personal variability. Validation with respect to online signature verification demonstrates the advantage of our framework over existing techniques that use either handcrafted or learned feature representations.
Xiaomeng Wu, Akisato Kimura, Brian Kenji Iwana, Seiichi Uchida, Kunio Kashino
ICDAR5
2019 Robust Learning for Deep Monocular Depth Estimation
abstract
Existing methods for deep monocular depth estimation are often trained with basic loss functions such as mean absolute error (MAE) or reverse Huber (BerHu). We revisit several basic loss functions to explore possibilities for improvement and show that the final depth estimation accuracy is dominated by pixels with small errors, which account for the vast majority. Based on this observation, we propose a new robust loss function that suppresses the contributions of pixels with higher errors by taking their square root. The loss function called the square root Huber (Ruber) is designed to be first-order differentiable on ℝ>0, so it can be directly applied to the end-to-end learning of a general type of neural network. Unlike the widely-used robust loss function called Huber, the Ruber loss function facilitates further refinement of the pixels with smaller errors by giving larger gradient values to the errors that are close to zero. Moreover, we show that the estimation accuracy can be further improved by introducing a second training step based on an edge-preserving loss. Experimental results with public indoor scene datasets demonstrate that our method outperforms major loss functions and can yield better accuracies than existing approaches in terms of root mean square error (RMSE).
Go Irie, Takahito Kawanishi, Kunio Kashino
ICIP3
2019 Understanding community structure in layered neural networks
Chihiro Watanabe, Kaoru Hiramatsu, Kunio Kashino
Neurocomputing3
2018 Generative Adversarial Image Synthesis With Decision Tree Latent Controller
abstract
This paper proposes the decision tree latent controller generative adversarial network (DTLC-GAN), an extension of a GAN that can learn hierarchically interpretable representations without relying on detailed supervision. To impose a hierarchical inclusion structure on latent variables, we incorporate a new architecture called the DTLC into the generator input. The DTLC has a multiple-layer tree structure in which the ON or OFF of the child node codes is controlled by the parent node codes. By using this architecture hierarchically, we can obtain the latent space in which the lower layer codes are selectively used depending on the higher layer ones. To make the latent codes capture salient semantic features of images in a hierarchically disentangled manner in the DTLC, we also propose a hierarchical conditional mutual information regularization and optimize it with a newly defined curriculum learning method that we propose as well. This makes it possible to discover hierarchically interpretable representations in a layer-by-layer manner on the basis of information gain by only using a single DTLC-GAN model. We evaluated the DTLC-GAN on various datasets, i.e., MNIST, CIFAR-10, Tiny ImageNet, 3D Faces, and CelebA, and confirmed that the DTLC-GAN can learn hierarchically interpretable representations with either unsupervised or weakly supervised settings. Furthermore, we applied the DTLC-GAN to image-retrieval tasks and showed its effectiveness in representation learning.
Takuhiro Kaneko, Kaoru Hiramatsu, Kunio Kashino
CVPR3
2018 Generating Sound Words from Audio Signals of Acoustic Events with Sequence-to-Sequence Model
abstract
Representing various sounds in language, such as sound words, or onomatopoeias, is not only useful as an auxiliary means for automatic speech recognition, but also essential in emerging fields such as natural human-machine communication, searching audio archives for acoustic events, and abnormality detection based on sounds. This paper proposes a novel method for sound word generation from audio signals. The method is based on an end-to-end, sequence-to-sequence framework to solve the audio segmentation problem to find an appropriate segment of audio signals along time that corresponds to a sequence of phonemes, and the ambiguity problem, where multiple words may correspond to the same sound, depending on the situations or listeners. Our tests show that the method worked efficiently and achieved a 2.8% mean phoneme error rate (MPER) and a 7.2% word error rate (WER) in a sound word generation task.
Shota Ikawa, Kunio Kashino
ICASSP2
2018 Statistical Phrase/Accent Command Estimation Algorithm Utilizing Linguistic Information
abstract
The importance of extracting non-linguistic information has been highlighted in a growing variety of applications of speech signal processing. Among the audio features carrying such information, fundamental frequency (F0) contours are considered primarily important. The Fujisaki model is a physical model that describes a F0contour with only a small number of parameters, namely, the timings and magnitudes of the phrase and accent commands, and a stochastic formulation and estimation algorithm have recently been proposed for it. However, the use of linguistic information has so far been limited, while it is known that accent commands are strongly related to linguistic information in many languages, and linguistic information could be obtained from the input audio signals by using speech recognition techniques. Against this background, this paper introduces a novel F0command parameter estimation method that incorporates linguistic information with the stochastic framework. Experiments using real speech data show that when linguistic information is appropriately utilized, the estimation accuracy of accent command parameters is improved by 43% under the proposed criteria.
Ryotaro Sato, Kunio Kashino
ICASSP2
2018 Query Expansion with Diffusion On Mutual Rank Graphs
abstract
In query expansion for object retrieval, there is substantial danger of query drift, where irrelevant information is inferred from pseudo-relevant images to enrich the query. To address this issue, we propose a query expansion method from the viewpoint of diffusion. It explores the structure of highly ranked images in a topological space, assuming that false positives reside on different manifolds from the query. For this purpose, a mutual rank graph is defined on pseudo-relevant images, and their distribution is learned by diffusing their query similarities through the graph. The relevance of a database image can thus be obtained by marginalizing over the learned distribution. The mutual rank graph accounts for varying local density in the image space, leading to great robustness as regards query drift and high generalization ability. The proposed method experimentally shows a consistent boost in the performance of object retrieval with handcrafted features on standard benchmarks.
Xiaomeng Wu, Go Irie, Kaoru Hiramatsu, Kunio Kashino
ICASSP4
2018 Weighted Generalized Mean Pooling for Deep Image Retrieval
abstract
Spatial pooling over convolutional activations (e.g., max pooling or sum pooling) has been shown to be successful in learning deep representations for image retrieval. However, most pooling techniques assume that every activation is equally important, and as a result they suffer from the presence of uninformative image regions that play a negative role as regards matching or lead to the confusion of particular visual instances. To address this issue, we propose a trainable building block that steers pooling to local information important to the task at hand. The method formulates pooling as a weighted generalized mean (wGeM), in which weights are learned on activations, reflecting the discriminative power of each activation in image matching. Embedding wGeM in a deep network improves image representation and boosts retrieval performance on standard benchmarks. wGeM does not require any bounding box annotations, but instead learns the latent probabilities of activations from scratch. It even goes beyond objectness, and learns to look at important visual details rather than the whole region of the object of interest.
Xiaomeng Wu, Go Irie, Kaoru Hiramatsu, Kunio Kashino
ICIP4
2018 Label Propagation with Ensemble of Pairwise Geometric Relations: Towards Robust Large-Scale Retrieval of Object Instances
Xiaomeng Wu, Kaoru Hiramatsu, Kunio Kashino
Int. J. Comput. Vis.3
2018 Modular representation of layered neural networks
Chihiro Watanabe, Kaoru Hiramatsu, Kunio Kashino
Neural Networks3
2017 Generative Attribute Controller with Conditional Filtered Generative Adversarial Networks
abstract
We present a generative attribute controller (GAC), a novel functionality for generating or editing an image while intuitively controlling large variations of an attribute. This controller is based on a novel generative model called the conditional filtered generative adversarial network (CFGAN), which is an extension of the conventional conditional GAN (CGAN) that incorporates a filtering architecture into the generator input. Unlike the conventional CGAN, which represents an attribute directly using an observable variable (e.g., the binary indicator of attribute presence) so its controllability is restricted to attribute labeling (e.g., restricted to an ON or OFF control), the CFGAN has a filtering architecture that associates an attribute with a multi-dimensional latent variable, enabling latent variations of the attribute to be represented. We also define the filtering architecture and training scheme considering controllability, enabling the variations of the attribute to be intuitively controlled using typical controllers (radio buttons and slide bars). We evaluated our CFGAN on MNIST, CUB, and CelebA datasets and show that it enables large variations of an attribute to be not only represented but also intuitively controlled while retaining identity. We also show that the learned latent space has enough expressive power to conduct attribute transfer and attribute-based image retrieval.
Takuhiro Kaneko, Kaoru Hiramatsu, Kunio Kashino
CVPR3
2017 Recursive Extraction of Modular Structure from Layered Neural Networks Using Variational Bayes Method
Chihiro Watanabe, Kaoru Hiramatsu, Kunio Kashino
DS3
2017 Generative adversarial network-based postfilter for statistical parametric speech synthesis
abstract
We propose a postfilter based on a generative adversarial network (GAN) to compensate for the differences between natural speech and speech synthesized by statistical parametric speech synthesis. In particular, we focus on the differences caused by over-smoothing, which makes the sounds muffled. Over-smoothing occurs in the time and frequency directions and is highly correlated in both directions, and conventional methods based on heuristics are too limited to cover all the factors (e.g., global variance was designed only to recover the dynamic range). To solve this problem, we focus on “spectral texture”, i.e., the details of the time-frequency representation, and propose a learning-based postfilter that captures the structures directly from the data. To estimate the true distribution, we utilize a GAN composed of a generator and a discriminator. This optimizes the generator to produce samples imitating the dataset according to the adversarial discriminator. This adversarial process encourages the generator to fit the true data distribution, i.e., to generate realistic spectral texture. Objective evaluation of experimental results shows that the GAN-based postfilter can compensate for detailed spectral structures including modulation spectrum, and subjective evaluation shows that its generated speech is comparable to natural speech.
Takuhiro Kaneko, Hirokazu Kameoka, Nobukatsu Hojo, Yusuke Ijima, Kaoru Hiramatsu, Kunio Kashino
ICASSP6
2017 Deep salience map guided arbitrary direction scene text recognition
abstract
Irregular scene text such as curved, rotated or perspective texts commonly appear in natural scene images due to different camera view points, special design purposes etc. In this work, we propose a text salience map guided model to recognize these arbitrary direction scene texts. We train a deep Fully Convolutional Network (FCN) to calculate the precise salience map for texts. Then we estimate the positions and rotations of the text and utilize this information to guide the generation of CNN sequence features. Finally the sequence is recognized with a Recurrent Neural Network (RNN) model. Experiments on various public datasets show that the proposed approach is robust to different distortions and performs superior or comparable to the state-of-the-art techniques.
Xinhao Liu 0001, Takahito Kawanishi, Xiaomeng Wu, Kaoru Hiramatsu, Kunio Kashino
ICASSP5
2017 Fast algorithm for statistical phrase/accent command estimation based on generative model incorporating spectral features
abstract
An important challenge in speech processing involves extracting non-linguistic information from a fundamental frequency (F0) contour of speech. We propose a fast algorithm for estimating the model parameters of the Fujisaki model, namely, the timings and magnitudes of the phrase and accent commands. Although a powerful parameter estimation framework based on a stochastic counterpart of the Fujisaki model has recently been proposed, it still had room for improvement in terms of both computational efficiency and parameter estimation accuracy. This paper describes our two contributions. First, we propose a hard expectation-maximization (EM) algorithm for parameter inference where the E step of the conventional EM algorithm is replaced with a point estimation procedure to accelerate the estimation process. Second, to improve the parameter estimation accuracy, we add a generative process of a spectral feature sequence to the generative model. This makes it possible to use linguistic or phonological information as an additional clue to estimate the timings of the accent commands. The experiments confirmed that the present algorithm was approximately 16 times faster and estimated parameters about 3% more accurately than the conventional algorithm.
Ryotaro Sato, Hirokazu Kameoka, Kunio Kashino
ICASSP3
2017 Edited film alignment via selective Hough transform and accurate template matching
abstract
Edited film alignment is the post-production process of finding small parts of unedited footage that temporally and spatially match an edited film. The huge amount of data to be processed makes significant downsampling of the videos essential in real-life applications. Simultaneously, professional users demand that the task be achieved with frame and pixel-level accuracy. We propose a novel selective Hough transform (SHT) and an accurate template matching method to address the difficult trade-off between accuracy and scalability. For robust temporal alignment, SHT investigates the selectivity of frame-level similarities and advantageously reduces the weights of mismatches. The template matching method encompasses spatial Hough transform and sum of squared differences (SSD) minimization. SSD is efficiently approximated by exploiting the second-order derivative of image intensity. Experiments conducted on real-world data show the superiority of our methods.
Xiaomeng Wu, Takahito Kawanishi, Minoru Mori, Kaoru Hiramatsu, Kunio Kashino
ICASSP5
2017 Contrast-accumulated histogram equalization for image enhancement
abstract
Among image enhancement methods, histogram equalization (HE) has received the most attention because of its intuitive implementation quality, high efficiency, and the monotonicity of its intensity mapping function. However, HE is indiscriminate and overemphasizes the contrast around intensities with large pixel populations but little visual importance. To address this issue, we propose an HE-based method that adaptively controls the contrast gain according to the potential visual importance of intensities and pixels. Observing that in natural scenes image details are usually hidden in darker regions that have noticeable local differences, we formulate the potential visual importance on the basis of the multi-resolution, dark-pass filtered gradients in the image. Experiments show that our method is highly discriminating in terms of noises and trivial image gradients, and it guarantees great global contrast preservation.
Xiaomeng Wu, Xinhao Liu 0001, Kaoru Hiramatsu, Kunio Kashino
ICIP4
2017 Sequence-to-Sequence Voice Conversion with Similarity Metric Learned Using Generative Adversarial Networks
Takuhiro Kaneko, Hirokazu Kameoka, Kaoru Hiramatsu, Kunio Kashino
INTERSPEECH4
2017 Visualizing Video Sounds With Sound Word Animation to Enrich User Experience
abstract
Sound information in videos plays an important role in shaping the user experience. When sound is not accessible in videos, text captions can provide sound information. However, conventional text captions are not very expressive for nonverbal sounds because they are designed to visualize speech sounds. Here, we present a framework to automatically transform nonverbal video sounds into animated sound words and position them near the sound source objects in the video for visualization. This provides natural visual representation of nonverbal sounds with rich information about the sound category and dynamics. To evaluate how the animated sound words generated by our framework affect the user experience, we implemented an experimental system and conducted a user study involving over 300 participants from an online crowdsourcing service. The results of the user study show that the animated sound words can effectively and naturally visualize the dynamics of sound while clarifying the position of the sound source as well as contribute to making video-watching more enjoyable and increasing the visual impact of videos.
Hidehisa Nagano, Kunio Kashino, Takeo Igarashi
IEEE Trans. Multim.3
2016 Scene text recognition with high performance CNN classifier and efficient word inference
abstract
The recognition of text in natural scene images is a practical yet challenging task due to the large variations in backgrounds, textures, fonts, and illumination conditions. In this paper, we propose a highly accurate character recognition model by utilizing the representational power of a specially designed Convolutional Neural Network (CNN). Based on the recognition model, we also develop an efficient post processing approach for error correction and hypothesis re-verification. Character and word image recognition experiments on two public datasets, namely the ICDAR 2003 Robust Reading dataset and the Street View Text (SVT) dataset both show that the proposed approach provides superior or comparable results to the state-of-the-art techniques.
Xinhao Liu 0001, Takahito Kawanishi, Xiaomeng Wu, Kunio Kashino
ICASSP4
2016 Scene text recognition with CNN classifier and WFST-based word labeling
abstract
Natural scene text recognition has proved to be challenging due to the unconstrained wild conditions. In this paper, to solve this problem we propose a method which first detects and recognizes characters by utilizing the high performance Convolutional Neural Network (CNN). Then for post-processing, inspired by its success in speech recognition, we employ the efficient and flexible Weight Finite State Transducer (WFST) based word labeling model for incorporation with a lexicon or high order language model. In the experiments we show that the proposed approach can correctly and robustly recognize the text in the scene images and the results for serveral public datasets (ICDAR 2003, SVT and IIIT5K) show comparable or superior performance to the state-of-the-art algorithms.
Xinhao Liu 0001, Takahito Kawanishi, Xiaomeng Wu, Kunio Kashino
ICPR4
2016 Adaptive Visual Feedback Generation for Facial Expression Improvement with Multi-task Deep Neural Networks
abstract
While many studies in computer vision and pattern recognition have been actively conducted to recognize people's current states, few studies have tackled the problem of generating feedback on how people can improve their states, although there are many real-world applications such as in sports, education, and health care. In particular, it has been challenging to develop such a system that can adaptively generate feedback for real-world situations, namely various input and target states, since it requires formulating various rules of feedback to do so. We propose a learning-based method to solve this problem. If we can obtain a large amount of feedback annotations, it is possible to explicitly learn the rules, but it is difficult to do so due to the subjective nature of the task. To mitigate this problem, our method implicitly learns the rules from training data consisting of input images, key-point annotations, and state annotations that do not require professional knowledge in feedback. Given such training data, we first learn a multi-task deep neural network with state recognition and key-point localization. Then, we apply a novel propagation method for extracting feedback information from the network. We evaluated our method in a facial expression improvement task using real-world data and clarified its characteristics and effectiveness.
Takuhiro Kaneko, Kaoru Hiramatsu, Kunio Kashino
ACM Multimedia3
2016 Unsupervised categorical shape reconstruction through manifolds
abstract
We tackle a new challenge of unsupervised categorical 3D reconstruction using images from different instances without any manually assigned key point correspondence. We take advantage of the observations that objects are generally captured from ground-level viewpoints, and that images transform smoothly between some characteristic viewpoints despite the difference among instances. To estimate 3D models of a category, we first embed the images into a manifold to cluster viewpoints into degenerate and intermediate viewpoints. Then, we select adequate triplets of images that capture similarly shaped objects by aligning manifolds of the different viewpoint clusters. To establish correspondence between the images, we prepare finer clusters from the original manifold, and obtain a common set of features that are at locations consistent with the average flow between neighboring clusters. Under the assumption of orthographical projection, camera parameters are estimated using the correspondence among views. Finally, visual hull for each image triplet is calculated using the silhouettes. Our method is capable of automatically reconstructing an approximate categorical model even without supervision.
Kent Fujiwara, Minoru Mori, Kunio Kashino
WACV3
2015 Robust Spatial Matching as Ensemble of Weak Geometric Relations
Xiaomeng Wu, Kunio Kashino
BMVC2
2015 Trademark Image Retrieval Using Inverse Total Feature Frequency and Multiple Detectors
Minoru Mori, Xiaomeng Wu, Kunio Kashino
CAIP (1)3
2015 A fast audio search method based on skipping irrelevant signals by similarity upper-bound calculation
abstract
In this paper, we describe an approach to accelerate fingerprint techniques by skipping the search for irrelevant sections of the signal and demonstrate its application to the divide and locate (DAL) audio fingerprint method. The search result for the applied method, DAL3, is the same as that of DAL mathematically. Experimental results show that DAL3 can reduce the computational cost of DAL to approximately 25% for the task of music signal retrieval.
Hidehisa Nagano, Ryo Mukai, Takayuki Kurozumi, Kunio Kashino
ICASSP4
2015 Adaptive Dither Voting for Robust Spatial Verification
abstract
Hough voting in a geometric transformation space allows us to realize spatial verification, but remains sensitive to feature detection errors because of the inflexible quantization of single feature correspondences. To handle this problem, we propose a new method, called adaptive dither voting, for robust spatial verification. For each correspondence, instead of hard-mapping it to a single transformation, the method augments its description by using multiple dithered transformations that are deterministically generated by the other correspondences. The method reduces the probability of losing correspondences during transformation quantization, and provides high robustness as regards mismatches by imposing three geometric constraints on the dithering process. We also propose exploiting the non-uniformity of a Hough histogram as the spatial similarity to handle multiple matching surfaces. Extensive experiments conducted on four datasets show the superiority of our method. The method outperforms its state-of-the-art counterparts in both accuracy and scalability, especially when it comes to the retrieval of small, rotated objects.
Xiaomeng Wu, Kunio Kashino
ICCV2
2015 Visualizing video sounds with sound word animation
abstract
Text captions are important means to provide sound information in videos when the sound is not accessible. However, conventional text captions are far less expressive for non-verbal sounds since they are designed to visualize speech sound. To address this problem, we propose a method for automatically transforming non-verbal video sounds to animated sound words, and positioning them near the sound source objects in the video for visualization. This provides natural visual representation of non-verbal sounds with rich information about the sound category and dynamics. We conducted a user study with over 300 participants using an online crowdsourcing service. The results showed that animated sound words could not only effectively and naturally visualize the dynamics of sound while clarify the position of the sound source, but also contribute to making video watching more enjoyable and increasing the visual impact of the video.
Hidehisa Nagano, Kunio Kashino, Takeo Igarashi
ICME3
2015 Data-driven taxonomy forest for fine-grained image categorization
abstract
Fine-grained image categorization must handle huge cross-class ambiguities and a large number of classes. Inspired by the success of rigid hierarchical classification, we propose a new flexible hierarchical classification method, called a data-driven taxonomy forest. It constructs a multitude of taxonomies, each of which converts a complex multi-class problem to a more easily tractable path-finding problem. We demonstrate how a stochastic representation of local classification hypotheses incorporated in multiple taxonomies deals skillfully with error propagation and over-fitting. Various strategies for instance space decomposition are investigated from the viewpoint of taxonomy complexity. We comprehensively evaluate our data-driven taxonomy forest using Oxford Flower 102 and Oxford Pet benchmarks and show its superiority in effectiveness and generality to rigid hierarchical classification in fine-grained image categorization tasks.
Xiaomeng Wu, Minoru Mori, Kunio Kashino
ICME3
2015 Visual Attention Driven by Auditory Cues - Selecting Visual Features in Synchronization with Attracting Auditory Events
Jiro Nakajima, Akisato Kimura, Akihiro Sugimoto, Kunio Kashino
MMM (2)4
2015 Interest point selection by topology coherence for multi-query image retrieval
Xiaomeng Wu, Kunio Kashino
Multim. Tools Appl.2
2015 Generative Modeling of Voice Fundamental Frequency Contours
abstract
This paper introduces a generative model of voice fundamental frequency (F0) contours that allows us to extract prosodic features from raw speech data. The present F0contour model is formulated by translating the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration, into a probabilistic model described as a discrete-time stochastic process. There are two motivations behind this formulation. One is to derive a general parameter estimation framework for the Fujisaki model that allows the introduction of powerful statistical methods. The other is to construct an automatically trainable version of the Fujisaki model that we can incorporate into statistical-model-based text-to-speech synthesizers in such a way that the Fujisaki-model parameters can be learned from a speech corpus in a unified manner. It could also be useful for other speech applications such as emotion recognition, speaker identification, speech conversion and dialogue systems, in which prosodic information plays a significant role. We quantitatively evaluated the performance of the proposed Fujisaki model parameter extractor using real speech data. Experimental results revealed that our method was superior to a state-of-the-art Fujisaki model parameter extractor.
Hirokazu Kameoka, Kota Yoshizato, Tatsuma Ishihara, Kento Kadowaki, Yasunori Ohishi, Kunio Kashino
IEEE ACM Trans. Audio Speech Lang. Process.6
2015 Second-Order Configuration of Local Features for Geometrically Stable Image Matching and Retrieval
abstract
Local features offer high repeatability, which supports efficient matching between images, but they do not provide sufficient discriminative power. Imposing a geometric coherence constraint on local features improves the discriminative power but makes the matching sensitive to anisotropic transformations. We propose a novel feature representation approach to solve the latter problem. Each image is abstracted by a set of tuples of local features. We revisit affine shape adaptation and extend its conclusion to characterize the geometrically stable feature of each tuple. The representation thus provides higher repeatability with anisotropic scaling and shearing than found in previous research. We develop a simple matching model by voting in the geometrically stable feature space, where votes arise from tuple correspondences. To make the required index space linear as regards the number of features, we propose a second approach called a centrality-sensitive pyramid to select potentially meaningful tuples of local features on the basis of their spatial neighborhood information. It achieves faster neighborhood association and has a greater robustness to errors in interest point detection and description. We comprehensively evaluated our approach using Flickr Logos 32, Holiday, Oxford Buildings, and Flickr 100 K benchmarks. Extensive experiments and comparisons with advanced approaches demonstrate the superiority of our approach in image retrieval tasks.
Xiaomeng Wu, Kunio Kashino
IEEE Trans. Circuits Syst. Video Technol.2
2014 Tri-Map Self-Validation Based on Least Gibbs Energy for Foreground Segmentation
Xiaomeng Wu, Kunio Kashino
BMVC2
2014 Mondrian hidden Markov model for music signal processing
abstract
This paper discusses a new extension of hidden Markov models that can capture clusters embedded in transitions between the hidden states. In our model, the state-transition matrices are viewed as representations of relational data reflecting a network structure between the hidden states. We specifically present a nonparametric Bayesian approach to the proposed state-space model whose network structure is represented by a Mondrian Process-based relational model. We show an application of the proposed model to music signal analysis through some experimental results.
Masahiro Nakano, Yasunori Ohishi, Hirokazu Kameoka, Ryo Mukai, Kunio Kashino
ICASSP5
2014 Mixture of Gaussian process experts for predicting sung melodic contour with expressive dynamic fluctuations
abstract
We present a generative model for predicting the sung melodic contour, i.e., F0contour, with expressive dynamic fluctuations, such as vibrato and portamento, for a given musical score. Although several studies have attempted to characterize such fluctuations, no systematic method has been developed for generating the F0contour with them in connection with musical notes. In our model, the relationship between a musical note sequence and F0contour is directly learned by a mixture of Gaussian process experts. This approach allows us to automatically characterize the fluctuations by utilizing the kernel function for each Gaussian process expert and predict the F0contour for an arbitrary musical note sequence. Experimental results show that our model can better predict the F0contour than a baseline method can. Additionally, we discuss the effective musical contexts and the amount of training data for the prediction.
Yasunori Ohishi, Daichi Mochihashi, Hirokazu Kameoka, Kunio Kashino
ICASSP4
2014 Image retrieval based on spatial context with Relaxed Gabriel Graph pyramid
abstract
Imposing the coherence of the spatial context on local features is becoming a necessity for object retrieval and recognition. Motivated by the success of proximity graphs in topological decomposition, clustering, and gradient estimation, we introduce a variation on and a generalization of Delaunay Triangulation, called a Relaxed Gabriel Graph (RGG), as the apex of spatial neighborhood association and design a Centrality-Sensitive Pyramid (CSP) model for hierarchical spatial context modeling. RGG is parameterized, and so allows the tuning of various applications and datasets. CSP achieves better neighborhood association and is more robust as regards feature description error than other related work. Our method is evaluated on Flickr Logos 32, Holiday, and Oxford Buildings benchmarks. Experimental results and comparisons demonstrate the superiority of our method in an image retrieval scenario.
Xiaomeng Wu, Kunio Kashino
ICASSP2
2014 Experimental Evaluation of Chromostereopsis with Varying Center Wavelength and FWHM of Spectral Power Distribution
Masaru Tsuchida, Kunio Kashino, Junji Yamato
ICISP2
2014 Video Content Detection with Single Frame Level Accuracy Using Dynamic Thresholding Technique
abstract
This paper proposes a video retrieval method that detects frame sections that correspond to shots in a query (video segment) with single frame level accuracy. The method adopts the coarse-to-fine strategy to decrease the processing time and the memory consumption, dynamic threshold with initial ranges for small segments is proposed to detect the exact beginning and end of each corresponding frame section to each shot in a query. Experiments on real videos show that our method can achieve accurate video detection with exact frame position while reducing processing time and memory consumption.
Minoru Mori, Takayuki Kurozumi, Hidehisa Nagano, Kunio Kashino
ICPR4
2014 Spatial People Density Estimation from Multiple Viewpoints by Memory Based Regression
abstract
Crowd analysis using cameras has attracted much attention for public safety and marketing. Among techniques of the crowd analysis, we focus on spatial people density estimation which estimates the number of people for each small area in a floor region. However, spatial people density cannot be estimated accurately for an area far from the camera because of the occlusion by people in a closer area. Therefore, we propose a method using a memory based regression method with images captured from cameras from multiple viewpoints. This method is realized by looking up a table that consists of correspondences between people density maps and crowd appearances. Since the crowd appearances include situations where various occlusions occur, an estimation robust to occlusion should be realized. In an experiment, we examined the effectiveness of the proposed method.
Yoshimune Tabuchi, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takayuki Kurozumi, Kunio Kashino
ICPR7
2014 Image Retrieval Based on Anisotropic Scaling and Shearing Invariant Geometric Coherence
abstract
Imposing a spatial coherence constraint on image matching is becoming a necessity for local feature based object retrieval. We tackle the affine invariance problem of the prior spatial coherence model and propose a novel approach for geometrically stable image retrieval. Compared with related studies focusing simply on translation, rotation, and isotropic scaling, our approach can deal with more significant transformations including anisotropic scaling and shearing. Our contribution consists of revisiting the first-order affine adaptation approach and extending its application to represent the geometric coherence of a second-order local feature structure. We comprehensively evaluated our approach using Flickr Logos 32, Holiday, and Oxford Buildings benchmarks. Extensive experimentation and comparisons with state-of-the-art spatial coherence models demonstrate the superiority of our approach in image retrieval tasks.
Xiaomeng Wu, Kunio Kashino
ICPR2
2014 BM25 With Exponential IDF for Instance Search
abstract
This paper deals with a novel concept of an exponential IDF in the BM25 formulation and compares the search accuracy with that of the BM25 with the original IDF in a content-based video retrieval (CBVR) task. Our video retrieval method is based on a bag of keypoints (local visual features) and the exponential IDF estimates the keypoint importance weights more accurately than the original IDF. The exponential IDF is capable of suppressing the keypoints from frequently occurring background objects in videos, and we found that this effect is essential for achieving improved search accuracy in CBVR. Our proposed method is especially designed to tackle instance video search, one of the CBVR tasks, and we demonstrate its effectiveness in significantly enhancing the instance search accuracy using the TRECVID2012 video retrieval dataset.
Masaya Murata, Hidehisa Nagano, Ryo Mukai, Kunio Kashino, Shin'ichi Satoh 0001
IEEE Trans. Multim.4
2013 Bayesian semi-supervised audio event transcription based on Markov indian buffet process
abstract
We present a novel generative model for audio event transcription that recognizes “events” on audio signals including multiple kinds of overlapping sounds. In the proposed model, firstly, the overlapping audio events are modeled based on nonnegative matrix factorization into which Bayesian nonparametric approaches: the Markov Indian buffet process and the Chinese restaurant process, are incorporated. This approach allows us to automatically transcribe the events while avoiding the model selection problem by assuming a countably infinite number of possible audio events in the input signal. Then, Bayesian logistic regression annotates the audio frames with the multiple event labels in a semi-supervised learning setup. Experimental results show that our model can better annotate an audio signal in comparison with a baseline method. Additionally, we verify that our infinite generative model is also able to detect unknown audio events that are not included in the training data.
Yasunori Ohishi, Daichi Mochihashi, Tomoko Matsui, Masahiro Nakano, Hirokazu Kameoka, Tomonori Izumitani, Kunio Kashino
ICASSP7
2013 Generative modeling of speech F0 contours
abstract
This paper introduces our ongoing work on generative modeling of speech fundamental frequency (F0) contours for estimating prosodic features from raw speech data. The present F0 contour model is formulated by translating the Fujisaki model, a well-founded mathematical model representing the control mechanism of vocal fold vibration, into a probabilistic model described as a discrete-time stochastic process. The motivation behind this formulation is two fold. One is to derive a general parameter estimation framework for the Fujisaki model, allowing for the introduction of powerful statistical methods. The other is to construct an automatically trainable version of the Fujisaki model so that in future it can be used to develop a statistical speaking style conversion system or incorporated into existing text-to-speech synthesis systems to improve the naturalness and intelligibility of computer-generated speech. We also briefly introduce a generative model of F0 contours of singing voice developed under the same spirit. Index Terms: speech F0 contour, Fujisaki model, generative model, hidden Markov model, EM algorithm
Hirokazu Kameoka, Kota Yoshizato, Tatsuma Ishihara, Yasunori Ohishi, Kunio Kashino, Shigeki Sagayama
INTERSPEECH5
2012 Constrained and regularized variants of non-negative matrix factorization incorporating music-specific constraints
abstract
Music spectrograms typically have many structural regularities that can be exploited to help solve the problem of decomposing a given spectrogram into distinct musically meaningful components. In this paper, we introduce new variants of the non-negative matrix factorization concept that incorporate music-specific constraints.
Hirokazu Kameoka, Masahiro Nakano, Kazuki Ochiai, Yutaka Imoto, Kunio Kashino, Shigeki Sagayama
ICASSP5
2012 Bayesian nonparametric music parser
abstract
This paper proposes a novel representation of music that can be used for similarity-based music information retrieval, and also presents a method that converts an input polyphonic audio signal to the proposed representation. The representation involves a 2-dimensional tree structure, where each node encodes the musical note and the dimensions correspond to the time and simultaneous multiple notes, respectively. Since the temporal structure and the synchrony of simultaneous events are both essential in music, our representation reflects them explicitly. In the conventional approaches to music representation from audio, note extraction is usually performed prior to structure analysis, but accurate note extraction has been a difficult task. In the proposed method, note extraction and structure estimation is performed simultaneously and thus the optimal solution is obtained with a unified inference procedure. That is, we propose an extended 2-dimensional infinite probabilistic context-free grammar and a sparse factor model for spectrogram analysis. An efficient inference algorithm, based on Markov chain Monte Carlo sampling and dynamic programming, is presented. The experimental results show the effectiveness of the proposed approach.
Masahiro Nakano, Yasunori Ohishi, Hirokazu Kameoka, Ryo Mukai, Kunio Kashino
ICASSP5
2012 A Stochastic Model of Singing Voice F0 Contours for Characterizing Expressive Dynamic Components
abstract
We present a novel stochastic model of singing voice fundamental frequency (F0) contours for characterizing expressive dynamic components, such as vibrato and portamento. Although dynamic components can be important features for any singing voice applications, modeling and extracting these components from a raw F0 contour have yet to be accomplished. Therefore, we describe a process for generating dynamic components explicitly and represent the process as a stochastic model. Then we develop an algorithm for estimating the model parameters based on statistical techniques. Experimental results show that our method successfully extracts the expressive components from raw F0 contours.
Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Kunio Kashino
INTERSPEECH4
2011 Automatic video annotation via Hierarchical Topic Trajectory Model considering cross-modal correlations
abstract
We propose a new statistical model, named Hierarchical Topic Trajectory Model (HTTM), for acquiring a dynamically changing topic model that represents the relationship between video frames and associated text labels. Model parameter estimation, annotation and retrieval can be executed within a unified framework with a few computation. It is also easy to add new modals such as audio signal and geotags. Preliminary experiments on video annotation task with manually annotated video dataset indicate that our proposed method can improve the annotation accuracy.
Takuho Nakano, Akisato Kimura, Hirokazu Kameoka, Shigeki Miyabe, Shigeki Sagayama, Nobutaka Ono, Kunio Kashino, Takuya Nishimoto
ICASSP7
2010 Statistical modeling of F0 dynamics in singing voices based on Gaussian processes with multiple oscillation bases
abstract
We present a novel statistical model for dynamics of various singing behaviors, such as vibrato and overshoot, in a fundamental frequency (F0) contour. These dynamics are the important cues for perceiving individuality of a singer, and can be a useful measure for various applications, such as singing skill evaluation and singing voice synthesis. While most previous studies have modeled the dynamics using a second-order linear system, the automatic and accurate estimation of model parameters has yet to be accomplished. In this paper, we first develop a complete stochastic representation of the second-order system with Gaussian processes from parametric discretization, and propose a complete, efficient scheme for parameter estimation using the Expectation-Maximization (EM) algorithm. Experimental results show that the proposed method can decompose an F0 contour into a musical component and a dynamics component. Finally, we discuss estimating singing styles from the model parameters for each singer.
Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Hidehisa Nagano, Kunio Kashino
INTERSPEECH5
2010 Fast Template Matching Based on Normalized Cross Correlation Using Adaptive Block Partitioning and Initial Threshold Estimation
abstract
This paper proposes a fast template matching method based on normalized cross correlation (NCC). NCC is more robust against image variations such as illumination changes than the widely-used sum of absolute difference (SAD). A problem with NCC has been its high computation cost. To deal with this problem, we use adaptive block partitioning and initial threshold estimation to extend the multilevel successive elimination algorithm. Adaptive block partitioning provides efficient sub-block partitioning and tighter boundaries. Initial threshold estimation yields a larger boundary threshold. They greatly suppress the number of search points at an earlier level from the beginning of search. The proposed method is exhaustive and robust with respect to template position and size. Experiments show that our method is up to 400 times faster than the brute force method and is significantly faster than conventional methods.
Minoru Mori, Kunio Kashino
ISM2
2009 Complex NMF: A new sparse representation for acoustic signals
abstract
This paper presents a new sparse representation for acoustic signals which is based on a mixing model defined in the complex-spectrum domain (where additivity holds), and allows us to extract recurrent patterns of magnitude spectra that underlie observed complex spectra and the phase estimates of constituent signals. An efficient iterative algorithm is derived, which reduces to the multiplicative update algorithm for non-negative matrix factorization developed by Lee under a particular condition.
Hirokazu Kameoka, Nobutaka Ono, Kunio Kashino, Shigeki Sagayama
ICASSP3
2009 Composite Autoregressive System for Sparse Source-filter Representation of speech
abstract
This paper presents a new generative model for speech signals called a ldquocomposite autoregressive systemrdquo. This model consists of a composite dictionary incorporating a set of the power spectral densities (PSDs) of excitation sources and a set of all-pole filters where the gain of each pair of excitation and filter elements is allowed to vary over time. We use this model to develop a computationally efficient scheme for generating a sparse mixture representation of speech based on the Expectation-Maximization algorithm. The algorithm iteratively updates the excitation PSDs and the gains through the update formulae, which reduce under a particular condition to the multiplicative update rule for non-negative matrix factorization with the Itakura-Saito distance criterion, and the all-pole parameters using the Levinson-Durbin algorithm.
Hirokazu Kameoka, Kunio Kashino
ISCAS2
2008 A background music detection method based on robust feature extraction
abstract
We propose a music segment detection method for audio signals. Unlike many existing methods, ours specifically focuses on a background-music detection task, that is, detecting music used in background of main sounds. This task is important because music is almost always overlapped by speech or other environmental sounds in visual materials such as TV programs. Our method consists of feature extraction, dimension reduction, and statistical discrimination steps. For each step, we analyzed a set of methods to maximize the detection accuracy. With a simple post processing step, we achieved a framewise error rate as low as 8 % even when the mixed speech was louder than the target music by 10dB.
Tomonori Izumitani, Ryo Mukai, Kunio Kashino
ICASSP3
2008 A stochastic model of selective visual attention with a dynamic Bayesian network
abstract
Recent studies in signal detection theory suggest that the human responses to the stimuli on a visual display are nondeterministic. People may attend to different locations on the same visual input at the same time. To predict the likelihood of where humans typically focus on a video scene, we propose a new stochastic model of visual attention by introducing a dynamic Bayesian network. Our model simulates and combines the visual saliency response and the cognitive state of a person to estimate the most probable attended regions. Experimental results have demonstrated that our model performs significantly better in predicting human visual attention compared to the previous deterministic model.
Derek Pang, Akisato Kimura, Tatsuto Takeuchi, Junji Yamato, Kunio Kashino
ICME5
2008 Dynamic Markov random fields for stochastic modeling of visual attention
abstract
This report proposes a new stochastic model of visual attention to predict the likelihood of where humans typically focus on a video scene. The proposed model is composed of a dynamic Bayesian network that simulates and combines a person’s visual saliency response and eye movement patterns to estimate the most probable regions of attention. Dynamic Markov random field (MRF) models are newly introduced to include spatiotemporal relationships of visual saliency responses. Experimental results have revealed that the propose model outperforms the previous deterministic model and the stochastic model without dynamic MRF in predicting human visual attention.
Akisato Kimura, Derek Pang, Tatsuto Takeuchi, Junji Yamato, Kunio Kashino
ICPR5
2008 Parameter estimation method of F0 control model for singing voices
abstract
In this paper, we propose a novel representation of F0 contours that provides a computationally efficient algorithm for automatically estimating the parameters of a F0 control model for singing voices. Although the best known F0 control model, based on a second-order system with a piece-wise constant function as its input, can generate F0 contours of natural singing voices, this model has no means of learning the model parameters from observed F0 contours automatically. Therefore, by modeling the piece-wise constant function by Hidden Markov Models (HMM) and approximating the second order differential equation by the difference equation, we estimate model parameters optimally based on iteration of Viterbi training and an LPC-like solver. Our representation is a generative model and can identify both the target musical note sequence and the dynamics of singing behaviors included in the F0 contours. Our experimental results show that the proposed method can separate the dynamics from the target musical note sequence and generate the F0 contours using estimated model parameters.
Yasunori Ohishi, Hirokazu Kameoka, Kunio Kashino, Kazuya Takeda
INTERSPEECH3
2008 A Quick Search Method for Audio Signals Based on a Piecewise Linear Representation of Feature Trajectories
abstract
This paper presents a new method for a quick similarity-based search through long unlabeled audio streams to detect and locate audio clips provided by users. The method involves feature-dimension reduction based on a piecewise linear representation of a sequential feature trajectory extracted from a long audio stream. Two techniques enable us to obtain a piecewise linear representation: the dynamic segmentation of feature trajectories and the segment-based Karhunen-L\'{o}eve (KL) transform. The proposed search method guarantees the same search results as the search method without the proposed feature-dimension reduction method in principle. Experiment results indicate significant improvements in search speed. For example the proposed method reduced the total search time to approximately 1/12 that of previous methods and detected queries in approximately 0.3 seconds from a 200-hour audio database.
Akihiro Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
IEEE Trans. Speech Audio Process.2
2007 Robust Search Methods for Music Signals Based on Simple Representation
abstract
Signal similarity search is an important technique for music information retrieval. A basic task is finding identical signal segments on unlabeled music-signal archives, given a short music signal fragment as a query. In such a task, the search must be fast and sufficiently robust against possible signal fluctuations due to noise and distortions. In this special session paper, we describe a search method designed to cope with additive interfering sounds by spectral partitioning. Then, we introduce another method designed to be robust under multiplicative noise or distortion based on binary area representation.
Kunio Kashino, Akisato Kimura, Hidehisa Nagano, Takayuki Kurozumi
ICASSP (4)1
2007 A Musical Audio Search Method Based on Self-Similarity Features
abstract
We propose a method for musical audio search based on signal matching. A major problem in the signal matching approach to musical audio search has been key variation; if the key of a query signal is significantly different from the one in the stored database, the search will fail. To cope with this problem, our method newly employs self-similarity as the feature for signal matching. The self-similarity proposed here is similarity of the power spectrum defined between two time points within an audio signal. We show that the method increases the robustness of musical audio search with respect to key variation. In our experimentation, for example, the proposed method yields precision and recall rates of around 0.75 even when the pitches in queries and stored signals differ from each other by seven semitones, whereas a conventional signal matching method does not produce meaningful results in such a case.
Tomonori Izumitani, Kunio Kashino
ICME2
2007 A Computational Model of Saliency Depletion/Recovery Phenomena for the Salient Region Extraction of Videos
abstract
This report proposes a new algorithm for extracting salient regions of videos by introducing two important properties of the early human visual system: (1) Instantaneous saliency depletion with gradual recovery, whereby saliency is instantaneously suppressed and gradually recovered in previously attended regions. (2) Gradual saliency depletion with instantaneous recovery, whereby saliency is gradually decreased over time in non-surprising regions and at the same time recovered in surprising locations. With the introduction of these properties, redundant information in videos can be suppressed and important information is eventually enhanced.
Clement Leung, Akisato Kimura, Tatsuto Takeuchi, Kunio Kashino
ICME4
2006 Frequency Component Restoration for Music Sounds using a Markov Random Field and Maximum Entropy Learning
abstract
We propose a method that estimates frequency component structures from music signals with noise and restores them. Restoring frequency components hidden by other interfering sounds is a difficult problem but has become important in various music information processing systems for melody extraction and audio retrieval. The proposed method is based on a probabilistic model of a frequency component structure represented as a Markov random field. To design an appropriate model, we introduce a supervised learning technique based on the maximum entropy model. We tested the method using musical audio signals generated from notes played by real instruments and noises. For four of six instruments, the proposed method achieves F-measures greater than 0.44 even in periods where signals are replaced by noises. We also evaluated the method in terms of feature distortion recovery in audio fingerprint matching tasks. The results show that the proposed method clearly reduces the effect of noise on the similarity values
Tomonori Izumitani, Kunio Kashino
ICASSP (5)2
2004 Bayesian estimation of simultaneous musical notes based on frequency domain modelling
abstract
The paper proposes a Bayesian method for polyphonic music description. The method first divides an input audio signal into a series of sections called snapshots, and then estimates parameters such as fundamental frequencies and amplitudes of the notes contained in each snapshot. The parameter estimation process is based on a frequency domain modelling and Gibbs sampling. Experimental results obtained from audio signals of test note patterns are encouraging; the accuracy is better than 80% for the estimation of fundamental frequencies in terms of semitones and instrument names when the number of simultaneous notes is two.
Kunio Kashino, Simon J. Godsill
ICASSP (4)1
2004 Similarity-based partial image retrieval guaranteeing same accuracy as exhaustive matching
abstract
We propose a new framework for quick and accurate partial image retrieval from a huge number of images based on a predefined distance measure. Finding partial similarities generally requires a huge amount of storage space for indexes due to the large number of portions of images. The proposed method extracts portions from each database image at a constant spacing, while it extracts all possible portions from a query image. In this way, the proposed method can greatly reduce the size of indexes while theoretically guaranteeing the same accuracy as exhaustive matching.
Akisato Kimura, Takahito Kawanishi, Kunio Kashino
ICME3
2003 Dynamic-segmentation-based feature dimension reduction for quick audio/video searching
abstract
We propose a new feature dimension reduction method for multimedia search. The main technique in the method is dynamic segmentation that partitions sequential feature trajectories dynamically. While dynamic segmentation reduces the average dimensionality and accelerates the search, it requires huge amount of calculation. Thus, our method quickly executes suboptimal partitioning of the trajectories by using the discreteness of dimension changes. This guarantees the optimal amount of calculation to derive the suboptimal partitioning under the condition that the dimension monotonously increases as the segment length increases. The experiment shows that our method is over 10 times faster than a straightforward dynamic segmentation method.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP (3)2
2003 A fast search algorithm for background music signals based on the search for numerous small signal components
abstract
The paper proposes a method for detecting and locating a known music signal in a long audio stream. Unlike existing methods, ours assumes that the music is used as background music (BGM) and overlapped by another sound such as speech and that the interfering sound is typically louder than the target music. The proposed method is based on time-series active search, which is a quick signal search method reported earlier (Kashino, K. et al., Proc. ICASSP-99, vol.VI, 1999). To realize the BGM search, however, a novel extension is introduced. That is, the music signal is first decomposed into a number of small time-frequency regions, and the search is carried out for each of those components. The results of the search are then integrated based on a voting scheme to find the target music locations. Experiments show that an accurate search is possible when SNR is -5 dB and that the search completes in about 8 s for a 30 min stored signal.
Hidehisa Nagano, Kunio Kashino, Hiroshi Murase
ICASSP (5)2
2003 Dynamic-segmentation-based feature dimension reduction for quick audio/video searching
abstract
We propose a new feature dimension reduction method for multimedia search. The main technique in the method is dynamic segmentation that partitions sequential feature trajectories dynamically. While dynamic segmentation reduces the average dimensionality and accelerates the search, it requires huge amount of calculation. Thus, our method quickly executes suboptimal partitioning of the trajectories by using the discreteness of dimension changes. This guarantees the optimal amount of calculation to derive the suboptimal partitioning under the condition that the dimension monotonously increases as the segment length increases. The experiment shows that our method is over 10 times faster than a straightforward dynamic segmentation method.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICME2
2003 A fast search algorithm for background music signals based on the search for numerous small signal components
abstract
This paper proposes a method for detecting and locating a known music signal in a long audio stream. Unlike existing methods, ours assumes that the music is used as background music (BGM) and overlapped by another sound such as speech and that the interfering sound is typically louder than the target music. The proposed method is based on time-series active search, which is a quick signal search method reported earlier. To realize the BGM search, however, a novel extension is introduced. That is, the music signal is firstly decomposed into a number of small time-frequency regions, and the search is carried out for each of those components. The results of the search are then integrated based on a voting scheme to find the target music locations. Experiments show that accurate search is possible when SNR is -5 dB and that the search completes in about 8 s for a 30-m stored signal.
Hidehisa Nagano, Kunio Kashino, Hiroshi Murase
ICME2
2003 A quick search method for audio and video signals based on histogram pruning
abstract
This paper proposes a quick method of similarity-based signal searching to detect and locate a specific audio or video signal given as a query in a stored long audio or video signal. With existing techniques, similarity-based searching may become impractical in terms of computing time in the case of searching through long-running (several-days' worth of) signals. The proposed algorithm, which is referred to as time-series active search, offers significantly faster search with sufficient accuracy. The key to the acceleration is an effective pruning algorithm introduced in the histogram matching stage. Through the pruning, the actual number of matching calculations can be reduced by 200 to 500 times compared with exhaustive search while guaranteeing exactly the same search result. Experiments show that the proposed method can correctly detect and locate a 15-s signal in a 48-h recording of TV broadcasts within 1 s, once the feature vectors are calculated and quantized. As extentions of the basic algorithm, efficient AND/OR search methods for searching for multiple query signals and a feature dithering method for coping with signal distortion are also discussed.
Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
IEEE Trans. Multim.1
2002 A quick search method for multimedia signals using feature compression based on piecewise linear maps
abstract
We propose a quick algorithm for multimedia signal search. The algorithm comprises two techniques: feature compression based on piecewise linear maps and distance bounding to efficiently limit the search space. When compared with existing multimedia search techniques, they greatly reduce the computational cost required in searching. Although feature compression is employed in our method, our bounding technique mathematically guarantees the same recall rate as the search based on the original features; no segment to be detected is missed. Experiments indicate that the proposed algorithm is approximately 10 times faster than and as accurate as an existing fast method maitaining the same search accuracy.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP2
2002 Fast music retrieval using polyphonic binary feature vectors
abstract
We propose a method for retrieving similar music from a polyphonic-music audio database using a polyphonic audio signal as a query. In this task, we must consider similarities among polyphonic signals of the music, and achieve quick retrieval. Therefore, we first introduce the polyphonic binary feature vector to represent the presence of multiple notes. This feature is suitable for the search based on the similarities among polyphonic audio signals. Then, we propose a new search method, which is quicker than the exhaustive use of DP matching. The search is accelerated using a "similarity matrix" to limit the search space. Experiments using a test database containing 216 music pieces show that the search accuracy of the proposed feature is 89%, which is approximately 26% higher than that of the conventional spectrum feature. It is also shown that the new search method retrieves similar music without significant accuracy degradation as well as the exhaustive search does and the computational complexity of the new search method is about 1/4 that of exhaustive search.
Hidehisa Nagano, Kunio Kashino, Hiroshi Murase
ICME (1)2
2001 Very quick audio searching: introducing global pruning to the Time-Series Active Search
abstract
Previously, we proposed a histogram-based quick signal search method called Time-Series Active Search (TAS). TAS is a method of searching through long audio or video recordings for a specified segment, based on signal similarity. TAS is fast; it can search through a 24-hour recording in 1 second after a query-independent preprocessing. However, an even faster method is required when we consider a huge amount of audio archives, for example a month's worth of recordings. Thus, we propose a preprocessing method that significantly accelerates TAS. The core part of this method comprises a global histogram clustering of long signals and a pruning scheme using those clusters. Tests using broadcast recording indicate that the proposed algorithm achieves a search speed approximately 3 to 30 times faster than TAS. In these tests, the search results are exactly the same as with TAS.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP2
2001 A method for robust and quick video searching using probabilistic dither-voting
abstract
We propose a quick and accurate search method for detecting a query signal from long video recordings The method is based on the time-series active search, which is a quick searching method for audio and video signals that we previously proposed. Time-series active search is based on a histogram matching scheme and an efficient pruning mechanism, and therefore, it was very quick. We found, however, that the accuracy sometimes deteriorates when it is applied to searches through long video archives that are composed of many similar video images or those containing feature distortions caused by video dubbing or low-bit-rate compression. The problem arises from (1) insufficient capability of representing features and (2) feature distortions. Thus, the method proposed here uses LBG-based VQ to improve the capacity to represent features and probabilistic dither-voting to improve robustness with respect to feature distortions. The experiments prove the effects of the proposed method.
Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICIP (2)1
2000 Feature Fluctuation Absorption for a Quick Audio Retrieval from Long Recordings
abstract
Kashino et al. proposed (1999) a histogram-based quick signal search method called time-series active search (TAS). TAS has only been effective in the exact matching case, where the segments to be detected are assumed to be exactly same as the reference signal. Here, we extend the method so that it is applicable even if the features fluctuate. In addition to the feature modification, feature dithering is discussed to absorb feature fluctuations. Efficient time-scaled search is also investigated to cope with variations of the reference signal duration. Tests using broadcast recordings show that the extended method improves the accuracy in nonexact-matching tasks such as hand-clap detection and word spotting in a single-speaker's narration. The tests also show the speed-ups by pruning introduced in the time-scaled search.
Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICPR1
1999 Time-series active search for quick retrieval of audio and video
abstract
This paper proposes a search method that can quickly detect and locate known sound (video) in a long audio (video) stream. The method is based on active search. Active search reduces the number of candidate matches between reference and input signals by approximately 10 to 100 times compared to exhaustive search, while guaranteeing the same retrieval accuracy. We proposed a quick search method in Smith et al. (1998), and here we focus on improvement of the accuracy. Thus the feature used has been extended to the audio power spectrum and temporal division of the histogram windows has been introduced to incorporate time information. Tests carried out under practical circumstances clearly show the accuracy improvement. The proposed method is still so fast that it can correctly retrieve a 15-s commercial in a 6-h recording of TV broadcasting within 2 s, once the features are calculated.
Kunio Kashino, Gavin Smith, Hiroshi Murase
ICASSP1
1999 A sound source identification system for ensemble music based on template adaptation and music stream extraction
Kunio Kashino, Hiroshi Murase
Speech Commun.1
1998 Music recognition using note transition context
abstract
As a typical example of sound-mixture recognition, the recognition of ensemble music is addressed. Here music recognition is defined as recognizing the pitch and the name of an instrument for each musical note in monaural or stereo recordings of real music performances. The first key part of the proposed method is adaptive template matching that can cope with variability in musical sounds. This is employed in the hypothesis-generation stage. The second key part of the proposed method is musical context integration based on the probabilistic networks. This is employed in the hypothesis-verification stage. The evaluation results clearly show the advantages of these two processes.
Kunio Kashino, Hiroshi Murase
ICASSP1
1998 Quick audio retrieval using active search
abstract
This paper discusses a method to search quickly through broadcast audio data to detect and locate known sounds using reference templates, based on the active search algorithm and histogram modeling of zero-crossing features. Active search reduces the number of candidate matches between reference and test template by up to 36 times compared to exhaustive search, while still remaining optimal. Computation is further reduced by using computationally inexpensive zero-crossing features. The method is robust against white noise addition down to 20 dB signal-to-noise ratios and digitization noise.
Gavin Smith, Hiroshi Murase, Kunio Kashino
ICASSP3
1997 A Music Stream Segregation System Based on Adaptive Multi-Agents
Kunio Kashino, Hiroshi Murase
IJCAI1
1996 A music scene analysis system with the MRF-based information integration scheme
abstract
This paper describes the process model for a system that recognizes the rhythm, chords, and source-separated musical notes in monaural music signals. The model consists of multiple processing modules and a MRF (Markov random field)-based hypothesis network for integration of multiple sources of information. Because the MRP enables information to be integrated on a multiply connected hypothesis network, the results of evaluation experiments show that the present system recognizes notes better than does a system based on a singly connected Bayesian hypothesis network.
Kunio Kashino, Norihiro Hagita
ICPR1
1995 Organization of Hierarchical Perceptual Sounds: Music Scene Analysis with Autonomous Processing Modules and a Quantitative Information Integration Mechanism
Kunio Kashino, Kazuhiro Nakadai, Tomoyoshi Kinoshita, Hidehiko Tanaka
IJCAI1