VLDB 2026 Research / reviewers in the wild / expert
Ryuki Tachibana
dblp:66/6844
· DBLP profile ↗
28ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 14 · 3 first-authorSystems, architecture and hardware · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 48% Language models and text generation · 32% Robot manipulation · 12% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.9 | 2 | 2020 | Q-learning with Language Model for Edit-based Unsupervised Summarization · EMNLP (1) 2020 Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based Games · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.4 | 1 | 2020 | Q-learning with Language Model for Edit-based Unsupervised Summarization · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text summarization
unsupervised summarization |
0.4 | 1 | 2020 | Q-learning with Language Model for Edit-based Unsupervised Summarization · EMNLP (1) 2020 |
Robotics › Robot manipulation › assembly › assembly task
assembly task execution |
0.3 | 1 | 2018 | MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning · ICRA 2018 |
Human-robot interaction
human-robot collaboration |
0.3 | 1 | 2018 | MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning · ICRA 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
rule and ontology reasoning |
0.1 | 1 | 2018 | MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level Reasoning · ICRA 2018 |
Digital forensics and information hiding › watermarking
audio watermarking |
0.1 | 1 | 2009 | Watermarked Movie Soundtrack Finds the Position of the Camcorder in a Theater · IEEE Trans. Multim. 2009 |
Methods — techniques the papers use, named apart from their topics
symbolic planning · 0.7ontology · 0.7natural language understanding · 0.7q-learning · 0.4observation pruning · 0.4language model · 0.4edit-based generation · 0.4context relevance · 0.4stochastic detection model · 0.2spread-spectrum watermarking · 0.1spread spectrum watermarking · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Data-Efficient Framework for Real-World Multiple Sound Source 2d LocalizationabstractDeep neural networks have recently led to promising results for the task of multiple sound source localization. Yet, they require a lot of training data to cover a variety of acoustic conditions and micro-phone array layouts. One can leverage acoustic simulators to inexpensively generate labeled training data. However, models trained on synthetic data tend to perform poorly with real-world recordings due to the domain mismatch. Moreover, learning for different microphone array layouts makes the task more complicated due to the infinite number of possible layouts. We propose to use adversarial learning methods to close the gap between synthetic and real do-mains. Our novel ensemble-discrimination method significantly improves the localization performance without requiring any label from the real data. Furthermore, we propose a novel explicit transformation layer to be embedded in the localization architecture. It enables the model to be trained with data from specific microphone array layouts while generalizing well to unseen layouts during inference. Guillaume Le Moing, Phongtharin Vinayavekhin, Don Joven Agravante, Tadanobu Inoue, Jayakorn Vongkulbhisal, Asim Munawar, Ryuki Tachibana |
ICASSP | 7 |
| 2020 | Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based GamesabstractSubhajit Chaudhury, Daiki Kimura, Kartik Talamadupula, Michiaki Tatsubori, Asim Munawar, Ryuki Tachibana. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Subhajit Chaudhury, Daiki Kimura, Kartik Talamadupula, Michiaki Tatsubori, Asim Munawar, Ryuki Tachibana |
EMNLP (1) | 6 |
| 2020 | Q-learning with Language Model for Edit-based Unsupervised SummarizationabstractUnsupervised methods are promising for abstractive textsummarization in that the parallel corpora is not required.However, their performance is still far from being satisfied, therefore research on promising solutions is on-going.In this paper, we propose a new approach based on Q-learning with an edit-based summarization.The method combines two key modules to form an Editorial Agent and Language Model converter (EALM).The agent predicts edit actions (e.t., delete, keep, and replace), and then the LM converter deterministically generates a summary on the basis of the action signals.Qlearning is leveraged to train the agent to produce proper edit actions.Experimental results show that EALM delivered competitive performance compared with the previous encoderdecoder-based methods, even with truly zero paired data (i.e., no validation set).Defining the task as Q-learning enables us not only to develop a competitive method but also to make the latest techniques in reinforcement learning available for unsupervised summarization.We also conduct qualitative analysis, providing insights into future study on unsupervised summarizers. 1 Ryosuke Kohita, Akifumi Wachi, Ryuki Tachibana |
EMNLP (1) | 4 |
| 2020 | Adversarial Discriminative Attention for Robust Anomaly DetectionabstractExisting methods for visual anomaly detection predominantly rely on global level pixel comparisons for anomaly score computation without emphasizing on unique local features. However, images from real-world applications are susceptible to unwanted noise and distractions, that might jeopardize the robustness of such anomaly score. To alleviate this problem, we propose a self-supervised masking method that specifically focuses on discriminative parts of images to enable robust anomaly detection. Our experiments reveal that discriminator's class activation map in adversarial training evolves in three stages and finally fixates on the foreground location in the images. Using this property of the activation map, we construct a mask that suppresses spurious signals from the background thus enabling robust anomaly detection by focusing on local discriminative attributes. Additionally, our method can further improve the accuracy by learning a semi-supervised discriminative classifier in cases where a few samples from anomaly classes are available during the training. Experimental evaluations on four different types of datasets demonstrate that our method outperforms previous state-of-the-art methods for each condition and in all domains. Daiki Kimura, Subhajit Chaudhury, Minori Narita, Asim Munawar, Ryuki Tachibana |
WACV | 5 |
| 2019 | Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports VideosabstractImage-based sports analytics enable automatic retrieval of key events in a game to speed up the analytics process for human experts. However, most existing methods focus on structured television broadcast video datasets with a straight and fixed camera having minimum variability in the capturing pose. In this paper, we study the case of event detection in sports videos for unstructured environments with arbitrary camera angles. The transition from structured to unstructured video analysis produces multiple challenges that we address in our paper. Specifically, we identify and solve two major problems: unsupervised identification of players in an unstructured setting and generalization of the trained models to pose variations due to arbitrary shooting angles. For the first problem, we propose a temporal feature aggregation algorithm using person re-identification features to obtain high player retrieval precision by boosting a weak heuristic scoring method. Additionally, we propose a data augmentation technique, based on multi-modal image translation model, to reduce bias in the appearance of training samples. Experimental evaluations show that our proposed method improves precision for player retrieval from 0.78 to 0.86 for obliquely angled videos. Additionally, we obtain an improvement in F1 score for rally detection in table tennis videos from 0.79 in case of global frame-level features to 0.89 using our proposed player-level features. Please see the supplementary video submission at https://ibm.biz/BdzeZA. Subhajit Chaudhury, Hiroki Ozaki, Daiki Kimura, Phongtharin Vinayavekhin, Asim Munawar, Ryuki Tachibana, Koji Ito, Yuki Inaba, Minoru Matsumoto, Shuji Kidokoro |
ISM | 6 |
| 2019 | Injective State-Image Mapping facilitates Visual Adversarial Imitation LearningabstractThe growing use of virtual autonomous agents in applications like games and entertainment demands better control policies for natural-looking movements and actions. Unlike the conventional approach of hard-coding motion routines, we propose a deep learning method for obtaining control policies by directly mimicking raw video demonstrations. Previous methods in this domain rely on extracting low-dimensional features from expert videos followed by a separate hand-crafted reward estimation step. We propose an imitation learning framework that reduces the dependence on hand-engineered reward functions by jointly learning the feature extraction and reward estimation steps using Generative Adversarial Networks (GANs). Our main contribution in this paper is to show that under injective mapping between low-level joint state (angles and velocities) trajectories and corresponding raw video stream, performing adversarial imitation learning on video demonstrations is equivalent to learning from the state trajectories. Experimental results show that the proposed adversarial learning method from raw videos produces a similar performance to state-of-the-art imitation learning techniques while frequently outperforming existing hand-crafted video imitation methods. Furthermore, we show that our method can learn action policies by imitating video demonstrations on YouTube with similar performance to learned agents from true reward signal. Please see the supplementary video submission at https://ibm.biz/BdzzNA. Subhajit Chaudhury, Daiki Kimura, Asim Munawar, Ryuki Tachibana |
MMSP | 4 |
| 2019 | Learning Multiple Sound Source 2D LocalizationabstractIn this paper, we propose novel deep learning based algorithms for multiple sound source localization. Specifically, we aim to find the 2D Cartesian coordinates of multiple sound sources in an enclosed environment by using multiple microphone arrays. To this end, we use an encoding-decoding architecture and propose two improvements on it to accomplish the task. In addition, we also propose two novel localization representations which increase the accuracy. Lastly, new metrics are developed relying on resolution-based multiple source association which enables us to evaluate and compare different localization approaches. We tested our method on both synthetic and real world data. The results show that our method improves upon the previous baseline approach for this problem. Guillaume Le Moing, Phongtharin Vinayavekhin, Tadanobu Inoue, Jayakorn Vongkulbhisal, Asim Munawar, Ryuki Tachibana, Don Joven Agravante |
MMSP | 6 |
| 2018 | Focusing on What is Relevant: Time-Series Learning and Understanding using AttentionabstractThis paper is a contribution towards interpretability of the deep learning models in different applications of time-series. We propose a temporal attention layer that is capable of selecting the relevant information to perform various tasks, including data completion, key-frame detection and classification. The method uses the whole input sequence to calculate an attention value for each time step. This results in more focused attention values and more plausible visualisation than previous methods. We apply the proposed method to three different tasks. Experimental results show that the proposed network produces comparable results to a state of the art. In addition, the network provides better interpretability of the decision, that is, it generates more significant attention weight to related frames compared to similar techniques attempted in the past. Phongtharin Vinayavekhin, Subhajit Chaudhury, Asim Munawar, Don Joven Agravante, Giovanni De Magistris, Daiki Kimura, Ryuki Tachibana |
ICPR | 7 |
| 2018 | MaestROB: A Robotics Framework for Integrated Orchestration of Low-Level Control and High-Level ReasoningabstractThis paper describes a framework called MaestROBe It is designed to make the robots perform complex tasks with high precision by simple high-level instructions given by natural language or demonstration. To realize this, it handles a hierarchical structure by using the knowledge stored in the forms of ontology and rules for bridging among different levels of instructions. Accordingly, the framework has multiple layers of processing components; perception and actuation control at the low level, symbolic planner and Watson APIs for cognitive capabilities and semantic understanding, and orchestration of these components by a new open source robot middleware called Project Intu at its core. We show how this framework can be used in a complex scenario where multiple actors (human, a communication robot, and an industrial robot) collaborate to perform a common industrial task. Human teaches an assembly task to Pepper (a humanoid robot from SoftBank Robotics) using natural language conversation and demonstration. Our framework helps Pepper perceive the human demonstration and generate a sequence of actions for UR5 (collaborative robot arm from Universal Robots), which ultimately performs the assembly (e.g. insertion) task. Asim Munawar, Giovanni De Magistris, Tu-Hoa Pham, Daiki Kimura, Michiaki Tatsubori, Takao Moriyama, Ryuki Tachibana, Grady Booch |
ICRA | 7 |
| 2018 | OptLayer - Practical Constrained Optimization for Deep Reinforcement Learning in the Real WorldabstractWhile deep reinforcement learning techniques have recently produced considerable achievements on many decision-making problems, their use in robotics has largely been limited to simulated worlds or restricted motions, since unconstrained trial-and-error interactions in the real world can have undesirable consequences for the robot or its environment. To overcome such limitations, we propose a novel reinforcement learning architecture, OptLayer, that takes as inputs possibly unsafe actions predicted by a neural network and outputs the closest actions that satisfy chosen constraints. While learning control policies often requires carefully crafted rewards and penalties while exploring the range of possible actions, OptLayer ensures that only safe actions are actually executed and unsafe predictions are penalized during training. We demonstrate the effectiveness of our approach on robot reaching tasks, both simulated and in the real world. Tu-Hoa Pham, Giovanni De Magistris, Ryuki Tachibana |
ICRA | 3 |
| 2017 | Effective joint training of denoising feature space transforms and Neural Network based acoustic modelsabstractNeural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improve speech recognition performance. In this paper, we propose an effective algorithm to jointly train such a denoising feature space transform and a NN based acoustic model with various kinds of data. Our proposed method first pretrains a Convolutional Neural Network (CNN) based denoising frontend and then jointly trains this frontend with a NN backend acoustic model. In the unsupervised pretraining stage, the frontend is designed to estimate clean log Mel-filterbank features from noisy log-power spectral input features. A subsequent multi-stage training of the proposed frontend, with the dropout technique applied only at the joint layer between the frontend and backend NNs, leads to significant improvements in the overall performance. On the Aurora-4 task, our proposed system achieves an average WER of 9.98%. This is a 9.0% relative improvement over one of the best reported speaker independent baseline system's performance. A final semi-supervised adaptation of the frontend NN, similar to feature space adaptation, reduces the average WER to 7.39%, a further relative WER improvement of 25%. Takashi Fukuda, Osamu Ichikawa, Gakuto Kurata, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2017 | Deep reinforcement learning for high precision assembly tasksabstractThe high precision assembly of mechanical parts requires precision that exceeds that of robots. Conventional part-mating methods used in the current manufacturing require numerous parameters to be tediously tuned before deployment. We show how a robot can successfully perform a peg-in-hole task with a tight clearance through training a recurrent neural network with reinforcement learning. In addition to reducing manual effort, the proposed method also shows a better fitting performance with a tighter clearance and robustness against positional and angular errors for the peg-in-hole task. The neural network learns to take the optimal action by observing the sensors of a robot to estimate the system state. The advantages of our proposed method are validated experimentally on a 7-axis articulated robot arm. Tadanobu Inoue, Giovanni De Magistris, Asim Munawar, Tsuyoshi Yokoya, Ryuki Tachibana |
IROS | 5 |
| 2016 | Convolutional neural network pre-trained with projection matrices on linear discriminant analysisabstractRecently, the hybrid architecture of a neural network (NN) and a hidden Markov model (HMM) has shown significant improvement on automatic speech recognition (ASR) over the conventional Gaussian mixture model (GMM)-based system. The convolutional neural network (CNN), a successful NN-based system, can represent local spectral variations spanning the time-frequency space. Meanwhile, spectro-temporal features have been widely studied to make ASR more robust. Typically, the spectro-temporal features are extracted from acoustic spectral patterns using a 2D filtering process. Convolutional layers in CNN that have various local windows can also be regarded as an efficient feature extractor to capture 2D spectral variations. In a standard procedure, the local windows in CNN are initialized randomly before the pre-training and are iteratively updated with a back propagation algorithm in the pre-training and fine-tuning steps. In this paper, we explore using projection matrices composed of eigenvectors estimated by linear discriminant analysis (LDA) objective function as initial weights for the first convolutional layer in CNN. From analysis of the local windows trained by the proposed method, we can see the eigenvectors of LDA has desirable properties as initial weights of CNN. The proposed method yielded a 8.1% relative improvement compared to CNN with local weights initialized randomly. Takashi Fukuda, Osamu Ichikawa, Ryuki Tachibana |
ICASSP | 3 |
| 2016 | Speech recognition robust against speech overlapping in monaural recordings of telephone conversationsabstractMonaural (single-channel) recording is sometimes used for telephone conversations in call centers. Generally speaking, the accuracy of automatic speech recognition of a monaural recording is worse than that of the multi-channel recording of the same conversation where each speaker's voice is separately recorded. The major reason is that the recognition system fails not only at the overlapping segments where the voices of the multiple speakers overlap, but also at the neighboring segments surrounding the overlapping segments. In this paper, we tackle this problem by using a combination of garbage modeling and noise-robust monaural acoustic modeling. Our proposed method trains the models by making use of multi-channel recordings and transcripts, which are relatively easy to prepare than monaural recordings and transcripts. We present experimental results where the proposed methods reduced the error rates by approximately 3% relative to the baseline methods for both of GMM-HMM and CNN-HMM cases. Because the proposed method is quite simple, the proposed method is easy to deploy to wide range of ASR systems for monaural speech transcription. Masayuki Suzuki, Gakuto Kurata, Tohru Nagano, Ryuki Tachibana |
ICASSP | 4 |
| 2016 | Domain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings
Masayuki Suzuki, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 2 |
| 2015 | A metric for evaluating speech recognizer output based on human-perception modelabstractWord error rate or character error rate are usually used as the metrics for evaluating the accuracy of speech recognition. These are naturally-defined objective metrics and are helpful for comparing recognition methods fairly. However the overall performance of the recognition systems and the usefulness of the results are not necessarily considered. To address this problem, we study and propose a metric which replicates human-annotated scores using their perception to the recognition results. The features that we use are the numbers of insertion errors, deletion errors, and substitution errors in the characters and the syllables. In addition we studied the numbers of consecutive errors, the misrecognized keywords, and the locations of errors. We created models using linear regression and random forest, predicted human-perceived scores, and compared them with the actual scores using Spearman’s rank-based correlation. According to our experiments the correlation of human perceived scores with character error rates is 0.456, while those with the predicted scores by using a random forest of 10 features is 0.715. The latter is close to the averaged correlation between the scores of the human subjects, 0.765, which suggests that we can predict the human-perceived scores using those features and that we can leverage human perception model for evaluating speech recognition performance. The important factors (features) for the prediction are the numbers of substitution errors and consecutive errors. Nobuyasu Itoh, Gakuto Kurata, Ryuki Tachibana, Masafumi Nishimura |
INTERSPEECH | 3 |
| 2012 | Constructing ensembles of dissimilar acoustic models using hidden attributes of training dataabstractOne of the objectives in acoustic modeling is to realize robust statistical models against the wide variety of acoustic conditions that are present in real world environments. As large amounts of training data become available, modeling subsets of the data with similar acoustic qualities can be done accurately and multiple acoustic models are jointly used as a form of system combination or model selection. In this paper, we propose a method to partition the training data for constructing ensembles of acoustic models using metadata attributes such as SNR, speaking rate, and duration via a binary tree. The metadata attribute used at each binary split in the decision tree is obtained using a metric proposed in this paper that is cosine-similarity based. The resulting multiple models are combined using voting techniques such as n-best ROVER. The proposed method improved the recognition accuracy by up to 4% relative over the state-of-the-art system on a large vocabulary continuous speech recognition voice search task. Takashi Fukuda, Ryuki Tachibana, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
ICASSP | 2 |
| 2011 | Frame-level AnyBoost for LVCSR with the MMI CriterionabstractThis paper propose a variant of AnyBoost for a large vocabulary continuous speech recognition (LVCSR) task. AnyBoost is an efficient algorithm to train an ensemble of weak learners by gradient descent for an objective function.We present a novel training procedure that trains acoustic models via the MMI criterion using data that is weighted proportional to the summation of the posterior functions of previous round of weak learners. Optimized for system combination by n-best ROVER at runtime, data weights for a new weak learner are computed as a weighted summation of posteriors of previous weak learners. We compare a frame-based version and a sentence-based version of our proposed algorithm with a frame-based AdaBoost algorithm. We will present results on a voice search task trained with different amounts of data with gains of 5.1% to 7.5% relative in WER can be obtained by three rounds of boosting. Ryuki Tachibana, Takashi Fukuda, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
ASRU | 1 |
| 2011 | Agglomerative Hierarchical Clustering of Emotions in Speech Based on Subjective Relative Similarity
Ryoichi Takashima, Tohru Nagano, Ryuki Tachibana, Masafumi Nishimura |
INTERSPEECH | 3 |
| 2010 | Speech synthesis by modeling harmonics structure with multiple functionabstractIn this paper, we present a new approach for the speech synthesis, in which speech utterances are synthesized using the parameters of spectro-modeling function (Multiple function). With this approach, only harmonic-parts are extracted from the phoneme spectrum, and the time-varying spectrum corresponding to the harmonics or sinusoidal components is modeled using the Multiple function. We introduce two types of the functions, and present the method to estimate the parameters of each function using the observed phoneme spectrum. In the synthesis stage, speech signals are generated from the parameters of the Multiple function. The advantage of this method is that it only requires a few speech synthesis parameters. We discuss the effectiveness of our proposed method through experimental results. Toru Nakashika, Ryuki Tachibana, Masafumi Nishimura, Tetsuya Takiguchi, Yasuo Ariki |
INTERSPEECH | 2 |
| 2009 | Efficient gradient F0 tree model for prosody modeling and unit-selection, applied for the embedded US English concatenative TTSabstractModeling of pitch dynamics in addition to absolute pitch modeling is highly desirable for robust pitch curve prediction and unit selection in concatenative TTS systems. Transition prosody models have been reported to improve consistency and naturalness for pitch-accent and tonal languages, like Japanese and Mandarin. In the current work we revise a Gradient F0 tree model, originally developed for Japanese, and adjust it for American English. The resultant model requires few computational resources at a runtime that makes it highly suitable for embedded TTS applications. We report encouraging results of applying it for an embedded concatenative TTS system for American English. Slava Shechtman, Ryuki Tachibana |
ICASSP | 2 |
| 2009 | Dynamic features in the linear domain for robust automatic speech recognition in a reverberant environment
Osamu Ichikawa, Takashi Fukuda, Ryuki Tachibana, Masafumi Nishimura |
INTERSPEECH | 3 |
| 2009 | Japanese pitch conversion for voice morphing based on differential modelingabstractAbstract In this paper, we convert the pitch contours predicted by a TTSsystem that models a source speaker to resemble the pitch con-toursofatargetspeaker. Whenthespeakingstylesofthespeak-ers are very different, complex conversions such as adding ordeletingpitchpeaksmayberequired. Ourmethoddoesthecon-versions by modeling the direct pitch features and differentialpitch features at the same time based on linguistic features. Thedifferential pitch features are calculated from matched pairs ofsource and target pitch values. We show experimental results inwhich the target speaker’s characteristics are successfully mod-eledbasedonaverylimitedtrainingcorpus. TheproposedpitchconversionmethodstretchesthepossibilitiesofTTScustomiza-tion for various speaking styles. Index Terms : pitch conversion, voice conversion, voice mor-phing, speech synthesis, differential modeling. 1. Introduction Voice conversion changes the characteristics of the voice of anSPS (SPeaker Source) to those of an SPT (SPeaker Target) forvarious applications. One important application is to build cus-tomized text-to-speech (TTS) systems for different companies,soaTTSsystemwitheachcompany’sfavoritevoicecanbecre-ated quickly and inexpensively by modifying the speech corpusof some original speaker.Spectra and prosody are the two major characteristics ofvoice. For spectral conversion, recent work such as [1, 2]has achieved significant improvements in the naturalness andsimilarity of the voices converted using only a limited amountof training data. However, not much research has been doneinto prosody conversion. Most spectral conversion researchuses simple linear transformations for the prosody. It is truethat the detailed prosody difference is sometimes difficult forlisteners to distinguish [3], especially when the speakers aremonotonously reading for TTS corpus recording. However, togenerate TTS voices with a lively colloquial speaking style, re-production of the detailed prosody characteristics is important.Our objective is to reproduce the SPT’s speaking style ofpitch contours based on limited training data. We assume 100sentences as training data is a reasonable amount to require forthe SPT’s speech corpus. A speech corpus with that size caneasily be recorded in a thirty-minute recording session. We donot assume the existence of a parallel corpus, which can be adifficult condition to satisfy. We focus on the pitch changesaround the syllable level, because the important pitch changesin Japanese are mainly at the syllable level.Figure 1 illustrates examples of pitch contour pairs that theproposed pitch conversion method can handle. They are (a)asymmetrical slope changes, (b) adding or deleting peaks, and(c) adding or deleting phrase-final rises. Ryuki Tachibana, Zhiwei Shuang, Masafumi Nishimura |
INTERSPEECH | 1 |
| 2009 | Watermarked Movie Soundtrack Finds the Position of the Camcorder in a TheaterabstractIn recent years, the problem of camcorder piracy in theaters has become more serious due to technical advances in camcorders. In this paper, as a new deterrent to camcorder piracy, we propose a system for estimating the recording position from which a camcorder recording is made. The system is based on spread-spectrum audio watermarking for the multichannel movie soundtrack. It utilizes a stochastic model of the detection strength, which is calculated in the watermark detection process. Our experimental results show that the system estimates recording positions in an actual theater with a mean estimation error of 0.44 m. The results of our MUSHRA subjective listening tests show the method does not significantly spoil the subjective acoustic quality of the soundtrack. These results indicate that the proposed system is applicable for practical uses. Yuta Nakashima, Ryuki Tachibana, Noboru Babaguchi |
IEEE Trans. Multim. | 2 |
| 2008 | Improving phoneme and accent estimation by leveraging a dictionary for a stochastic TTS front-endabstractDetermining the correct phonemes and pitch accents is important for creating natural Japanese speech. We implemented a TTS frontend system based on an n-gram model. However, the vocabulary of the word n-gram model is limited to the list of the words found in the training corpus, and collecting a very large training corpus is not an easy task. In this paper, we propose using an additional class n-gram model to incorporate not only the words found in the training corpus, but the words found in the dictionary to further improve the accuracy. In our experiments, our proposed model relatively improves the accuracy for estimating accents by 16.9% and the accuracy for estimating phonemes by 21.6% compared to the word n-gram model. Tohru Nagano, Ryuki Tachibana, Nobuyasu Itoh, Masafumi Nishimura |
ICASSP | 2 |
| 2007 | Determining Recording Location Based on Synchronization Positions of AudiowatermarkingabstractIn this paper, we propose a novel application of digital watermarking, determination of recording locations. This application enables us to determine the seat location in an auditorium where a recording was made. Precisely measured synchronization positions of the spread-spectrum watermarks are used for the determination. To avoid use of mismeasured synchronization positions, the algorithm discards synchronization positions with the corresponding normalized correlation values below a threshold. The experiments with our implementation resulted in accurate determinations; almost all of the locations can be determined within the error of 0.5 m. These experimental results successfully show the potential applicability of our application. Yuta Nakashima, Ryuki Tachibana, Masafumi Nishimura, Noboru Babaguchi |
ICASSP (2) | 2 |
| 2007 | Preliminary experiments toward automatic generation of new TTS voices from recorded speech alone
Ryuki Tachibana, Tohru Nagano, Gakuto Kurata, Masafumi Nishimura, Noboru Babaguchi |
INTERSPEECH | 1 |
| 2002 | An audio watermarking method using a two-dimensional pseudo-random array
Ryuki Tachibana, Shuichi Shimizu, Seiji Kobayashi, Taiga Nakamura |
Signal Process. | 1 |