EDBT 2026 Demo / reviewers in the wild / expert
Kazuya Takeda
dblp:38/3483
· DBLP profile ↗
192ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0002-0330-1787ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 129 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 111 · 6 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Estimation of Mobile Robot Waiting Locations using Fluid Simulation without Prior ObservationabstractThe social implementation of small autonomous mobile robots is advancing in delivery and accompanying services. However, inappropriate waiting positions can obstruct pedestrian flow and impair facility operations. Conventional waiting location estimation methods require long-term observation of actual pedestrian walking history, resulting in high introduction costs and environmental dependence. This research proposes a fluid simulation-based method that estimates waiting locations using only floor maps without prior observation. By approximating pedestrian flow as a two-dimensional incompressible fluid, the method extracts waiting candidates from velocity fields determined by environmental geometry. Through comparative verification with agent-based simulation and human subject experiments, we confirmed that the proposed method achieves estimation accuracy equivalent to history-based approaches while requiring no prior observation. The method demonstrates environmental adaptability across simple corridors to complex spaces, with a practical threshold of 0.4 times maximum flow velocity effectively distinguishing appropriate waiting locations. The approach enables rapid deployment and adaptation to layout changes, addressing key challenges in mobile robot social implementation. Ryoya Sakaki, Yoshio Ishiguro, Kento Ohtani, Takanori Nishino, Kazuya Takeda |
HRI | 5 |
| 2025 | Estimating Counterfactual Treatment Outcomes Over Time in Complex Multiagent ScenariosabstractEvaluation of intervention in a multiagent system, for example, when humans should intervene in autonomous driving systems and when a player should pass to teammates for a good shot, is challenging in various engineering and scientific fields. Estimating the individual treatment effect (ITE) using counterfactual long-term prediction is practical to evaluate such interventions. However, most of the conventional frameworks did not consider the time-varying complex structure of multiagent relationships and covariate counterfactual prediction. This may lead to erroneous assessments of ITE and difficulty in interpretation. Here, we propose an interpretable, counterfactual recurrent network in multiagent systems to estimate the effect of the intervention. Our model leverages graph variational recurrent neural networks (GVRNNs) and theory-based computation with domain knowledge for the ITE estimation framework based on long-term prediction of multiagent covariates and outcomes, which can confirm the circumstances under which the intervention is effective. On simulated models of an automated vehicle and biological agents with time-varying confounders, we show that our methods achieved lower estimation errors in counterfactual covariates and the most effective treatment timing than the baselines. Furthermore, using real basketball data, our methods performed realistic counterfactual predictions and evaluated the counterfactual passes in shot scenarios. Keisuke Fujii 0001, Koh Takeuchi 0001, Atsushi Kuribayashi, Naoya Takeishi, Yoshinobu Kawahara, Kazuya Takeda |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Audio Difference Learning for Audio CaptioningabstractThis study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a feature representation space that preserves the relationship between audio, enabling the generation of captions that detail intricate audio information. This method employs a reference audio along with the input audio, both of which are transformed into feature representations via a shared encoder. Captions are then generated from these differential features to describe their differences. Furthermore, a unique technique is proposed that involves mixing the input audio with additional audio, and using the additional audio as a reference. This results in the difference between the mixed audio and the reference audio reverting back to the original input audio. This allows the original input’s caption to be used as the caption for their difference, eliminating the need for additional annotations for the differences. In the experiments using the Clotho and ESC50 datasets, the proposed method demonstrated an improvement in the SPIDEr score by 7% compared to conventional methods. Tatsuya Komatsu, Yusuke Fujita, Kazuya Takeda, Tomoki Toda |
ICASSP | 3 |
| 2024 | Contrasting Disentangled Partial Observations for Pedestrian Action PredictionabstractData-driven approaches have been recently proven effective in pedestrian action prediction by extensive works. However, frame-level annotations of pedestrian actions require a significant amount of manpower and time. In this paper, we propose a simple yet effective contrastive learning framework that enables pedestrian action prediction models to be trained on data without action labels. First of all, we regard disentangled visual observations, such as appearance, motion and trajectories, as multiple modalities. Then we construct a joint latent space where multimodal features from the same sample are encouraged to be close, whereas features from different samples are encouraged to be far from each other. Since most existing models use a similar architecture composed of separate feature extractors and fusion modules, our proposed framework can be applied directly to existing methods to boost the feature extractors. We pretrained state-of-the-art models on datasets without action labels, nuScenes and BDD100k, and evaluated these models on PIE, JAAD and TITAN. Quantitative results show that the pretrained with only the fusion parameters fine-tuned can compete with or even outperform models that are completely trained the one dataset. Alexander Carballo, Yingjie Niu, Kazuya Takeda |
IV | 4 |
| 2024 | RSG-Search Plus: An Advanced Traffic Scene Retrieval Methods based on Road Scene GraphabstractCurrently, with the rapid growth of training datasets for autonomous driving systems, we are faced with a challenge: how to efficiently retrieve specific traffic scenes from massive amount of scene in multiple datasets. This challenge primarily stems from the heterogeneity of existing datasets, meaning these datasets contain different types of data, follow different data formats, and use different sensors for data collection. To address this issue, we present RSG-Search Plus, a universal traffic scene searching method based on Road Scene Graph and Large Language Models (LLMs). Our approach first transform datasets into scene graphs to exclude irrelevant details, then efficiently retrieving specific configurations among thousands of traffic scenes by matching isomorphic sub-graphs between input graph and road scene graph. Experimental results demonstrate that our graph searching method can accurately match the scenes described by input condition. Additionally, this method is easily adaptable to different datasets, significantly simplifying the scene search process. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Simon Thompson 0002, Kazuya Takeda |
IV | 5 |
| 2024 | 360 LiDAR + 360 RGB + 360 Thermal: Multimodal Targetless CalibrationabstractNowadays, using LiDARs, RGB cameras and Thermal cameras for automatic systems, in particular self-driving cars, has become the common approach in multiple deployments. Each kind of sensor has distinct advantages, leading to the fact that using multiple sensors can help autonomous systems improve their performance. Calibration between sensors is the precondition of fusing multiple sensors. This paper presents a novel way to register extrinsic parameters for LiDAR, 360 RGB camera and 360 Thermal camera automatically based on features information. To evaluate the method, we use our dataset around Nagoya University. Khanh Bao Tran, Alexander Carballo, Kazuya Takeda |
IV | 3 |
| 2024 | Estimation of control area in badminton doubles with pose information from top and back view drone videosabstractAbstract The application of visual tracking to the performance analysis of sports players in dynamic competitions is vital for effective coaching. In doubles matches, coordinated positioning is crucial for maintaining control of the court and minimizing opponents’ scoring opportunities. The analysis of such teamwork plays a vital role in understanding the dynamics of the game. However, previous studies have primarily focused on analyzing and assessing singles players without considering occlusion in broadcast videos. These studies have relied on discrete representations, which involve the analysis and representation of specific actions (e.g., strokes) or events that occur during the game while overlooking the meaningful spatial distribution. In this work, we present the first annotated drone dataset from top and back views in badminton doubles and propose a framework to estimate the control area probability map, which can be used to evaluate teamwork performance. We present an efficient framework of deep neural networks that enables the calculation of full probability surfaces. This framework utilizes the embedding of a Gaussian mixture map of players’ positions and employs graph convolution on their poses. In the experiment, we verify our approach by comparing various baselines and discovering the correlations between the score and control area. Additionally, we propose a practical application for assessing optimal positioning to provide instructions during a game. Our approach offers both visual and quantitative evaluations of players’ movements, thereby providing valuable insights into doubles teamwork. The dataset and related project code is available at https://github.com/Ning-D/Drone_BD_ControlArea Ning Ding 0002, Kazuya Takeda, Wenhui Jin, Yingjiu Bei, Keisuke Fujii 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Runner re-identification from single-view running video in the open-world setting
Kazushi Tsutsui, Kazuya Takeda, Keisuke Fujii 0001 |
Multim. Tools Appl. | 3 |
| 2024 | Decentralized policy learning with partial observation and mechanical constraints for multiperson modeling
Keisuke Fujii 0001, Naoya Takeishi, Yoshinobu Kawahara, Kazuya Takeda |
Neural Networks | 4 |
| 2023 | Uncertainty Aware Task Allocation for Human-Automation Cooperative Recognition in Autonomous Driving SystemsabstractCooperative recognition, a method to achieve human-automation cooperation in the recognition phase of the autonomous driving system, has been proposed to address the challenges in the conventional control phase cooperation, e.g., taking over vehicle control. In cooperative recognition, the operator intervenes in recognition tasks that are difficult for the automated system alone to improve driving efficiency and safety. The challenge is the integration of both human and automated systems while both participants have different characteristics, processing capabilities, and uncertainty in the decisions (recognition results). The objectives of this study are task allocation (i.e., when and for which targets the operator should intervene) taking into account the intervention efficiency and human state. And also combine the human intervention and recognition result of the automated systems to solve the uncertainties in both participants. We formulated this problem with a Partially Observable Markov Decision Process (POMDP). The simulator experiment indicated that the recognition result of the automated system and the operator’s intervention were stochastically combined. The intervention requests to the operator adapted to the operator state and could be reduced while maintaining driving efficiency and minimizing risk omissions. Atsushi Kuribayashi, Eijiro Takeuchi, Alexander Carballo, Yoshio Ishiguro, Kazuya Takeda |
IV | 5 |
| 2023 | Expert-driven Rule-based Refinement of Semantic Segmentation Maps for Autonomous VehiclesabstractSemantic segmentation aims at assigning labels to every pixel of a given image. In the context of autonomous vehicles, semantic segmentation models should be trained with data collected from the traffic network through which vehicles are expected to circulate. Road regulation, weather conditions, and other context features may differ between regions, making local semantic segmentation datasets extremely valuable. However, the high ground truth annotation costs represent a hindrance to the development of such models. The upsurge of powerful feature learning architectures leaves room for semantic segmentation models trained on an unsupervised fashion. This observation vertebrates the purpose of this work: to produce coarse segmentation maps for scene understanding without the need of annotated data. We depart from an unsupervised model that yields low-quality results. The proposed methodology establishes a set of guidelines for the enhancement of segmentation maps. Obtained results expose an improvement of the segmentation quality thanks to the application of our devised guidelines, paving the way for the automatic generation of semantic segmentation datasets. Eric Manibardo, Ibai Lana, Javier Del Ser, Alexander Carballo, Kazuya Takeda |
IV | 5 |
| 2023 | Open-world driving scene segmentation via multi-stage and multi-modality fusion of vision-language embeddingabstractIn this study, a pixel-text level multi-stage multi-modality fusion segmentation method is proposed to make the open-world driving scene segmentation more efficient. It can be used for different semantic perceptual needs of autonomous driving scenarios for real-world driving situations. The method can finely segment unseen labels without additional corresponding semantic segmentation labels, only using the existing semantic segmentation data. The proposed method consists of 4 modules. A visual representation embedding module and a segmentation command embedding module are used to extract the driving scene and the segmentation category command. A multi-stage multi-modality fusion module is used to fuse the driving scene visual information and segmentation command text information for different sizes at the pixel-text level. Finally, a cascade segmentation head is used to ground the segmentation command text to the driving scene for encouraging the model to generate corresponding high-quality semantic segmentation results. In the experiment, we first verify the effectiveness of the method for zero-shot segmentation using a popular driving scene segmentation dataset. We also confirm the effectiveness of synonyms unseen label and hierarchy unseen label for the open-world semantic segmentation. Yingjie Niu, Ming Ding 0002, Maoning Ge, Hanting Yang, Kazuya Takeda |
IV | 6 |
| 2023 | Real-Time Graph-Based Optimization for GNSS-Doppler Integrated RTK-GNSS/IMU/DR Positioning System in Urban AreaabstractAutonomous driving of vehicles and robots requires highly accurate position information, and RTK-GNSS is expected to be utilized for this purpose. In this paper, we propose a robust and real-time operation method by introducing graph optimization into the integrated RTK-GNSS/IMU method. The proposed method is an extension of a method using vehicle trajectories that can estimate positions with lane-level accuracy even in urban areas. The position is estimated by removing GNSS multipaths from the shape of a vehicle trajectory of several hundred meters and averaging the remaining GNSS results. This method does not take into account the errors in the vehicle trajectory and cannot fully benefit from the high accuracy positioning solution of RTK-GNSS. To solve this problem, we introduce graph optimization to the base method, which treats the error state as a probabilistic model. However, general graph optimization methods have problems with processing time and outlier elimination. The proposed method solves these problems by restricting the time series data to be optimized and using a two-step optimization structure. Evaluations show that the proposed method is effective because it satisfies the requirements for real-time operation and improves accuracy compared to conventional methods. Aoki Takanose, Eijiro Takeuchi, Alexander Carballo, Junichi Meguro, Kazuya Takeda |
IV | 5 |
| 2023 | RSG-Search: Semantic Traffic Scene Retrieval Using Graph-Based Scene RepresentationabstractBrowsing specific traffic scene in large-scale dataset is an increasing demand from researchers, self-driving community and insurance companies. It is easy to search scenes with specific tags such as "rain", "snow", or "on highway". However, searching specific scene configurations, like "two vehicles waiting for a person crossing the road", is still an open problem. In this paper, we provide RSG-search, a scene-graph based traffic scene retrieval method, based on our previous research on traffic scene-graph generation. By previously translating open datasets to scene graphs, we can ignore irrelevant details, and efficiently search specific scene configuration among thousands of traffic scenes. Experiment results shows that our graph searching method is able to retrieve results for a given query with high accuracy. Our method simplifies the task of scene retrieval, opening opportunities for new applications. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Kazuya Takeda |
IV | 4 |
| 2023 | Synthesizing Realistic Snow Effects in Driving Images Using GANs and Real Data with Semantic GuidanceabstractIntelligent vehicle perception algorithms often have difficulty accurately analyzing and interpreting images in adverse weather conditions. Snow is a corner case that not only reduces visibility and contrast but also affects the stability of the road environment. While it is possible to train deep learning models on real-world driving datasets in snow weather, obtaining such data can be challenging. Synthesizing snow effects on existing driving datasets is a viable alternative. In this work, we propose a method based on Cycle Consistent Generative Adversarial Networks (CycleGANs) that utilizes additional semantic information to generate snow effects. We apply deep supervision by using intermediate outputs from the last two convolutional layers in the generator as multi-scale supervision signals for training. We collect a small set of driving image data captured under heavy snow as the translation source. We compare the generated images with those produced by various network architectures and evaluate the results qualitatively and quantitatively on the Cityscapes and EuroCity Persons datasets. Experiment results indicate that our model can synthesize realistic snow effects in driving images. Hanting Yang, Ming Ding 0002, Alexander Carballo, Kento Ohtani, Yingjie Niu, Maoning Ge, Kazuya Takeda |
IV | 9 |
| 2023 | LiDAR Point Cloud Translation Between Snow and Clear Conditions Using Depth Images and GANsabstractSnow corrupts LiDAR point clouds with scattered noise points and false objects, posing a serious threat to the perception of autonomous driving systems. Existing effective point cloud de-snow methods are mainly based on outlier filters that rigidly remove isolated points. There are deep-learning and algorithm-based weather models that can handle adverse conditions such as rain and fog, but snow conditions are rarely considered. In this study, we propose a LiDAR point cloud translation model based on refined generative adversarial networks (GANs) that is not only able to de-noise snow in point clouds but also to generate fake snow points on clear data. Our model is trained on depth image representations of point clouds from unpaired datasets, with a customized loss function for grayscale depth images that can maintain scale consistency. A pixel-wise discriminator structure is designed to improve the de-snowing effect around the ego vehicle. The proposed model expresses a better feature capture on snow in LiDAR point clouds, and experiment results show high-quality snow removal performance on both the scattered and clustered snow points, as well as satisfactory fake snow generation on clear road point clouds. Ming Ding 0002, Hanting Yang, Yingjie Niu, Maoning Ge, Alexander Carballo, Kazuya Takeda |
IV | 8 |
| 2022 | Improving Dense Representation Learning by Superpixelization and Contrasting Cluster Assignment
Robin Karlsson, Tomoki Hayashi, Keisuke Fujii 0001, Alexander Carballo, Kento Ohtani, Kazuya Takeda |
BMVC | 6 |
| 2022 | Estimating counterfactual treatment outcomes over time in multi-vehicle simulationabstractEvaluation of intervention in a multi-agent system, e.g., when humans should intervene in autonomous driving systems, is challenging in various engineering and scientific fields. Estimating the individual treatment effect (ITE) using counterfactual long-term prediction is practical to evaluate such interventions. However, most of the conventional frameworks did not consider the time-varying complex structure of multi-agent relationships and covariate counterfactual prediction. Here we propose an interpretable, counterfactual recurrent network in multi-agent systems to estimate the effect of the intervention. Our model leverages graph variational recurrent neural networks and theory-based computation with domain knowledge for the ITE estimation framework based on long-term prediction of multi-agent covariates and outcomes, which can confirm the circumstances under which the intervention is effective. On simulated models of an automated vehicle with time-varying confounders, we show that our methods achieved lower estimation errors in counterfactual covariates. Keisuke Fujii 0001, Koh Takeuchi 0001, Atsushi Kuribayashi, Naoya Takeishi, Yoshinobu Kawahara, Kazuya Takeda |
SIGSPATIAL/GIS | 6 |
| 2022 | Driving Risk and Intervention: Subjective Risk Lane Change DatasetabstractWhen developing truly driverless mobility for the future, one key index used to measure the matureness of a particular self-driving technology is the driver intervention rate. One method which has proven to be effective for decreasing intervention rates is the use of personalized driving models that can mimic the driving style and preferences of a targeted user, so that autonomous driving feels safer and more natural to them. To create such models, quantitative data should be collected from users in order to determine the style of driving that a particular user, or type of user, prefers. In this paper, we introduce the Subjective Risk Lane Change (SRLC) Dataset, which includes ego vehicle driving behavior data, surrounding vehicle location information, and the subjective risk scores of users, collected during both safe and risky lane change scenarios encountered in CARLA simulators, as well as demographic information for our 30 participants. Furthermore, user intervention data for all of our participants was collected from Personalized Model Predictive Controllers during the generated lane change maneuvers. As far as the authors are able to determine, no other public dataset provides driving behavior signal and intervention timing information collected during driver interventions. Our dataset can be used to gain insights into a variety of personal driving styles, allowing the improvement of adaptive autonomous driving systems, and leading to safer and more widely accepted driverless technology. Naren Bao, Alexander Carballo, Kazuya Takeda |
IV | 3 |
| 2022 | Auditory and visual warning information generation of the risk object in driving scenes based on weakly supervised learningabstractIn this research, a two-stage risk object warning method is proposed to generate the auditory and visual warning information simultaneously from the driving scene. The auditory warning module (AWM) is designed as a classification task by combining the rough location and type information as warning sentences and treating each sentence as one class. The visual warning module (VWM) is designed as a weakly supervised method to save the labor-intensive bounding box marking of risk objects. To confirm the effectiveness of the proposed method, we also create a linguistic risk notification (LRN) dataset by describing the driving scenario as several different sentences. The average accuracy of auditory warning is 96.4% for generating the warning sentences. The average accuracy of the weakly supervised visual warning algorithm is 81.3% for getting the risk vehicle localization without any supervisory information. Yingjie Niu, Ming Ding 0002, Kento Ohtani, Kazuya Takeda |
IV | 5 |
| 2022 | An enhanced driver's risk perception modeling based on gate recurrent unit networkabstractRisk perception is one of the most important driving skills for drivers to detect potential traffic accident and make correct risk avoidance behaviors. Accurate model and evaluation of risk perception can effectively identify driver’s perception deficiencies and serve as an important human factor for the design of advance driver assistance systems. Most traditional perception ability assessing methods are based on macroscopic statistical results and lack effective mathematical models, thus making it difficult to evaluate risk perception quantitatively. To this end, this paper first obtains semantic understanding of traffic scenes through a semantic segmentation method based on deep learning. Then, the semantic understanding information is fused with the driver’s risk perception results to form the time series data reflecting the risk perception ability. Finally, we use gate recurrent unit (GRU) network-based learning framework to learn the time series data features and thus obtain the behavioral model which responds to the driver’s risk perception ability. By verifying the classification performance on multiple drivers, the risk perception model based on traffic scene understanding can accurately predict the actual cognitive behavior of individual driver. Besides, the proposed GRU-based method outperforms the traditional machine learning algorithm in modeling the driver’s risk perception behavior. Peng Ping, Weiping Ding 0001, Yongkang Liu 0005, Kazuya Takeda |
IV | 4 |
| 2022 | Real-to-Synthetic: Generating Simulator Friendly Traffic Scenes from Graph RepresentationabstractReproducing real-world traffic scenes in the simulator is fundamental to training self-driving systems. Creating a simulation scenario is a complex task, generally done manually: the ego-vehicle and other entities are placed and their trajectories defined, trying to recreate some situation found in real traffic. To reduce the manual burden, here we propose the Real-to-Synthetic toolset. This toolset provides synthetic traffic scene in openDrive format, which can be directly simulated in many simulators such as SUMO or CARLA. Also, we provide a scene generator which generates near-realistic scene from minimum user effort. To maintain the similarity between real-world scene and generated one, here we introduce the concept “Road Scene Graph”(RSG). In this graph, nodes represent entities while edges stand for pairwise relationships. These relationships could be maintained in the scene generation process while the actor is generated according to the distribution sampled from real-world data. Experiments proved that by using “Road Scene Graph”, our scene generator proposes a much more convenient way to conFigure traffic scenes rather than manually defining every actor’s initial status and trajectories. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Kazuya Takeda |
IV | 4 |
| 2022 | Disentangled Bad Weather Removal GAN for Pedestrian DetectionabstractBad weather such as rain and haze will lower the visibility and contrast of captured images and often occur at the same time, which makes the situation worse. Existing CNN-based methods can achieve impressive results when processing each condition individually. But few works consider removing rain and haze under a unified framework. Besides, these models will experience degraded performance when the distribution of the data changes. In this work, we attempt to find a method that can remove all bad weather without switching between single task models and paired training data for driving scene applications like pedestrian detection. In specific, we adopt a disentanglement strategy to obtain weather layer from subtraction. The statistic distance is calculated between two weather layers from each pipeline. In addition, the weather layer serves as information guidance and input to the reconstruct generator. Experiments on mixed dataset from RainCityscapes and Foggy Cityscapes show that the effectiveness of pedestrian detector is improved. Finally, we contribute a new dataset called Realistic Driving Scene under Bad Weather (RDSBW), which contains over 70K real bad weather images. Hanting Yang, Alexander Carballo, Kazuya Takeda |
VTC Spring | 3 |
| 2021 | Automatic Generation of Road Trip Summary Video for Reminiscence and Entertainment using Dashcam VideoabstractVehicle dashboard cameras are becoming an increasingly popular kind of automotive accessory. While it is easy to obtain the high-definition video data recorded by dashcams using Secure Digital memory cards, this data is rarely used except for safety purposes because it takes substantial time and effort to review or edit many hours of such recorded videos. In this paper, we propose a new usage for this data through the automatic video editing system we have developed that can create enjoyable video summaries of road trips utilizing video and other data from the vehicle. We also report the results of comparisons between automatically edited videos created by the proposed system and manually edited videos created by study participants. The prototype developed in this study and the findings from our experiments will contribute to improving the driving experience by providing entertainment for automobile users after road trips, and by memorializing their travels. Kana Bito, Itiro Siio, Yoshio Ishiguro, Kazuya Takeda |
AutomotiveUI | 4 |
| 2021 | Prediction of Personalized Driving Behaviors via Driver-Adaptive Deep Generative ModelsabstractHuman drivers have complex and unique driving characteristics, even when driving in common, well-defined scenarios such as lane changes. In this study, we propose using probabilistic, deep generative models to predict personalized driving behavior including velocity, acceleration, and steering angle sequence. Probabilistic approaches are applied to model uncertainty in the driving behavior of individual drivers as a distribution, while surrounding vehicle and driver ID information are considered as given conditions in the distribution. We train individual driver models using real-world driving data, and use them to predict sequences of future driving behavior in dynamic environments, using historical data to take personal driving styles into account. Our results show that the proposed driver behavior modeling method is able to learn from a driver's vehicle operation data and their interactions with surrounding vehicles to reproduce their specific driving style. Naren Bao, Alexander Carballo, Kazuya Takeda |
IV | 3 |
| 2021 | How to monitor multiple autonomous vehicles remotely with few observers: An active management methodabstractIn this research, we proposed an active management method to tele-monitor and tele-operate more autonomous vehicles (AVs) with few observers by adjusting the movement of the AVs actively. A management system is created to get the status from the AVs and separate the monitoring requirement to the observers optimally. When the requirements might be intensive, the management system can also adjust the movements of the AVs actively to distribute the monitoring time, which can make the observers monitor more vehicles. We implement and verify the management system in an autonomous driving simulator - CARLA for the limited number of AVs and observers. Based on the data acquired from the driving simulator, we also create a numerical simulator and tested our method with the pairs of a large number of AVs and observers. The result of both simulators shows that our method can reduce the utilization degree of observers and make them monitor more AVs. Ming Ding 0002, Eijiro Takeuchi, Yoshio Ishiguro, Yoshiki Ninomiya, Nobuo Kawaguchi, Kazuya Takeda |
IV | 6 |
| 2021 | A Comparison of Methods for Sharing Recognition Information and Interventions to Assist Recognition in Autonomous Driving SystemabstractAs research and development related to the practical use of autonomous driving systems (ADS) continue to advance, one of the remaining challenges is achieving both safe and natural autonomous driving. In part, this is due to the difficulty of achieving flawless, automated perception and understanding of the driving environment. In our previous study, we proposed a recognition assistance interface to solve this problem by sharing ADS recognition information with the passenger, allowing them to assist in the recognition stage of the autonomous driving process. In this study, we incorporate our recognition assistance interface into Autoware and test it in a simulated driving environment, using scenarios in which the ADS must recognize the intent of pedestrians and respond to the presence of trash in the road. Natural driving was dened as driving that avoids significant, unnecessary deceleration when performing these challenging recognition tasks. The results of our experiment with 11 participants showed that sharing recognition information with passengers is effective for avoiding unnecessary deceleration and achieving only minor variations in speed when encountering obstacles flagged in error by the recognition system. Atsushi Kuribayashi, Eijiro Takeuchi, Alexander Carballo, Yoshio Ishiguro, Kazuya Takeda |
IV | 5 |
| 2021 | RSG-Net: Towards Rich Sematic Relationship Prediction for Intelligent Vehicle in Complex EnvironmentsabstractBehavioral and semantic relationships play a vital role on intelligent self-driving vehicles and ADAS systems. Different from other research focused on trajectory, position, and bounding boxes, relationship data provides a human understandable description of the object's behavior, and it could describe an object's past and future status in an amazingly brief way. Therefore it is a fundamental method for tasks such as risk detection, environment understanding, and decision making. In this paper, we propose RSG-Net (Road Scene Graph Net): a graph convolutional network designed to predict potential semantic relationships from object proposals, and to produce a graph-structured result, called “Road Scene Graph”. The experimental results indicate that this network, trained on Road Scene Graph dataset, could efficiently predict potential semantic relationships among objects around the ego-vehicle. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Kazuya Takeda |
IV | 4 |
| 2021 | Motion Analysis and Performance Improved Method for 3D LiDAR Sensor Data CompressionabstractContinuous point cloud data is being used more and more widely in practical applications such as mapping, localization and object detection in autonomous driving systems, but due to the huge volume of data involved, sharing and storing this data is currently expensive and difficult. One possible solution is the development of more efficient methods of compressing the data. Other researchers have proposed converting 3D point cloud data into 2D images, or using tree structures to store the data. In a previous study targeting streaming point cloud data, we proposed an MPEG-like compression method which utilizes simultaneous localization and mapping (SLAM) results to simulate LiDAR’s operating process. In this paper, instead of imitating MPEG, we propose new strategy for more efficient reference frame distribution and more natural frame prediction, and use a different algorithm to encode the residual, greatly improving the algorithm’s performance and its stability in different scenarios. We also discuss how various parameters affect compression performance. Using our proposed method, streaming point cloud data collected by LiDAR sensors can be compressed to 1/50th of its original size, with only 2 cm of Root Mean Square Error for each detected point. We evaluate our proposed method by comparing its performance with several other existing point cloud compression methods in three different driving scenarios, demonstrating that our proposed method outperforms them. Chenxi Tu, Eijiro Takeuchi, Alexander Carballo, Chiyomi Miyajima, Kazuya Takeda |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech ToolkitabstractThis paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit supports state-of- the-art E2E-TTS models, including Tacotron 2, Transformer TTS, and FastSpeech, and also provides recipes inspired by the Kaldi automatic speech recognition (ASR) toolkit. The recipes are based on the design unified with the ESPnet ASR recipe, providing high reproducibility. The toolkit also provides pre-trained models and samples of all of the recipes so that users can use it as a baseline. Furthermore, the unified design enables the integration of ASR functions with TTS, e.g., ASR-based objective evaluation and semi- supervised learning with both ASR and TTS models. This paper describes the design of the toolkit and experimental evaluation in comparison with other toolkits. The experimental results show that our models can achieve state-of-the-art performance comparable to the other latest toolkits, resulting in a mean opinion score (MOS) of 4.25 on the LJSpeech dataset. The toolkit is publicly available at https://github.com/espnet/espnet. Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe 0001, Tomoki Toda, Kazuya Takeda, Yu Zhang 0033, Xu Tan 0003 |
ICASSP | 7 |
| 2020 | Weakly-Supervised Sound Event Detection with Self-AttentionabstractIn this paper, we propose a novel sound event detection (SED) method that incorporates a self-attention mechanism of the Transformer for a weakly-supervised learning scenario. The proposed method utilizes the Transformer encoder, which consists of multiple self-attention modules, allowing to take both local and global context information of the input feature sequence into account. Furthermore, inspired by the great success of BERT in the natural language processing field, the proposed method introduces a special tag token into the input sequence for weak label prediction, which enables the aggregation of the whole sequence information. To demonstrate the performance of the proposed method, we conduct the experimental evaluation using the DCASE2019 Task4 dataset. The experimental results demonstrate that the proposed method outperforms the DCASE2019 Task4 baseline method, which is based on the convolutional recurrent neural network, and the self-attention mechanism effectively works for SED. Koichi Miyazaki, Tatsuya Komatsu, Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Kazuya Takeda |
ICASSP | 6 |
| 2020 | End-to-End Automatic Speech Recognition Integrated with CTC-Based Voice Activity DetectionabstractThis paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification (CTC) and its extension of CTC/attention architectures. As opposed to an attention-based architecture, input-synchronous label prediction can be performed based on a greedy search with the CTC (pre-)softmax output. This prediction includes consecutive long blank labels, which can be regarded as a non-speech region. We use the labels as a cue for detecting speech segments with simple thresholding. The threshold value is directly related to the length of a non-speech region, which is more intuitive and easier to control than conventional VAD hyperparameters. Experimental results on unsegmented data show that the proposed method outperformed the baseline methods using the conventional energy-based and neural-network-based VAD methods and achieved an RTF less than 0.2. The proposed method is publicly available. Takenori Yoshimura, Tomoki Hayashi, Kazuya Takeda, Shinji Watanabe 0001 |
ICASSP | 3 |
| 2020 | Intelligibility Enhancement Based on Speech Waveform Modification Using Hearing Impairment
Shu Hikosaka, Shogo Seki, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, Hideki Banno, Tomoki Toda |
INTERSPEECH | 5 |
| 2020 | LIBRE: The Multiple 3D LiDAR DatasetabstractIn this work, we present LIBRE: LiDAR Benchmarking and Reference, a first-of-its-kind dataset featuring 10 different LiDAR sensors, covering a range of manufacturers, models, and laser configurations. Data captured independently from each sensor includes three different environments and configurations: static targets, where objects were placed at known distances and measured from a fixed position within a controlled environment; adverse weather, where static obstacles were measured from a moving vehicle, captured in a weather chamber where LiDARs were exposed to different conditions (fog, rain, strong light); and finally, dynamic traffic, where dynamic objects were captured from a vehicle driven on public urban roads, multiple times at different times of the day, and including supporting sensors such as cameras, infrared imaging, and odometry devices. LIBRE will contribute to the research community to (1) provide a means for a fair comparison of currently available LiDARs, and (2) facilitate the improvement of existing self-driving vehicles and robotics-related software, in terms of development and tuning of LiDAR-based perception algorithms. Alexander Carballo, Jacob Lambert 0001, Abraham Monrroy Cano, David Robert Wong, Patiphon Narksri, Yuki Kitsukawa, Eijiro Takeuchi, Shinpei Kato, Kazuya Takeda |
IV | 9 |
| 2020 | Point Grid Map-Based Mid-To-Mid Driving without Object DetectionabstractTeaching autonomous vehicles to imitate human driving in complex, urban traffic scenarios is a difficult task. “End-to-end” autonomous driving systems, based on “imitation learning”, are an expecting approach. A model learns the relationships between sensing input and vehicle control signal outputs. These methods can successfully achieve driving in simple scenarios such as lane keeping. In contrast, the “mid-to-mid” autonomous driving methods now being proposed. In such framework, the model learns the relationships between pre-processed feature maps from the model-based system as input and the future position of the ego vehicle as the output. Mid-to-mid driving methods can direct vehicles more robustly than end-to-end driving methods in some complex driving environments. However, mid-to-mid driving methods use the results of the object detection module to create the feature map. If object detection fails, or detection performance is poor due to changes in the driving environment, prediction performance may also be degraded. Our proposed method uses a prediction module that outputs point grid maps directly, without the use of an object detection module, which are then incorporated into the feature map. Point grid maps represent the locations of surrounding vehicles and obstacles directly, based on LiDAR point cloud data. Since the results of object detection are not used by the prediction module, detection performance does not affect prediction performance. In this study we conduct two experiments, an off-line evaluation using a Lyft dataset, and an on-line evaluation using the CARLA simulator. The results show that our model can achieve the same level of ego-vehicle position prediction performance as a model using annotated object location information. Shunya Seiya, Alexander Carballo, Eijiro Takeuchi, Kazuya Takeda |
IV | 4 |
| 2019 | Attention-Based Speech Recognition Using Gaze InformationabstractWe assume that there is a correlation between an utterance and a corresponding gaze object, and propose a new paradigm of multi-modal end-to-end speech recognition using multimodal information, namely, utterances and corresponding gaze points. In our method, the system extracts acoustic features and corresponding images around gaze points, and inputs the information into the proposed attention-based multiple encoder-decoder networks. This makes it possible to integrate the two different modalities, and the performance of speech recognition is improved. To evaluate the proposed method, we prepared a simulation task of power-line control operations, and built a corpus that contains utterances and corresponding gaze points in the operations. We conducted an experimental evaluation using this corpus, and the results showed the reduction in the CER, suggesting the effectiveness of the proposed method in which acoustic features and gaze information are integrated. Osamu Segawa, Tomoki Hayashi, Kazuya Takeda |
ASRU | 3 |
| 2019 | Scene-dependent Anomalous Acoustic-event Detection Based on Conditional Wavenet and I-vectorabstractThis paper proposes a scene-dependent anomalous acoustic-event detection based on conditional WaveNet and i-vector. The WaveNet builds normal acoustic event models by exhaustive learning of time-domain signals in the public space to provide scene-independent anomaly detection. I-vectors are used as additional features to describe acoustic scenes, where the input signals are observed, to complement the WaveNet. The proposed method can detect anomalous acoustic-events in environments whose acoustic scenes vary depending on time, location, and surrounding environment. Evaluations with data recorded from the real environment demonstrate that the proposed method achieved as much as 15 pt higher F-measure than LSTM and AE. The difference in F-measure by the WaveNet with and without i-vector turned out to be 1.5 pt. Tatsuya Komatsu, Tomoki Hayashi, Reishi Kondo, Tomoki Toda, Kazuya Takeda |
ICASSP | 5 |
| 2019 | A Predictive Reward Function for Human-Like Driving Based on a Transition Model of Surrounding EnvironmentabstractDriving is a complex task that requires the perception of the surrounding environment, decision making and control of the vehicle. Human drivers predict how surrounding objects move and decide an appropriate driving behavior. As with human drivers, autonomous driving vehicles should consider the condition of the surrounding environment and behave naturally so as not to disturb the traffic flow. We propose a reward function for learning how natural the driving is based on the hypothesis that the movement of surrounding vehicles becomes unpredictable when the ego vehicle takes an unnatural driving behavior. The reward function is based on the prediction error of a deep predictive network that models the transition of the surrounding environment. Occupancy grid image is used to perceive the surrounding environment and the predictions up to two seconds are used to calculate the reward function. We evaluated the reward function using both simulated and the real world data. We trained the prediction network using real driving data and trained a reinforcement learning agent based on the reward function. Then we compared the speed planned by the agent and a human driver, which showed a correlation of 0.52. We also confirmed the benefit of taking prediction into account by observing the behavior of the agent in a specific traffic scenario. Daiki Hayashi, Takashi Bando, Kazuya Takeda |
ICRA | 4 |
| 2019 | Point Cloud Compression for 3D LiDAR Sensor using Recurrent Neural Network with Residual BlocksabstractThe use of 3D LiDAR, which has proven its capabilities in autonomous driving systems, is now expanding into many other fields. The sharing and transmission of point cloud data from 3D LiDAR sensors has broad application prospects in robotics. However, due to the sparseness and disorderly nature of this data, it is difficult to compress it directly into a very low volume. A potential solution is utilizing raw LiDAR data. We can rearrange the raw data from each frame losslessly in a 2D matrix, making the data compact and orderly. Due to the special structure of 3D LiDAR data, the texture of the 2D matrix is irregular, in contrast to 2D matrices of camera images. In order to compress this raw, 2D formatted LiDAR data efficiently, in this paper we propose a method which uses a recurrent neural network and residual blocks to progressively compress one frame's information from 3D LiDAR. Compared to our previous image compression based method and generic octree point cloud compression method, the proposed approach needs much less volume while giving the same decompression accuracy. Potential application scenarios for point cloud compression are also considered in this paper. We describe how decompressed point cloud data can be used with SLAM (simultaneous localization and mapping) as well as for localization using a given map, illustrating potential uses of the proposed method in real robotics applications. Chenxi Tu, Eijiro Takeuchi, Alexander Carballo, Kazuya Takeda |
ICRA | 4 |
| 2019 | Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis
Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, Karen Livescu |
INTERSPEECH | 4 |
| 2019 | Robustness of Statistical Voice Conversion Based on Direct Waveform Modification Against Background Sounds
Yusuke Kurita, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
INTERSPEECH | 3 |
| 2018 | Multi-Head Decoder for End-to-End Speech RecognitionabstractThis paper presents a new network architecture called multihead decoder for end-to-end speech recognition as an extension of a multi-head attention model.In the multi-head attention model, multiple attentions are calculated, and then, they are integrated into a single attention.On the other hand, instead of the integration in the attention level, our proposed method uses multiple decoders for each attention and integrates their outputs to generate a final output.Furthermore, in order to make each head to capture the different modalities, different attention functions are used for each head, leading to the improvement of the recognition performance with an ensemble effect.To evaluate the effectiveness of our proposed method, we conduct an experimental evaluation using Corpus of Spontaneous Japanese.Experimental results demonstrate that our proposed method outperforms the conventional methods such as locationbased and multi-head attention models, and that it can capture different speech/linguistic contexts within the attention-based encoder-decoder framework. Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Kazuya Takeda |
INTERSPEECH | 4 |
| 2018 | Back-Translation-Style Data Augmentation for end-to-end ASRabstractIn this paper we propose a novel data augmentation method for attention-based end-to-end automatic speech recognition (E2E-ASR), utilizing a large amount of text which is not paired with speech signals. Inspired by the back-translation technique proposed in the field of machine translation, we build a neural text-to-encoder model which predicts a sequence of hidden states extracted by a pre-trained E2E-ASR encoder from a sequence of characters. By using hidden states as a target instead of acoustic features, it is possible to achieve faster attention learning and reduce computational cost, thanks to sub-sampling in E2E-ASR encoder, also the use of the hidden states can avoid to model speaker dependencies unlike acoustic features. After training, the text-to-encoder model generates the hidden states from a large amount of unpaired text, then E2E-ASR decoder is retrained using the generated hidden states as additional training data. Experimental evaluation using LibriSpeech dataset demonstrates that our proposed method achieves improvement of ASR performance and reduces the number of unknown words without the need for paired data. Tomoki Hayashi, Shinji Watanabe 0001, Yu Zhang 0033, Tomoki Toda, Takaaki Hori, Ramón Fernandez Astudillo, Kazuya Takeda |
SLT | 7 |
| 2018 | Learning How to Drive in Blind Intersections from Human DataabstractIn this paper we present a method to learn how to drive in different types of blind intersections using expert driving data. We cluster different intersections based on the velocity of how drivers approach them, and train a linear SVM classifier for each class of intersection. Through clustering we found that there were three different classes of intersections in typical residential areas in Japan. We used inverse reinforcement learning (IRL) to build a driving model for each type of intersection. The models were trained from 308 trajectories traversed by 5 different drivers. The models and policies were implemented and evaluated in a ROS simulator where the agent is provided a global path, and upon it reaching an intersection, it selects the appropriate trained policy. By doing this, the simulated autonomous vehicle can perform proactive safe driving behaviors when approaching blind intersections. Kyle Sama, Luis Yoichi Morales Saiki, Naoki Akai, Eijiro Takeuchi, Kazuya Takeda |
SMC | 5 |
| 2017 | An investigation of multi-speaker training for wavenet vocoderabstractIn this paper, we investigate the effectiveness of multi-speaker training for WaveNet vocoder. In our previous work, we have demonstrated that our proposed speaker-dependent (SD) WaveNet vocoder, which is trained with a single speaker's speech data, is capable of modeling temporal waveform structure, such as phase information, and makes it possible to generate more naturally sounding synthetic voices compared to conventional high-quality vocoder, STRAIGHT. However, it is still difficult to generate synthetic voices of various speakers using the SD-WaveNet due to its speaker-dependent property. Towards the development of speaker-independent WaveNet vocoder, we apply multi-speaker training techniques to the WaveNet vocoder and investigate its effectiveness. The experimental results demonstrate that 1) the multispeaker WaveNet vocoder still outperforms STRAIGHT in generating known speakers' voices but it is comparable to STRAIGHT in generating unknown speakers' voices, and 2) the multi-speaker training is effective for developing the WaveNet vocoder capable of speech modification. Tomoki Hayashi, Akira Tamamori, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
ASRU | 4 |
| 2017 | BLSTM-HMM hybrid system combined with sound activity detection network for polyphonic Sound Event DetectionabstractThis paper presents a new hybrid approach for polyphonic Sound Event Detection (SED) which incorporates a temporal structure modeling technique based on a hidden Markov model (HMM) with a frame-by-frame detection method based on a bidirectional long short-term memory (BLSTM) recurrent neural network (RNN). The proposed BLSTM-HMM hybrid system makes it possible to model sound event-dependent temporal structures and also to perform sequence-by-sequence detection without having to resort to thresholding such as in the conventional frame-by-frame methods. Furthermore, to effectively reduce insertion errors of sound events, which often occurs under noisy conditions, we additionally implement a binary mask post-processing using a sound activity detection (SAD) network to identify segments with any sound event activity. We conduct an experiment using the DCASE 2016 task 2 dataset to compare our proposed method with typical conventional methods, such as non-negative matrix factorization (NMF) and a standard BLSTM-RNN. Our proposed method outperforms the conventional methods and achieves an F1-score 74.9 % (error rate of 44.7 %) on the event-based evaluation, and an F1-score of 80.5 % (error rate of 33.8 %) on the segment-based evaluation, most of which also outperforms the best reported result in the DCASE 2016 task 2 challenge. Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Takaaki Hori, Jonathan Le Roux, Kazuya Takeda |
ICASSP | 6 |
| 2017 | Music staging AIabstractThrough smartphones, user enables to download/listen music anytime and anywhere. As a concept of a future audio player, we propose a framework of "music staging artificial intelligence (AI)". In that framework, audio object signals, e.g. vocal, guitar, bass, drums and keyboards, are assumed to be extracted from stereo music signals. To visualize music as if live performance is virtually conducted, playing motion sequence is estimated by using separated signals. After adjusting the spatial arrangement of audio objects so as to each user prefers it, audio/visual rendering is conducted. We constructed two types of demonstration systems for music staging AI. In the smartphone-based implementation, each user enables to change the spatial arrangement through sliderbar dragging. Since information of user preferable spatial arrangement can be sent from each smartphone to server, it would enable to predict/recommend the user preferable spatial arrangement. In another implementation, head mount display (HMD) was utilized to dive into virtual music live performance. Each user enables to walk/teleport anywhere and audio is then changing corresponding to the user view. Kenta Niwa, Kento Ohtani, Kazuya Takeda |
ICASSP | 3 |
| 2017 | Speaker-Dependent WaveNet Vocoder
Akira Tamamori, Tomoki Hayashi, Kazuhiro Kobayashi, Kazuya Takeda, Tomoki Toda |
INTERSPEECH | 4 |
| 2017 | Continuous point cloud data compression using SLAM based predictionabstractInterest in continuous point cloud data has increased rapidly as it becomes more widely used in practical applications, such as the development of autonomous driving systems. Sharing and storing continuous point cloud data is currently expensive and difficult due to the huge volume of data involved, however. As a result, developing efficient methods of compressing cloud point data has become an urgent task. Current methods compress the data directly using tree structures or height map-based methods. Taking into consideration the specialized requirements of autonomous driving, in [1] we proposed a method of compressing raw point cloud data using image compression methods and obtained superior results. Image compression-based methods are not able to utilize the 3D characteristics of point clouds, however. Therefore, in this paper we propose a new compression method using location and orientation information from Simultaneous Localization and Mapping (SLAM). Proposed method can take advantage of the 3D characteristic of point cloud by a predicting process which simulating the working procedure of 3D LiDAR. Strategies from MPEG and DPCM compression are utilized and several issues which can affect performance are discussed. In addition, based on proposed method, this paper discusses the ways to segment steam point cloud for compression task. We call this segmentation process adaptive sequence decision. We then compare the proposed SLAM-based method with image compression-based methods in various situations. Our experimental results show that, while requiring a small margin of error, the SLAM-based method outperforms image compression-based methods. Chenxi Tu, Eijiro Takeuchi, Chiyomi Miyajima, Kazuya Takeda |
Intelligent Vehicles Symposium | 4 |
| 2017 | Duration-Controlled LSTM for Polyphonic Sound Event DetectionabstractThis paper presents a new hybrid approach called duration-controlled long short-term memory (LSTM) for polyphonic sound event detection (SED). It builds upon a state-of-the-art SED method that performs frame-by-frame detection using a bidirectional LSTM recurrent neural network (BLSTM), and incorporates a duration-controlled modeling technique based on a hidden semi-Markov model. The proposed approach makes it possible to model the duration of each sound event precisely and to perform sequence-by-sequence detection without having to resort to thresholding, as in conventional frame-by-frame methods. Furthermore, to effectively reduce sound event insertion errors, which often occur under noisy conditions, we also introduce a binary-mask-based postprocessing that relies on a sound activity detection network to identify segments with any sound event activity, an approach inspired by the well-known benefits of voice activity detection in speech recognition systems. We conduct an experiment using the DCASE2016 task 2 dataset to compare our proposed method with typical conventional methods, such as nonnegative matrix factorization and standard BLSTM. Our proposed method outperforms the conventional methods both in an event-based evaluation, achieving a 75.3% F1 score and a 44.2% error rate, and in a segment-based evaluation, achieving an 81.1% F1 score, and a 32.9% error rate, outperforming the best results reported in the DCASE2016 task 2 Challenge. Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Takaaki Hori, Jonathan Le Roux, Kazuya Takeda |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Robust Example Search Using Bottleneck Features for Example-Based Speech Enhancement
Atsunori Ogawa, Shogo Seki, Keisuke Kinoshita, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, Kazuya Takeda |
INTERSPEECH | 7 |
| 2016 | Integrating driving behavior and traffic context through signal symbolizationabstractThis paper presents a novel method for integrating driving behavior and traffic context through signal symbolization in order to summarize driving semantics from sensor outputs. The method has been applied to risky lane change detection. Language models (nested Pitman-Yor language model) and speech recognition algorithms (hidden Markov Model) have been utilized for converting continuous sensor signals into a sequence of non-uniform segments (chunks). After symbolization, Latent Dirichlet Allocation (LDA) is used to integrate the symbolized driving behavior and the surrounding vehicle information for establishing the semantics of the driving scene. 988 lane changes of real-world highway driving are used for the evaluation. Risk level of each lane change rated by 10 subjects are used as ground truth. Best results have been obtained when driving behavior and surrounding vehicle information are integrated through co-occurrence chunking after independent symbolization of behavior and context signals. Suguru Yamazaki, Chiyomi Miyajima, Ekim Yurtsever, Kazuya Takeda, Masataka Mori, Kentarou Hitomi, Masumi Egawa |
Intelligent Vehicles Symposium | 4 |
| 2016 | Accelerated Deformable Part Models on GPUsabstractObject detection is a fundamental challenge facing intelligent applications. Image processing is a promising approach to this end, but its computational cost is often a significant problem. This paper presents schemes for accelerating the deformable part models (DPM) on graphics processing units (GPUs). DPM is a well-known algorithm for image-based object detection, and it achieves high detection rates at the expense of computational cost. GPUs are massively parallel compute devices designed to accelerate data-parallel compute-intensive workload. According to an analysis of execution times, approximately 98 percent of DPM code exhibits loop processing, which means that DPM could be highly parallelized by GPUs. In this paper, we implement DPM on the GPU by exploiting multiple parallelization schemes. Results of an experimental evaluation of this GPU-accelerated DPM implementation demonstrate that the best scheme of GPU implementations using an NVIDIA GPU achieves a speed up of 8.6x over a naive CPU-based implementation. Manato Hirabayashi, Shinpei Kato, Masato Edahiro, Kazuya Takeda, Seiichi Mita |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | Exploring multi-channel features for denoising-autoencoder-based speech enhancementabstractThis paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Although multi-channel speech enhancement usually outperforms single channel approaches, there has been little research on the use of multi-channel processing in the context of DAE. In this paper, we explore the use of several multi-channel features as DAE input to confirm whether multi-channel information can improve performance. Experimental results show that certain multi-channel features outperform both a monaural DAE and a conventional time-frequency-mask-based speech enhancement method. Shoko Araki, Tomoki Hayashi, Marc Delcroix, Masakiyo Fujimoto, Kazuya Takeda, Tomohiro Nakatani |
ICASSP | 5 |
| 2015 | Integration of deep bottleneck features for audio-visual speech recognition
Hiroshi Ninomiya, Norihide Kitaoka, Satoshi Tamura, Yurie Iribe, Kazuya Takeda |
INTERSPEECH | 5 |
| 2015 | Analyzing driver gaze behavior and consistency of decision making during automated drivingabstractWe investigate a possible method for detecting a driver's negative adaptation to an automated driving system by analyzing consistency of driver decision making and driver gaze behavior during automated driving. We focus on an automated driving system equivalent to Level 2 automation per the NHTSA's definition. At this level of automation, drivers must be ready to take control of the vehicle in critical situations by monitoring the driving environment and vehicle behavior. Since drivers are not required to operate the pedals or steering wheel during automated driving, a driver's negative adaptation to an automated system needs to be detected from behavior other than vehicle operation. In this study, we focus on driver gaze behavior. We conduct a simulator study to compare the gaze behavior of fifteen drivers during conventional and automated driving. We also analyze the consistency of driver decision making when changing lanes during conventional and automated driving. Experimental results show that drivers who pay less attention to the road ahead during automated driving tend to be less sensitive to risk factors in the surrounding environment and also tend to make inconsistent lane change decisions during automated driving. Chiyomi Miyajima, Suguru Yamazaki, Takashi Bando, Kentarou Hitomi, Hitoshi Terai, Hiroyuki Okuda, Takatsugu Hirayama, Masumi Egawa, Tatsuya Suzuki 0001, Kazuya Takeda |
Intelligent Vehicles Symposium | 10 |
| 2015 | Automatic lane change extraction based on temporal patterns of symbolized driving behavioral dataabstractThis paper proposes a method of automatically extracting lane change situations from large-scale driving corpora. Naturalistic driving data stored in large-scale corpora has a potential of contributing for developing novel advanced driver-assistance systems based on estimated information about driver's intent and/or potential risk of accidents. However, direct estimation of such kind of information from stream data is difficult. To address the issue, we apply an unsupervised symbolization method and topic representation to driving data. Driving stream data is converted to sequences of discrete symbols by a non-parametric symbolization method, and then the symbols are characterized by topics which represent typical distribution of driving behavior observed during the symbols. Because these symbols are separated on changing points of driving behavior, similar driving situations are effectively retrieved from sequences of the symbols. For evaluating effectiveness of the symbolization approach, we extract lane change situations based on the topic proportions and their temporal patterns. Distinctive elements of topic proportions and their temporal patterns for lane change situations are extracted by AdaBoost classifier. As a result, proposed approach outperforms baselines with neither topic proportions nor their temporal patterns in terms of extracting lane change situations. This result shows effectiveness of symbols with topic proportions for representing characteristics of driving situations. Masataka Mori, Kazuhito Takenaka, Takashi Bando, Tadahiro Taniguchi, Chiyomi Miyajima, Kazuya Takeda |
Intelligent Vehicles Symposium | 6 |
| 2015 | Traffic trajectory history and drive path generation using GPS data cloudabstractThis paper proposes a novel approach for extracting the traffic trajectory history, with the use of GPS data collected over a certain period of time, to be used as an input for driver models. In this approach, driving curvature is distinguished from actual road shape curvature with the use of real driving data. After sufficient amount of drive data has been collected, high degree polynomials are fitted to GPS point cloud. Traffic trajectory history is the tangential unit vectors and curvature values that are calculated from these polynomials. Then a single drivers driving path has been predicted with using traffic trajectory history and road shape curvature for comparison and validation. Experimental results show that the predictions made with categorized traffic trajectory history have less errors than the predictions made with road shape curvature. Ekim Yurtsever, Kazuya Takeda, Chiyomi Miyajima |
Intelligent Vehicles Symposium | 2 |
| 2015 | Modeling of Physical Characteristics of Speech under StressabstractThis letter presents a method to perform the classification of speech under stress based on physical characteristics. A physical model is proposed to model airflow patterns in the physiological system in order to represent the process of speech production under psychological stress, and physical parameters characterizing airflow variations in the vocal folds, the vocal tract, and laryngeal ventricle are explored. Experimental evaluations show that the physical parameters are effective for the classification of stressed speech. Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
IEEE Signal Process. Lett. | 5 |
| 2014 | Stochastic modeling and disaggregation of energy-consumption behaviorabstractThis paper focuses on stochastic modeling and energy dis-aggregation based on conditional random fields (CRFs) using real-world energy consumption data. Firstly, energy-consumption activities modeling aims at understanding and identifying energy-consumption activities using behavior models based on observed energy signals. Our ultimate goal is to suggest ways to modify human behavior activities in order to conserve energy by optimizing the use of energy. Preliminary analysis of energy consumption data clearly shows the potential effectiveness of activity behavior changes on the changing energy consumption behavior. Secondly energy disaggregation aims at breaking up the total energy signal into its component appliances. This is very useful since it can provide home owners with feedback about the way they use electrical energy, and can also motivate users to conserve significant amounts of energy. In the current study, we focus on activity/event disaggregation using the total energy-consumption signal. Panikos Heracleous, Pongtep Angkititrakul, Kazuya Takeda |
ICASSP | 3 |
| 2013 | A Discussion on the Consistency of Driving Behavior across Laboratory and Real Situational Studies
Hitoshi Terai, Kazuhisa Miwa, Hiroyuki Okuda, Yuichi Tazaki, Tatsuya Suzuki 0001, Kazuaki Kojima, Junya Morita, Akihiro Maehigashi, Kazuya Takeda |
CogSci | 9 |
| 2013 | Analysis and modeling of entrainment in chorus singingabstractThe dynamics of the contour of the fundamental frequency (F0) of singing voices in a chorus is analyzed from the view point of `entrainment' in singing behavior. The One-Mass-Two-Spring (OMTS) coupled system is used as the mathematical model of the contour of the F0of singing voices that are concurrently singing the same melody. Using this model, the characteristics of the F0dynamics of a voice singing in a chorus are parameterized by the mass, the coefficients of friction, and the spring factors of an OMTS system. It is experimentally confirmed that a steepest decent method can estimate the four model parameters, so that the model can generate the F0contour of a voice singing in a chorus with a less than 44.4 cents of RMS error. Preliminary experiments also show that experienced and novice singers can be correctly identified using the parameters of our model, because their entrainment behaviors are significantly different. Motonari Kawagishi, Shota Kawabuchi, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 5 |
| 2013 | Modeling head-related transfer functions via spatial-temporal Gaussian processabstractWe propose a novel application of a family of non-parametric statistical models to estimate head-related transfer functions (HRTFs) using spatial-temporal Gaussian processes (GPs). In this approach, we model the head-related impulse response (HRIR) utilizing non-parametric regression via a GP. The challenge posed by this problem involves accurate modeling of the spatial correlation structure jointly with the temporal correlation structure at each spatial location for the HRIR. We solve this problem by constructing a joint spatial-temporal kernel characterizing the GP regression model. To perform inference, we estimate the hyper-parameters of the GP regression kernel via maximum signal-to-deviation-ratio estimation on the basis of a real experimental setup in which we collected observations of the HRIR using two head-and-torso simulators (HATSs): KEMAR and B&K. We also perform cross validation of the model by training on the KEMAR system and assessing the generalization of our model and its out-of-sample predictive power for HRIRs at any locations that we predict by the model assessed on the B&K system. The corresponding HRTFs are obtained as the Fourier transform of the HRIRs. In the experiments, we show that our method is robust against variation in the azimuth interval needed to perform high-accuracy interpolation and has the expressive power to handle the individual characteristics of each HATS. Tatsuya Komatsu, Takanori Nishino, Gareth W. Peters, Tomoko Matsui, Kazuya Takeda |
ICASSP | 5 |
| 2013 | Computationally efficient single channel dereverberation based on complementary wiener filterabstractA single-channel dereverberation method with low computational complexity is proposed. We introduce the complementary Wiener filter which can suppress a late reverberation during silence intervals via theoretical analysis and numerical calculation. An implementation represents reductions both of memory consumption and operative calculations compared to a conventional method; each reduction is almost by half. Dereverberation performance is evaluated by an experimental simulation using speech signals and measured room impulse responses. The performance under several hundreds of msec of the reverberation time is similar to the conventional method: 6 [dB] reverberation reduction and 3 [dB] improvement of target-to-interference ratio. Kazunobu Kondo, Yu Takahashi, Tatsuya Komatsu, Takanori Nishino, Kazuya Takeda |
ICASSP | 5 |
| 2013 | Estimation of vocal tract parameters for the classification of speech under stressabstractIn this work, we propose a method for the classification of speech under stress that is based on a physical model. Using this method, the characteristics of the vocal folds and the vocal tract are taken into consideration, based on the process of speech production. In addition to vocal fold parameters, we estimate parameters of the vocal tract representing cross-sectional areas and vocal tract length, by fitting a two-mass model to real speech. Results show that calculation of vocal tract length for each speaker can improve the accuracy of the estimation of other physical parameters. Analysis is performed under vowel-dependent and vowel-independent conditions, showing that the proposed physical features are effective for the classification of neutral and stressed speech. Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 5 |
| 2013 | Classification of speech under stress by modeling the aerodynamics of the laryngeal ventricle
Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 5 |
| 2013 | Modeling and Analysis of Driving Behavior Based on a Probability-Weighted ARX ModelabstractThis paper proposes a probability-weighted autoregressive exogenous (PrARX) model wherein the multiple ARX models are composed of the probabilistic weighting functions. This model can represent both the motion-control and decision-making aspects of the driving behavior. As the probabilistic weighting function, a “softmax” function is introduced. Then, the parameter estimation problem for the proposed model is formulated as a single optimization problem. The “soft” partition defined by the PrARX model can represent the decision-making characteristics of the driver with vagueness. This vagueness can be quantified by introducing the “decision entropy.” In addition, it can be easily extended to the online estimation scheme due to its small computational cost. Finally, the proposed model is applied to the modeling of the vehicle-following task, and the usefulness of the model is verified and discussed. Hiroyuki Okuda, Norimitsu Ikami, Tatsuya Suzuki 0001, Yuichi Tazaki, Kazuya Takeda |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2012 | Multi-platform Experiment to Discuss Behavioral Consistency across Laboratory and Real Situational Studies
Hitoshi Terai, Kazuhisa Miwa, Hiroyuki Okuda, Yuichi Tazaki, Tatsuya Suzuki 0001, Kazuaki Kojima, Junya Morita, Akihiro Maehigashi, Kazuya Takeda |
CogSci | 9 |
| 2012 | Estimating sound source depth using a small-size arrayabstractA method for estimating the sound source depth, i.e., the distance between a source and receiver, using a small-size array is proposed. The proposed method uses the spatial distribution pattern of quasi-independent signal components obtained by the frequency-domain independent component analysis (FDICA) as the cue for depth estimation. The quasi-independent components are calculated by applying FDICA to array signals with very high redundancy, for example, 60 microphone signals for a pair of sources; therefore, signal components associated with reflection signals are obtained even though they are correlated with the direct signal. Experimental evaluation using a small-size microphone array with a large number of elements confirms that the average (RMS) estimation error of the proposed method is 0.33 m, which is sufficiently accurate for our applications. Satoshi Esaki, Kenta Niwa, Takanori Nishino, Kazuya Takeda |
ICASSP | 4 |
| 2012 | Physical characteristics of vocal folds during speech under stressabstractWe focus on variations in the glottal source of speech production, which is essential for understanding the generation of speech under psychological stress. In this paper, a two-mass vocal fold model is fitted to estimate the stiffness parameters of vocal folds during speech, and the stiffness parameters are then analyzed in order to classify recorded samples into neutral and stressed speech. Mechanisms of vocal folds under stress are derived from the experimental results. We propose using a Muscle Tension Ratio (MTR) to identify speech under stress. Our results show that MTR is more effective than a conventional method of stress measurement. Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 5 |
| 2012 | Classification of Stressed Speech Using Physical Parameters Derived from Two-Mass Model
Takatoshi Jitsuhiro, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 5 |
| 2012 | An improved driver-behavior model with combined individual and general driving characteristicsabstractIn this paper, we propose a stochastic driver-behavior modeling framework which takes into account both individual and general driving characteristics as one aggregate model. Patterns of individual driving styles are modeled using Dirichlet process mixture model, a nonparametric Bayesian approach which automatically selects the optimal number of model components to fit sparse observations of each particular driver's behavior. In addition, general or background driving patterns are also captured with a Gaussian mixture model using a reasonably large amount of development observed data from several drivers. By combining both probability distributions, the aggregate driver-dependent model can better emphasize driving characteristics of each particular driver, while also backing off to exploit general driving behavior in cases of unmatched parameter spaces from individual training observations. The proposed driver-behavior model was employed to anticipate pedal-operation behavior during car-following maneuvers involving several drivers on the road. The experimental results showed advantages of the combined model over the adapted model previously proposed. Pongtep Angkititrakul, Chiyomi Miyajima, Kazuya Takeda |
Intelligent Vehicles Symposium | 3 |
| 2012 | Causal analysis of task completion errors in spoken music retrieval interactions
Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
LREC | 3 |
| 2012 | Self-Coaching System Based on Recorded Driving Data: Learning From One's ExperiencesabstractThis paper describes the development of a self-coaching system to improve driving behavior by allowing drivers to review a record of their own driving activity. By employing stochastic driver-behavior modeling, the proposed system is able to detect a wide range of potentially hazardous situations, which conventional event data recorders are not able to capture, including those involving latent risks, of which drivers themselves are unaware. By utilizing these automatically detected hazardous situations, our web-based system offers a user-friendly interface for drivers to navigate and review each hazardous situation in detail (e.g., driving scenes are categorized into different types of hazardous situations and are displayed with corresponding multimodal driving signals). Furthermore, the system provides feedback on each risky driving behavior and suggests how users can safely respond to such situations. The proposed system establishes a cooperative relationship between the driver, the vehicle, and the driving environment, leading to the development of the next generation of safety systems and paving the way for an alternative form of driving education that could further reduce the number of fatal accidents. The system's potential benefits are demonstrated through preliminary extensive evaluation of an on-road experiment, showing that safe-driving behavior can be significantly improved when drivers use the proposed system. Kazuya Takeda, Chiyomi Miyajima, Tatsuya Suzuki 0001, Pongtep Angkititrakul, Kenji Kurumida, Yuichi Kuroyanagi, Hiroaki Ishikawa, Ryuta Terashima, Toshihiro Wakita, Masato Oikawa, Yuichi Komada |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2011 | Robust seed model training for speaker adaptation using pseudo-speaker features generated by inverse CMLLR transformationabstractIn this paper, we propose a novel acoustic model training method which is suitable for speaker adaptation in speech recognition. Our method is based on feature generation from a small amount of speakers' data. For decades, speaker adaptation methods have been widely used. Such adaptation methods need some amount of adaptation data and if the data is not sufficient, speech recognition performance degrade significantly. If the seed models to be adapted to a specific speaker can widely cover more speakers, speaker adaptation can perform robustly. To make such robust seed models, we adopt inverse maximum likelihood linear regression (MLLR) transformation-based feature generation, and then train our seed models using these features. First we obtain MLLR transformation matrices from a limited number of existing speakers. Then we extract the bases of the MLLR transformation matrices using PCA. The distribution of the weight parameters to express the MLLR transformation matrices for the existing speakers is estimated. Next we generate pseudo-speaker MLLR transformations by sampling the weight parameters from the distribution, and apply the inverse of the transformation to the normalized existing speaker features to generate the pseudo-speakers' features. Finally, using these features, we train the acoustic seed models. Using this seed models, we obtained better speaker adaptation results than using simply environmentally adapted models. Arata Itoh, Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
ASRU | 4 |
| 2011 | Driver risk evaluation based on acceleration, deceleration, and steering behaviorabstractWe propose a driver risk evaluation method based on the analysis of driving data captured with drive recorders. To evaluate the acceleration behavior of each driver we plot the maximum acceleration per minute to velocity on a two dimensional plane and approximate the distribution by linear regression. We assume that the higher the y-intercept of the line, the quicker the driver accelerates from a stop, and the higher the x-intercept, the higher the preferred speed of travel. To evaluate deceleration behavior, brake pedal operation patterns are classified into four types, based on how the brake is depressed and released. We evaluate deceleration risk levels based on these four braking pattern categories. Steering behavior is evaluated based on the relationship between the radius of road curvature and road design speed as defined in the road construction ordinance. Some correlation is observed between our evaluation results and those manually scored by risk consultants. Chiyomi Miyajima, Hiroki Ukai, Atsumi Naito, Hideomi Amata, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 6 |
| 2011 | Improving head-related impulse response measured in noisy environments with spatio-temporal frequency analysisabstractA new noise reduction method based on spatio-temporal frequency analysis is proposed that can be applied to head-related impulse response (HRIR), which is an impulse response between the sound source and the ear canal entrance. HRIR measurement is some times conducted in spaces which are not free fields, such as sound proof chambers with low reverberation. Conventional noise reduction methods such as synchronous averaging need multiple measurements under an identical measurement condition. However, taking many measurements requires much time, and subjects must be patient. Therefore, a noise reduction method which requires fewer measurements is needed. In our proposed method, noises are sup pressed using a two-dimensional filter that is designed in the spatio temporal frequency domain. Two experiments were conducted to improve the signal-to-noise ratio (SNR). The maximum SNR improvement achieved 1.28 dB, and the HRIRs measured in high reverberation were also improved, indicating that the proposed method effectively improved HRIRs without multiple measurements. Takanori Nishino, Kazuya Takeda |
ICASSP | 2 |
| 2011 | Detection of Task-Incomplete Dialogs Based on Utterance-and-Behavior Tag N-Gram for Spoken Dialog SystemsabstractWe propose a method of detecting “task incomplete” dialogs in spoken dialog systems using N-gram-based dialog models. We used a database created during a field test in which inexperienced users used a client-server music retrieval system with a spoken dialog interface on their own PCs. In this study, the dialog for a music retrieval task consisted of a sequence of user and system tags that related their utterances and behaviors. The dialogs were manually classified into two classes: the dialog either completed the music retrieval task or it didn’t. We then detected dialogs that did not complete the task, using N-gram probability models or a Support Vector Machine with N-gram feature vectors trained using manually classified dialogs. Off-line and on-line detection experiments were conducted on a large amount of real data, and the results show that our proposed method achieved good classification performance. Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 3 |
| 2011 | Modeling and adaptation of stochastic driver-behavior model with application to car followingabstractIn this paper, we present our recently developed stochastic driver-behavior model based on Gaussian mixture model (GMM) framework. The proposed driver-behavior modeling is employed to anticipate car-following behavior in terms of pedal control operations in response to the observable driving signals, such as the own vehicle velocity and the following distance to the leading vehicle. In addition, the proposed driver modeling allows adaptation scheme to enhance the model capability to better represent particular driving characteristics of interest (i.e., individual driving style) from the observed driving data themselves. Validation and comparison of the proposed driver-behavior models on realistic car-following data of several drivers showed the promising results. Furthermore, the adapted driver models showed consistent improvement over the unadapted driver models in both short-term and long-term predictions. Pongtep Angkititrakul, Chiyomi Miyajima, Kazuya Takeda |
Intelligent Vehicles Symposium | 3 |
| 2011 | Analysis of Real-World Driver's FrustrationabstractThis paper investigates a method for estimating a driver's spontaneous frustration in the real world. In line with a specific definition of emotion, the proposed method integrates information about the environment, the driver's emotional state, and the driver's responses in a single model. Driving data are recorded using an instrumented vehicle on which multiple sensors are mounted. While driving, drivers also interact with an automatic speech recognition (ASR) system to retrieve and play music. Using a Bayesian network, we combine knowledge on the driving environment assessed through data annotation, speech recognition errors, the driver's emotional state (frustration), and the driver's responses measured through facial expressions, physiological condition, and gas- and brake-pedal actuation. Experiments are performed with data from 20 drivers. We discuss the relevance of the proposed model and features of frustration estimation. When all of the available information is used, the overall estimation achieves a true positive rate of 80% and a false positive rate of 9% (i.e., the system correctly estimates 80% of the frustration and, when drivers are not frustrated, makes mistakes 9% of the time). Lucas Malta, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2011 | International Large-Scale Vehicle Corpora for Research on Driver Behavior on the RoadabstractThis paper considers a comprehensive and collaborative project to collect large amounts of driving data on the road for use in a wide range of areas of vehicle-related research centered on driving behavior. Unlike previous data collection efforts, the corpora collected here contain both human and vehicle sensor data, together with rich and continuous transcriptions. While most efforts on in-vehicle research are generally focused within individual countries, this effort links a collaborative team from three diverse regions (i.e., Asia, American, and Europe). Details relating to the data collection paradigm, such as sensors, driver information, routes, and transcription protocols, are discussed, and a preliminary analysis of the data across the three data collection sites from the U.S. (Dallas), Japan (Nagoya), and Turkey (Istanbul) is provided. The usability of the corpora has been experimentally verified with a Cohen's kappa coefficient of 0.74 for transcription reliability, as well as being successfully exploited for several in-vehicle applications. Most importantly, the corpora are publicly available for research use and represent one of the first multination efforts to share resources and understand driver characteristics. Future work on distributing the corpora to the wider research community is also discussed. Kazuya Takeda, John H. L. Hansen, Pinar Boyraz Baykas, Lucas Malta, Chiyomi Miyajima, Hüseyin Abut |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2010 | A small dodecahedral microphone array for blind source separationabstractA sound source separation method based on frequency-domain independent component analysis (FD-ICA) is proposed. This method fully utilizes the dodecahedral microphone array (DHMA), which has several merits: 1) the size of the array is very small and thus easy to handle; 2) the amplitude difference among microphones on the different surfaces is large; and 3) it is less affected by spatial aliasing in the higher frequency region. In the proposed method, in order to solve the permutation problem in FD-ICA through clustering acoustic transfer functions, amplitude and phase differences are optimally combined as a function of frequency. A DHMA of 8 cm in diameter with 60 microphones is used for the experiment, where up to twelve sound sources (speech/musical instruments) are separated using the proposed algorithm. The separation performance of the proposed method attains 24 dB in the signal-to-interference ratio (SIR) improvement score for the case of twelve sources. Since the performance is better by up to 10 dB in comparison to the conventional method, our results confirm the effectiveness of the proposed method. Motoki Ogasawara, Takanori Nishino, Kazuya Takeda |
ICASSP | 3 |
| 2010 | Analyzing grasping for inferring cognitive states of usersabstractWe study the effect of cognitive states, feelings about tasks, on grasping behavior to estimate user's feelings from their motion. Since people solve the inverse kinematics problem of grasping based on their cognition for the task, when they grasp an object, the way to grasp the object reflects their cognitive states. We are analyzing the way of grasping a cup depending on whether a user is stressed. The physical properties of grasping, volume and entropy of Grasp Jacobian ellipsoids are analyzed. The volume of Grasp Jacobian ellipsoids, which indicates the possible size of object movement, was shrunk after learning the grasp motion. Also the volumes between the relaxed and the stressed cognitive conditions were significantly different. These results show that the user's cognition for tasks reflects the grasp forms and the possible size of object movement. Kotaro Ogino, Takatoshi Jitsuhiro, Chiyomi Miyajima, Kazuya Takeda |
ICASSP | 4 |
| 2010 | Automatic detection of task-incompleted dialog for spoken dialog system based on dialog act n-gramabstractIn this paper, we propose a method of detecting task- incompleted users for a spoken dialog system using an N-gram- based dialog history model. We collected a large amount of spoken dialog data accompanied by usability evaluation scores by users in real environments. The database was made by a field test in which naive users used a client-server music retrieval system with a spoken dialog interface on their own PCs. An N-gram model was trained from sequences that consist of user dialog acts and/or system dialog acts for two dialog classes, that is, the dialog completed the music retrieval task or the dialog incompleted the task. Then the system detects unknown dialogs that is not completed the task based on the N-gram likelihood. Experiments were conducted on large real data, and the results show that our proposed method achieved good classification performance. When the classifier correctly detected all of the task-incompleted dialogs, our proposed method achieved a false detection rate of 6%. Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 3 |
| 2010 | A browsing and retrieval system for driving dataabstractWith the increased presence and recent advances of drive recorders, rich driving data that include video, vehicle acceleration signals, driver speech, GPS data, and several sensor signals can be continuously recorded and stored. These advances enable researchers to study driving behavior more extensively for traffic safety. However, increasing the variety and the amount of driving data complicates the simultaneous browsing of various data and finding desired data from large databases. In this study, we develop a browsing and retrieval system for driving data that provides a multi-modal data browser, query- and similarity-based retrieval functions, and a fast browsing function that skips redundant scenes. For sharing data with several users, this system can be used via networks from PCs or smartphones, This system uses a time-series active search, which has been successfully used for fast search of audio and video data, as its retrieval function algorithm. In a few seconds, this system can retrieve driving scenes that are similar to an input scene from 80,000 scenes. Retrieval performance was compared in various retrieval conditions by changing the codebook size of the vector quantization for the histogram features and a combination of driving signals. Experimental results showed that more than 97% retrieval performance was achieved for driving behaviors of left/right turns and curves using a combination of such complementary information as steering angles and lateral acceleration. We also compared the proposed method to a conventional image-based retrieval method using subjective similarity scores of driving scenes. Our proposed system retrieved similar scenes with about a 75% retrieval performance that was five points higher than a conventional image-based retrieval method. It is because image-based method is sensitive to changes of image in the area except in the region of interest for driving data retrieval. The fast browsing function also skipped scenes that could not be skipped by an image-based method. Masashi Naito, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Kazuya Takeda |
Intelligent Vehicles Symposium | 5 |
| 2010 | Estimation Method of User Satisfaction Using N-gram-based Dialog History Model for Spoken Dialog System
Sunao Hara, Norihide Kitaoka, Kazuya Takeda |
LREC | 3 |
| 2009 | Spoken dialog strategy based on understanding graph searchabstractWe regarded information retrieval as a graph search problem and proposed several novel dialog strategies that can recover from misrecognition through a spoken dialog that traverses the graph. To recover from misrecognition without seeking confirmation, our system kept multiple understanding hypotheses at each turn and searched for a globally optimal hypothesis in the graph whose nodes express understanding states across user utterances in a whole dialog. As for a dialog strategy, we introduced a new criterion based on efficiency in information retrieval and consistency with understanding hypotheses to select an appropriate system response. Using such criterion, the system removes the ambiguity so that users do not feel that a response that conflicts with the actual user intent is unnatural. We developed a spoken dialog system using these techniques and showed dialog examples in which misrecognition was naturally corrected. We also showed that our strategy was efficient in terms of the number of turns. Yuji Kinoshita, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 4 |
| 2009 | Stochastic modeling of vehicle trajectory during lane-changingabstractA signal processing approach for modeling vehicle trajectory during lane changing driving is discussed. Because individual driving habits are not a deterministic process, we developed a stochastic method. The proposed model consists of two parts: a dynamic system represented by a hidden Markov model and a cognitive distance space derived from the range distance distribution. The first part models the local dynamics of vehicular movements and generates a set of probable trajectories. The second part selects an optimal trajectory by stochastically evaluating the distances from surrounding vehicles. From experimental evaluation, we show that the model can predict the vehicle trajectory at given traffic conditions with 17.6 m prediction error for two different drivers. Yoshihiro Nishiwaki, Chiyomi Miyajima, Hidenori Kitaoka, Kazuya Takeda |
ICASSP | 4 |
| 2009 | Feature transformation based on discriminant analysis preserving local structure for speech recognitionabstractTo improve speech recognition performance, a feature transformation based on discriminant analysis has been widely used to reduce redundant dimensions of features. Linear discriminant analysis (LDA) and heteroscedastic discriminant analysis (HDA) are often used for this purpose, and a generalization method for LDA and HDA called power LDA (PLDA) has been proposed. However, these methods may result in unexpected dimensionality reduction for multimodal data. It is important to preserve the local structure of the data in reducing the dimensionality of multimodal data. In this paper we introduce two methods, locality preserving HDA and locality preserving PLDA. We also give an efficient calculation scheme to obtain an optimal projection. Makoto Sakai, Norihide Kitaoka, Kazuya Takeda |
ICASSP | 3 |
| 2009 | A multimedia corpus of driving behaviorsabstractIn this paper we present our multimedia corpus of real-world driving data (NUDrive), built with the primary objective of firming foundations for applying digital signal processing technologies in the vehicular environment. NUDrive is a content rich corpus composed of driving, speech, video, and physiological signals. So far, we have collected data from 250 drivers, who drove an instrumented vehicle under very similar conditions. In order to provide a more meaningful description of the situations drivers experience, a comprehensive data annotation protocol is proposed. We also briefly present a multimedia processing system, which uses information from various sources in NUDrive to implement a context-dependent estimation of a driver's spontaneous frustration. Results are encouraging and stress the relevance of content rich driving corpora to driver behavior modeling. Lucas Malta, Akira Ozaki, Chiyomi Miyajima, Norihide Kitaoka, Kazuya Takeda |
MMSP | 5 |
| 2009 | A Study of Driver Behavior Under Potential Threats in Vehicle TrafficabstractAlthough, in recent years, significant developments have been made in road safety, traffic statistics indicate that we still need significant improvements in the field. Since traffic accidents usually reflect human factors, in this paper, we focus on clarifying the understanding of driver behaviors under hazardous scenarios. Brake pedal signals or driver speech, or both, are utilized to detect incidents from a real-world driving database of 373 drivers. Results are then analyzed to address the individuality in driver behaviors, the multimodality of driver reactions, and the detection of potentially dangerous locations. All of the existing 25 potentially hazardous scenes in the database are hand labeled and categorized. Based on the joint histograms of behavioral signals and their time derivatives, a detection feature is proposed and satisfactorily applied to the indication of anomalies in driving behavior. Seventeen scenes, where a reaction utilizing the brake pedal was observed, are detected with a true positive (TP) rate of 100% and a false positive (FP) rate of 4.1%. We demonstrate the relevance of considering behavior individuality. During 11 scenes, the drivers verbally reacted. Scenes that included high-energy words are adequately detected by the speech-based method, which achieved a TP rate of 54% for an FP rate of 6.4%. The integration of different behavior modalities satisfactorily boosts the detection of the most subjectively hazardous situations, which suggests the importance of considering multimodal reactions. Finally, a strong relationship is presented between locations where potentially hazardous situations occurred and areas of frequent strong braking. Lucas Malta, Chiyomi Miyajima, Kazuya Takeda |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2009 | Driving Profile Modeling and Recognition Based on Soft Computing ApproachabstractAdvancements in biometrics-based authentication have led to its increasing prominence and are being incorporated into everyday tasks. Existing vehicle security systems rely only on alarms or smart card as forms of protection. A biometric driver recognition system utilizing driving behaviors is a highly novel and personalized approach and could be incorporated into existing vehicle security system to form a multimodal identification system and offer a greater degree of multilevel protection. In this paper, detailed studies have been conducted to model individual driving behavior in order to identify features that may be efficiently and effectively used to profile each driver. Feature extraction techniques based on Gaussian mixture models (GMMs) are proposed and implemented. Features extracted from the accelerator and brake pedal pressure were then used as inputs to a fuzzy neural network (FNN) system to ascertain the identity of the driver. Two fuzzy neural networks, namely, the evolving fuzzy neural network (EFuNN) and the adaptive network-based fuzzy inference system (ANFIS), are used to demonstrate the viability of the two proposed feature extraction techniques. The performances were compared against an artificial neural network (NN) implementation using the multilayer perceptron (MLP) network and a statistical method based on the GMM. Extensive testing was conducted and the results show great potential in the use of the FNN for real-time driver identification and verification. In addition, the profiling of driver behaviors has numerous other potential applications for use by law enforcement and companies dealing with buses and truck drivers. Abdul Wahab 0001, Hiok Chai Quek, Chin Keong Tan, Kazuya Takeda |
IEEE Trans. Neural Networks | 4 |
| 2008 | Encoding large array signals into a 3D sound field representation for selective listening point audio based on blind source separationabstractABSTRACT A sound field reproduction method which uses blind source separation and head-related transfer function is proposed. In the proposed system, multichannel acoustic signals captured at the distant microphones are encoded to a set of location/signal pairs of virtual sound sources based on frequency-domain ICA. After estimating the locations and the signals of the virtual sources, by convolving the controlled acoustic transfer functions with each signal, the spatial sound at the selected point is constructed. In the evaluation, the sound field made by 6 sound sources is captured using 48 distant microphones and is encoded into set of virtual sound sources. Subjective evaluation shows that there is no significant difference between natural and reconstructed sound when more than 6 virtual sources are used. Therefore the effectiveness of the encoding algorithm as well as the virtual source representation is confirmed. Kenta Niwa, Takanori Nishino, Kazuya Takeda |
ICASSP | 3 |
| 2008 | An integrative recognition method for speech and gesturesabstractWe propose an integrative recognition method of speech accompanied with gestures such as pointing. Simultaneously generated speech and pointing complementarily help the recognition of both, and thus the integration of these multiple modalities may improve recognition performance. As an example of such multimodal speech, we selected the explanation of a geometry problem. While the problem was being solved, speech and fingertip movements were recorded with a close-talking microphone and a 3D position sensor. To find the correspondence between utterance and gestures, we propose probability distribution of the time gap between the starting times of an utterance and gestures. We also propose an integrative recognition method using this distribution. We obtained approximately 3-point improvement for both speech and fingertip movement recognition performance with this method. Madoka Miki, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Kazuya Takeda |
ICMI | 5 |
| 2008 | CENSREC-4: development of evaluation framework for distant-talking speech recognition under reverberant environmentsabstractIn this paper, we newly introduce a collection of databases and evaluation tools called CENSREC-4, which is an evaluation framework for distant-talking speech under hands-free conditions. Distant-talking speech recognition is crucial for a handsfree speech interface. Therefore, we measured room impulse responses to investigate reverberant speech recognition in various environments. The data contained in CENSREC-4 are connected digit utterances, as in CENSREC-1. Two subsets are included in the data: basic data sets and extra data sets. The basic data sets are used for the evaluation environment for the room impulse response-convolved speech data. The extra data sets consist of simulated and recorded data. An evaluation framework is only provided for the basic data sets as evaluation tools. The results of evaluation experiments proved that CENSREC-4 is an effective database for evaluating the new dereverberation method because the traditional dereverberation process had difficulty sufficiently improving the recognition performance. Index Terms: Various environments, Impulse response, Convolution, Real recorded data, Evaluation framework Masato Nakayama, Takanobu Nishiura, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Tetsuji Ogawa, Shigeki Matsuda, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
INTERSPEECH | 15 |
| 2008 | Parameter estimation method of F0 control model for singing voicesabstractIn this paper, we propose a novel representation of F0 contours that provides a computationally efficient algorithm for automatically estimating the parameters of a F0 control model for singing voices. Although the best known F0 control model, based on a second-order system with a piece-wise constant function as its input, can generate F0 contours of natural singing voices, this model has no means of learning the model parameters from observed F0 contours automatically. Therefore, by modeling the piece-wise constant function by Hidden Markov Models (HMM) and approximating the second order differential equation by the difference equation, we estimate model parameters optimally based on iteration of Viterbi training and an LPC-like solver. Our representation is a generative model and can identify both the target musical note sequence and the dynamics of singing behaviors included in the F0 contours. Our experimental results show that the proposed method can separate the dynamics from the target musical note sequence and generate the F0 contours using estimated model parameters. Yasunori Ohishi, Hirokazu Kameoka, Kunio Kashino, Kazuya Takeda |
INTERSPEECH | 4 |
| 2008 | Building and combining document and music spaces for music query-by-webpage systemabstractBuilding and combining document and music spaces of songs are discussed for a new music recommendation applica-tion, which uses commonly read texts such as Web log as query input. The most important application of this flexible recom-mendation system is its music query-by-Webpage, from which a song that appropriately matches Webpage is automatically played. The key idea of the proposed system is to train a lin-ear transformation between document and music spaces so that query documents can be mapped onto a music space in which similarities based on acoustic characteristics is represented. The basic system has been trained using 2,650 pairs of song and review texts. Through experimental evaluations, we show the effectiveness of the system, which is three times better than the previous system. Web text as a training corpus and a bigram representation for the document vector are also investigated for the purpose of improving the system, and their effectiveness is also confirmed. Index Terms: Music, information retrieval, music similarity, latent semantic analysis, multimedia databases Ryoei Takahashi, Yasunori Ohishi, Norihide Kitaoka, Kazuya Takeda |
INTERSPEECH | 4 |
| 2008 | Evaluation Framework for Distant-talking Speech Recognition under Reverberant Environments: newest Part of the CENSREC Series -
Takanobu Nishiura, Masato Nakayama, Yuki Denda, Norihide Kitaoka, Kazumasa Yamamoto, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
LREC | 13 |
| 2008 | In-car Speech Data Collection along with Various Multimodal Signals
Akira Ozaki, Sunao Hara, Takashi Kusakawa, Chiyomi Miyajima, Takanori Nishino, Norihide Kitaoka, Katunobu Itou, Kazuya Takeda |
LREC | 8 |
| 2008 | 3DAV integrated system featuring arbitrary listening-point and viewpoint generationabstractIn this paper, we propose two novel methods for arbitrary listening-point generation for 3D audio-video (3DAV) integration in a large-scale multipoint cameras and microphones system with abilities to process, and display information of any recorded 3D scene in realtime. With this system, users are able to control their own viewpoint/listening-point position, freely. Arbitrary listening-point can be generated by either (i) ray-space representation of sound wave field (i.e. source sound independent) for multi frequency layers, or (ii) acoustic transfer function estimation (i.e. source sound dependent) and blind separation of sources of sounds. Arbitrary viewpoint generation is based on ray-space method, which is enhanced by using multipass dynamic programming for geometry compensation. Integration is done by either (i) ray-space representation of sound wave and image together, or (ii) integrating each camera video signal and acoustic transfer function of the same location as integrated 3DAV data. The prototype system of integrated audio-visual viewer achieves both good image and sound qualities with 15 frames/second. Mehrdad Panahpour Tehrani, Kenta Niwa, Norishige Fukushima, Yasushi Hirano, Toshiaki Fujii, Masayuki Tanimoto, Kazuya Takeda, Kenji Mase, Akio Ishikawa, Shigeyuki Sakazawa, Atsushi Koike |
MMSP | 7 |
| 2007 | Development of VAD evaluation framework CENSREC-1-C and investigation of relationship between VAD and speech recognition performanceabstractVoice activity detection (VAD) plays an important role in speech processing including speech recognition, speech enhancement, and speech coding in noisy environments. We developed an evaluation framework for VAD in such environments, called corpus and environment for noisy speech recognition 1 concatenated (CENSREC-1-C). This framework consists of noisy continuous digit utterances and evaluation tools for VAD results. By adoptiong two evaluation measures, one for frame-level detection performance and the other for utterance-level detection performance, we provide the evaluation results of a power-based VAD method as a baseline. When using VAD in speech recognizer, the detected speech segments are extended to avoid the loss of speech frames and the pause segments are then absorbed by a pause model. We investigate the balance of an explicit segmentation by VAD and an implicit segmentation by a pause model using an experimental simulation of segment extension and show that a small extension improves speech recognition. Norihide Kitaoka, Kazumasa Yamamoto, Tomohiro Kusamizu, Seiichi Nakagawa, Takeshi Yamada, Satoru Tsuge, Chiyomi Miyajima, Takanobu Nishiura, Masato Nakayama, Yuki Denda, Masakiyo Fujimoto, Tetsuya Takiguchi, Satoshi Tamura, Shingo Kuroiwa, Kazuya Takeda, Satoshi Nakamura 0001 |
ASRU | 15 |
| 2007 | Statistical segmentation and recognition of fingertip trajectories for a gesture interfaceabstractThis paper presents a virtual push button interface created by drawing a shape or line in the air with a fingertip. As an example of such a gesture-based interface, we developed a four-button interface for entering multi-digit numbers by pushing gestures within an invisible 2x2 button matrix inside a square drawn by the user. Trajectories of fingertip movements entering randomly chosen multi-digit numbers are captured with a 3D position sensor mounted on the the forefinger's tip. We propose a statistical segmentation method for the trajectory of movements and a normalization method that is associated with the direction and size of gestures. The performance of the proposed method is evaluated in HMM-based gesture recognition. The recognition rate of 60.0% was improved to 91.3% after applying the normalization method. Kazuhiro Morimoto, Chiyomi Miyajima, Norihide Kitaoka, Katunobu Itou, Kazuya Takeda |
ICMI | 5 |
| 2007 | Driver Modeling Based on Driving Behavior and Its Evaluation in Driver IdentificationabstractAll drivers have habits behind the wheel. Different drivers vary in how they hit the gas and brake pedals, how they turn the steering wheel, and how much following distance they keep to follow a vehicle safely and comfortably. In this paper, we model such driving behaviors as car-following and pedal operation patterns. The relationship between following distance and velocity mapped into a two-dimensional space is modeled for each driver with an optimal velocity model approximated by a nonlinear function or with a statistical method of a Gaussian mixture model (GMM). Pedal operation patterns are also modeled with GMMs that represent the distributions of raw pedal operation signals or spectral features extracted through spectral analysis of the raw pedal operation signals. The driver models are evaluated in driver identification experiments using driving signals collected in a driving simulator and in a real vehicle. Experimental results show that the driver model based on the spectral features of pedal operation signals efficiently models driver individual differences and achieves an identification rate of 76.8% for a field test with 276 drivers, resulting in a relative error reduction of 55% over driver models that use raw pedal operation signals without spectral analysis Chiyomi Miyajima, Yoshihiro Nishiwaki, Koji Ozawa, Toshihiro Wakita, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
Proc. IEEE | 6 |
| 2006 | Multichannel Speech Enhancement Based on Speech Spectral Magnitude Estimation Using Generalized Gamma Prior DistributionabstractWe present multichannel speech enhancement method based on MAP speech spectral magnitude estimation using a generalized gamma model of speech prior distribution, where the model parameters are adapted from actual noisy speech in a frame-by-frame manner. The utilization of a more general prior distribution with its online estimation is shown to be effective for speech spectral estimation. We tested the proposed algorithm in an in-car speech database and obtained significant improvements on the speech recognition performance, particularly under nonstationary noise conditions such as music, air-conditioner and open window. Tran Huy Dat, Kazuya Takeda, Fumitada Itakura |
ICASSP (4) | 2 |
| 2006 | Development of Micro-Dodecahedral Loudspeaker for Measuring Head-Related Transfer Functions in The Proximal regionabstractThis paper describes new equipment for measuring head-related transfer functions (HRTFs) near a listener's head. 3D sounds in headphones are generated by the convolution of sound signals and an HRTF, which is defined as the acoustical transfer function between a point sound source and the entrance to the ear canal. A loudspeaker is usually used for HRTF measurements, and a distance of more than 1 m separates the loudspeaker and the subject. The region within 1 m of the head is called the 'proximal region,' where a small loudspeaker is needed for accurately measuring HRTF, that is, a conventional loudspeaker cannot be used. In our study, a micro-dodecahedral loudspeaker with twelve piezoelectric ceramic devices is used for HRTF measurements at a diameter is 38 mm. Our experiments examined the characteristics of this loudspeaker. From the results, our developed loudspeaker provides similar performance to a point source, and it is very effective for measuring the HRTFs in the proximal region. Seiichiro Hosoe, Takanori Nishino, Katunobu Itou, Kazuya Takeda |
ICASSP (5) | 4 |
| 2006 | Adaptive Regression Based Framework for In-Car Speech RecognitionabstractWe address issues for improving hands-free speech recognition performance in different car environments using a single distant microphone. In our previous work, we proposed a regression based enhancement method for in-car speech recognition. In this paper, we describe recent improvements and propose a data-driven adaptive regression based speech recognition system, in which both feature enhancement and model compensation are performed. Based on isolated word recognition experiments conducted in 15 real car environments, the proposed adaptive regression approach shows an advantage in average relative word error rate (WER) reductions of 52.5% and 14.8%, compared to original noisy speech and ETSI advanced front-end, respectively. Weifeng Li 0001, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 3 |
| 2006 | Cepstral Analysis of Driving Behavioral Signals for Driver IdentificationabstractSpectral analysis is applied to such driving behavioral signals as gas and brake pedal operation signals for extracting drivers' characteristics while accelerating or decelerating. Cepstral features of each driver obtained through spectral analysis of driving signals are modeled with a Gaussian mixture model (GMM). A GMM driver model based on cepstral features is evaluated in driver identification experiments using driving signals collected in a driving simulator and in a real vehicle on a city road. Experimental results show that the driver model based on cepstral features achieves a driver identification rate of 89.6% for driving simulator and 76.8% for real vehicle, resulting in 61 % and 55 % error reduction, respectively, over a conventional driver model that uses raw driving signals without spectral analysis Chiyomi Miyajima, Yoshihiro Nishiwaki, Koji Ozawa, Toshihiro Wakita, Katunobu Itou, Kazuya Takeda |
ICASSP (5) | 6 |
| 2006 | Arbitrary Listening-Point Generation Using Sub-Band Representation of Sound Wave Ray-SpaceabstractThis paper proposes arbitrary listening-point generation of sound without source localization, and a theory based on the ray-space representation of light rays using sub-band signal processing. An array of beam-formed dynamic microphone-arrays (MAs), are set and each MA generates a multi frequency layer sound-image (SImage) by scanning the viewing range of a camera. Each layer is captured by active microphones in MA for a given frequency range which makes the correct interval to capture that frequency layer, and a sound wave ray-space. Arbitrary listening-point generation of each layer is done by geometry compensation of corresponding images in the location of each MA or SImages or their combination. Each layer sound of an SIamge is generated by averaging the sound wave in each pixel or group of pixel. The listening-point is generated after combining each layer sound in frequency domain and inverse transformation from frequency to time domain Mehrdad Panahpour Tehrani, Yasushi Hirano, Toshiaki Fujii, Shoji Kajita, Kazuya Takeda, Kenji Mase |
ICASSP (5) | 5 |
| 2006 | Multipoint Measuring System for Video and Sound - 100-camera and microphone systemabstractWe developed a novel multipoint measurement system capable of acquiring video and sound at more than 100 points in a "synchronized" manner. In this paper, we first describe the specification of the system and how the system works in detail. Then we report some experimental results that confirm the performance of the system. We also describe test data set we provided for MPEG (moving picture experts group) multi-viewpoint video coding activities. Using this system, we are planning to conduct projects to measure humans and their activities, collect a large volume of real-world data of video and sound, and release them to the public Toshiaki Fujii, Kensaku Mori, Kazuya Takeda, Kenji Mase, Masayuki Tanimoto, Yasuhito Suenaga |
ICME | 3 |
| 2006 | CENSREC2: corpus and evaluation environments for in car continuous digit speech recognition
Satoshi Nakamura 0001, Masakiyo Fujimoto, Kazuya Takeda |
INTERSPEECH | 3 |
| 2006 | Statistical Analysis for Thesaurus Construction using an Encyclopedic Corpus
Yasunori Ohishi, Katunobu Itou, Kazuya Takeda, Atsushi Fujii |
LREC | 3 |
| 2006 | On-line Gaussian mixture modeling in the log-power domain for signal-to-noise ratio estimation and speech enhancement
Tran Huy Dat, Kazuya Takeda, Fumitada Itakura |
Speech Commun. | 2 |
| 2005 | Generalized gamma modeling of speech and its online estimation for speech enhancementabstractGeneralized gamma modeling and its online method of parameter estimation of speech spectral magnitude are proposed for MAP based speech enhancement systems. Generalized gamma modeling is shown to be a natural extension of the Gaussian modeling of speech spectral component distribution, and is therefore, able to fit the prior distribution better than the conventional method. An online parameter estimation method for the gamma distribution, based on a moment matching method, is then proposed. The effectiveness of the proposed methods are confirmed by improvement in both SNR and ASR using the AURORA2 standard database, where about 4 dB improvement in SNR and 20% improvement in relative ASR performance are obtained. Tran Huy Dat, Kazuya Takeda, Fumitada Itakura |
ICASSP (4) | 2 |
| 2005 | Analysis of a large in-car speech corpus and its application to the multimodel ASRabstractIn-car ASR performance improvement, utilizing a large in-car speech corpus, consisting of the utterances of more than five hundred drivers under real driving conditions, is discussed. A subset design method for efficient cross validations in large-scale speech recognition experiments is proposed. The factor analysis of the results of the recognition experiments show the relationship between word accuracy and utterance characteristics, i.e., SNR, entropy and speaking rates. Based on the factor analysis results, a multimodel approach which uses the utterance duration and subband SNRs as the model selection measures for acoustic and language models, respectively, is proposed. By the proposed multimodel approach, a relative error reduction of 16% is obtained. Hiroshi Fujimura, Chiyomi Miyajima, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 4 |
| 2005 | Spatial coding based on the extraction of moving sound sources in wavefield synthesisabstractSince a sound field reproduction system based on wavefield synthesis usually needs a great number of channel signals, the amount of data transmitted should be reduced. This paper therefore proposes a spatial coding method that is based on the extraction of moving sound sources and reduces the amount of data transmitted from an amount proportional to the number of channels to an amount proportional to the number of sound sources. A coding experiment was performed for a reverberant sound field which was simulated with an image method. The effect of the proposed method on the perceptual quality was evaluated by subjective assessment. Toshiyuki Kimura, Kazuya Takeda, Fumitada Itakura |
ICASSP (3) | 3 |
| 2005 | Two-stage Noise Spectra Estimation and Regression based In-car Speech Recognition using Single Distant MicrophoneabstractWe present a two-stage noise spectra estimation approach. After the first-stage noise estimation using the improved minima controlled recursive averaging (IMCRA) method, the second-stage noise estimation is performed by employing a maximum a posteriori (MAP) noise amplitude estimator. We also develop a regression-based speech enhancement system by approximating the clean speech with the estimated noise and the original noisy speech. Evaluation experiments show that the proposed two-stage noise estimation method results in lower estimation error for all test noise types. Compared to the original noisy speech, the proposed regression-based approach obtains an average relative word error rate (WER) reduction of 65% in our isolated word recognition experiments conducted in 12 real car environments. Weifeng Li 0001, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 3 |
| 2005 | SNR and Local Noise Power Estimations Based on Gaussian Mixture Modeling on the Log-Power DomainabstractWe propose a flexible and robust SNR estimation method for the real conditions, when neither clean reference signal nor speech activity is available. This method is based on Gaussian mixture modeling on the log-power domain of the noisy speech and use the estimated subspace distribution parameters to derive the SNR measures. The experimental results show better performance in estimating both the segmental and global SNR compared to the conventional method based on voice activity detection (VAD). The second application presented in this work is local noise power estimation, where the same model is applied to each frequency bin. Furthermore, an empirical MAP solution using second order statistics is applied to estimate the local noise powers in order to implement a Wiener filtering system. The evaluation experiments show the improvements of the proposed speech enhancement method in both segmental SNR and automatic speech recognition (ASR) performance. Kazuya Takeda, Tran Huy Dat, Hiroshi Fujimura, Fumitada Itakura |
ICASSP (1) | 1 |
| 2005 | The sound wave ray-spaceabstractThis paper addresses the problem of 3D sound representation without sound source localization and proposes a theory based on the ray-space representation of light rays, which is independent of object's specifications. An array of beam-formed microphone-arrays (MAs), are set and each MA generates a sound-image (SImage) by scanning the viewing range of a camera in the same location. SImage has the same size of an image and contains of blocks of sound wave with duration of one image-frame. Captured SImages with the array of MAs generate the sound wave ray-space. To make a dense SImage ray-space, we propose to use the geometry compensation of corresponding images in the location of each MA. By a dense sound ray-space, any virtual SImage, which corresponds to an arbitrary listening-point, can be generated. The listening-point sound is generated by averaging the sound wave in each pixel or group of pixel of the virtual SImage. Mehrdad Panahpour Tehrani, Yasushi Hirano, Toshiaki Fujii, Shoji Kajita, Kazuya Takeda, Masayuki Tanimoto, Kenji Mase |
ICME | 5 |
| 2005 | Subjective and objective quality assessment of regression-enhanced speech in real car environments
Weifeng Li 0001, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 3 |
| 2005 | Discrimination between singing and speaking voicesabstractDiscriminating between singing and speaking voices by using the local and global characteristics of voice signals is discussed. From the results of subjective experiments, we show that human beings can discriminate singing and speaking voices with more than 70 % and 95 % accuracy from 300 ms and one second long signals, respectively. From the subjective experiment results, assuming that different features are effective for shortterm and long-term signals, we designed two measures using a spectral envelope (MFCC) and the fundamental frequency (F0, perceived as pitch) contour. Experimental results show that the F0 measure performs better than the spectral envelope measure when the input voice signals are longer than one second. Particularly, it can discriminate singing and speaking voices with more than 80 % accuracy with two-second signals. On the other hand, when the input signals are shorter than one second, the spectral envelope measure performs better than the F0 measure. Finally, by simply combining the two measures, more than 90% accuracy is obtained for two-second signals. 1. Yasunori Ohishi, Masataka Goto, Katunobu Itou, Kazuya Takeda |
INTERSPEECH | 4 |
| 2005 | Data collection and evaluation of speech recognition for motorbike ridersabstractAbstract Speech recognition should be as an eyes-free and hands-free interface. To realise this technology, we need to clar-ify acoustics in a helmet and determine how much high-level riding noise affects captured speech data. This pa-per describes the acoustics in a helmet and transfer func-tions of the microphone position. We constructed a datacollection system and collected the speech data of motor-bike riders on city roads and express highways. Speechrecognition experiments were conducted and we obtaineda recognition rate high of 83.1%. 1. Introduction Riding a motorbike requires more care than driving a car.Even when a rider idles his/her motorbike, such as whenhe/she waits at a red light, button operations are incon-venient because the rider needs to remove his/her glovesto push the buttons. Therefore, for motorbike riders, aneyes-free and hands-free interface is required for operat-ing information appliances, such as a cellular phone anda route navigation system. Thus, speech recognition is avery important technology.This study investigated the feasibility of speechrecognition for motorbike riders. On a motorbike, rid-ers are exposed directly to high-level noises such as windnoise, engine noise, and road noise. It is known thatexposed noise level is varied by various factors such asspeed, riding position, and helmets[1, 2]. In order to re-alize speech recognition on a motorbike, we first need toinvestigate how much such factors degrade conventionalspeech recognition performance.Riders must put on a helmet when they ride motor-bikes. Helmets are designed to reduce noise level, how-ever, we need to clarify how much this reduction con-tributes to speech recognition using microphones insidethe side of helmets. Moreover, we need to investigateacoustics in a helmet, because a helmet has a very smallcavity.In this paper, we measured acoustics in a helmetto determine microphone positions for collecting riders’speech corpus. We then collected the speech data utteredby motorbike riders riding on a highway. We also providean analysis of the corpus and the results of our speechrecognition evaluation. Hiroshi Fujimura, Chiyomi Miyajima, Takanori Nishino, Katunobu Itou, Kazuya Takeda |
INTERSPEECH | 6 |
| 2005 | Speaker verification using Gaussian mixture models within changing real car environments
Xianxian Zhang, John H. L. Hansen, Pongtep Angkititrakul, Kazuya Takeda |
INTERSPEECH | 4 |
| 2005 | Performance Evaluation of H.264 Video Streaming over Inter-Vehicular 802.11 Ad Hoc NetworksabstractThis paper evaluates the performance of video streaming in inter-vehicular environments using the 802.11 ad hoc network protocol. We performed transmission experiments while driving two cars equipped with 802.11b standard devices in urban and highway scenarios. Different sequences, bitrates and packctization policies have been tested. The experiments show that each scenario presents peculiar characteristics in terms of average link availability and SNR, which can be exploited to develop more efficient applications. In this paper we also determine the best packetization policies for the two scenarios, showing that large packets lead to better performance in the highway scenario and vice versa. Perceptual quality results indicate that the best packetization policy achieves consistent gains in terms of PSNR values (up to 5 dB), and reduced quality variations, with respect to a fixed-policy transmission technique Paolo Bucciol, Enrico Masala, Nobuo Kawaguchi, Kazuya Takeda, Juan Carlos De Martin |
PIMRC | 4 |
| 2005 | Analysis and recognition of whispered speech
Taisuke Ito, Kazuya Takeda, Fumitada Itakura |
Speech Commun. | 2 |
| 2005 | Adaptive log-spectral regression for in-car speech recognition using multiple distributed microphonesabstractThis letter addresses issues in improving hands-free speech recognition performance in different car environments. We propose a new speech-enhancement approach based on optimizing regression of the log-spectra, which is used to estimate the log-spectra of speech at a close-talking microphone by using multiple spatially distributed microphones. The regression weights can be adapted automatically for different noise environments. Compared to the nearest distant microphone and adaptive beamformer generalized sidelobe canceller (GSC), the proposed approach shows an advantage in the average relative word error rate (WER) reduction of 58.5 and 10.3%, respectively, for isolated word recognition under 15 real-car environments. Weifeng Li 0001, Kazuya Takeda, Fumitada Itakura |
IEEE Signal Process. Lett. | 2 |
| 2004 | Biometric identification using driving behavioral signalsabstractWe investigate the uniqueness of driver behavior in vehicles and the possibility of using it for personal identification with the objectives of achieving safer driving, of assisting the driver in case of emergencies, and of being a part of a multi-mode biometric signature for driver identification. We use Gaussian mixture models (GMM) for modeling the individualities of the accelerator and brake pedal pressures, and focus on not only the static features, but also the dynamics of the pedal pressures. Experimental results show that the dynamic features significantly improve the performance of driver identification. Kei Igarashi, Chiyomi Miyajima, Katunobu Itou, Kazuya Takeda, Fumitada Itakura, Hüseyin Abut |
ICME | 4 |
| 2004 | Speech recognition using synchronization between speech and finger tapping
Hiromitsu Ban, Chiyomi Miyajima, Katunobu Itou, Fumitada Itakura, Kazuya Takeda |
INTERSPEECH | 5 |
| 2004 | Analysis of in-car speech recognition experiments using a large-scale multi-mode dialogue corpusabstractThe dependency of conversational utterances on themode of dialogue is analyzed. A speech corpus of 800 speak-ers collected under three different modes, i.e., talking to a human operator, an WOZ system and an ASR system, is used for analysis. Some characteristics such as sen-tence complexity loudness of the voice and speaking-rate are found to be significantly different among the dialogue modes. Linear regression analysis results also clarify the relative importance of those characteristics on speech recognition accuracy. 1. Hiroshi Fujimura, Katunobu Itou, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 3 |
| 2004 | CIAIR in-car speech database
Nobuo Kawaguchi, Shigeki Matsubara, Yukiko Yamaguchi, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 4 |
| 2004 | Recent progress of open-source LVCSR engine julius and Japanese model repositoryabstractContinuous Speech Recognition Consortium (CSRC) was founded for further enhancement of Japanese Dictation Toolkit that had been developed by the support of a Japanese agency. Overview of its product software is reported in this paper. The open-source LVCSR (large vocabulary continuous speech recognition) engine Julius has been improved both in performance and functionality, and it is also ported to Microsoft Windows in compliance with SAPI (Speech API). The software is now used for not a few languages and plenty of applications. For plug-and-play speech recognition in various applications, we have also compiled a repository of acoustic and language models for Japanese. Especially, the acoustic model set realizes wider coverage of user generations and speech-input environments. Tatsuya Kawahara, Akinobu Lee, Kazuya Takeda, Katunobu Itou, Kiyohiro Shikano |
INTERSPEECH | 3 |
| 2004 | Optimizing regression for in-car speech recognition using multiple distributed microphonesabstractIn this paper, we address issues in improving handsfree speech recognition performance in different car environments using multiple spatially distributed microphones. In previous work, we proposed multiple regression of the log-spectra (MRLS) for estimating the logspectra of speech at a close-talking microphone. In this paper, the idea is extended to nonlinear regressions. Isolated word recognition experiments under real car environments show that, compared to the nearest distant microphone, recognition accuracies could be improved by about 40% for very noisy driving conditions by using the optimizing regression method, The proposed approach outperforms linear regression methods and also outperforms adaptive beamformer by 8% and 3% respectively in terms of averaged recognition accuracies. Weifeng Li 0001, Fumitada Itakura, Kazuya Takeda |
INTERSPEECH | 3 |
| 2004 | Speech enhancement based on magnitude estimation using the gamma priorabstractIn this paper, we propose a speech enhancement method based on spectral magnitude estimation. We modify the noise estimation from the minimum statistics method and combine with a maximum a posterior (MAP) decomposition, using the Rice-conditional probability and a non-Gaussian statistic model of the speech. We derive two versions of magnitude decomposition and magnitude-phase decomposition and compare to spectral subtraction and other MAP methods based on the Gaussian statistic (MMSE, LSA). The experiments show the advantage of the proposed method in the improvement of both SNR (up to 12 dB) and recognition accuracy rate (up to 21 % to base line). Weifeng Li 0001, Kazuya Takeda, Fumitada Itakura, Tran Huy Dat |
INTERSPEECH | 2 |
| 2004 | Example-based spoken dialogue system with online example augmentationabstractIn this paper, we propose a new method to expand an examplebased spoken dialogue system to handle context dependent utterances. The dialogue system refers to the dialogue examples to find an example that is suitable to promote dialogue. Here, the dialogue contexts are expressed in the form of dialogue slots. By constructing dialogue examples with the text of utterances and the dialogue slots, the system handle context dependent dialogue. And we also propose a new framework of spoken dialogue, named “GROW architecture” that consists of the dialogue system and a Wizard-of-OZ (WOZ) system. By using the WOZ system to add dialogue examples via network, it becomes efficient to augment dialogue examples. Hiroya Murao, Nobuo Kawaguchi, Shigeki Matsubara, Yukiko Yamaguchi, Kazuya Takeda, Yasuyoshi Inagaki |
INTERSPEECH | 5 |
| 2004 | Audio-visual SPeaker localization for car navigation systemsabstractHuman-computer interaction for in-vehicle information and navigation systems is a challenging problem because of the diverse and changing acoustic environments. It is proposed that the integration of video and audio information can significantly improve dialog system performance, since the visual modality is not impacted by acoustic noise. In this paper, we propose a robust audio-visual integration system for source tracking and speech enhancement for an in-vehicle speech dialog system. The proposed system integrates both audio and visual information to locate the desired speaker source. Using real data collected in car environments, the proposed system can improve speech accuracy by up to 40.75% compared with audio data alone. Xianxian Zhang, Kazuya Takeda, John H. L. Hansen, Toshiki Maeno |
INTERSPEECH | 2 |
| 2003 | In-car speech recognition using distributed microphones-adapting to automatically detected driving conditionsabstractIn this paper, we describe a multichannel method of noisy speech recognition that can adapt to various in-car noise situations during driving. The method allows us to estimate the log spectrum of speech at a close-talking microphone based on the multiple regression of the log spectra (MRLS) of noisy signals captured by multiple distributed microphones. Through clustering of the spatial noise distributions under various driving conditions, the regression weights for MRLS are effectively adapted to the driving conditions. The experimental evaluation shows an average error rate reduction of 43 % in isolated word recognition under 15 different driving conditions. Hideki Banno, Tetsuya Shinde, Kazuya Takeda, Fumitada Itakura |
ICASSP (1) | 3 |
| 2003 | In-car speech recognition using distributed microphones: adapting to automatically detected driving conditionsabstractIn this paper, we describe a multichannel method of noisy speech recognition that can adapt to various in-car noise situations during driving. The method allows as to estimate the log spectrum of speech at a close-talking microphone based on the multiple regression of the log spectra (MRLS) of noisy signals captured by multiple distributed microphones. Through clustering of the spatial noise distributions under various driving conditions, the regression weights for MRLS are effectively adapted to the driving conditions. The experimental evaluation shows average error rate reductions of 43% in isolated word recognition under 15 different driving conditions. Hideki Banno, Tetsuya Shinde, Kazuya Takeda, Fumitada Itakura |
ICME | 3 |
| 2003 | A study on domain recognition of spoken dialogue systems
Toshihiro Isobe, Shoji Hayakawa, Hiroya Murao, Tatsuji Mizutani, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 5 |
| 2003 | Integration of noise reduction algorithms for Aurora2 taskabstractTo achieve high recognition performance for a wide variety of noise and for a wide range of signal-to-noise ratios, this paper presents the integration of four noise reduction algorithms: spectral subtraction with smoothing of time direction, temporal domain SVD-based speech enhancement, GMM-based speech estimation and KLT-based comb-filtering. Recognition results on the Aurora2 task show that the effectiveness of these algorithms and their combinations strongly depends on noise conditions, and excessive noise reduction tends to degrade recognition performance in multicondition training. Takeshi Yamada, Jiro Okada, Kazuya Takeda, Norihide Kitaoka, Masakiyo Fujimoto, Shingo Kuroiwa, Kazumasa Yamamoto, Takanobu Nishiura, Mitsunori Mizumachi, Satoshi Nakamura 0001 |
INTERSPEECH | 3 |
| 2002 | Synthesis of car noise based on a composition of engine noise and friction noiseabstractThis paper describes a method for generation of car noise based on the engine noise and the “friction noise.” The engine noise is modeled by composition of a stationary background noise that depends on the size of the engine and nonstationary noise that depends rotational speed of the engine. The friction noise is modeled as a white noise with ranging power. Based on these models, methods for synthesis of these components are developed. Subjective assessment of the car noise synthesis method shows that it is fairly similar to the actual noise. Yoshihide Ban, Hideki Banno, Kazuya Takeda, Fumitada Itakura |
ICASSP | 3 |
| 2002 | Acoustic analysis and recognition of whispered speechabstractIn this paper, the acoustic properties and recognition of whispered speech is discussed, A whispered speech database that consists of whispered speech, nonnal speech and their corresponding facial video images of more than 6,000 sentences from 100 speakers was prepared. The comparison between whispered and nonnal utterances show that 1) the cepstrum distance between them is 4 dB for voiced and 2 dB for unvoiced phonemes, respectively, 2) the spectral tilt of the whispered speech is less sloped than the nonnal speech and 3) the frequency of the lower formants (below 1.5 kHz) is higher than that of the nonnal speech, Acoustic models (HMM) trained by the whispered speech database attain an accuracy of 68% in word recognition experiments. This accuracy can be improved to 78% when MLLR adaptation is applied, while the nonnal speech HMM adapted with the whispered speech attain only 62 % word accuracy. Taisuke Ito, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 2002 | Spatial compression of multi-channel audio signals using inverse filtersabstractA large number of transmission audio channels are necessary for reproduction of a sound field based on Huygens principle. A method is proposed for spatial compression of the multi-channel audio signals. The signals are compressed by convolving them with the inverse filter of the room impulse response to reduce the number of transmission channels to the number of source signals. Then the source signals are transmitted to reproduce the sound field by convolving them with the impulse response. The compression method is evaluated using the signal-to-noise ratio and a subjective assessment. Experimental results show that SNR between a source signal and an extracted signal is more than 40dB and that there is no significant difference of the directional perception due to compression. Toshiyuki Kimura, Kazuya Takeda, Fumitada Itakura |
ICASSP | 3 |
| 2002 | Recognition of continuous speech segments of monophone units using support vector machines
Weifeng Lee, Chellu Chandra Sekhar, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 3 |
| 2002 | Multiple regression of log-spectra for in-car speech recognition
Tetsuya Shinde, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 2 |
| 2002 | Experiments on recognition of lavalier microphone speech and whispered speech in real world environments
Kiyoshi Tatara, Taisuke Ito, Parham Zolfaghari, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 4 |
| 2002 | Multi-Dimensional Data Acquisition for Integrated Acoustic Information Research
Nobuo Kawaguchi, Shigeki Matsubara, Kazuya Takeda, Fumitada Itakura |
LREC | 3 |
| 2002 | The Present Status of Speech Database in Japan: Development, Management, and Application to Speech Research
Hisao Kuwabara, Shuichi Itahashi, Mikio Yamamoto, Toshiyuki Takezawa, Satoshi Nakamura 0001, Kazuya Takeda |
LREC | 6 |
| 2002 | Continuous Speech Recognition Consortium an Open Repository for CSR Tools and Models
Akinobu Lee, Tatsuya Kawahara, Kazuya Takeda, Masato Mimura, Atsushi Yamada, Akinori Ito, Katunobu Itou, Kiyohiro Shikano |
LREC | 3 |
| 2001 | Recognition of consonant-vowel utterances using Support Vector Machines
Chellu Chandra Sekhar, Kazuya Takeda, Fumitada Itakura |
ESANN | 2 |
| 2001 | Close-Class-Set Discrimination Method for Recognition of Stop_Consonant-Vowel Utterances Using Support Vector Machines
Chellu Chandra Sekhar, Kazuya Takeda, Fumitada Itakura |
ICANN | 2 |
| 2001 | A study on perceptual distance measure for phase spectrum of stimuliabstractThis paper describes a perceptual distance measure for phase spectrum based on results from a subjective experiment using stimuli. The stimuli have flat amplitude spectrum and, in a particular frequency band, have a certain group delay value. The experiment was performed using stimuli with different group delay peak values where the group delay center frequencies are fixed, and their associated group delay bandwidths are also fixed. It was found that when the peak values of stimuli are between -1 ms and 2 ms, they are perceived to be zero phase regardless of their center frequencies and bandwidths. Moreover, when the peak values are less than -8 ms or more than 10 ms and the bandwidths are less than 1 ERB, each of the stimuli are perceived to be similar. Based on these perceptual similarity results, we introduce an ellipsoidal function to estimate the similarity scores with a simple equation. It is found that the estimated similarity scores well approximate the subjective similarity scores. Hideki Banno, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 2001 | Direction of arrival estimation based on nonlinear microphone arrayabstractThis paper describes a new method for estimating the direction of arrival (DOA) using a nonlinear microphone array based on complementary beamforming. Complementary beamforming is based on two types of beamformers designed to obtain complementary directivity patterns with each other. In this system, since the resultant directivity pattern is proportional to the product of these directivity patterns, the proposed method can be used to estimate DOAs even when the number of sound sources is equal to or exceeds that of microphones. First, DOA-estimation experiments are performed using actual devices in real acoustic environments. The results clarify that DOA estimation for two sound sources can be accomplished by the proposed method with only two microphones. Also, by comparing the resolutions of DOA estimation by the proposed method and by the conventional minimum variance method, we can show that the performance of the proposed method is superior to that of the conventional method. Hidekazu Kamiyanagida, Hiroshi Saruwatari, Kazuya Takeda, Fumitada Itakura |
ICASSP | 3 |
| 2001 | Blind source separation combining frequency-domain ICA and beamformingabstractWe describe a new method of blind source separation (BSS) on a microphone array combining subband independent component analysis (ICA) and beamforming. The proposed array system consists of the following three sections: (1) subband-ICA-based BSS section with direction-of-arrival (DOA) estimation; (2) null beamforming section based on the estimated DOA information; and (3) integration of (1) and (2) based on the algorithm diversity. Using this technique, we can resolve the low-convergence problem through optimization in ICA. The results of the signal separation experiments reveal that a noise reduction rate (NRR) of about 18 dB is obtained under the nonreverberant condition, and NRR of 8 dB and 6 dB are obtained in the case that the reverberation times are 150 msec and 300 msec. These performances are superior to those of both simple ICA-based BSS and simple beamforming method. Hiroshi Saruwatari, Satoshi Kurita, Kazuya Takeda |
ICASSP | 3 |
| 2001 | Continuous speech recognition without end-point detectionabstractA continuous speech recognition method that does not need explicit speech end-point detection is proposed. A one-pass decoding algorithm is modified to decode input speech of infinite length so that, with appropriate nonspeech models for silence and ambient noises, continuous speech recognition can be executed without explicit endpoint detection. The basic algorithm: 1) decodes a processing block of predetermined length, 2) traces back and finds the boundaries of the processing blocks where the word history in the preceding processing block is merged into one, and 3) restarts decoding from the boundary frame with the merged word history. The effectiveness of the method is verified by the two dictating experiments. With 100 consecutive sentences of utterances from a newspaper, the degradation of the recognition accuracy due to the modification of the decoder is about 5% compared with the results when the correct end-point is given. With a 30 minutes dialogue in a moving car, 75% correct and 69% accuracy score is obtained. Osamu Segawa, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 2001 | Multimedia data collection of in-car speech communicationabstractThis paper reports the details of the collection of the multimedia data such as audio, video and auxiliary information of the vehicle during a spoken dialogue in a moving car. The system specially built in a Data CollectionVehicle (DCV) supports synchronous recording of multi-channel audio data from 16 microphones, 3-channel video data and the vehicle related data. Multimedia data has been collected for three sessions of spoken dialogue in about a 60-minute drive by each of 200 subjects. Data has been collected for two dialogue modes:(1) prompted dialogue between the driver and an accompanying operator and (2) natural dialogue between the driver and a telephone operator for information access over a cellular phone while driving a car. The corpus can be used for analysis of multimedia data in a moving car environment and also for modeling spoken dialogue in scenarios such as information access while driving a car. Nobuo Kawaguchi, Shigeki Matsubara, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 3 |
| 2001 | Robust speech recognition based on selective use of missing frequency band HMMs
Takayoshi Kawamura, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 2 |
| 2000 | Evaluation of blind signal separation method using directivity pattern under reverberant conditionsabstractThis paper describes a new blind signal separation method using the directivity patterns of a microphone array. In this method, to deal with the arriving lags among each microphone, the inverses of the mixing matrices are calculated in the frequency domain so that the separated signals are mutually independent. Since the calculations are carried out in each frequency independently, the following problems arise: (1) permutation of each sound source, (2) arbitrariness of each source gain. In this paper, we propose a new solution that directivity patterns are explicitly used to estimate each sound source direction. As the results of signal separation experiments, it is shown that the proposed method improves the SNR of degraded speech by about 16 dB under non-reverberant condition. Also, the proposed method improves the SNR by 8.7 dB when the reverberation time is 184 ms, and by 5.1 dB when the reverberation time is 322 ms. Satoshi Kurita, Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 4 |
| 2000 | A new phonetic tied-mixture model for efficient decodingabstractA phonetic tied-mixture (PTM) model for efficient large vocabulary continuous speech recognition is presented. It is synthesized from context-independent phone models with 64 mixture components per state by assigning different mixture weights according to the shared states of triphones. Mixtures are then re-estimated for optimization. The model achieves a word error rate of 7.0% with a 20000-word dictation of newspaper corpus, which is comparable to the best figure by the triphone of much higher resolutions. Compared with conventional PTMs that share Gaussians by all states, the proposed model is easily trained and reliably estimated. Furthermore, the model enables the decoder to perform efficient Gaussian pruning. It is found out that computing only two out of 64 components does not cause any loss of accuracy. Several methods for the pruning are proposed and compared, and the best one reduced the computation to about 20%. Akinobu Lee, Tatsuya Kawahara, Kazuya Takeda, Kiyohiro Shikano |
ICASSP | 3 |
| 2000 | Speech enhancement using nonlinear microphone array with noise adaptive complementary beamformingabstractThis paper describes an improved complementary beamforming microphone array with a new noise adaptation. Complementary beamforming is based on two types of beamformers designed to obtain complementary directivity patterns. In this system, two directivity patterns of the beamformers are adapted to the noise directions so that the expectation values of each noise power spectrum are minimized. Using this technique, we can realize the directional nulls for each noise even when the number of sound sources exceeds that of microphones. To evaluate the effectiveness, speech enhancement experiments are performed based on computer simulations with a two-element array and three sound sources. Compared with the conventional spectral subtraction method cascaded with the adaptive beamformer, it is shown that the proposed array improves the signal-to-noise ratio of degraded speech by more than 6 dB and performs more than 18% better in word recognition rates when the interfering noise is two speakers. Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 3 |
| 2000 | Speech recognition based on space diversity using distributed multi-microphoneabstractThis paper proposes space diversity speech recognition technique using distributed multi-microphones in a room, as a new paradigm of speech recognition. The key technology to realize the system is (1) distant-talking speech recognition and (2) the integration method of multiple inputs. In this paper, we propose the use of a distant speech model for distant-talking speech recognition, and feature-based and likelihood-based integration methods for multimicrophones distributed in the room. The distant speech model is a set of HMMs learned using speech data convolved with the impulse responses measured at several points in the room. The experimental results of simulated distant-talking speech recognition show that the proposed space diversity speech recognition system can attain about 80% in accuracy, while the performances of conventional HMMs using close-talking microphones are less than 50%. These results indicate that the space diversity approach is promising for robust speech recognition under a real acoustic environment. Yasuhiro Shimizu, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 3 |
| 2000 | An acoustic measure for predicting recognition performance degradationabstractAn acoustic measure for predicting the degradation of speech recognition performance due to noise contamination is developed. The merits of the proposed measure over using conventional SNR are that (1) the measure does not require the original clean signal as a reference signal (2) the measure takes the spectral shape of the noise into account and, (3) the measure can predict recognition performance directly. The basic idea of the measure is to utilize the dynamic range of the sub-band signals as an estimate of SNR in the corresponding subband and, to predict the degradation of the recognition performance by taking a product of the recognition accuracy of each sub-band. The proposed measure is tested through experimental evaluation using white Gaussian and human speech like (HSL) noise. In the experiment, the correlation between the predicted and the actual recognition accuracies are 0.96 and 0.99 for white and HSL noise respectively. From the results, the effectiveness of the proposed measure is confirmed. Kazuya Takeda, Masaaki Kondo, Fumitada Itakura |
ICASSP | 1 |
| 2000 | Construction of speech corpus in moving car environmentabstractThe Center for Integrated Acoustic Information Research (CIAIR) at Nagoya University has been collecting speech corpora in moving cars which are made available as resources to advance the research and development of robust ASRs and spoken dialogue systems under high-noise conditions. The speech corpus consists of (1) phonetically balanced sentences, (2) digit strings, (3) discrete words and (4) transcribed spoken dialogues between drivers and information systems for navigation and information retrieval. These data are collected in vehicles under both idling and driving situations. The language of the corpus is currently Japanese. The number of subjects is currently about 300, total recording time is over 200 hours and total corpus size is about 160GByte. We have also been recording video images from three different angles, vehicle-control signals, and vehicle location, all synchronized with the speech recording. We report the objective of the speech corpus, the recording methods and the recording vehicle developed. Nobuo Kawaguchi, Shigeki Matsubara, Hiroyuki Iwa, Shoji Kajita, Kazuya Takeda, Fumitada Itakura, Yasuyoshi Inagaki |
INTERSPEECH | 5 |
| 2000 | Free software toolkit for Japanese large vocabulary continuous speech recognitionabstractA sharable software repository for Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) is introduced. It is designed as a baseline platform for research and developed by researchers of different academic institutes under a governmental support. The repository consists of a recognition engine (Julius), Japanese acoustic models and statistical language models as well as Japanese morphological analysis tools. These modules can be easily integrated and replaced under a plug-and-play framework, which makes it possible to fairly evaluate components and to develop specific application systems. Assessment of these modules and systems in a 20000-word dictation task is reported. The software repository is freely available to the public. Tatsuya Kawahara, Akinobu Lee, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Shigeki Sagayama, Katunobu Itou, Akinori Ito, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
INTERSPEECH | 4 |
| 2000 | Blind source separation based on subband ICA and beamformingabstractThis paper describes a new blind source separation (BSS) method on microphone array using the subband independent component analysis (ICA) and beamforming. The proposed array system consists of the following three sections: (1) subband-ICA-based BSS section, (2) null beamforming section, and (3) integration of (1) and (2) based on the algorithm diversity. Using this technique, we can resolve the low-convergence problem on optimization in ICA. Signal separation and speech recognition experiments clarify that the noise reduction rate (NRR) of about 18 dB is obtained under the nonreverberant condition, and NRRs of 8 dB and 6 dB are obtained in the case that the reverberation times are 150 msec and 300 msec. These performances are superior to those of both simple ICA-based BSS and simple beamforming method. Also, the improvements of the proposed method in word recognition rates are superior to those of the conventional ICA-based BSS method under all reverberant conditions. Hiroshi Saruwatari, Satoshi Kurita, Kazuya Takeda, Fumitada Itakura, Kiyohiro Shikano |
INTERSPEECH | 3 |
| 2000 | Vector space representation of language probabilities through SVD of n-gram matrix
Shiro Terashima, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 2 |
| 2000 | IPA Japanese Dictation Free Software Project
Katunobu Itou, Kiyohiro Shikano, Tatsuya Kawahara, Kazuya Takeda, Atsushi Yamada, Akinori Ito, Takehito Utsuro, Tetsunori Kobayashi, Nobuaki Minematsu, Mikio Yamamoto, Shigeki Sagayama, Akinobu Lee |
LREC | 4 |
| 1999 | Audio data hiding by use of band-limited random sequencesabstractThis paper proposes the use of band-limited random sequences to introduce further flexibility in the spread spectrum based audio data hiding. To realize the sub-band data hiding, a systematic method is developed in order to generate band-limited and orthonormal random sequences of any length. In experiments, we evaluated the selective use of frequency channels to be used for information embedding, and the robustness against the MPEG1 layer 3 encoding and decoding. From the results, it is clarified that the proposed method is robust against more than 160 kbps MPEG1 coding and decoding when the center frequency of the sub-band is lower than 11 kHz. Mikio Ikeda, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 1999 | Compensating of room acoustic transfer functions affected by change of room temperatureabstractThis paper proposes an efficient compensation method using a first-order approximation of time axis scaling for the variations of the room acoustic transfer function. The time axis scaling model is based on the fact that the change of the sound velocity due to the change of room temperature is a dominant factor for the variations of room impulse response affected by environmental conditions. In this paper, the effectiveness of the compensation method is evaluated using room impulse responses measured in the real environment. As the results, it is clarified that the variations of room impulse response can be modeled by the first-order approximated time axis scaling when the successive re-estimation is performed every small change of temperature. Furthermore, it is shown that the compensation method applied to an inverse filtering based dereverberation approach improves the intelligibility and speech recognition rates dramatically. Michiaki Omura, Motohiko Yada, Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 5 |
| 1999 | Speech enhancement using nonlinear microphone array with complementary beamformingabstractThis paper describes an improved spectral subtraction method by using the complementary beamforming microphone array to enhance noisy speech signals for speech recognition. The complementary beamforming is based on two types of beamformers designed to obtain complementary directivity patterns with respect to each other. It is shown that the nonlinear subtraction processing with complementary beamforming can result in a kind of the spectral subtraction without the need for speech pause detection. In addition, the design of the optimization algorithm for the directivity pattern is also described. To evaluate the effectiveness, speech enhancement experiments and speech recognition experiments are performed based on computer simulations. In comparison with the optimized conventional delay-and-sum array, it is shown that the proposed array improves the signal-to-noise ratio of degraded speech by about 2 dB and performs about 10% better in word recognition rates under heavy noisy conditions. Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 3 |
| 1999 | Speaker conversion through non-linear frequency warping of straight spectrumabstractThis paper presents our approach to automatically detect tone nuclei, and to use their features for recognizing lexical tones of Chinese continuous speech. We have suggested that Fundamental frequency (F0) contour of a syllable usually consists of three segments: onset course, tone nucleus and offset course. Among them, only tone nucleus contains key features for tone discrimination, hence the tone critical segment of a syllable. The other two segments result from physiological transition effect of human vocal cords, and are affected largely by adjacent tones in continuous speech. The tone nucleus can be detected out by a two-process scheme; the first process segments a syllable F0 contour by Segmental Clustering algorithm, and the second one finds tone nucleus according to knowledge rules on suprasegmental features. Tone recognition performance can be improved by using tone nucleus features and discarding others. Tone recognition experimental results proved the advantage of our method over the conventional ones. Noriyasu Maeda, Hideki Banno, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
EUROSPEECH | 4 |
| 1999 | Speech enhancement using nonlinear microphone array under nonstationary noise conditions
Hiroshi Saruwatari, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
EUROSPEECH | 3 |
| 1998 | Spectral weighting of SBCOR for noise robust speech recognitionabstractSubband-autocorrelation (SBCOR) analysis is a noise robust acoustic analysis based on filter bank and autocorrelation analysis, and aims to extract the periodicities associated with the inverse of the center frequency in a subband. In this paper, it is derived that SBCOR results in the lateral inhibitive weighting (LIW) processing of the power spectrum, and it is shown that the LIW is significantly effective for noise robust acoustic analysis using a DTW word recognizer. An interpretation of the LIW is also described. A flattening technique of the noise spectral envelope using an LPC inverse filter is applied to speech degraded with noise, and DTW word recognition is performed. The idea of this inverse filtering technique comes from weakening the strong periodic components included in noise. The experimental results using a 32th order LPC inverse filter show that the recognition performance of SBCOR (or LIW) is improved for computer room noise. Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 1998 | Balancing acoustic and linguistic probabilitiesabstractThe length of the word sequence is not taken into account under language modeling of n-gram local probability modeling. Due to this property the optimal values of the language weight and word insertion penalty for balancing acoustic and linguistic probabilities is affected by the length of word sequence. To deal with this problem, a new language model is developed based on the Bernoulli trial model taking the length of the word sequence into account. Not only better recognition accuracy but also more robust balancing with acoustic probability compared with the normal n-gram model of the proposed method is confirmed through recognition experiments. Atsunori Ogawa, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 1998 | The design of the newspaper-based Japanese large vocabulary continuous speech recognition corpusabstractIn this paper we present the first public Japanese speech corpus for large vocabulary continuous speech recognition (LVCSR) technology, which we have titled JNAS (Japanese Newspaper Article Sentences). We designed it to be comparable to the corpora used in the American and European LVCSR projects. The corpus contains speech recordings (60 hrs.) and their orthographic transcriptions for 306 speakers (153 males and 153 females) reading excerpts from the newspaper's articles and phonetically balanced (PB) sentences. This corpus contains utterances of about 45,000 sentences as a whole with each speaker reading about 150 sentences. JNAS is being distributed on 16 CD-ROMs. Katunobu Itou, Mikio Yamamoto, Kazuya Takeda, Toshiyuki Takezawa, Tatsuo Matsuoka, Tetsunori Kobayashi, Kiyohiro Shikano, Shuichi Itahashi |
ICSLP | 3 |
| 1998 | Sharable software repository for Japanese large vocabulary continuous speech recognitionabstractThe project of Japanese LVCSR (Large Vocabulary Continuous Speech Recognition) platform is introduced. It is a collaboration of researchers of different academic institutes and intended to develop a sharable software repository of not only databases but also models and programs. The platform consists of a standard recognition engine, Japanese phone models and Japanese statistical language models. A set of Japanese phone HMMs are trained with ASJ (Acoustic Society of Japan) databases of 20K sentence utterances per each gender. Japanese word N-gram (2-gram and 3-gram) models are constructed with a corpus of Mainichi newspaper of four years. The recognition engine JULIUS is developed for assessment of both acoustic and language models. The modules are integrated as a Japanese LVCSR system and evaluated on 5000-word dictation task. The software repository is available to the public. Tatsuya Kawahara, Tetsunori Kobayashi, Kazuya Takeda, Nobuaki Minematsu, Katunobu Itou, Mikio Yamamoto, Atsushi Yamada, Takehito Utsuro, Kiyohiro Shikano |
ICSLP | 3 |
| 1998 | Estimating entropy of a language from optimal word insertion penalty
Kazuya Takeda, Atsunori Ogawa, Fumitada Itakura |
ICSLP | 1 |
| 1997 | A binaural speech processing method using subband-cross correlation analysis for noise robust recognitionabstractThis paper describes an extended subband-cross-correlation (SBXCOR) analysis to improve the robustness against noise. The SBXCOR analysis, which has been already proposed, is a binaural speech processing technique using two input signals and extracts the periodicities associated with the inverse of the center frequency (CF) in each subband. In this paper, by taking an exponentially weighted sum of crosscorrelation at the integral multiples of the inverse of CF, SBXCOR is extended so as to capture more periodicities included in two input signals. The experimental results using a DTW word recognizer showed that the processing improves the performance of SBXCOR for both that of the white noise and a computer room noise. For white noise, the extended SBXCOR performed significantly better than the smoothed group delay spectrum and the mel-frequency cepstral coefficient (MFCC) extracted from both monaural and binaural signals. However, for the computer room noise, it outperformed only at SNR 0 dB. Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICASSP | 2 |
| 1997 | Voice activity detection using source separation techniques
Tomohiko Taniguchi, Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
EUROSPEECH | 3 |
| 1996 | Subband-crosscorrelation analysis for robust speech recognition
Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICSLP | 2 |
| 1996 | Extracting speech features from human speech-like noise
Shoji Kajita, Kazuya Takeda, Fumitada Itakura |
ICSLP | 3 |
| 1996 | Variability of lombard effects under different noise conditions
Atsushi Wakao, Kazuya Takeda, Fumitada Itakura |
ICSLP | 2 |
| 1995 | A prototype of a Japanese-Korean realtime speech translation system
Masami Suzuki, Naomi Inoue, Fumihiro Yato, Kazuya Takeda, Seiichi Yamamoto |
EUROSPEECH | 4 |
| 1995 | Top-down speech detection and n-best meaning search in a voice activated telephone extension system
Kazuya Takeda, Shingo Kuroiwa, Masaki Naito, Seiichi Yamamoto |
EUROSPEECH | 1 |
| 1994 | A trellis-based implementation of minimum error rate training
Kazuya Takeda, Tetsunori Murakami, Shingo Kuroiwa, Seiichi Yamamoto |
ICSLP | 1 |
| 1993 | A voice-activated extension telephone exchange system
Shingo Kuroiwa, Kazuya Takeda, Naomi Inoue, Izuru Nogaito, Seiichi Yamamoto, Makoto Shozakai, Kunihiko Owa, Masahiko Takahashi, Ryuuji Matsumoto |
EUROSPEECH | 2 |
| 1993 | Improving robustness of network grammar by using class HMM
Kazuya Takeda, Naomi Inoue, Shingo Kuroiwa, Tomohiro Konuma, Seiichi Yamamoto |
EUROSPEECH | 1 |
| 1992 | Architecture and algorithms of a real-time word recognizer for telephone input
Shingo Kuroiwa, Kazuya Takeda, Fumihiro Yato, Seiichi Yamamoto, Kunihiko Owa, Makoto Shozakai, Ryuuji Matsumoto |
ICSLP | 2 |
| 1990 | Statistical analysis for segmental duration rules in Japanese speech synthesis
Nobuyoshi Kaiki, Kazuya Takeda, Yoshinori Sagisaka |
ICSLP | 2 |
| 1990 | A large-scale Japanese speech database
Yoshinori Sagisaka, Kazuya Takeda, M. Abel, Shigeru Katagiri, T. Umeda, Hisao Kuwabara |
ICSLP | 2 |
| 1990 | On the unit search criteria and algorithms for speech synthesis using non-uniform units
Kazuya Takeda, Katsuo Abe, Yoshinori Sagisaka |
ICSLP | 1 |
| 1990 | ATR Japanese speech database as a tool of speech recognition and synthesis
Akira Kurematsu, Kazuya Takeda, Yoshinori Sagisaka, Shigeru Katagiri, Hisao Kuwabara, Kiyohiro Shikano |
Speech Commun. | 2 |
| 1989 | Construction of a large-scale Japanese speech database and its management systemabstractA large-scale Japanese speech database is described. It consists of (1) an isolated-word speech database and (2) a continuous-speech database. To facilitate usage for speech research, multiple transcriptions have been made in five different layers from a simple phonemic description to fine acoustic-phonetic transcriptions. A database management system has also been developed to establish an effective link between speech data and label data for easy access to both data.> Hisao Kuwabara, Kazuya Takeda, Yoshinori Sagisaka, Shigeru Katagiri, S. Morikawa |
ICASSP | 2 |
| 1989 | Adaptive manipulation of non-uniform synthesis units using multi-level unit transcription
Kazuya Takeda, Katsuo Abe, Yoshinori Sagisaka, Hisao Kuwabara |
EUROSPEECH | 1 |