VLDB 2026 Research / reviewers in the wild / expert
Hiromitsu Nishizaki
dblp:87/1050
· DBLP profile ↗
49ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-7717-8312ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 30 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 10 · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reference-Free Handwritten Japanese Character Generation via CLIP-Conditioned Diffusion Models
Koki Fujita, Hideaki Yajima, Chee Siang Leow, Hiromitsu Nishizaki |
ICDAR (3) | 4 |
| 2025 | Design and Training of a Sound Classification Model on a Resource-Constrained Edge Device for Fruit Theft PreventionabstractIn recent years, fruit theft has become a serious social problem in Japan, where conventional surveillance measures are limited due to the unique challenges in agricultural areas, such as limited power supply and large areas. To address this, previous studies have developed battery-powered, microphone-based theft detection systems that operate on microcontrollers; however, achieving high classification accuracy under severe resource constraints remains a challenge. This paper presents a design and training methodology for a compact, high-accuracy sound classification model for fruit theft prevention. Our approach combines Depthwise Separable Convolution (DSC) for efficient feature extraction with Ensemble Distillation (EnD), which transfers knowledge from two teacher models with distinct feature extraction approaches. Experiments on a three-class task (environmental sounds, speech, and footsteps) demonstrated that the proposed model achieved an F1-score of 88.61%, an improvement of 12.00 percentage points over the baseline. The inference memory usage was limited to 49 kB, and including the entire system for processes other than deep learning inference, the total memory usage remained within 160 kB, well below the 256 kB RAM capacity of the target hardware. Moreover, the model achieved an inference time of approximately 0.4519 seconds per 1-second audio segment on a Seeeduino XIAO nRF52840 micro-controller. These results confirm the practical feasibility of the proposed system as a robust AI-based surveillance solution under strict resource constraints for real-world orchard applications. Haruki Endo, Hideaki Yajima, Chee Siang Leow, Tsutomu Tanzawa, Koji Makino, Kazuyoshi Ishida, Hiromitsu Nishizaki |
IECON | 7 |
| 2025 | VCapAV: A Video-Caption Based Audio-Visual Deepfake Detection Dataset
Yikang Wang, Qishan Zhang, Hiromitsu Nishizaki, Ming Li 0026 |
INTERSPEECH | 4 |
| 2025 | Non-invasive estimation of Shine Muscat grape color and sensory evaluation from standard camera imagesabstractAbstract This study proposes a non-invasive method to estimate both color and sensory attributes of Shine Muscat grapes from standard camera images. First, we focus on color estimation by integrating a Vision Transformer (ViT) feature extractor with interquartile range (IQR)-based outlier removal. Experimental results show that our approach achieves 97.2% accuracy, significantly outperforming Convolutional Neural Network (CNN) models. This improvement underscores the importance of capturing global contextual information to differentiate subtle color variations in grape ripeness. Second, we address human sensory evaluation by collecting questionnaire responses on 13 attributes (e.g., “Sweetness,” “Overall taste rating”), each rated on a five-point scale. Because these ratings tend to cluster around midrange values (labels “2,” “3,” and “4”), we initially limit the dataset to the extreme labels “1” (“lowest grade”) and “5” (“highest grade”) for binary classification. Three attributes—“Overall color,” “Sweetness,” and “Overall taste rating”—exhibit relatively high classification accuracies of 79.9%, 75.1%, and 75.7%, respectively. By contrast, the other 10 attributes reach only 50%–66%, suggesting that subjective variations and limited visual cues pose significant challenges. Overall, the proposed approach demonstrates the feasibility of an image-based system that integrates color estimation and sensory evaluation to support more objective, data-driven harvest timing decisions for Shine Muscat grapes. Ryosuke Shimazu, Chee Siang Leow, Prawit Buayai, Xiaoyang Mao, Wan-Young Chung, Hiromitsu Nishizaki |
Vis. Comput. | 6 |
| 2024 | The dynamic load-induced bioelectric potentials of knee jointabstractThe monitoring and prediction of joint load based on physical activity is important but challenging in the medical field due to the limited neural perception of cartilage and technical obstacles in real-time monitoring of dynamic daily activities. This study explores a novel, non-invasive method to measure load-generated potentials in knee cartilage using surface electrodes. Twenty subjects performed both active and passive knee extensions while potentials were recorded from seven electrodes around the knee. The signals were analyzed across five phases: initial flat, activity, post-activity, return, and post-return. Results showed significantly higher mean amplitudes in active extensions compared to passive ones, particularly in the post-activity and post-return phases. This study identifies specific bioelectric signals correlating with knee kinematics, offering potential for improved assessment of joint health and rehabilitation strategies. Jae Hyun Lee, Ye-Seul Jang, Hiromitsu Nishizaki, Won-Du Chang |
CW | 3 |
| 2024 | Analysis of Acute Stress Reaction with EEG using MI-mRMR EEG Channel Selection and Ensemble Learning with Majority VotingabstractIndividuals facing challenging and threatening situations significantly strain their minds and bodies. As a result, they may experience emotional, physical, or psychological stress. However, people perceive stress differently depending on how long or how severe their exposition to traumatic events is. Prolonged stress can cause significant and often inexplicable damage to both the body and mind. Because of that, early detection of human stress levels has become a prominent means of early diagnosis. Therefore, the major goal of this study is to investigate how human stress detection using Electroencephalography (EEG) which has been applied to EEG channel selection using Mutual Information (MI) with Minimum Redundancy Maximum Relevance (mRMR) (Mi$m R M R$) and the ensemble learning algorithm mainly using the method of bagging, boosting and stacking works in classifying different stress levels. On the other hand, the features were extracted using time-domain, frequency domain and timefrequency domain analysis containing relevant features related to stress data. Based on the experiment results, the ensemble learning method bagging performs better compared to the boosting method and can classify the data with the highest accuracy, highest F1 score, and highest precision, of 88 percent and highest precision of 94 percent. Hence, this shows the capability of the ensemble learning algorithms to classify stress EEG data that has been selected using (Mi-mRMR) algorithms as a means of early stress detection and diagnosis model. Muhammad Rasydan Mazlan, Abdul Syafiq Abdull Sukor, Abdul Hamid Adom, Latifah Kamarudin, Hiromitsu Nishizaki, Norasmadi Abdul Rahim |
CW | 5 |
| 2024 | High Quality Color Estimation of Shine Muscat Grape Using Vision TransformerabstractCurrently, skilled farmers judge the ripeness of the Shine Muscat grape variety by looking at the color on the surface of the grapes. However, the color of Shine Muscat grapes does not change much as they grow, and there are individual differences in the way the color is perceived. Furthermore, the same color can look very different depending on the exposure to sunlight and shadows. Therefore, there is a need for a system that can quantitatively determine the color of Shine Muscat grapes to pass on the harvesting techniques of experienced farmers to amateurs and inexperienced farmers. This research aims to improve the accuracy of the color estimation of Shine Muscat grapes using deep learning. We propose a method to estimate the color of individual grapes using a color estimation model with a self-attention mechanism, from which the color of the whole bunch is estimated. A Vision Transformer model with a self-attention mechanism was found to improve the color estimation accuracy to $96.9 \%$. Furthermore, by eliminating outliers using the interquartile range, a color estimation accuracy of $97.2 \%$ could be achieved, demonstrating the effectiveness of the new color estimation model. Ryosuke Shimazu, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki |
CW | 6 |
| 2024 | Prediction Model with Penalized Hyperparameter Optimization for mRNA Vaccine Degradation Based on Tetra-nitrogenous-base AnalysisabstractMessenger ribonucleic acid (mRNA) vaccines, despite their rapid degradation, play a critical role in pandemic response due to their high efficacy and swift production capabilities. Accurate prediction of mRNA vaccine degradation rates is vital for determining their shelf life and maintaining efficacy. This study presents a tetra-nitrogenous-base label encoding approach (4-ntb-lbA) integrated with a novel hyperparameter optimization (HPO) technique, named the Hyperparameter Optimization Penalizer (HOPeR), aimed at enhancing prediction precision. The state-of-the-art hybrid Dense-BiGRU-BiLSTM-BiLSTM (Hybrid_LGSS) model underwent rigorous testing to assess both the model’s performance and the proposed methodologies, broadening its interdisciplinary applications. Results indicate that the 4 -ntblbA method, which leverages bioinformatic data, substantially improves prediction reliability and accuracy while reducing error rates (Set_I: training loss $={0. 0 9 0 4}$, validation loss $=$ 0.0938; Set_II: training loss $=0.0141$, validation loss $=0.0145$), assessed with mean column-wise root mean square error (MCRMSE). The HOPeR approach demonstrated effectiveness across various models and HPO algorithms, including Particle Swarm Optimization (PSO), Bayesian Optimization with Gaussian Process (BOGP), and the RIME optimization algorithm (RIME), by minimizing the risk of suboptimal hyperparameter configurations while promoting fast convergence through penalization strategies. These findings validate the proposed approaches’ flexibility, applicability, and robustness across different algorithms. Additionally, beyond advancing mRNA vaccine degradation prediction, this research introduces a versatile framework for HPO with broad applicability. By fostering interdisciplinary integration and innovation, this paper significantly enhances the precision and efficiency of predictive models in bioinformatics, machine learning (ML), and beyond, thereby paving the way for future advancements across various scientific and engineering disciplines. Hwai Ing Soon, Abdullah Azian Azamimi, Hiromitsu Nishizaki, Latifah Kamarudin |
CW | 3 |
| 2024 | Development of a Fruit Theft Reporting System Using a Compact Microcontroller with Deep Learning Based on Suspicious SoundsabstractIn Japanese fruit-growing regions, fruit theft is a significant issue. Traditional anti-theft methods are fraught with various shortcomings and are often ineffective. To address this, we have developed a new approach for preventing fruit theft: a device that detects suspicious sounds by integrating a sound sensor with a compact, low-power microcontroller. This paper presents a suspicious sound detection system designed for a fruit theft alert device. It details the development of a deep learning model for detecting suspicious sounds, covering aspects from data collection to model training and evaluation. Despite operating on a microcontroller with limited memory and computational power, the system under development has achieved a classification accuracy of up to 78.5% in F1-score for footsteps, speech, and other environmental sounds. Chee Siang Leow, Tsutomu Tanzawa, Tze Yaw Bong, Koji Makino, Kazuyoshi Ishida, Hiromitsu Nishizaki |
IECON | 6 |
| 2024 | The Database and Benchmark For the Source Speaker Tracing Challenge 2024abstractVoice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/ Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026 |
SLT | 8 |
| 2023 | Automatic Exploration of Optimal Data Processing Operations for Sound Data Augmentation Using Improved Differentiable Automatic Data Augmentation
Toki Sugiura, Hiromitsu Nishizaki |
INTERSPEECH | 2 |
| 2023 | A new speech corpus of super-elderly Japanese for acoustic modelingabstractThe development of accessible speech recognition technology will allow the elderly to more easily access electronically stored information. However, the necessary level of recognition accuracy for elderly speech has not yet been achieved using conventional speech recognition systems, due to the unique features of the speech of elderly people. To address this problem, we have created a new speech corpus named EARS (Elderly Adults Read Speech), consisting of the recorded read speech of 123 super-elderly Japanese people (average age: 83.1), as a resource for training automated speech recognition models for the elderly. In this study, we investigated the acoustic features of super-elderly Japanese speech using our new speech corpus. In comparison to the speech of less elderly Japanese speakers, we observed a slower speech rate and extended vowel duration for both genders, a slight increase in fundamental frequency for males, and a slight decrease in fundamental frequency for females. To demonstrate the efficacy of our corpus, we also conducted speech recognition experiments using two different acoustic models (DNN-HMM and transformer-based), trained with a combination of data from our corpus and speech data from three conventional Japanese speech corpora. When using the DNN-HMM trained with EARS and speech data from existing corpora, the character error rate (CER) was reduced by 7.8% (to just over 9%), compared to a CER of 16.9% when using only the baseline training corpora. We also investigated the effect of training the models with various amounts of EARS data, using a simple data expansion method. The acoustic models were also trained for various numbers of epochs without any modifications. When using the Transformer-based end-to-end speech recognizer, the character error rate was reduced by 3.0% (to 11.4%) by using a doubled EARS corpus with the baseline data for training, compared to a CER of 13.4% when only data from the baseline training corpora were used. Meiko Fukuda, Ryota Nishimura, Hiromitsu Nishizaki, Koharu Horii, Yurie Iribe, Kazumasa Yamamoto, Norihide Kitaoka |
Comput. Speech Lang. | 3 |
| 2022 | Peer Collaborative Learning for Polyphonic Sound Event DetectionabstractThis paper describes how semi-supervised learning, called peer collaborative learning (PCL), can be applied to the polyphonic sound event detection (PSED) task, which is one of the tasks in the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge. Many deep learning models have been studied to determine what kind of sound events occur where and for how long in a given audio clip. The characteristic of PCL used in this paper is the combination of ensemble-based knowledge distillation into sub-networks and student-teacher model-based knowledge distillation, which can train a robust PSED model from a small amount of strongly labeled data, weakly labeled data, and a large amount of unlabeled data. We evaluated the proposed PCL model using the DCASE 2019 Task 4 dataset, and achieved an F1-score improvement of about 8.2 points compared with the baseline model. Hayato Endo, Hiromitsu Nishizaki |
ICASSP | 2 |
| 2022 | Handwritten Character Generation using Y-Autoencoder for Character Recognition Model TrainingabstractIt is well-known that the deep learning-based optical character recognition (OCR) system needs a large amount of data to train a high-performance character recognizer. However, it is costly to collect a large amount of realistic handwritten characters. This paper introduces a Y-Autoencoder (Y-AE)-based handwritten character generator to generate multiple Japanese Hiragana characters with a single image to increase the amount of data for training a handwritten character recognizer. The adaptive instance normalization (AdaIN) layer allows the generator to be trained and generate handwritten character images without paired-character image labels. The experiment shows that the Y-AE could generate Japanese character images then used to train the handwritten character recognizer, producing an F1-score improved from 0.8664 to 0.9281. We further analyzed the usefulness of the Y-AE-based generator with shape images, out-of-character (OOC) images, which have different character images styles in model training. The result showed that the generator could generate a handwritten image with a similar style to that of the input character. Tomoki Kitagawa, Chee Siang Leow, Hiromitsu Nishizaki |
LREC | 3 |
| 2022 | Appropriate grape color estimation based on metric learning for judging harvest timingabstractAbstract The color of a bunch of grapes is a very important factor when determining the appropriate time for harvesting. However, judging whether the color of the bunch is appropriate for harvesting requires experience and the result can vary by individuals. In this paper, we describe a system to support grape harvesting based on color estimation using deep learning. To estimate the color of a bunch of grapes, bunch detection, grain detection, removal of pest grains, and color estimation are required, for which deep learning-based approaches are adopted. In this study, YOLOv5, an object detection model that considers both accuracy and processing speed, is adopted for bunch detection and grain detection. For the detection of diseased grains, an autoencoder-based anomaly detection model is also employed. Since color is strongly affected by brightness, a color estimation model that is less affected by this factor is required. Accordingly, we propose multitask learning that uses metric learning. The color estimation model in this study is based on AlexNet. Metric learning was applied to train this model. Brightness is an important factor affecting the perception of color. In a practical experiment using actual grapes, we empirically selected the best three image channels from RGB and CIELAB (L*a*b*) color spaces and we found that the color estimation accuracy of the proposed multi-task model, the combination with “L” channel from L*a*b color space and “GB” from RGB color space for the grape image (represented as “LGB” color space), was 72.1%, compared to 21.1% for the model which used the normal RGB image. In addition, it was found that the proposed system was able to determine the suitability of grapes for harvesting with an accuracy of 81.6%, demonstrating the effectiveness of the proposed system. Tatsuyoshi Amemiya, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki |
Vis. Comput. | 6 |
| 2021 | Development of a Support System for Judging the Appropriate Timing for Grape HarvestingabstractThe color of grape bunches is a significant factor when harvesting grapes at the appropriate timing. Judging the suitable color for shipment requires experience and varies from one person to another. We herein describe a support system for grape harvesting based on color estimation. To estimate the color of a bunch of grapes, bunch detection, grain detection, removal of diseased grains, and color estimation should be performed. Models based on deep learning are employed for this series of processes. Since color is strongly affected by sunlight, we propose a multitask model that considers sunlight exposure to achieve a robust color estimation model that exhibits decreased sensitivity to sunlight. Our results show that the color estimation accuracy of the model is 76% when sunlight exposure is not considered and 81% when sunlight exposure is considered. In addition, we performed a practical field test of the developed harvest support system in an actual grape field. The results show that our support system can determine the appropriateness of grape harvest with an accuracy of 90%, demonstrating the effectiveness of the system. Tatsuyoshi Amemiya, Kodai Akiyama, Chee Siang Leow, Prawit Buayai, Koji Makino, Xiaoyang Mao, Hiromitsu Nishizaki |
CW | 7 |
| 2021 | End-to-End Inflorescence Measurement for Supporting Table Grape Trimming with Augmented RealityabstractInflorescence trimming is a crucial process to produce high-quality table grapes. It can eliminate nutrient competition in a bunch and makes it less vulnerable to disease development. After trimming, the remaining part of the inflorescence should have a target length decided by the grape variety. This is challenging for novice farmers because of the time constraint. The farmer needs to finish trimming the inflorescence before the berries develop. This paper proposes a novel end-to-end inflorescence length measurement method for supporting a trimming process with augmented reality technology. The proposed technique makes use of the state-of-the-art deep neural network model for detecting the inflorescence area, as well as the scissors from the images captured with a camera installed on an optical see-through head-mounted display. A new algorithm is designed to estimate the length of the remaining inflorescence with the screw of the scissors loop as the calibrator. The estimated length is then visualized on the head-mounted display to support the farmer in performing the trimming correctly and efficiently. The experiment, conducted with real inflorescence trimming tasks, shows that the mean absolute error of the length estimation is only 0.19 cm, which is small enough for use in real applications. Prawit Buayai, Kabin Yok-In, Daisuke Inoue 0004, Chee Siang Leow, Hiromitsu Nishizaki, Koji Makino, Xiaoyang Mao |
CW | 5 |
| 2021 | Supporting Vine Vegetation Status Observation Using ARabstractAugmented reality (AR) is a technology that expands information by superimposing digital information, such as virtual objects, on the real world using smartphones, smart glasses, and head-mounted displays. It is used in a variety of situations. In this paper, we propose a system that allows vine farmers to investigate effectively the vegetation condition of trellising-style vineyards using a head mounted display and AR technology. The experiment results show that by using a hybrid navigation approach include showing the whole vineyard in a small window and showing the details only, when necessary, the proposed system enable the farmers to move to a location with concern accurately. Daisuke Inoue 0004, Prawit Buayai, Hiromitsu Nishizaki, Koji Makino, Xiaoyang Mao |
CW | 3 |
| 2021 | Semi-Supervised Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis is a rapidly growing domain in natural language processing which is a fine-grained study. Within this broad field, most existing studies use large amounts of labeled data by deep learning methods. However, obtaining massive quantities of labeled data to train a deep neural network model is frequently time-consuming and laborious. In this paper, we focus on semi-supervised learning based on ACSA with few labeled data in restaurant reviews and scholarly paper reviews. In order to leverage information from unlabeled data, the semi-supervised learning method-Ladder network is proposed to fix the problem. Furthermore, the pre-trained language models BERT, ALBERT and Longformer are used for text pre-processing and feature extraction. Extensive experiments on both datasets demonstrate the superiority of the Longformer based Ladder Network compared with supervised learning methods and other semi-supervised learning methods including$\Gamma$-Model and VAT. Yoshimi Suzuki, Fumiyo Fukumoto, Hiromitsu Nishizaki |
CW | 5 |
| 2021 | Language and Speaker-Independent Feature Transformation for End-to-End Multilingual Speech Recognition
Tomoaki Hayakawa, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki |
Interspeech | 5 |
| 2021 | Voice Activity Detection for Live Speech of Baseball Game Based on Tandem Connection with Speech/Noise Separation Model
Yuto Nonaka, Chee Siang Leow, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki |
Interspeech | 5 |
| 2020 | Sentiment analysis using semi-supervised learning with few labeled dataabstractSentiment analysis has been widely explored in many text domains, including tweets, movie reviews, shop/restaurant reviews, product reviews, and peer reviews for scholarly papers. However, it is very costly to manually label the training data for sentiment analysis. We focus on the problem and presents an approach for leveraging contextual features from unlabeled movie and restaurant reviews with a neural-network-based learning model, Ladder network. The experimental results by using two benchmark datasets, IMDb and YelpNYC, show that our model outperforms the baseline models including LSTM and SVM. Especially we verified that our model is better performance gaining on limited training datasets with 1% data labeled. Our source codes are available online.11Our source code can be obtained from https://github.com/jepyh/sentiment_analysis_few_labeled. Yuhao Pan, Zhiqun Chen, Yoshimi Suzuki, Fumiyo Fukumoto, Hiromitsu Nishizaki |
CW | 5 |
| 2020 | Automatic Fluency Evaluation of Spontaneous Speech Using Disfluency-Based FeaturesabstractThis paper describes an automatic fluency evaluation of spontaneous speech. Although we regularly observe a variety of different disfluencies in spontaneous speech, we focus on two types of phenomena, i.e., filled pauses and word fragments. This paper aims to reveal that these two types of disfluencies have effects on speech fluency evaluation differently. To this end, we conduct a series of SVM classification experiments on the Japanese spontaneous speech corpus. The experimental results show that the features derived from word fragments are effective in evaluating disfluent speech especially when combined with prosodic features such as speech rate and pauses/silence, while the features from filled pauses are not effective in evaluating fluency. Huaijin Deng, Youchao Lin, Takehito Utsuro, Akio Kobayashi, Hiromitsu Nishizaki, Junichi Hoshino |
ICASSP | 5 |
| 2020 | Integrating Disfluency-based and Prosodic Features with Acoustics in Automatic Fluency Evaluation of Spontaneous SpeechabstractThis paper describes an automatic fluency evaluation of spontaneous speech. In the task of automatic fluency evaluation, we integrate diverse features of acoustics, prosody, and disfluency-based ones. Then, we attempt to reveal the contribution of each of those diverse features to the task of automatic fluency evaluation. Although a variety of different disfluencies are observed regularly in spontaneous speech, we focus on two types of phenomena, i.e., filled pauses and word fragments. The experimental results demonstrate that the disfluency-based features derived from word fragments and filled pauses are effective relative to evaluating fluent/disfluent speech, especially when combined with prosodic features, e.g., such as speech rate and pauses/silence. Next, we employed an LSTM based framework in order to integrate the disfluency-based and prosodic features with time sequential acoustic features. The experimental evaluation results of those integrated diverse features indicate that time sequential acoustic features contribute to improving the model with disfluency-based and prosodic features when detecting fluent speech, but not when detecting disfluent speech. Furthermore, when detecting disfluent speech, the model without time sequential acoustic features performs best even without word fragments features, but only with filled pauses and prosodic features. Huaijin Deng, Youchao Lin, Takehito Utsuro, Akio Kobayashi, Hiromitsu Nishizaki, Junichi Hoshino |
LREC | 5 |
| 2020 | Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech RecognitionabstractIn an aging society like Japan, a highly accurate speech recognition system is needed for use in electronic devices for the elderly, but this level of accuracy cannot be obtained using conventional speech recognition systems due to the unique features of the speech of elderly people. S-JNAS, a corpus of elderly Japanese speech, is widely used for acoustic modeling in Japan, but the average age of its speakers is 67.6 years old. Since average life expectancy in Japan is now 84.2 years, we are constructing a new speech corpus, which currently consists of the utterances of 221 speakers with an average age of 79.2, collected from four regions of Japan. In addition, we expand on our previous study (Fukuda, 2019) by further investigating the construction of acoustic models suitable for elderly speech. We create new acoustic models and train them using a combination of existing Japanese speech corpora (JNAS, S-JNAS, CSJ), with and without our ‘super-elderly’ speech data, and conduct speech recognition experiments. Our new acoustic models achieve word error rates (WER) as low as 13.38%, exceeding the results of our previous study in which we used the CSJ acoustic model adapted for elderly speech (17.4% WER). Meiko Fukuda, Hiromitsu Nishizaki, Yurie Iribe, Ryota Nishimura, Norihide Kitaoka |
LREC | 2 |
| 2020 | Semi-Automatic Construction and Refinement of an Annotated Corpus for a Deep Learning Framework for Emotion ClassificationabstractIn the case of using a deep learning (machine learning) framework for emotion classification, one significant difficulty faced is the requirement of building a large, emotion corpus in which each sentence is assigned emotion labels. As a result, there is a high cost in terms of time and money associated with the construction of such a corpus. Therefore, this paper proposes a method of creating a semi-automatically constructed emotion corpus. For the purpose of this study sentences were mined from Twitter using some emotional seed words that were selected from a dictionary in which the emotion words were well-defined. Tweets were retrieved by one emotional seed word, and the retrieved sentences were assigned emotion labels based on the emotion category of the seed word. It was evident from the findings that the deep learning-based emotion classification model could not achieve high levels of accuracy in emotion classification because the semi-automatically constructed corpus had many errors when assigning emotion labels. In this paper, therefore, an approach for improving the quality of the emotion labels by automatically correcting the errors of emotion labels is proposed and tested. The experimental results showed that the proposed method worked well, and the classification accuracy rate was improved to 55.1% from 44.9% on the Twitter emotion classification task. Kyosuke Masuda, Hiromitsu Nishizaki, Fumiyo Fukumoto, Yoshimi Suzuki |
LREC | 3 |
| 2019 | Classification of Swing Motion of Tennis using a Recurrent-based Neural NetworkabstractAll generation person should enjoy playing sports to keep and improve health. It is important to instruct a beginner in the techniques of sports to enjoy a sport. However, instruction for the beginner is challenging because of the difficulty of evaluation of the motion. In this paper, the classification for the evaluation of the motion in the sport is addressed using the recurrent-based neural network that is one of the deep learning. Moreover, this paper deals with the swing motion of tennis, since the swing motion of the tennis is essential, and the beginner is often instructed in the swing motion at first. First, a developed measurement system for human motion is described. Next, the recurrent-based neural network to classify the swing motion is shown. Final, the classification results are discussed. The individual swing motion can be classified using deep learning framework. However, it is clear that the swing motion of the experienced player is not always the same. Therefore, we confirm that individual instruction is important to improve the motion of the sport. Koji Makino, Yudai Kitano, Hiromitsu Nishizaki |
HSI | 3 |
| 2019 | Audio Classification of Bit-Representation WaveformabstractThis study investigated the waveform representation for audio signal classification. Recently, many studies on audio waveform classification such as acoustic event detection and music genre classification have been published. Most studies on audio waveform classification have proposed the use of a deep learning (neural network) framework. Generally, a frequency analysis method such as Fourier transform is applied to extract the frequency or spectral information from the input audio waveform before inputting the raw audio waveform into the neural network. In contrast to these previous studies, in this paper, we propose a novel waveform representation method, in which audio waveforms are represented as a bit sequence, for audio classification. In our experiment, we compare the proposed bit representation waveform, which is directly given to a neural network, to other representations of audio waveforms such as a raw audio waveform and a power spectrum with two classification tasks: one is an acoustic event classification task and the other is a sound/music classification task. The experimental results showed that the bit representation waveform achieved the best classification performance for both the tasks. Masaki Okawa, Takuya Saito, Naoki Sawada, Hiromitsu Nishizaki |
INTERSPEECH | 4 |
| 2017 | Usability and Learning Effect Evaluations of an Electrical Note-Taking Support System with Speech Processing Technologies
Hiromitsu Nishizaki, Yosuke Narita |
ICCE | 1 |
| 2017 | Parallel Hierarchical Attention Networks with Shared Memory Reader for Multi-Stream Conversational Document Classification
Naoki Sawada, Ryo Masumura, Hiromitsu Nishizaki |
INTERSPEECH | 3 |
| 2016 | Recurrent Neural Network-Based Phoneme Sequence Estimation Using Multiple ASR Systems' Outputs for Spoken Term Detection
Naoki Sawada, Hiromitsu Nishizaki |
INTERSPEECH | 2 |
| 2015 | Two-step spoken term detection using SVM classifier trained with pre-indexed keywords based on ASR result
Kentaro Domoto, Takehito Utsuro, Naoki Sawada, Hiromitsu Nishizaki |
INTERSPEECH | 4 |
| 2012 | Designing an Evaluation Framework for Spoken Term Detection and Spoken Document Retrieval at the NTCIR-9 SpokenDoc Task
Tomoyosi Akiba, Hiromitsu Nishizaki, Kiyoaki Aikawa, Tatsuya Kawahara, Tomoko Matsui |
LREC | 2 |
| 2011 | Utterance verification using garbage words for a hospital appointment system with speech interfaceabstractOn a system that captures spoken dialog, users often use out-of-domain utterances to the system. The speech recognition component in the dialog system cannot correctly recognize such utterances, which causes fatal errors. This paper proposes a method to verify whether utterances are in-domain or out-of-domain. The proposed method trains systems with two language models: one that can accept both in-domain and out-of-domain utterances and the other that can accept only in-domain utterances. These models are installed into two speech recognition systems. A comparison of the recognizers' outputs provides a good verification of utterances. We installed our method in a hospital appointment system and evaluated it. The experimental results showed that the proposed method worked well. Mitsuru Takaoka, Hiromitsu Nishizaki, Yoshihiro Sekiguchi |
ASRU | 2 |
| 2010 | Constructing Japanese test collections for spoken term detectionabstractSpoken Document Retrieval (SDR) and Spoken Term Detection (STD) have been two of the most intensively investigated topics in spoken document processing research according to the establishment of the SDR and STD test collections by the Text REtrieval Conference (TREC) and NIST. Because Japanese spoken document processing researchers also requires such test collections for SDR and STD, we have established a working group to develop these collections in Special Interest Group-Spoken Language Processing (SIG-SLP) of the Information Processing Society of Japan. The working group has constructed and made available a test collection for SDR, and is now constructing new test collections for STD that will be open to researchers. The present paper introduces the policies, outline, and schedule of the new test collections. Then, the new test collections are compared with the NIST STD test collections. Index Terms: spoken term detection, test collection 1. Yoshiaki Itoh 0001, Hiromitsu Nishizaki, Xinhui Hu, Hiroaki Nanjo, Tomoyosi Akiba, Tatsuya Kawahara, Seiichi Nakagawa, Tomoko Matsui, Yoichi Yamashita, Kiyoaki Aikawa |
INTERSPEECH | 2 |
| 2010 | Japanese spoken term detection using syllable transition network derived from multiple speech recognizers' outputs
Satoshi Natori, Hiromitsu Nishizaki, Yoshihiro Sekiguchi |
INTERSPEECH | 2 |
| 2008 | Is a speech recognizer useful for characteristic analysis of classroom lecture speech?
Kenji Kobayashi, Mitsuhiro Somiya, Hiromitsu Nishizaki, Yoshihiro Sekiguchi |
INTERSPEECH | 3 |
| 2008 | Speech recognition performance of CJLC: corpus of Japanese lecture contentsabstractThis paper discusses the speech recognition of Japanese classroom lecture speech. In particular, we mention the influences of microphone differences and the language model differences on the speech recognition performance of classroom lectures. First, we collected actual classroom lecture contents from several universities in Japan. In this paper, we recorded the lecture speech using lapel microphones because lapel microphones are more commonly used to record lectures. LVCSR is one of the essential technologies for adding tag information to such lecture speech. Next, therefore, we researched the influence of the differences between microphones used for recording lecture on speech recognition performance. Finally, seven types of language models that were trained using three types of corpora were compared on the basis of their ability to lecture speech. Satoru Kogure, Hiromitsu Nishizaki, Masatoshi Tsuchiya, Kazumasa Yamamoto, Shingo Togashi, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2008 | Test Collections for Spoken Document Retrieval from Lecture Audio Data
Tomoyosi Akiba, Kiyoaki Aikawa, Yoshiaki Itoh 0001, Tatsuya Kawahara, Hiroaki Nanjo, Hiromitsu Nishizaki, Norihito Yasuda, Yoichi Yamashita, Katunobu Itou |
LREC | 6 |
| 2008 | Developing Corpus of Japanese Classroom Lecture Speech Contents
Masatoshi Tsuchiya, Satoru Kogure, Hiromitsu Nishizaki, Kengo Ohta, Seiichi Nakagawa |
LREC | 3 |
| 2007 | The effect of filled pauses in a lecture speech on impressive evaluation of listenersabstractThis paper examines and reports on how ”filled pauses” included when delivering speeches influence the understanding, and change the impression, of the speech as shown through the research and experiments we conducted on trial subjects. We conducted research about speeches and lectures given at classes at our university, and at academic meetings. A questionnaire related to filled pauses was given to audiences in university classrooms, and the speeches given where recorded. Then, we prepared a number of speeches that were manually altered to put emphasis on the frequency, position, and duration of filled pauses in the speeches. Comparing those speeches with the original speeches which were not processed in our listening experiments, we were able to estimate the effect of filled pauses in a lecture speech and how effective these were in altering the impressions of the audiences. We were able to find the best conditions related to the frequency, position, and duration of filled pauses, and how these conditions cleary changed a lecture or speech into a better one which is easy to understanding and listen to for the audience. Index Terms: Spoken language, Lecture speech, Evaluation of lectures, Filled pause Hiromitsu Nishizaki, Mitsuhiro Somiya, Kenji Kobayashi, Yoshihiro Sekiguchi |
INTERSPEECH | 1 |
| 2004 | Keyword recognition and extraction by multiple-LVCSRs with 60, 000 words in speech-driven WEB retrieval taskabstractThis paper presents speech-driven Web retrieval models which accepts spoken search topics (queries) in the NTCIR-3 Web retrieval task. We experimentally evaluate the techniques of combining outputs of multiple LVCSR models with a language model(LM) with a 60,000 vocabulary size in recognition of spoken queries. As model combination techniques, we use the SVM learning. We show that the techniques of multiple LVCSR model combination can achieve improvement both in speech recognition and retrieval accuracies in speech-driven text retrieval. Comparing with the retrieval accuracies when a LM with a 20,000/60,000 vocabulary size is used in LVCSRs, the LM that has larger size of the vocabulary improves also retrieval accuracies. Masahiko Matsushita, Hiromitsu Nishizaki, Seiichi Nakagawa, Takehito Utsuro |
INTERSPEECH | 2 |
| 2004 | Unsupervised speaker adaptation using high confidence portion recognition results by multiple recognition systems
Tomohiro Watanabe, Hiromitsu Nishizaki, Takehito Utsuro, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2003 | Confidence of agreement among multiple LVCSR models and model combination by SVMabstractFor many practical applications of speech recognition systems, it is quite desirable to have an estimate of confidence for each hypothesized word. Unlike previous works on confidence measures, we have proposed features for confidence measures that are extracted from outputs of more than one LVCSR models. For further analysis of the proposed confidence measure, this paper examines the correlation between each word's confidence and the word's features such as its part-of-speech and syllable length. We then apply SVM learning technique to the task of combining outputs of multiple LVCSR models, where, as features of SVM learning, information such as the pairs of the models which output the hypothesized word are useful for improving the word recognition rate. Experimental results show that the combination results achieve a relative word error reduction of up to 72 % against the best performing single model and that of up to 36 % against ROVER. Takehito Utsuro, Yasuhiro Kodama, Tomohiro Watanabe, Hiromitsu Nishizaki, Seiichi Nakagawa |
ICASSP (1) | 4 |
| 2003 | Evaluating multiple LVCSR model combination in NTCIR-3 speech-driven web retrieval taskabstractThis paper studies speech-driven Web retrieval models which accepts spoken search topics (queries) in the NTCIR-3 Web retrieval task. The major focus of this paper is on improving speech recognition accuracy of spoken queries and then improving retrieval accuracy in speech-driven Web retrieval. We experimentally evaluate the techniques of combining outputs of multiple LVCSR models in recognition of spoken queries. As model combination techniques, we compare the SVM learning technique and conventional voting schemes such as ROVER. We show that the techniques of multiple LVCSR model combination can achieve improvement both in speech recognition and retrieval accuracies in speech-driven text retrieval. We also show that model combination by SVM learning outperforms conventional voting schemes both in speech recognition and retrieval accuracies. Masahiko Matsushita, Hiromitsu Nishizaki, Takehito Utsuro, Yasuhiro Kodama, Seiichi Nakagawa |
INTERSPEECH | 2 |
| 2002 | Comparing isolately spoken keywords with spontaneously spoken queries for Japanese spoken document retrieval
Hiromitsu Nishizaki, Seiichi Nakagawa |
INTERSPEECH | 1 |
| 2002 | A confidence measure based on agreement among multiple LVCSR models - correlation between pair of acoustic models and confidenceabstractFor many practical applications of speech recognition systems, it is quite desirable to have an estimate of confidence for each hypothesized word. Unlike previous works on confidence measures, this paper studies features for confidence measures that are extracted from outputs of more than one LVCSR models. More specifically, this paper experimentally evaluates the agreement among the outputs of multiple Japanese LVCSR models, with respect to whether it is effective as an estimate of confidence for each hypothesized word. The results of experimental evaluation show that the agreement between the outputs with two LVCSR models with different decoders and acoustic models can achieve quite reliable confidence. Furthermore, among various features of acoustic models based on Gaussian mixture HMMs, it is concluded that ones such as whether or not to have short pause models, as well as different units in HMMs (e.g., triphone model or syllable model) are the most effective in achieving highly reliable confidence. Takehito Utsuro, Tetsuji Harada, Hiromitsu Nishizaki, Seiichi Nakagawa |
INTERSPEECH | 3 |
| 2001 | Experimental evaluation on confidence of agreement among multiple Japanese LVCSR modelsabstractFor many practical applications of speech recognition systems, it is quite desirable to have an estimate of confidence for each hypothesized word. Unlike previous works on confidence measures, this paper studies features for confidence measures that are extracted from outputs of more than one LVCSR models. More specifically, this paper experimentally evaluates the agreement among the outputs of multiple Japanese LVCSR models, with respect to whether it is effective as an estimate of confidence for each hypothesized word. The results of experimental evaluation show that the agreement between the outputs with two acoustic models which have different units in HMMs, such as phonemes and syllables, can achieve quite reliable confidence. 1. Yasuhiro Kodama, Takehito Utsuro, Hiromitsu Nishizaki, Seiichi Nakagawa |
INTERSPEECH | 3 |
| 2000 | A system for retrieving broadcast news speech documents using voice input keywords and similarity between words
Hiromitsu Nishizaki, Seiichi Nakagawa |
INTERSPEECH | 1 |