VLDB 2026 Research / reviewers in the wild / expert
Yusuke Shinohara
dblp:49/407
· DBLP profile ↗
42ranked-venue papers
10as first author
18since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 5 since 2021Computer networks · 8 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Self-Supervised Speech Models Via Text-Based LLMsabstractSelf-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyper-parameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task. Takashi Maekaku, Keita Goto, Jinchuan Tian, Yusuke Shinohara, Shinji Watanabe 0001 |
ASRU | 4 |
| 2025 | ViT-PQC: Vision Transformer-Based Patch Quality Controller for Edge-Assisted Visual-SLAMabstractEdge-based visual simultaneous localization and mapping (Visual-SLAM) enables real-time indoor localization of robots by using substantial computing resources of an edge server. However, this system can place a burden on resource-constrained wireless networks because robots transmit their video data to the edge server via wireless networks. Given the sharing of the limited network capacity among multiple devices, the video bitrate needs to be reduced so as not to overshoot the available bandwidth. Although lowering the video quality can reduce the video bitrate, it also results in a loss of detailed image features and can negatively impact localization accuracy. In this paper, we propose a Vision Transformer-based patch quality controller (ViT-PQC) that reduces the video bitrate below the available bandwidth while preserving localization accuracy. ViT-PQC segments the video frame into patches and assigns an optimal image quality to each patch by taking into account both the image features of the recent frame and time shift features between the two most recent frames. Evaluation results demonstrate that ViT-PQC reduces the number of bitrate overshoots by 94% while preserving localization accuracy. Yuma Katsuki, Hayato Itsumi, Yusuke Shinohara, Koichi Nihei, Anan Sawabe, Takanori Iwai |
CCNC | 3 |
| 2025 | Wireless Multi-Connectivity Management with Packet-level Delay Gradient AnalysisabstractWireless multi-connectivity solutions are essential for reliable low-latency wireless communication services enabling delay-sensitive applications such as industrial robotics. However, non-expert industry vertical players in networking seek low-installation-cost and high-quality solutions to use wireless multi-connectivity. This paper proposes a wireless multi-connectivity management gateway (WMC-GW) to effectively utilize multiple wireless networks without modifying applications and wireless network systems. We install a WMC-GW at each mobile robot and another WMC-GW at the application server. There are two features. The first is real-time radio access technology (RAT) selection, where the WMC-GW select an appropriate RAT based on predicted delay trends using IP packet-level delay gradients to follow sensitive delay variations. The second is the flexibility of flow-level policy control, where we develop multiple RAT selection policies based on the delay gradients. Through performance evaluation, our approach effectively works in low-latency RAT selection. Anan Sawabe, Yusuke Shinohara, Yuma Katsuki, Takanori Iwai |
CCNC | 2 |
| 2025 | Traffic Pattern Re-Arrangement by Hierarchical Calendar Queueing-Based Packet SchedulerabstractCommunication traffic patterns, representing network and application behavior, are beneficial for the transport and application layers to estimate the states of black-box network systems. Masking noisy traffic features helps high-quality communication by reducing the misestimation of network states due to unstable delay behavior for delay-sensitive applications. Motivated by the effects, we propose TrafficArranger, a traffic pattern re-arrangement system using packet scheduling to provide a target quality of service (QoS), including throughput, delay, and packet interval, as required by each flow. The main component of TrafficArranger is Hierarchical Calendar Queueing (HCQ) for delay-based packet scheduling. We introduce a number of techniques: (i) stepped dequeue to control the dequeue level for flexible jitter control for improving robustness to traffic load while keeping control accuracy and (ii) rush dequeue and reenqueue skipping to address out-of-order issues in HCQ. Through performance comparison with some leaves of Linux qdisc and state-of-the-art HCQ method, i.e., Gearbox, only TrafficArranger controls QoS as required for all QoS classes (i.e., throughput, delay, and packet interval). Also, we demonstrate that TrafficArranger effectively works for performance stabilization in a cooperative adaptive cruise control system. Anan Sawabe, Yusuke Shinohara, Yuma Katsuki, Takanori Iwai |
ICC | 2 |
| 2025 | Detection and Mitigation of False Data Injection Attacks for MEC-based Leader-Follower CACCabstractCountermeasures against false data injection (FDI) attacks are necessary for developing cooperative adaptive cruise control (CACC) systems. Fixed redundant path selection (FRPS) based on majority voting has been used to detect and mitigate FDI attacks in leader–follower CACC systems. However, the applications of FRPS are limited to distributed CACC architectures. This study proposes the application of FRPS to a centralized CACC architecture based on multiaccess edge computing (MEC). Simulations confirm that the proposed method achieves stable and safe MEC-based leader–follower CACC while receiving FDI attacks. Naoya Sato, Yuma Katsuki, Anan Sawabe, Yusuke Shinohara, Ryogo Kubo |
ICCCN | 5 |
| 2025 | OpusLM: A Family of Open Unified Speech Language Models
Jinchuan Tian, Yifan Peng 0003, Jiatong Shi, Siddhant Arora, Shikhar Bharadwaj, Takashi Maekaku, Yusuke Shinohara, Keita Goto, Xiang Yue, Chao-Han Huck Yang, Shinji Watanabe 0001 |
INTERSPEECH | 8 |
| 2024 | Congestion State Estimation via Packet-Level RTT Gradient Analysis with Gradual RTT SmoothingabstractAccurately estimating congestion states of the bottleneck link in mobile networks (e.g., LTE and 5G) from round-trip times (RTTs) is an important task for congestion control algorithms (CCAs). In mobile networks, RTTs fluctuate easily because they are sensitive to stochastic behavior, e.g., congestion and radio quality, and deterministic behavior, e.g., scheduling, making it difficult to estimate the congestion state accurately. In this paper, we propose an RTT-gradient analysis method for accurately estimating the congestion state (i.e., overuse/normal/underuse) by filtering stochastic noise while reducing misestimation caused by deterministic delay behavior in mobile networks. The proposed method consists of two features. The first is gradual RTT smoothing, a two-stage Kalman filter designed in series for filtering noisy RTT variations. The second is adaptive threshold clipping, which eliminates estimator instability caused by deterministic delay variations in the mobile network. Our experimental results in a commercial 5G network show that our method is more robust than a conventional method (the state estimator of Google Congestion Control (GCC)). Furthermore, simulation results using an open dataset for mobile networks show that our method improves throughput by 1.6% on average for 13 scenarios out of 16 scenarios total from GCC. Anan Sawabe, Yusuke Shinohara, Takanori Iwai |
CCNC | 2 |
| 2024 | Revisiting TCP Pacing for Throughput Performance Enhancement Over TDD Band in Private Mobile NetworksabstractPrivate mobile networks, such as local 5G, have attracted the attention of industry players who expect flexible radio resource allocation by methods such as time-division duplex (TDD) scheduling based on the uplink and downlink traffic demand of their solutions. However, there are two challenges when communicating using TCP congestion control algorithms (CCAs) over the TDD link: TDD-induced ACK-waiting time and misestimating congestion states due to deterministic delay variation caused by TDD scheduling. In this paper, we propose a TDD-aware TCP pacing method for improving TCP throughput by pacing the sending time between two consecutive segments within the ACK-waiting time. We determine the pacing rate on the basis of the TDD-induced delay variation for sending TCP segments within allocated TDD slots while reducing round-trip time (RTT). We evaluate the performance of our method by using a network simulator (ns-3). TCP pacing improves throughput by about 10–70% compared with when there is no pacing, especially for TCP Illinois. We also verify that our TDD-aware pacing improves throughput by about 10% compared to the default pacing rate on the Linux kernel. Anan Sawabe, Yusuke Shinohara, Takanori Iwai |
CCNC | 2 |
| 2024 | Rethinking Delay Behavior in Mobile Networks as a Lifeline of Industrial ApplicationsabstractUnderstanding packet-level communication delay behavior in mobile networks is becoming increasingly important with the rise of delay-sensitive applications such as industrial mobile robots. Although attention has traditionally focused on queueing delays due to congestion, in mobile networks, variable communication delays due to multiple delay factors degrade the performance of delay-sensitive applications that send packets with a high packet rate. This paper contributes to identifying the relationship between packet rate and delay components. We first categorize delays in end-to-end communication into four categories: transmission, propagation, queueing, and processing delays. We then formulate the relationship between the packet transmission interval and the delay factors, which shows an interesting trend that the ratio of delay components differs with transmission interval time. Also, when the packet transmission interval is short, the impact of deterministic delay jitter due to processing delay is more significant than that of stochastic queueing delay. We examine the trend of delay component ratio through experiments in an operational 5G network in Japan. Anan Sawabe, Yusuke Shinohara, Takanori Iwai |
GLOBECOM | 2 |
| 2023 | Domain Adaptation by Data Distribution Matching Via Submodularity For Speech RecognitionabstractWe study the problem of building a domain-specific speech recognition model given some text from the target domain. One of the most popular approaches to this problem is shallow fusion, which incorporates a domain-specific language model build from the given text. However, shallow fusion significantly increases the model size and inference cost, which makes its deployment harder. In this paper, we propose domain adaptation by data distribution matching, where a subset is selected from an existing multi-domain training data to match the target-domain distribution, and a model is fine-tuned on the subset. A submodular optimization algorithm with a novel extension is employed for the subset selection. Experiments on LibriSpeech, a corpus of audiobooks, where we treat each book as a domain, show that the proposed distribution-matching approach achieves WERs equivalent with the conventional shallow-fusion approach, without any increase in the model size and inference cost. Yusuke Shinohara, Shinji Watanabe 0001 |
ASRU | 1 |
| 2023 | Recognition-aware Bitrate Allocation for AI-Enabled Remote Video SurveillanceabstractWith a growing number of deployed video surveillance cameras, deep learning based video recognition becomes increasingly important for analysing the large amounts of generated video content. For use-cases with computationally intensive recognition applications and with moving cameras or frequently changing camera layouts, video content is often streamed in real-time over wireless networks to cloud environments where video recognition technologies operate. However, it is in general difficult to transmit a large number of video streams and also achieve a high video recognition accuracy with only the limited radio frequency resources of wireless networks, because of video compression artefacts appearing at low video bitrates. To allow transmission of more streams and to increase recognition accuracy, we propose a dynamic prediction-based bitrate allocation method that balances encoder bitrates among cameras in a way that approximately minimizes the total recognition error. For predicting the number of recognition errors at given bitrate and video content, our method uses a neural network operating in a distributed fashion at both edge and cloud locations. In the experiments, we show that our predictions are highly correlated with the actual numbers of recognition errors (Pearson correlation of 0.872). Furthermore, an evaluation of our allocation approach in a construction site surveillance scenario with ten video streams shows that at low average video stream bitrates under 1 Mbps, our method can reduce the false negative rate of an action recognition engine by 23% over the baseline approach (evenly distributed allocation) on average. Florian Beye, Yusuke Shinohara, Hayato Itsumi, Koichi Nihei |
CCNC | 2 |
| 2023 | Multiple Cars Remote Monitoring System using AI-based Video Streaming and AlertabstractAutonomous driving is attracting attention. Though related technologies are evolving, it is forecasted that Level 5 autonomous driving (full driving automation) would be ordinary in the 2040s to 2060s. Until the epoch, remote human support is necessary to deal with the cases in which an autonomous car cannot decide by itself. For the remote human support, live video streaming and user interface for effective monitoring are significant challenges. This paper proposes a live video streaming method that minimizes video bitrate while keeping the required video quality (video analytics accuracy) and AI-assisted monitoring GUI (graphical user interface) to monitor multiple cars effectively. First, the authors evaluated the GUI through simulation. The result shows that it improves efficiency and reduces mental and physical loads. Then, the authors conducted field tests in which an operator monitored two cars. The operators said that the proposed system is much better than the conventional systems they usually use and can monitor and control at least two cars simultaneously employing the system. Koichi Nihei, Hayato Itsumi, Yusuke Shinohara, Tomonao Araki, Takanori Iwai |
VTC2023-Spring | 3 |
| 2022 | Delay Jitter Modeling for Low-Latency Wireless Communications in Mobility ScenariosabstractUnderstanding the delay jitter of mobile communications becomes important because of widely spreading delay-sensitive applications such as remote control of mobile robots with high-frequency communications via wireless networks. Prior studies on delay jitter modeling have proposed using a single probability distribution (e.g., Gamma and Laplace distributions). However, mobility-induced wireless quality fluctuations form a mixture of probability patterns, e.g., several peaks and a heavy tail. This paper proposes a method to estimate delay jitter accurately in high-frequency and mobile communications. Our method has two features. The first is to model the delay jitter by a mixture of multiple Laplace distributions by taking into account the probability patterns. For quick convergence of model training, the model is trained with access manner-aware initialization in each Wi-Fi and mobile network. The second is to construct a likelihood-based observation segmentation for estimating model parameters accurately against mobility. Performance evaluation through experiments in indoor Wi-Fi and outdoor 5G scenarios shows that our proposed method improves modeling accuracy by 28.7% compared with the case of the prior studies. Anan Sawabe, Yusuke Shinohara, Takanori Iwai |
GLOBECOM | 2 |
| 2022 | Minimum latency training of sequence transducers for streaming end-to-end speech recognition
Yusuke Shinohara, Shinji Watanabe 0001 |
INTERSPEECH | 1 |
| 2021 | A Study of Transducer Based End-to-End ASR with ESPnet: Architecture, Auxiliary Loss and Decoding StrategiesabstractIn this study, we present recent developments of models trained with the RNN-T loss in ESPnet. It involves the use of various archi-tectures such as recently proposed Conformer, multi-task learning with different auxiliary criteria and multiple decoding strategies, in-cluding our own proposition. Through experiments and benchmarks, we show that our proposed systems can be competitive against other state-of-art systems on well-known datasets such as LibriSpeech and AISHELL-1. Additionally, we demonstrate that these models are promising against other already implemented systems in ESPnet in regards to both performance and decoding speed, enabling the pos-sibility to have powerful systems for a streaming task. With these additions, we hope to expand the usefulness of the ESPnet toolkit for the research community and also give tools for the ASR industry to deploy our systems in realistic and production environments. Florian Boyer, Yusuke Shinohara, Takaaki Ishii, Hirofumi Inaguma, Shinji Watanabe 0001 |
ASRU | 2 |
| 2021 | DCM: Delay as Component Model based on Hidden Striping Structure in Mobile NetworksabstractUnderstanding communication delay in mobile networks is becoming more important as delay-sensitive scenarios become more prevalent. Round-trip time measurement, the conventional technique to measure communication delay, e.g., ping, outputs network-induced delay for each packet but is insufficient in identifying specific delay factors. We propose a model called Delay Component Model (DCM) to aid in clearly visualizing communication delay. We construct the DCM on the basis of our measurement study on packet-receipt intervals with packet transmission at a constant interval via commercial mobile networks in Japan. We find that the measured receipt-time intervals form striped patterns due to the combination of two components: a constant scheduler (ConstSched) and probability scheduler (ProbSched). We use the principle of forming striped patterns and develop a method of estimating the DCM structure. Finally, we evaluate our method by analyzing delay patterns measured in Long Term Evolution (LTE) and fifth-generation mobile (5G) networks. The results indicate that delay patterns in these networks are due to the combination of three components, i.e., a ConstSched with 20-ms intervals, ProbSched with 8-ms delay for LTE and 6-ms delay for 5G, and ProbSched with 1-ms delay. Anan Sawabe, Shinya Yasuda, Yusuke Shinohara, Takanori Iwai, Akihiro Nakao |
GLOBECOM | 3 |
| 2021 | Simpleflat: A Simple Whole-Network Pre-Training Approach for RNN Transducer-Based End-to-End Speech RecognitionabstractRecurrent neural network-transducer (RNN-T) is promising for building time-synchronous end-to-end automatic speech recognition (ASR) systems, in part because it does not need frame-wise alignment between input features and target labels in the training step. Although training without alignment is beneficial, it makes it difficult to discern the relation between input features and output token sequences. This, in effect, degrades RNN-T performance. Our solution is SimpleFlat (SF), a novel and simple whole-network pretraining approach for RNN-T. SF extracts frame-wise alignments on-the-fly from the training dataset, and does not require any external resources. We distribute equal numbers of target tokens to each frame following RNN-T encoder output lengths by repeating each token. The frame-wise tokens so created are shifted, and also used as the prediction network inputs. Therefore, SF can be implemented by cross entropy loss computation as in autoregressive model training. Experiments on Japanese and English ASR tasks demonstrate that SF can effectively improve various RNN-T architectures. Takafumi Moriya, Takanori Ashihara, Tomohiro Tanaka, Tsubasa Ochiai, Hiroshi Sato 0002, Atsushi Ando, Yusuke Ijima, Ryo Masumura, Yusuke Shinohara |
ICASSP | 9 |
| 2021 | Edge Cloud Ensemble with Motion Vectors for Object Detection in Wireless EnvironmentsabstractCloud offloading enables smaller edge devices which contain fewer computing resources to be used for computer vision. This contributes to mobile applications of computer vision such as robotics and augmented reality. However, cloud offloading is impacted by bandwidth fluctuations in wireless networks. When the available bandwidth is restricted, it is difficult to offload workloads (i.e. frame) to cloud instances. Instead of offloading workloads to the cloud, existing approaches send a result of edge prediction or use a past cloud prediction to cover non-offloaded frames, which result in low accuracy. In this paper, we propose an Edge Cloud Ensemble method for object detection to improve accuracy in low bandwidth environments. In our method, the edge sends an inaccurate prediction and a motion vector of a detected object’s region to the cloud while maintaining low transmission overhead. The cloud corrects the inaccurate prediction by using the motion vectors to shift a past, accurate cloud prediction. The results of our experiments demonstrate that our approach can improve accuracy in low bandwidth environments compared with existing methods, especially in moving cameras. Hayato Itsumi, Florian Beye, Yusuke Shinohara, Charvi Vitthal, Takanori Iwai |
ICC | 3 |
| 2020 | Sequence-Level Consistency Training for Semi-Supervised End-to-End Automatic Speech RecognitionabstractThis paper presents a novel semi-supervised end-to-end automatic speech recognition (ASR) method that employs consistency training with the use of unlabeled data. In consistency training, unlabeled data can be utilized for constraining a model such that it becomes invariant to small deformation. In fact, considering consistency can make the model robust to a variety of input examples. While previous studies have applied consistency training to primitive classification problems, no studies have employed consistency training to tackle sequence-to-sequence generation problems including end-to- end ASR. One problem is that existing consistency training schemes cannot take sequence-level generation consistency into consideration. In this paper, we propose a sequence-level consistency training scheme specialized to handle sequence-to-sequence generation problems. Our key idea is to consider the consistency of the generation function by utilizing beam search decoding results. For semi- supervised learning, we adopt Transformer as the end-to-end ASR model, and SpecAugment as the deformation function in consistency training. Our experiments show that our semi-supervised learning proposal with sequence-level consistency training can efficiently improve ASR performance using unlabeled speech data. Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Atsushi Ando, Yusuke Shinohara |
ICASSP | 6 |
| 2020 | Distilling Attention Weights for CTC-Based ASR SystemsabstractWe present a novel training approach for connectionist temporal classification (CTC) -based automatic speech recognition (ASR) systems. CTC models are promising for building both a conventional acoustic model and an end-to-end (E2E) ASR model. However, CTC models make it difficult to capture the correct timing of each output label because timing is not given explicitly in the training data. In this paper, we propose a new auxiliary task with frame-wise targets for CTC model enhancement. We utilize attention weights generated by an attention-based encoder-decoder model (S2S) for making the targets, called the attention matrix. The attention matrix is the sum of the products of the attention weights (spike timing information) and the corresponding target vectors (probability information), and used for S2S-to-CTC knowledge distillation loss computation. Therefore, the attention matrix makes the CTC models jointly train-able as regards spike timings and their posteriors. Experiments on Japanese ASR tasks demonstrate that our proposal is effective for CTC model training; it achieves a 10.2% (E2E) / 9.4% (acoustic model) relative reduction in the character/kana-syllable error rates compared to models trained using only CTC loss. Takafumi Moriya, Hiroshi Sato 0002, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara |
ICASSP | 6 |
| 2020 | Training With Cache: Specializing Object Detectors From Live Streams Without OverfittingabstractOnline distillation can dynamically adapt to changes of distribution in a target domain by continuously updating a smaller student model from a live video stream. This ensures the accuracy of the student model even if distribution changes occur due to a change of content. However, online distillation degrades the overall accuracy because it causes overfitting to “current” distribution not to “recent” distribution. The student model is trained on sequential incoming data, and its model parameters are overwritten with the “current” distribution. As a result, the student model forgets the “recent” distribution. To overcome this problem, we propose a new training framework using cache. Our framework temporarily stores incoming frames and teacher model outputs in a cache and trains a student model with data selected from the cache. Since our approach trains the student model with not only incoming data but also past data, it can improve the overall accuracy while adapting to changes of distribution without overfitting. To use limited cache size efficiently, we also propose a loss-aware cache algorithm that chooses training data prioritized by its loss value. Our experiments show that training with cache improves the accuracy compared with online distillation, and the loss-aware cache algorithm outperforms a cache algorithm modeled on traditional offline training. Hayato Itsumi, Florian Beye, Yusuke Shinohara, Takanori Iwai |
ICIP | 3 |
| 2020 | Self-Distillation for Improving CTC-Transformer-Based ASR Systems
Takafumi Moriya, Tsubasa Ochiai, Shigeki Karita, Hiroshi Sato 0002, Tomohiro Tanaka, Takanori Ashihara, Ryo Masumura, Yusuke Shinohara, Marc Delcroix |
INTERSPEECH | 8 |
| 2020 | Video Compression Estimating Recognition Accuracy for Remote Site Object DetectionabstractCurrent video compression algorithms are designed to achieve a smaller data size and enable higher human perception for the real time streaming of video applications. However, with the recent explosive technical progress of deep neural network (DNNs), video surveillance is increasingly being performed not by humans but by computer vision systems. In this work, we propose a video compression method for object detection by computer vision algorithms. Our method detects the ROI in an image and differentiates the image quality between the ROI and other areas. It also optimizes the image quality in the ROI by estimating the recognition accuracy of the object detection model. The results of an experimental evaluation demonstrate that our proposed method can achieve high-quality encoding in terms of data size and successfully estimate the recognition accuracy of the object detection model. Yusuke Shinohara, Hayato Itsumi, Florian Beye, Takanori Iwai |
IWCMC | 1 |
| 2019 | Towards Accurate and Scalable Performance Prediction for Automated Service Design in NFVabstractAutomatizing the process of designing communication services in network function virtualization (NFV) is important because it may reduce provisioning time and lead to more efficient designs. The design process involves solving performance constraints imposed by service level agreements (SLAs), which in turn requires accurate and fast performance prediction. However, effects such as resource contention make performance prediction in virtualized environments challenging when large numbers of possible combinations of software and hardware are considered. The key to scalability lies in finding a componentized approach that reduces the number of model degrees of freedom while still allowing high accuracy. In this work, we propose a componentized approach based on feed-forward networks that are composited from software and hardware models. Model parameter data is obtained from a machine learning technique which is fed using data generated from automatized offline performance measurements. An evaluation showed that our technology achieves a prediction accuracy close to 95% and prediction evaluation times of a few milliseconds. Florian Beye, Yusuke Shinohara, Hideyuki Shimonishi |
CCNC | 2 |
| 2019 | Large Context End-to-end Automatic Speech Recognition via Extension of Hierarchical Recurrent Encoder-decoder ModelsabstractThis paper describes a novel end-to-end automatic speech recognition (ASR) method that takes into consideration long-range sequential context information beyond utterance boundaries. In spontaneous ASR tasks such as those for discourses and conversations, the input speech often comprises a series of utterances. Accordingly, the relationships between the utterances should be leveraged for transcribing the individual utterances. While most previous end-to-end ASR methods only focus on utterance-level ASR that handles single utterances independently, the proposed method (which we call "large-context end-to-end ASR") can explicitly utilize relationships between a current target utterance and all preceding utterances. The method is modeled by combining an attention-based encoder-decoder model, which is one of the most representative end-to-end ASR models, with hierarchical recurrent encoder-decoder models, which are effective language models for capturing long-range sequential contexts beyond the utterance boundaries. Experiments on Japanese discourse speech tasks demonstrate the proposed method yields significant ASR performance improvements compared with the conventional utterance-level end-to-end ASR system. Ryo Masumura, Tomohiro Tanaka, Takafumi Moriya, Yusuke Shinohara, Takanobu Oba, Yushi Aono |
ICASSP | 4 |
| 2019 | Neural Whispered Speech Detection with Imbalanced Learning
Takanori Ashihara, Yusuke Shinohara, Hiroshi Sato 0002, Takafumi Moriya, Kiyoaki Matsui, Takaaki Fukutomi, Yoshikazu Yamaguchi, Yushi Aono |
INTERSPEECH | 2 |
| 2019 | Joint Maximization Decoder with Neural Converters for Fully Neural Network-Based Japanese Speech Recognition
Takafumi Moriya, Tomohiro Tanaka, Ryo Masumura, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono |
INTERSPEECH | 5 |
| 2018 | Adversarial Training for Multi-task and Multi-lingual Joint Modeling of Utterance Intent ClassificationabstractThis paper proposes an adversarial training method for the multi-task and multi-lingual joint modeling needed for utterance intent classification.In joint modeling, common knowledge can be efficiently utilized among multiple tasks or multiple languages.This is achieved by introducing both languagespecific networks shared among different tasks and task-specific networks shared among different languages.However, the shared networks are often specialized in majority tasks or languages, so performance degradation must be expected for some minor data sets.In order to improve the invariance of shared networks, the proposed method introduces both language-specific task adversarial networks and task-specific language adversarial networks; both are leveraged for purging the task or language dependencies of the shared networks.The effectiveness of the adversarial training proposal is demonstrated using Japanese and English data sets for three different utterance intent classification tasks. Ryo Masumura, Yusuke Shinohara, Ryuichiro Higashinaka, Yushi Aono |
EMNLP | 2 |
| 2018 | Multi-task Learning with Augmentation Strategy for Acoustic-to-word Attention-based Encoder-decoder Speech Recognition
Takafumi Moriya, Sei Ueno, Yusuke Shinohara, Marc Delcroix, Yoshikazu Yamaguchi, Yushi Aono |
INTERSPEECH | 3 |
| 2018 | Encoder Transfer for Attention-based Acoustic-to-word Speech Recognition
Sei Ueno, Takafumi Moriya, Masato Mimura, Shinsuke Sakai, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono, Tatsuya Kawahara |
INTERSPEECH | 5 |
| 2018 | Automatic DNN Node Pruning Using Mixture Distribution-based Group Regularization
Tsukasa Yoshida, Takafumi Moriya, Kazuho Watanabe, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono |
INTERSPEECH | 4 |
| 2018 | Efficient Building Strategy with Knowledge Distillation for Small-Footprint Acoustic ModelsabstractIn this paper, we propose a novel training strategy for deep neural network (DNN) based small-footprint acoustic models. The accuracy of DNN-based automatic speech recognition (ASR) systems can be greatly improved by leveraging large amounts of data to improve the level of expression. DNNs use many parameters to enhance recognition performance. Unfortunately, resource-constrained local devices are unable to run complex DNN-based ASR systems. For building compact acoustic models, the knowledge distillation (KD) approach is often used. KD uses a large, well-trained model that outputs target labels to train a compact model. However, the standard KD cannot fully utilize the large model outputs to train compact models because the soft logits provide only rough information. We assume that the large model must give more useful hints to the compact model. We propose an advanced KD that uses mean squared error to minimize the discrepancies between the final hidden layer outputs. We evaluate our proposal on recorded speech data sets assuming car-and home-use scenarios, and show that our models achieve lower character error rates than the conventional KD approach or from-scratch training on computation resource-constrained devices. Takafumi Moriya, Hiroki Kanagawa, Kiyoaki Matsui, Takaaki Fukutomi, Yusuke Shinohara, Yoshikazu Yamaguchi, Manabu Okamoto, Yushi Aono |
SLT | 5 |
| 2016 | Adversarial Multi-Task Learning of Deep Neural Networks for Robust Speech Recognition
Yusuke Shinohara |
INTERSPEECH | 1 |
| 2014 | A submodular optimization approach to sentence set selectionabstractA new method for selecting a sentence set with a desired phoneme distribution is presented. Selection of a sentence set for speech corpus recording is a fundamental step in speech processing research. The problem of designing phonetically-balanced sentence sets has been studied extensively in the past. One of the popular approaches is to select a sentence set so that its phoneme distribution gets close to a given (desired) distribution. Several methods have been proposed in the literature to realize this approach. However, these methods were designed by heuristics, which means they are not optimal. In this paper, we propose a near-optimal method for selecting sentence sets along this approach. We first define our objective function, and show it to be a submodular function. Then, we show that a greedy algorithm is near-optimal for this problem, according to the submodular optimization theory. We also show that a significant speedup is possible by exploiting the submodularity of the objective function. Our experimental result on Japanese phonetically-balanced sentence set selection shows the effectiveness of the proposed method. Yusuke Shinohara |
ICASSP | 1 |
| 2013 | Tying rotations of covariance matrices via riemannian subspace clusteringabstractThe use of full covariance matrices in acoustic modeling is getting popular, but its huge computational burden in likelihood calculation is a major issue. Semi-tied covariance matrices are commonly used to speed-up the computation, where global or phone-based tying of transforms, or “rotations”, is usually used. However, such tyings are heuristic, and not necessarily optimal. In this paper, we propose a Riemannian-geometric approach to optimally tying rotations of covariance matrices. We first introduce a tangent space of the Riemannian manifold of covariance matrices, which has an excellent distance for measuring dissimilarity between covariance matrices. We then show that covariance matrices having the same rotation to each other lie on the same subspace in the tangent space. Exploiting this property, we fit subspaces to samples (covariances) in the tangent space for finding out clusters of samples that have similar rotations, and tie them together. By doing so, an optimal tying that minimizes the sum of “distortions” of covariance matrices can be found. Experimental results on the Wall Street Journal corpus show a superior performance of the proposed tying over the conventional ones. Yusuke Shinohara |
ICASSP | 1 |
| 2013 | N-best rescoring by phoneme classifiers using subclass adaboost algorithm
Hiroshi Fujimura, Yusuke Shinohara, Takashi Masuko |
INTERSPEECH | 2 |
| 2011 | N-Best rescoring by adaboost phoneme classifiers for isolated word recognitionabstractThis paper proposes a novel technique to exploit generative and discriminative models for speech recognition. Speech recognition using discriminative models has attracted much attention in the past decade. In particular, a rescoring framework using discriminative word classifiers with generative-model-based features was shown to be effective in small-vocabulary tasks. However, a straightforward application of the framework to large-vocabulary tasks is difficult because the number of classifiers increases in proportion to the number of word pairs. We extend this framework to exploit generative and discriminative models in large-vocabulary tasks. N-best hypotheses obtained in the first pass are rescored using AdaBoost phoneme classifiers, where generative-model-based features, i.e. difference-of-likelihood features in particular, are used for the classifiers. Special care is taken to use context-dependent hidden Markov models (CDHMMs) as generative models, since most of the state-of-the-art speech recognizers use CDHMMs. Experimental results show that the proposed method reduces word errors by 32.68% relatively in a one-million-vocabulary isolated word recognition task. Hiroshi Fujimura, Masanobu Nakamura, Yusuke Shinohara, Takashi Masuko |
ASRU | 3 |
| 2010 | Covariance clustering on Riemannian manifolds for acoustic model compressionabstractA new method of covariance clustering for acoustic model compression is proposed. Since covariance matrices do not form a Euclidean vector space, standard vector clustering algorithms cannot be used effectively for covariance clustering. In this paper, we propose a novel clustering algorithm based on a Riemannian framework, where the covariance space is considered as a Riemannian manifold equipped with the Fisher information metric, and notions of distance and mean are defined on the manifold. The LBG clustering algorithm is naturally extended to the covariance space under the Riemannian framework. Experimental results show the effectiveness of the proposed method, reducing the acoustic model size nearly to the half without noticeable loss in recognition performance. Yusuke Shinohara, Takashi Masuko, Masami Akamine |
ICASSP | 1 |
| 2010 | Source flow: handling millions of flows on flow-based nodesabstractFlow-based networks such as OpenFlow-based networks have difficulty handling a large number of flows in a node due to the capacity limitation of search engine devices such as ternary content-addressable memory (TCAM). One typical solution of this problem would be to use MPLS-like tunneling, but this approach spoils the advantage of flow-by-flow path selection for load-balancing or QoS. We demonstrate a method named "Source Flow" that allows us to handle a huge amount of flows without changing the granularity of flows. By using our method, expensive and power consuming search engine devices can be removed from the core nodes, and the network can grow pretty scalable. In our demo, we construct a small network that consists of small number of OpenFlow switches, a single OpenFlow controller, and end-hosts. The hosts generate more than one million flows simultaneously and the flows are controlled on a per-flow-basis. All active flows are monitored and visualized on a user interface and the user interface allows audiences to confirm if our method is feasible and deployable. Yasunobu Chiba, Yusuke Shinohara, Hideyuki Shimonishi |
SIGCOMM | 2 |
| 2009 | Bayesian feature enhancement using a mixture of unscented transformation for uncertainty decoding of noisy speechabstractA new parameter estimation method for the model-Based feature enhancement (MBFE) is presented. The conventional MBFE uses the vector Taylor series to calculate the parameters of non-linearly transformed distributions, though the linearization leads to a degraded performance. We use the unscented transformation to estimate the parameters, where a minimal number of samples propagated through the nonlinear transformation are used. By avoiding the linearization, the parameters are estimated more accurately. Experimental results on Aurora2 show that the proposed method reduces the word error rate by 8.48% relatively, while the computational cost is just modestly higher, compared with the conventional MBFE. Yusuke Shinohara, Masami Akamine |
ICASSP | 1 |
| 2009 | Programmable and Scalable Per-Flow Traffic Management Scheme Using a Control ServerabstractWe propose a programmable and scalable traffic management scheme. Programmable traffic management at high-speed routers is difficult because programmability and high-speed packet processing have involved a serious tradeoff. To attain both, the new scheme combines control programs at a control server and simple packet handling functions, such as sampling packet headers and discarding packets, at routers. Therefore, by installing appropriate control programs into the server, a variety of active queue management schemes, per-flow bandwidth management schemes, DoS mitigation schemes, and so on, are achieved. One of the main contributions of this paper is its proposal of a statistical scheme for handling flows. As only a fraction of complete flow information stored at the control server is loaded into the router's flow table and it is replaced cyclically, the proposed scheme scales more than the router's flow table capacity. Our simulation results indicate that the scheme provides efficient traffic management, per-flow WFQ emulation in our example, even with very small flow tables compared to the number of concurrently active flows. Furthermore, we discuss implementation issues with the proposed scheme and reveal that the processing cost at the server and router is sufficiently small for use with 10 Gbps links. Yusuke Shinohara, Hideyuki Shimonishi, Hideki Tode, Koso Murakami |
ICC | 1 |
| 2008 | Feature enhancement by speaker-normalized splice for robust speech recognitionabstractThe SPLICE method of feature enhancement is known for its powerful performance. It learns a mapping from noisy to clean feature vectors given a set of stereo training data. However, feature vector variation caused by speaker changes conceals noise-induced variation, which is what we want to find in the SPLICE training. In this paper, an improvement of SPLICE by means of speaker-normalization is proposed. The training data is first normalized with respect to speaker variation, and a mapping is learned afterward. CMLLR with a GMM as its target is utilized for the speaker-normalization, where the GMM representing a standard speaker is learned via a novel variant of the speaker adaptive training. The proposed method was evaluated on Aurora2, and achieved a relative word error rate reduction of 38% over the conventional SPLICE. Yusuke Shinohara, Takashi Masuko, Masami Akamine |
ICASSP | 1 |