VLDB 2026 Research / reviewers in the wild / expert
Kazuhito Koishida
dblp:88/3464
· DBLP profile ↗
42ranked-venue papers
7as first author
19since 2021 · last 2025
0000-0002-3111-5375ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 25 · 3 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and CompressionabstractStructured pruning and quantization are fundamental techniques used to reduce the size of deep neural networks (DNNs), and typically are applied independently. Applying these techniques jointly via co-optimization has the potential to produce smaller, high-quality models. However, existing joint schemes are not widely used because of (1) engineering difficulties (complicated multi-stage processes), (2) black-box optimization (extensive hyperparameter tuning to control the overall compression), and (3) insufficient architecture generalization. To address these limitations, we present the framework GETA, which automatically and efficiently performs joint structured pruning and quantization- aware training on any DNN. GETA introduces three key innovations: (i) a quantization-aware dependency graph (QADG) that constructs a pruning search space for generic quantization-aware DNN, (ii) a partially projected stochastic gradient method that guarantees layerwise bit constraints are satisfied, and (iii) a new joint learning strategy that incorporates interpretable relationships between pruning and quantization. We present numerical experiments on both convolutional neural networks and transformer architectures that show that our approach achieves competitive (often superior) performance compared to existing joint pruning and quantization methods. Source code is available at https://github.com/microsoft/GETA. Xiaoyi Qu, David Aponte, Colby R. Banbury, Daniel P. Robinson, Tianyu Ding, Kazuhito Koishida, Ilya Zharkov |
CVPR | 6 |
| 2025 | CorrGAN: Simultaneous Learning of Speech Enhancement and Perceptual Quality Loss FunctionsabstractDeep-learning models have allowed effective end-to-end SE systems in the Speech Enhancement (SE) field. Most of these methods are trained using a fixed reconstruction loss in a supervised setting. Often these losses do not perfectly represent the desired perceptual quality metrics, resulting in sub-optimal performance. Recently, there have been efforts to learn the behavior of those metrics directly via neural nets for training SE models. However, an accurate estimation of the true metric function introduces statistical complexity for training because it attempts to capture the exact value of the metric. We propose an adversarial training strategy based on statistical correlation that avoids the complexity of estimating the SE metric while learning to mimic its overall behavior. We call this framework CorrGAN and show its significant improvement over standard losses of the SOTA baselines and achieve SOTA performance on the VoiceBank+DEMAND dataset. Vasily Zadorozhnyy, Saeed Amizadeh, Qiang Ye 0003, Kazuhito Koishida |
ICASSP | 4 |
| 2025 | VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web TasksabstractVideos are often used to learn or extract the necessary information to complete
tasks in ways different than what text or static imagery can provide. However, many
existing agent benchmarks neglect long-context video understanding, instead focus-
ing on text or static image inputs. To bridge this gap, we introduce VideoWebArena
(VideoWA), a benchmark for evaluating the capabilities of long-context multimodal
agents for video understanding. VideoWA consists of 2,021 web agent tasks based
on manually crafted video tutorials, which total almost four hours of content. For
our benchmark, we define a taxonomy of long-context video-based agent tasks with
two main areas of focus: skill retention and factual retention. While skill retention
tasks evaluate whether an agent can use a given human demonstration to complete
a task efficiently, the factual retention task evaluates whether an agent can retrieve
instruction-relevant information from a video to complete a task. We find that the
best model achieves a 13.3% success rate on factual retention tasks and 45.8% on
factual retention QA pairs—far below human success rates of 73.9% and 79.3%,
respectively. On skill retention tasks, long-context models perform worse with
tutorials than without, exhibiting a 5% performance decrease in WebArena tasks
and a 10.3% decrease in VisualWebArena tasks. Our work highlights performance
gaps in the agentic abilities of long-context multimodal models and provides as a
testbed for the future development of long-context video agents. Lawrence Jang, Yinheng Li, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, Kazuhito Koishida |
ICLR | 8 |
| 2025 | Windows Agent Arena: Evaluating Multi-Modal OS Agents at ScaleabstractLarge language models (LLMs) show potential as computer agents, enhancing productivity and software accessibility in multi-modal tasks.
However, measuring agent performance in sufficiently realistic and complex environments becomes increasingly challenging as:
(i) most benchmarks are limited to specific modalities/domains (e.g., text-only, web navigation, Q&A) and
(ii) full benchmark evaluations are slow (on order of magnitude of multiple hours/days) given the multi-step sequential nature of tasks.
To address these challenges, we introduce Windows Agent Arena: a general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real OS to use the same applications and tools available to human users when performing tasks.
We create 150+ diverse tasks across representative domains that require agentic abilities in planning, screen understanding, and tool usage.
Our benchmark is scalable and can be seamlessly parallelized for a full benchmark evaluation in as little as $20$ minutes.
Our work not only speeds up the development and evaluation cycle of multi-modal agents, but also highlights and analyzes existing shortfalls in the agentic abilities of several multimodal LLMs as agents within the Windows computing environment---with the best achieving only a 19.5\% success rate compared to a human success rate of 74.5\%. Rogerio Bonatti, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, Zheng Hui |
ICML | 9 |
| 2025 | Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale ProblemsabstractTransformers and their attention mechanism have been revolutionary in the field of Machine Learning. While originally proposed for the language data, they quickly found their way to the image, video, graph, etc. data modalities with various signal geometries. Despite this versatility, generalizing the attention mechanism to scenarios where data is presented at different scales from potentially different modalities is not straightforward. The attempts to incorporate hierarchy and multi-modality within transformers are largely based on ad hoc heuristics, which are not seamlessly generalizable to similar problems with potentially different structures. To address this problem, in this paper, we take a fundamentally different approach: we first propose a mathematical construct to represent multi-modal, multi-scale data. We then mathematically derive the neural attention mechanics for the proposed construct from the first principle of entropy minimization. We show that the derived formulation is optimal in the sense of being the closest to the standard Softmax attention while incorporating the inductive biases originating from the hierarchical/geometric information of the problem. We further propose an efficient algorithm based on dynamic programming to compute our derived attention mechanism. By incorporating it within transformers, we show that the proposed hierarchical attention mechanism not only can be employed to train transformer models in hierarchical/multi-modal settings from scratch, but it can also be used to inject hierarchical information into classical, pre-trained transformer models post training, resulting in more efficient models in zero-shot manner. Saeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito Koishida |
NeurIPS | 4 |
| 2024 | uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio MixturesabstractMasked Autoencoders (MAEs) learn rich low-level representations from unlabeled data but require substantial labeled data to effectively adapt to downstream tasks. Conversely, Instance Discrimination (ID) emphasizes high-level semantics, offering a potential solution to alleviate annotation requirements in MAEs. Although combining these two approaches can address downstream tasks with limited labeled data, naively integrating ID into MAEs leads to extended training times and high computational costs. To address this challenge, we introduce uaMix-MAE, an efficient ID tuning strategy that leverages unsupervised audio mixtures. Utilizing contrastive tuning, uaMix-MAE aligns the representations of pretrained MAEs, thereby facilitating effective adaptation to task-specific semantics. To optimize the model with small amounts of unlabeled data, we propose an audio mixing technique that manipulates audio samples in both input and virtual label spaces. Experiments in low/few-shot settings demonstrate that uaMix-MAE achieves 4 − 6% accuracy improvements over various benchmarks when tuned with limited unlabeled data, such as AudioSet-20K. Afrina Tabassum, Dung N. Tran, Trung Dang 0002, Ismini Lourentzou, Kazuhito Koishida |
ICASSP | 5 |
| 2024 | Learned Image Compression With Text Quality EnhancementabstractLearned image compression has gained widespread popularity for their efficiency in achieving ultra-low bit-rates. Yet, images containing substantial textual content, particularly screen-content images (SCI), often suffers from text distortion at such compressed levels. To address this, we propose to minimize a novel text logit loss designed to quantify the disparity in text between the original and reconstructed images, thereby improving the perceptual quality of the reconstructed text. Through rigorous experimentation across diverse datasets and employing state-of-the-art algorithms, our findings reveal significant enhancements in the quality of reconstructed text upon integration of the proposed loss function with appropriate weighting. Notably, we achieve a Bjontegaard delta (BD) rate of $-32.64 \%$ for Character Error Rate (CER) and $-28.03 \%$ for Word Error Rate (WER) on average by applying the text logit loss for two screenshot datasets. Additionally, we present quantitative metrics tailored for evaluating text quality in image compression tasks. Our findings underscore the efficacy and potential applicability of our proposed text logit loss function across various text-aware image compression contexts. Chih-Yu Lai, Dung N. Tran, Kazuhito Koishida |
ICIP | 3 |
| 2024 | Weakly-supervised Audio Separation via Bi-modal Semantic SimilarityabstractConditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop with multi-source training mixtures due to the lack of supervision signal for single source separation cases during training. However, in the case of language-conditional audio separation, we do have access to corresponding text descriptions for each audio mixture in our training data, which can be seen as (rough) representations of the audio samples in the language modality. That raises the curious question of how to generate supervision signal for single-source audio extraction by leveraging the fact that single-source sounding language entities can be easily extracted from the text description. To this end, in this paper, we propose a generic bi-modal separation framework which can enhance the existing unsupervised frameworks to separate single-source signals in a target modality (i.e., audio) using the easily separable corresponding signals in the conditioning modality (i.e., language), without having access to single-source samples in the target modality during training. We empirically show that this is well within reach if we have access to a pretrained joint embedding model between the two modalities (i.e., CLAP). Furthermore, we propose to incorporate our framework into two fundamental scenarios to enhance separation performance. First, we show that our proposed methodology significantly improves the performance of purely unsupervised baselines by reducing the distribution shift between training and test samples. In particular, we show that our framework can achieve 71% boost in terms of Signal-to-Distortion Ratio (SDR) over the baseline, reaching 97.5% of the supervised learning performance. Second, we show that we can further improve the performance of the supervised learning itself by 17% if we augment it by our proposed weakly-supervised framework. Our framework achieves this by making large corpora of unsupervised data available to the supervised learning model as well as utilizing a natural, robust regularization mechanism through weak supervision from the language modality, and hence enabling a powerful semi-supervised framework for audio separation. Code is released at https://github.com/microsoft/BiModalAudioSeparation. Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana Marculescu |
ICLR | 3 |
| 2024 | LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Trung Dang 0002, David Aponte, Dung N. Tran, Kazuhito Koishida |
INTERSPEECH | 4 |
| 2024 | ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
Yatong Bai, Trung Dang 0002, Dung N. Tran, Kazuhito Koishida, Somayeh Sojoudi |
INTERSPEECH | 4 |
| 2024 | Automatic Disfluency Detection From Untranscribed SpeechabstractSpeech disfluencies, such as filled pauses or repetitions, are disruptions in the typical flow of speech. All speakers experience disfluencies at times, and the rate at which we produce disfluencies may be increased by certain speaker or environmental characteristics. Modeling disfluencies has been shown to be useful for a range of downstream tasks, and as a result, disfluency detection has many potential applications. In this work, we investigate language, acoustic, and multimodal methods for frame-level automatic disfluency detection and categorization. Each of these methods relies on audio as an input. First, we evaluate several automatic speech recognition (ASR) systems in terms of their ability to transcribe disfluencies, measured using disfluency error rates. We then use these ASR transcripts as input to a language-based disfluency detection model. We find that disfluency detection performance is largely limited by the quality of transcripts and alignments. We find that an acoustic-based approach that does not require transcription as an intermediate step outperforms the ASR language approach. Finally, we present multimodal architectures which we find improve disfluency detection performance over the unimodal approaches. Ultimately, this work introduces novel approaches for automatic frame-level disfluency and categorization. In the long term, this will help researchers incorporate automatic disfluency detection into a range of applications. Amrit Romana, Kazuhito Koishida, Emily Mower Provost |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Toward A Multimodal Approach for Disfluency Detection and CategorizationabstractSpeech disfluencies, such as filled pauses, repetitions, or revisions, disrupt the typical flow of speech. Disfluency detection and categorization has gained traction as a research area because modeling disfluent events has been shown to be helpful for downstream tasks. However, the majority of work on disfluency detection and categorization has focused on language-based approaches that process manually transcribed text. While these methods have shown high accuracy, requiring manually transcribed text limits the scalability and practicality of these approaches. In this paper, we evaluate the impact of using automatic speech recognition (ASR) transcripts for disfluency detection and categorization. Additionally, we explore skipping transcription altogether with an acoustic approach. We assess the strengths and weaknesses of each modality, and in this paper we introduce a model fusion approach that combines the two modalities. We find that multimodal disfluency detection and categorization outperforms using either individual modality, and that the improvement in performance is especially significant when the language-based model processes ASR transcripts. Amrit Romana, Kazuhito Koishida |
ICASSP | 2 |
| 2023 | SCP-GAN: Self-Correcting Discriminator Optimization for Training Consistency Preserving Metric GAN on Speech Enhancement Tasks
Vasily Zadorozhnyy, Qiang Ye 0003, Kazuhito Koishida |
INTERSPEECH | 3 |
| 2023 | Progressive Ensemble Distillation: Building Ensembles for Efficient InferenceabstractKnowledge distillation is commonly used to compress an ensemble of models into a single model. In this work we study the problem of progressive ensemble distillation: Given a large, pretrained teacher model , we seek to decompose the model into an ensemble of smaller, low-inference cost student models . The resulting ensemble allows for flexibly tuning accuracy vs. inference cost, which can be useful for a multitude of applications in efficient inference. Our method, B-DISTIL, uses a boosting procedure that allows function composition based aggregation rules to construct expressive ensembles with similar performance as using much smaller student models. We demonstrate the effectiveness of B-DISTIL by decomposing pretrained models across a variety of image, speech, and sensor datasets. Our method comes with strong theoretical guarantees in terms of convergence as well as generalization. Don Kurian Dennis, Abhishek Shetty, Anish Prasad Sevekari, Kazuhito Koishida, Virginia Smith |
NeurIPS | 4 |
| 2022 | Training Robust Zero-Shot Voice Conversion Models with Self-Supervised FeaturesabstractUnsupervised Zero-Shot Voice Conversion (VC) aims to modify the speaker characteristic of an utterance to match an unseen target speaker without relying on parallel training data. Recently, self-supervised learning of speech representation has been shown to produce useful linguistic units without using transcripts, which can be directly passed to a VC model. In this paper, we showed that high-quality audio samples can be achieved by using a length resampling decoder, which enables the VC model to work in conjunction with different linguistic feature extractors and vocoders without requiring them to operate on the same sequence length. We showed that our method can outperform many baselines on the VCTK dataset. Without modifying the architecture, we further demonstrated that a) using pairs of different audio segments from the same speaker, b) adding a cycle consistency loss, and c) adding a speaker classification loss can help to learn a better speaker embedding. Our model trained on LibriTTS using these techniques achieves the best performance, producing audio samples transferred well to the target speaker’s voice, while preserving the linguistic content that is comparable with actual human utterances in terms of Character Error Rate. Trung Dang 0002, Dung N. Tran, Sang (Peter) Chin, Kazuhito Koishida |
ICASSP | 4 |
| 2022 | A Training Framework for Stereo-Aware Speech Enhancement Using Deep Neural NetworksabstractDeep learning-based speech enhancement has shown unprecedented performance in recent years. The most popular mono speech enhancement frameworks are end-to-end networks mapping the noisy mixture into an estimate of the clean speech. With growing computational power and availability of multichannel microphone recordings, prior work has aimed to incorporate spatial statistics along with spectral information to boost up performance. Despite an improvement in enhancement performance of mono output, the spatial image preservation and subjective evaluations have not gained much attention in the literature. This paper proposes a novel stereo-aware framework for speech enhancement, i.e., a training loss for deep learning-based speech enhancement to preserve the spatial image while enhancing the stereo mixture. The proposed framework is model independent, hence it can be applied to any deep learning based architecture. We provide an extensive objective and subjective evaluation of the trained models through a listening test. We show that by regularizing for an image preservation loss, the overall performance is improved, and the stereo aspect of the speech is better preserved. Bahareh Tolooshams, Kazuhito Koishida |
ICASSP | 2 |
| 2021 | Cascaded Time + Time-Frequency Unet For Speech Enhancement: Jointly Addressing Clipping, Codec Distortions, And GapsabstractSpeech enhancement aims to improve speech quality by eliminating noise and distortions. While most speech enhancement methods address signal independent additive sources of noise, several degradations to speech signals are signal dependent and non-additive, like speech clipping, codec distortions, and gaps in speech. In this work, we first systematically study and achieve state of the art results on each of these three distortions individually. Next, we demonstrate a neural network pipeline that cascades a time domain convolutional neural network with a time-frequency domain convolutional neural network to address all three distortions jointly. We observe that such a cascade achieves good performance while also keeping the action of each neural network component interpretable. Arun Asokan Nair, Kazuhito Koishida |
ICASSP | 2 |
| 2021 | Single-Channel Speech Enhancement Using Learnable Loss Mixup
Oscar Chang, Dung N. Tran, Kazuhito Koishida |
Interspeech | 3 |
| 2021 | INTERSPEECH 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP 2020. We open-sourced training and test datasets for the wideband scenario. We also open-sourced a subjective evaluation framework based on ITU-T standard P.808, which was also used to evaluate participants of the challenge. Many researchers from academia and industry made significant contributions to push the field forward, yet even the best noise suppressor was far from achieving superior speech quality in challenging scenarios. In this version of the challenge organized at INTERSPEECH 2021, we are expanding both our training and test datasets to accommodate full band scenarios. The two tracks in this challenge will focus on real-time denoising for (i) wide band, and(ii) full band scenarios. We are also making available a reliable non-intrusive objective speech quality metric called DNSMOS for the participants to use during their development phase. Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 3 |
| 2020 | MMTM: Multimodal Transfer Module for CNN FusionabstractIn late fusion, each modality is processed in a separate unimodal Convolutional Neural Network (CNN) stream and the scores of each modality are fused at the end. Due to its simplicity, late fusion is still the predominant approach in many state-of-the-art multimodal applications. In this paper, we present a simple neural network module for leveraging the knowledge from multiple modalities in convolutional neural networks. The proposed unit, named Multimodal Transfer Module (MMTM), can be added at different levels of the feature hierarchy, enabling slow modality fusion. Using squeeze and excitation operations, MMTM utilizes the knowledge of multiple modalities to recalibrate the channel-wise features in each CNN stream. Unlike other intermediate fusion methods, the proposed module could be used for feature modality fusion in convolution layers with different spatial dimensions. Another advantage of the proposed method is that it could be added among unimodal branches with minimum changes in the their network architectures, allowing each branch to be initialized with existing pretrained weights. Experimental results show that our framework improves the recognition accuracy of well-known multimodal networks. We demonstrate state-of-the-art or competitive performance on four datasets that span the task domains of dynamic hand gesture recognition, speech enhancement, and action recognition with RGB and body joints. Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito Koishida |
CVPR | 4 |
| 2020 | Low-Latency Single Channel Speech Enhancement Using U-Net Convolutional Neural NetworksabstractSingle-channel speech enhancement (SE) can be described, in its simplest terms, as learning a transformation from single-channel noisy speech to the clean speech. To do this, we propose a simple but effective U-Net convolutional neural network (CNN) based architecture with skip-connections with a focus on real-time applications which require low-latency processing. To that end, we choose to process relatively small temporal windows and apply time-frequency (T-F) featurization on it to achieve magnitude estimation. Two state-of-the-art systems are picked for bench-marking: One operating on spectral-domain [1] and the other on temporal-domain [2]. We evaluate the performance of the systems in terms of perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI). Experimental results show that in terms of PESQ measure the proposed method provides around 27% and 11% relative improvement over the baseline systems respectively and has significantly lower latency compared to them. We further investigate the trade-off between performance and overall latency of the proposed system. Ahmet Emin Bulut, Kazuhito Koishida |
ICASSP | 2 |
| 2020 | AV(SE)2: Audio-Visual Squeeze-Excite Speech EnhancementabstractThe goal of audio-visual speech enhancement (AVSE) is to supplement audio-only information with visual information, such as target speaker's lip movements, to improve the intelligibility and overall perceptual quality of noisy speech signals. We propose a new mechanism for audio-visual (AV) fusion that leverages a cross-modal squeeze-excitation (SE) block for speech enhancement: AV(SE)2. The fusion block is adaptable to any feature layer of the audio and visual networks and significantly reduces model parameters as compared to standard AV fusion methods of channel-wise concatenation without loss of performance. We show that AV(SE)2with time-based gating across multiple feature layers outperforms baseline methods of single-point, channel-wise concatenated AV fusion on objective evaluations. Michael L. Iuzzolino, Kazuhito Koishida |
ICASSP | 2 |
| 2020 | Geometrically Constrained Independent Vector Analysis for Directional Speech EnhancementabstractThis paper addresses the multichannel directional speech enhancement problem with geometrically constrained independent vector analysis (GCIVA), where we aim to combine the high separation performance from blind source separation and the capability of directional focus from beamforming. The proposed method exploits geometric constraints composed from the spatial information of sources to guide the target speech to the desired output channel. A convergence-guaranteed parameter estimation algorithm is derived from the framework of auxiliary function-based IVA (AuxIVA) to take advantage of fast convergence, low computational cost, and no step-size tuning. We propose a dual-microphone speech enhancement system based on the proposed method and investigate its effectiveness with objective metrics. The experimental evaluations revealed that the proposed system outperformed the conventional beamforming and the standard AuxIVA in a large margin in terms of source-to-distortion and source-to-interference ratios. Li Li 0063, Kazuhito Koishida |
ICASSP | 2 |
| 2020 | Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"
Saeed Amizadeh, Hamid Palangi, Oleksandr Polozov, Kazuhito Koishida |
ICML | 5 |
| 2020 | Online Directional Speech Enhancement Using Geometrically Constrained Independent Vector Analysis
Li Li 0063, Kazuhito Koishida, Shoji Makino |
INTERSPEECH | 2 |
| 2020 | Low-Latency Single Channel Speech Dereverberation Using U-Net Convolutional Neural Networks
Ahmet Emin Bulut, Kazuhito Koishida |
INTERSPEECH | 2 |
| 2020 | Robust Pitch Regression with Voiced/Unvoiced Classification in Nonstationary Noise Environments
Dung N. Tran, Uros Batricevic, Kazuhito Koishida |
INTERSPEECH | 3 |
| 2020 | Single-Channel Speech Enhancement by Subspace Affinity Minimization
Dung N. Tran, Kazuhito Koishida |
INTERSPEECH | 2 |
| 2019 | Speech Super Resolution Generative Adversarial NetworkabstractThe goal of speech super-resolution (SSR) or speech bandwidth expansion is to generate the missing high-frequency components for a given low-resolution speech signal. It has the potential to improve the quality of telecommunications. We propose a new method for SSR that leverages the generative adversarial networks (GANs) and a regularization method for stabilizing the GAN training. The generator network is a convolutional autoencoder with 1D convolution kernels, operating along time-axis and generating the high-frequency log-power spectra from the low-frequency log-power spectra input. We employ two recent deep neural network (DNN) based approaches to compare them with our proposed method, including both objective speech quality metrics and subjective perceptual tests. We show that our proposed method outperforms the baseline methods in terms of both objective and subjective evaluations. Sefik Emre Eskimez, Kazuhito Koishida |
ICASSP | 2 |
| 2019 | Sound Event Detection in Multichannel Audio Using Convolutional Time-Frequency-Channel Squeeze and ExcitationabstractIn this study, we introduce a convolutional time-frequency-channel Squeeze and Excitation (tfc-SE) module to explicitly model inter-dependencies between the time-frequency domain and multiple channels. The tfc-SE module consists of two parts: tf-SE block and c-SE block which are designed to provide attention on time-frequency and channel domain, respectively, for adaptively recalibrating the input feature map. The proposed tfc-SE module, together with a popular Convolutional Recurrent Neural Network (CRNN) model, are evaluated on a multi-channel sound event detection task with overlapping audio sources: the training and test data are synthesized TUT Sound Events 2018 datasets, recorded with microphone arrays. We show that the tfc-SE module can be incorporated into the CRNN model at a small additional computational cost and bring significant improvements on sound event detection accuracy. We also perform detailed ablation studies by analyzing various factors that may influence the performance of the SE blocks. We show that with the best tfc-SE block, error rate (ER) decreases from 0.2538 to 0.2026, relative 20.17\% reduction of ER, and 5.72\% improvement of F1 score. The results indicate that the learned acoustic embeddings with the tfc-SE module efficiently strengthen time-frequency and channel-wise feature representations to improve the discriminative performance. Kazuhito Koishida |
INTERSPEECH | 2 |
| 2018 | Text-Independent Speaker Verification Based on Triplet Convolutional Neural Network EmbeddingsabstractThe effectiveness of introducing deep neural networks into conventional speaker recognition pipelines has been broadly shown to benefit system performance. A novel text-independent speaker verification (SV) framework based on the triplet loss and a very deep convolutional neural network architecture (i.e., Inception-Resnet-v1) are investigated in this study, where a fixed-length speaker discriminative embedding is learned from sparse speech features and utilized as a feature representation for the SV tasks. A concise description of the neural network based speaker discriminative training with triplet loss is presented. An Euclidean distance similarity metric is applied in both network training and SV testing, which ensures the SV system to follow an end-to-end fashion. By replacing the final max/average pooling layer with a spatial pyramid pooling layer in the Inception-Resnet-v1 architecture, the fixed-length input constraint is relaxed and an obvious performance gain is achieved compared with the fixed-length input speaker embedding system. For datasets with more severe training/test condition mismatches, the probabilistic linear discriminant analysis (PLDA) back end is further introduced to replace the distance based scoring for the proposed speaker embedding system. Thus, we reconstruct the SV task with a neural network based front-end speaker embedding system and a PLDA that provides channel and noise variabilities compensation in the back end. Extensive experiments are conducted to provide useful hints that lead to a better testing performance. Comparison with the state-of-the-art SV frameworks on three public datasets (i.e., a prompt speech corpus, a conversational speech Switchboard corpus, and NIST SRE10 10 s-10 s condition) justifies the effectiveness of our proposed speaker embedding system. Kazuhito Koishida, John H. L. Hansen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | End-to-end text-independent speaker verification with flexibility in utterance durationabstractWe continue to investigate end-to-end text-independent speaker verification by incorporating the variability from different utterance durations. Our previous study [1] showed a competitive performance with a triplet loss based end-to-end text-independent speaker verification system. To normalize the duration variability, we provided fixed length inputs to the network by a simple cropping or padding operation. Those operations do not seem ideal, particularly for long duration where some amount of information is discarded, while an i-vector system typically has improved accuracy with an increase in input duration. In this study, we propose to replace the final max/average pooling layer with a Spatial Pyramid Pooling layer in the Inception-Resnet-v1 architecture, which allows us to relax the fixed-length input constraint and train the entire network with the arbitrary size of input in an end-to-end fashion. In this way, the modified network can map variable length utterances into fixed length embeddings. Experiments shows that the new end-to-end system with variable size input relatively reduces EER by 8.4% over the end-to-end system with fixed-length input, and 24.0% over the i-vector/PLDA baseline system. an end-to-end system with. Kazuhito Koishida |
ASRU | 2 |
| 2017 | End-to-End Text-Independent Speaker Verification with Triplet Loss on Short Utterances
Kazuhito Koishida |
INTERSPEECH | 2 |
| 2008 | Hybrid low bitrate audio coding using adaptive gain shape vector quantizationabstractAudio coding at low bitrates typically suffers from artifacts caused by bandwidth truncation. In this paper we present a novel scheme to code audio signals at low bitrates which uses a traditional scalar quantization followed by entropy coding to code some portions of the spectrum (typically the lower portion). The other portions (typically the higher portions) of the spectrum are coded at a low bitrate using an adaptive gain shape vector quantizer where the codebook for vector quantization is formed by unmodified or modified versions of the portions of the spectrum which have already been coded. Fixed pre-trained codebooks are also available for use in certain cases. The use of such a scheme results in an audio codec which has been shown to be among the best audio codecs available at low bitrates. In addition, the decoder complexity of this audio codec is significantly lower than any other codec of equal quality at low bitrates. Sanjeev Mehrotra, Wei-Ge Chen, Kazuhito Koishida, Naveen Thumpudi |
MMSP | 3 |
| 2000 | A 16-kbit/s bandwidth scalable audio coder based on the G.729 standardabstractThis paper proposes a bandwidth-scalable coding scheme based on the G.729 standard as a base layer coder. In the scheme, according to the channel conditions, the output speech of the decoder can be selected to be narrowband (4-kHz bandwidth) or wideband (8-kHz bandwidth). The proposed scheme consists of two layers: base and enhancement. The base coder uses the G.729 algorithm to encode narrowband speech. The enhancement coder is based on a full-band CELP model and it encodes wideband speech while making use of the available base layer information. Two bandwidth-scalable coders are designed: one is scalable with the 8 kbit/s G.729 base coder and another with the 6.4 kbit/s G.729 (Annex D) base coder. Subjective tests show that, for wideband speech, the proposed coders at 16 kbit/s achieve better performance than the 16 kbit/s MPEG-4 CELP with bandwidth scalability. Kazuhito Koishida, Vladimir Cuperman, Allen Gersho |
ICASSP | 1 |
| 2000 | A 1200 bps speech coder based on MELPabstractThis paper presents a 1.2 kbps speech coder based on the mixed excitation linear prediction (MELP) analysis algorithm. In the proposed coder, the MELP parameters of three consecutive frames are grouped into a superframe and jointly quantized to obtain a high coding efficiency. The interframe redundancy is exploited with distinct quantization schemes for different unvoiced/voiced (U/V) frame combinations in the superframe. Novel techniques for improving performance make use of the superframe structure. These include pitch vector quantization using pitch differentials, joint quantization of pitch and U/V decisions and LSF quantization with a forward-backward interpolation method. Subjective test results indicate that the 1.2 kbps speech coder achieves approximately the same quality as the proposed federal standard 2.4 kbps MELP coder. Tian Wang 0003, Kazuhito Koishida, Vladimir Cuperman, Allen Gersho, John S. Collura |
ICASSP | 2 |
| 1998 | A wideband CELP speech coder at 16 kbit/s based on mel-generalized cepstral analysisabstractThis paper proposes a wideband CELP coder using frequency warping. Instead of linear prediction, the proposed coder adopts the mel-generalized cepstral analysis, and encodes the fullband of the speech signal through a warped frequency scale. It is shown that the subjective quality of the proposed coder at 16 kbit/s is better than that of the ITU-T G.722 at 64 kbit/s. Furthermore, the proposed coder gives a much smaller difference in performance for male and female speakers than the conventional CELP coder. These results indicate that the frequency warping makes a large contribution to the improvement of the subjective quality for wideband speech coding. Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi |
ICASSP | 1 |
| 1998 | A 16 kbit/s wideband CELP coder using MEL-generalized cepstral analysis and its subjective evaluationabstractWe have proposed a wideband CELP coder, called MGC-CELP, which provides high quality speech by utilizing mel-generalized cepstral (MGC) analysis instead of linear prediction (LP). In this paper, we investigate the performance of the wideband MGCCELP coder at 16 kbit/s in terms of short-term predictor order, i.e., order of MGC analysis. Subjective tests show that the MGCCELP coder with a predictor of order 20 gives better performance than ITU-T G.722 at 64 kbit/s. It is also found that the MGCCELP coder with 12th order achieves comparable quality to the 64 kbit/s G.722, and outperforms the 16 kbit/s conventional CELP coder using 20th-order LP analysis under the same conditions. 1. INTRODUCTION Recently several schemes for high-quality wideband speech coding at low bit rates have been developed. Most of the work in this field uses either transform/subband coding or CELP (Code Excited Linear Prediction) coding. At the bit rates around 16 kbit/s, CELP coding has received much attention s... Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi |
ICSLP | 1 |
| 1997 | Efficient encoding of mel-generalized cepstrum for CELP codersabstractThe performance of several algorithms for the quantization of the mel-generalized cepstral coefficients is studied. First, the objective and subjective performance of two-stage vector quantization (VQ) is measured. It is shown that the subjective quality for the mel-generalized cepstral coefficients is higher than that for LSP. Secondly, interframe prediction is introduced in the encoding of mel-generalized cepstral coefficients. By utilizing interframe moving average (MA) prediction, the mel-generalized cepstral coefficients can be encoded more efficiently than LSP in terms of cepstral distortion. Finally, we implement a CELP coder based on mel-generalized cepstral analysis in which mel-generalized cepstral coefficients are quantized using MA prediction. This coder has a higher objective quality than conventional CELP. Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 1 |
| 1996 | CELP coding system based on mel-generalized cepstral analysis
Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICSLP | 1 |
| 1995 | CELP coding based on mel-cepstral analysisabstractWe propose a CELP coder based on mel-cepstral analysis. In the coder, since the transfer functions of perceptual weighting and postfiltering are defined through mel-cepstral coefficients, the effects of perceptual weighting and postfiltering should fit with the characteristics of the human auditory sensation. We use a basic CELP structure without adaptive codebook, and the subjective speech quality of the proposed coder in terms of the opinion equivalent Q is measured and compared with that of the conventional CELP coder. It is shown that the improvement of more than 1.8 dB is achieved by the proposed coder over the conventional CELP coder. Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICASSP | 1 |
| 1994 | Speech coding based on adaptive MEL-cepstral analysis for noisy channels
Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
ICSLP | 1 |