Noboru Harada

dblp:31/4910 · DBLP profile ↗
← Back
58ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0002-1759-4533ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorComputer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 FedPM: Federated Learning Using Second-order Optimization with Preconditioned Mixing of Local Parameters
abstract
We propose Federated Preconditioned Mixing (FedPM), a novel Federated Learning (FL) method that leverages second-order optimization. Prior methods - such as LocalNewton, LTDA, and FedSophia - have incorporated second-order optimization in FL by performing iterative local updates on clients and applying simple mixing of local parameters on the server. However, these methods often suffer from drift in local preconditioners, which significantly disrupts the convergence of parameter training, particularly in heterogeneous data settings. To overcome this issue, we refine the update rules by decomposing the ideal second-order update - computed using globally preconditioned global gradients - into parameter mixing on the server and local parameter updates on clients. As a result, our FedPM introduces preconditioned mixing of local parameters on the server, effectively mitigating drift in local preconditioners. We provide a theoretical convergence analysis demonstrating a superlinear rate for strongly convex objectives in scenarios involving a single local update. To demonstrate the practical benefits of FedPM, we conducted extensive experiments. The results showed significant improvements with FedPM in the test accuracy compared to conventional methods incorporating simple mixing, fully leveraging the potential of second-order optimization.
Hiro Ishii, Kenta Niwa, Hiroshi Sawada, Akinori Fujino, Noboru Harada, Rio Yokota
AAAI5
2026 Guided Masked Self-Distillation Modeling for Distributed Multimedia Sensor Event Analysis
abstract
This article addresses a new task: distributed multimedia sensor event analysis (DiMSEA). DiMSEA aims to analyze a series of human and machine activities (called “events” in this article) in complex and extensive real-world environments. Since an observation from a single sensor is often missing or fragmented in such an environment, observations from multiple locations and modalities should be integrated to analyze events comprehensively. However, a learning method has yet to be established to extract joint representations that effectively combine such distributed observations. Therefore, we propose guided masked self-distillation modeling (Guided-MELD) for inter-sensor relationship modeling. The basic idea of Guided-MELD is to learn to supplement the information from the masked sensor with information from other sensors needed to detect the event. Guided-MELD is expected to effectively distill fragmented target event information from sensors without over-relying on any specific sensors. To validate the effectiveness of the proposed method in DiMSEA, we recorded two new datasets: MM-Store and MM-Office. These datasets consist of human activities in a convenience store and an office, recorded using distributed cameras and microphones. Experimental results show that the proposed Guided-MELD improves event tagging and detection performance and outperforms conventional inter-sensor relationship modeling methods. Furthermore, the proposed method performed robustly even when sensors were reduced.
Masahiro Yasuda, Noboru Harada, Yasunori Ohishi, Shoichiro Saito, Akira Nakayama, Nobutaka Ono
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Stereo Downmix in 3GPP IVAS for EVS Compatibility
abstract
The 3GPP IVAS codec specifies an EVS-compatible stereo downmix as one of the key functionalities. This paper describes how this novel active downmix scheme has been devised to achieve high and stable quality from stereo input to EVS encoder/decoder with no additional algorithmic delay. An example of a network configuration for a multi-party conference using the proposed EVS-compatible downmix is also provided. This paper shows several fundamental schemes, of active processes, including adaptive weight and phase-compensated weight between two channels. Subjective listening test results show that the devised scheme’s quality is better than that of a passive downmix.
Takehiro Moriya, Stéphane Ragot, Arnaud Lefort, Alexandre Guérin, Noboru Harada, Ryosuke Sugiura, Yutaka Kamamoto
ICASSP5
2025 Collision-less and Balanced Sampling for Language-Queried Audio Source Separation
abstract
Language-queried audio source separation (LASS) is an emerging research field that has recently received increasing attention. This task aims to isolate individual sources from a mixture of signals using natural language descriptions, enabling applications in various areas such as automatic audio editing. While conventional methods focus on the system architecture, the important aspect of data processing has been overlooked. The data for training LASS are typically created by mixing various audio signals in the dataset to form a mixture. One signal is then used as the target, whereas the others are regarded as interference. However, sound events in the target signal could overlap with those in the interference signals, which may cause confusion that instructs the model to both retain and suppress the same sound events within a single training example. In addition, training LASS with large-scale datasets may suffer from the data imbalance problem, where some sound events appear too frequently while others are rare. In this paper, we address these problems by using data sampling techniques. Specifically, the interference signals are sampled so that their audio tags do not conflict with those of the target signal, where the tags are generated using an audio tagging model. To balance the data, we consider several balanced sampling approaches using tag or caption embedding. By leveraging their distribution information, we use either weighted or group sampling to boost the occurrence of underrepresented samples while reducing the presence of overrepresented ones. Experimental results show the superiority of the proposed method over state-of-the-art LASS systems in DCASE 2024 Challenge Task 9. Pre-trained model is available at: https://github.com/tucothien/LASS-CLBS.
Binh Thien Nguyen, Daiki Takeuchi, Masahiro Yasuda, Daisuke Niizumi, Noboru Harada
ICASSP5
2025 Sound Source Distance Estimation Utilizing Physics-informed Prior for Sound Event Localization and Detection
abstract
Sound Event Localization and Detection (SELD) is the combined task of detecting sound events and estimating their spatial locations. We propose a Sound source Distance Estimation (SDE) method for SELD that utilizes a physics-informed prior. The conventional data-driven approach of SDE for SELD can handle complex situations where sound sources move or overlap, thanks to multitask learning. Deep Neural Network (DNN)-based SELD systems are generally pre-trained with synthesized data to compensate for the lack of real data. However, in the context of recent SELD tasks, the performance of DNN models, which are pre-trained with synthetic data, significantly degrades on real data. One possible cause of this is that the DNN models overfit sound characteristics that do not exist in the real data, i.e., are not physically reasonable. Therefore, we propose a hybrid SDE method that utilizes a physics-informed prior in a data-driven SELD system. Focusing on the specificity of the expected sound PoWer Level (PWL) of the sound sources depending on the class, we set the typical PWL for each class as a prior. To improve real-world applicability, we also adopt the sound attenuation model as a physics-informed prior for explicitly utilizing physical laws in SDE. Experimental results suggested the effectiveness of utilizing a physics-informed prior in SDE for SELD to improve its applicability to the real world.
Nao Sato, Masahiro Yasuda, Shoichiro Saito, Noboru Harada
ICASSP4
2025 Spatial Annotation-free Training for Sound Event Localization and Detection
abstract
Sound Event Localization and Detection (SELD) is the task of estimating the class, duration, and direction of arrival (DOA) of sound events. State-of-the-art SELD systems use a data-driven approach based on Deep Neural Networks (DNNs) to deal with complex situations involving overlapping and moving sound sources. Such systems need to be trained on real-recording data to succeed in real-world situations. However, annotation of real-world data, especially spatial annotation, incurs huge costs, and the amount of labeled data is currently insufficiently small. Therefore, we are introducing spatial annotation-free training for SELD, which trains a SELD system using only sound class and duration labels. As a first attempt at this task, we propose Beam-based Multiple Instance Learning (Beam-MIL). Beam-MIL first instantiates acoustic signals for each DOA by beamforming. Then, the sound events contained in each instance are trained indirectly by MIL without using each instance’s ground truth information, i.e., spatial annotation. Experimental results show that Beam-MIL can effectively train a valid DOA estimator without spatial annotations. Moreover, when the small amount of the annotated data is available, enlarging data size by adding annotation-free data significantly improved the performance of the system.
Masahiro Yasuda, Shoichiro Saito, Nao Sato, Noboru Harada
ICASSP4
2025 Towards Pre-training an Effective Respiratory Audio Foundation Model
Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen, Yasunori Ohishi, Noboru Harada
INTERSPEECH6
2025 CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
Daiki Takeuchi, Binh Thien Nguyen, Masahiro Yasuda, Yasunori Ohishi, Daisuke Niizumi, Noboru Harada
INTERSPEECH6
2025 SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
abstract
Development of optical technology has enabled imaging of two-dimensional (2D) sound fields. This acoustooptic sensing enables understanding of the interaction between sound and objects such as reflection and diffraction. Moreover, it is expected to be used an advanced measurement technology for sonars in self-driving vehicles and assistive robots. However, the low sound-pressure sensitivity of the acousto-optic sensing results in high intensity of noise on images. Therefore, denoising is an essential task to visualize and analyze the sound fields. In addition to denoising, segmentation of sound and object silhouette is also required to analyze interactions between them. In this paper, we propose sound-field-images-with-object-silhouette denoising and segmentation (SoundSil-DS) that jointly perform denoising and segmentation for sound fields and object silhouettes on a visualized image. We developed a new model based on the current state-of-the-art denoising network. We also created a dataset to train and evaluate the proposed method through acoustic simulation. The proposed method was evaluated using both simulated and measured data. We confirmed that our method can applied to experimentally measured data. These results suggest that the proposed method may improve the post-processing for sound fields, such as physical model-based three-dimensional reconstruction since it can remove unwanted noise and separate sound fields and other object silhouettes. Our code is available at https://github.com/httcslab/soundsil-ds.
Risako Tanigawa, Kenji Ishikawa, Noboru Harada, Yasuhiro Oikawa
WACV3
2024 6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on Self-Motioning Human
abstract
We aim to perform sound event localization and detection (SELD) using wearable equipment for a moving human, such as a pedestrian. Conventional SELD tasks have dealt only with microphone arrays located in static positions. However, self-motion with three rotational and three translational degrees of freedom (6DoF) shall be considered for wearable microphone arrays. A system trained only with a dataset using microphone arrays in a fixed position would be unable to adapt to the fast relative motion of sound events associated with self-motion, resulting in the degradation of SELD performance. To address this, we designed 6DoF SELD Dataset1for wearable systems, the first SELD dataset considering the self-motion of microphones. Furthermore, we proposed a multi-modal SELD system that jointly utilizes audio and motion tracking sensor signals. These sensor signals are expected to help the system find useful acoustic cues for SELD on the basis of the current self-motion state. Experimental results on our dataset show that the proposed method effectively improves SELD performance with a mechanism to extract acoustic features conditioned by sensor signals.
Masahiro Yasuda, Shoichiro Saito, Akira Nakayama, Noboru Harada
ICASSP4
2024 Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan
abstract
With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often ineffective, which remains a challenging problem. In this paper, we found that the human ear cannot sensitively perceive the difference between a precise phase spectrum and a biased phase (BP) spectrum. Therefore, we propose an optimization method of phase reconstruction, allowing freedom on the global-phase bias instead of reconstructing the precise phase spectrum. We applied it to a Conformer-based Metric Generative Adversarial Networks (CMGAN) baseline model, which relaxes the existing constraints of precise phase and gives the neural network a broader learning space. Results show that this method achieves a new state-of-the-art performance without incurring additional computational overhead.
Zheng Qiu, Daiki Takeuchi, Noboru Harada, Shoji Makino
ICASSP4
2024 M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Masahiro Yasuda, Shunsuke Tsubaki, Keisuke Imoto
INTERSPEECH4
2024 Masked Modeling Duo: Towards a Universal Audio Pre-Training Framework
abstract
Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals. Unlike conventional methods, M2D obtains a training signal by encoding only the masked part, encouraging the two networks in M2D to model the input. While M2D improves general-purpose audio representations, a specialized representation is essential for real-world applications, such as in industrial and medical domains. The often confidential and proprietary data in such domains is typically limited in size and has a different distribution from that in pre-training datasets. Therefore, we propose M2D for X (M2D-X), which extends M2D to enable the pre-training of specialized representations for an application X. M2D-X learns from M2D and an additional task and inputs background noise. We make the additional task configurable to serve diverse applications, while the background noise helps learn on small data and forms a denoising task that makes representation robust. With these design choices, M2D-X should learn a representation specialized to serve various application needs. Our experiments confirmed that the representations for general-purpose audio, specialized for the highly competitive AudioSet and speech domain, and a small-data medical task achieve top-level performance, demonstrating the potential of using our models as a universal audio pre-training framework. Our code is available online for future studies.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
abstract
Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods learn representations directly by predicting representations of masked patches; however, we think using all patches to encode training signal representations is suboptimal. We propose a new method, Masked Modeling Duo (M2D), that learns representations directly while obtaining training signals using only masked patches. In the M2D, the online network encodes visible patches and predicts masked patch representations, and the target network, a momentum encoder, encodes masked patches. To better predict target representations, the online network should model the input well, while the target network should also model it well to agree with online predictions. Then the learned representations should better model the input. We validated the M2D by learning general-purpose audio representations, and M2D set new state-of-the-art performance on tasks such as UrbanSound8K, VoxCeleb1, AudioSet20K, GTZAN, and SpeechCommandsV2.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
ICASSP4
2023 Masked Modeling Duo for Speech: Specializing General-Purpose Audio Representation to Speech using Denoising Distillation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
INTERSPEECH4
2023 BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations
abstract
Pre-trained models are essential as feature extractors in modern machine learning systems in various domains. In this study, we hypothesize that representations effective for general audio tasks should provide multiple aspects of robust features of the input sound. For recognizing sounds regardless of perturbations such as varying pitch or timbre, features should be robust to these perturbations. For serving the diverse needs of tasks such as recognition of emotions or music genres, representations should provide multiple aspects of information, such as local and global features. To implement our principle, we propose a self-supervised learning method: Bootstrap Your Own Latent (BYOL) for Audio (BYOL-A, pronounced “viola”). BYOL-A pre-trains representations of the input sound invariant to audio data augmentations, which makes the learned representations robust to the perturbations of sounds. Whereas the BYOL-A encoder combines local and global features and calculates their statistics to make the representation provide multi-aspect information. As a result, the learned representations should provide robust and multi-aspect information to serve various needs of diverse tasks. We evaluated the general audio task performance of BYOL-A compared to previous state-of-the-art methods, and BYOL-A demonstrated generalizability with the best average result of 72.4% and the best VoxCeleb1 result of 57.6%. Extensive ablation experiments revealed that the BYOL-A encoder architecture contributes to most performance, and the final critical portion resorts to the BYOL framework and BYOL-A augmentations. Our code is available online for future studies.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Multi-View And Multi-Modal Event Detection Utilizing Transformer-Based Multi-Sensor Fusion
abstract
We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed sensors are utilized complementarily to capture events that are difficult to capture with a single sensor, such as a series of actions of people moving in an intricate room, or communication between people located far apart in a room. For sensors to cooperate effectively in such a situation, the system should be able to exchange information among sensors and combines information that is useful for identifying events in a complementary manner. For such a mechanism, we propose a Transformer-based multi-sensor fusion (MultiTrans) which combines multi-sensor data on the basis of the relationships between features of different viewpoints and modalities. In the experiments using a dataset1newly collected for this task, our proposed method using MultiTrans improved the event detection performance and outperformed comparatives.
Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada
ICASSP4
2022 Introducing Auxiliary Text Query-modifier to Content-based Audio Retrieval
Daiki Takeuchi, Yasunori Ohishi, Daisuke Niizumi, Noboru Harada, Kunio Kashino
INTERSPEECH4
2022 ConceptBeam: Concept Driven Target Speech Extraction
abstract
We propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speaker in a mixture. Typical approaches have been exploiting properties of audio signals, such as harmonic structure and direction of arrival. In contrast, ConceptBeam tackles the problem with semantic clues. Specifically, we extract the speech of speakers speaking about a concept, i.e., a topic of interest, using a concept specifier such as an image or speech. Solving this novel problem would open the door to innovative applications such as listening systems that focus on a particular topic discussed in a conversation. Unlike keywords, concepts are abstract notions, making it challenging to directly represent a target concept. In our scheme, a concept is encoded as a semantic embedding by mapping the concept specifier to a shared embedding space. This modality-independent space can be built by means of deep metric learning using paired data consisting of images and their spoken captions. We use it to bridge modality-dependent information, i.e., the speech segments in the mixture, and the specified, modality-independent concept. As a proof of our scheme, we performed experiments using a set of images associated with spoken captions. That is, we generated speech mixtures from these spoken captions and used the images or speech signals as the concept specifiers. We then extracted the target speech using the acoustic characteristics of the identified segments. We compare ConceptBeam with two methods: one based on keywords obtained from recognition systems and another based on sound source separation. We show that ConceptBeam clearly outperforms the baseline methods and effectively extracts speech based on the semantic representation.
Yasunori Ohishi, Marc Delcroix, Tsubasa Ochiai, Shoko Araki, Daiki Takeuchi, Daisuke Niizumi, Akisato Kimura, Noboru Harada, Kunio Kashino
ACM Multimedia8
2021 An Extension of Sparse Audio Declipper to Multiple Measurement Vectors
abstract
This paper proposes formulating declipping as a constrained multiple measurement vector (MMV) optimization problem that has a ${\ell _{2,0}}$ group norm as its cost function for further improving the state-of-the-art declipping method SParse Audio DEclipper (SPADE). This paper shows that the MMV optimization problem can be solved by extending the steps of alternating direction method of multipliers (ADMM) in the SPADE algorithm. The proposed method improved the signal-to-distortion ratio and reduced the cepstral distance for all clipping levels in a numerical simulation.
Satoru Emura, Noboru Harada
ICASSP2
2021 Asynchronous Decentralized Optimization With Implicit Stochastic Variance Reduction
abstract
A novel asynchronous decentralized optimization method that follows Stochastic Variance Reduction (SVR) is proposed. Average consensus algorithms, such as Decentralized Stochastic Gradient Descent (DSGD), facilitate distributed training of machine learning models. However, the gradient will drift within the local nodes due to statistical heterogeneity of the subsets of data residing on the nodes and long communication intervals. To overcome the drift problem, (i) Gradient Tracking-SVR (GT-SVR) integrates SVR into DSGD and (ii) Edge-Consensus Learning (ECL) solves a model constrained minimization problem using a primal-dual formalism. In this paper, we reformulate the update procedure of ECL such that it implicitly includes the gradient modification of SVR by optimally selecting a constraint-strength control parameter. Through convergence analysis and experiments, we confirmed that the proposed ECL with Implicit SVR (ECL-ISVR) is stable and approximately reaches the reference performance obtained with computation on a single-node using full data set.
Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn, Noboru Harada, Hiroshi Sawada, Akinori Fujino
ICML4
2021 BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation
abstract
Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio representation learning approach. We propose learning general-purpose audio representation from a single audio segment without expecting relationships between different time segments of audio samples. To implement this principle, we introduce Bootstrap Your Own Latent (BYOL) for Audio (BYOL-A, pronounced “viola”), an audio self-supervised learning method based on BYOL for learning general-purpose audio representation. Unlike most previous audio self-supervised learning methods that rely on agreement of vicinity audio segments or disagreement of remote ones, BYOL-A creates contrasts in an augmented audio segment pair derived from a single audio segment. With a combination of normalization and augmentation techniques, BYOL-A achieves state-of-the-art results in various downstream tasks. Extensive ablation studies also clarified the contribution of each component and their combinations.
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino
IJCNN4
2020 A Frequency-Domain BSS Method Based on ℓ1 Norm, Unitary Constraint, and Cayley Transform
abstract
We propose a frequency-domain blind source separation method that uses (a) the ℓ1norm of orthonormal vectors of estimated source signals as a sparsity measure and (b) Cayley transform for optimizing the objective function under the unitary constraint in the Riemannian geometry approach. The orthonormal vectors of estimated source signals, obtained by the sphering of observed mixed signals and the unitary constraint on the separation filters, enables us to use the ℓ1norm properly as a sparsity measure. The Cayley transform enables us to handle the geometrical aspects of the unitary constraint efficiently. According to the simulation of a two-channel case, the proposed method achieved a 20-dB improvement in the source-to-interference ratio in a room with a reverberation time of T60= 300ms.
Satoru Emura, Hiroshi Sawada, Shoko Araki, Noboru Harada
ICASSP4
2020 SPIDERnet: Attention Network For One-Shot Anomaly Detection In Sounds
abstract
We propose a similarity function for one-shot anomaly detection in sounds (ADS) called SPecific anomaly IDentifiER network (SPIDERnet). In ADS systems, since overlooking an anomaly may result in serious incidents, we need to update such systems using an (often only one) overlooked anomalous sample. A previous study proposed the use of memory-based one-shot learning. A problem with this previous method is that it can detect only short anomalous sounds such as collision sounds because its similarity function is based on a naive mean-squared-error between the input and memorized spectrogram. To detect various anomalous sounds, SPIDERnet consists of (i) a neural network-based feature extractor for measuring similarity in embedded space and (ii) attention mechanisms for absorbing time-frequency stretching. Experimental results on two public datasets indicate that SPIDERnet outperforms conventional methods and robustly detects various anomalous sounds.
Yuma Koizumi, Masahiro Yasuda, Shin Murata, Shoichiro Saito, Hisashi Uematsu, Noboru Harada
ICASSP6
2020 Subjective Quality Estimation Using PESQ For Hands-Free Terminals
abstract
Previous reports have mentioned the possibility that subjective quality of the echo-suppressed speech signal can be estimated based on perceptual evaluation of speech quality (PESQ), but there are few experimental results. We propose third-party listening and conversational test procedures to assess whether PESQ can be used for predicting the subjective quality of an acoustic echo canceler. In the proposed third-party listening test procedure, near-end and far-end signals are presented separately in the left and right channels of stereo playback and differential category rating evaluation is applied to those stimuli for obtaining differential mean opinion scores. In the proposed conversational test procedure, impaired and non-impaired reference signals are recorded during a conversation to make PESQ processing possible. Experimental results indicate that there is a strong correlation between PESQ and subjective scores.
Sachiko Kurihara, Masahiro Fukui, Suehiro Shimauchi, Noboru Harada
ICASSP4
2020 Phase Reconstruction Based On Recurrent Phase Unwrapping With Deep Neural Networks
abstract
Phase reconstruction, which estimates phase from a given amplitude spectrogram, is an active research field in acoustical signal processing with many applications including audio synthesis. To take advantage of rich knowledge from data, several studies presented deep neural network (DNN)–based phase reconstruction methods. However, the training of a DNN for phase reconstruction is not an easy task because phase is sensitive to the shift of a waveform. To overcome this problem, we propose a DNN-based two-stage phase reconstruction method. In the proposed method, DNNs estimate phase derivatives instead of phase itself, which allows us to avoid the sensitivity problem. Then, phase is recursively estimated based on the estimated derivatives, which is named recurrent phase unwrapping (RPU). The experimental results confirm that the proposed method outperformed the direct phase estimation by a DNN.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP5
2020 Real-Time Speech Enhancement Using Equilibriated RNN
abstract
We propose a speech enhancement method using a causal deep neural network (DNN) for real-time applications. DNN has been widely used for estimating a time-frequency (T-F) mask which enhances a speech signal. One popular DNN structure for that is a recurrent neural network (RNN) owing to its capability of effectively modelling time-sequential data like speech. In particular, the long short-term memory (LSTM) is often used to alleviate the vanishing/exploding gradient problem which makes the training of an RNN difficult. However, the number of parameters of LSTM is increased as the price of mitigating the difficulty of training, which requires more computational resources. For real-time speech enhancement, it is preferable to use a smaller network without losing the performance. In this paper, we propose to use the equilibriated recurrent neural network (ERNN) for avoiding the vanishing/exploding gradient problem without increasing the number of parameters. The proposed structure is causal, which requires only the information from the past, in order to apply it in real-time. Compared to the uni- and bi-directional LSTM networks, the proposed method achieved the similar performance with much fewer parameters.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP5
2020 Invertible DNN-Based Nonlinear Time-Frequency Transform for Speech Enhancement
abstract
We propose an end-to-end speech enhancement method with trainable time-frequency (T-F) transform based on invertible deep neural network (DNN). The resent development of speech enhancement is brought by using DNN. The ordinary DNN-based speech enhancement employs T-F transform, typically the short-time Fourier transform (STFT), and estimates a T-F mask using DNN. On the other hand, some methods have considered end-to-end networks which directly estimate the enhanced signals without T-F transform. While end-to-end methods have shown promising results, they are black boxes and hard to understand. Therefore, some end-to-end methods used a DNN to learn the linear T-F transform which is much easier to understand. However, the learned transform may not have a property important for ordinary signal processing. In this paper, as the important property of the T-F transform, perfect reconstruction is considered. An invertible nonlinear T-F transform is constructed by DNNs and learned from data so that the obtained transform is perfectly reconstructing filterbank.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP5
2020 Crossmodal Sound Retrieval Based on Specific Target Co-Occurrence Denoted with Weak Labels
Masahiro Yasuda, Yasunori Ohishi, Yuma Koizumi, Noboru Harada
INTERSPEECH4
2020 Edge-consensus Learning: Deep Learning on P2P Networks with Nonhomogeneous Data
abstract
An effective Deep Neural Network (DNN) optimization algorithm that can use decentralized data sets over a peer-to-peer (P2P) network is proposed. In applications such as medical data analysis, the aggregation of data in one location may not be possible due to privacy issues. Hence, we formulate an algorithm to reach a global DNN model that does not require transmission of data among nodes. An existing solution for this issue is gossip stochastic gradient descend (SGD), which updates by averaging node models over a P2P network. However, in practical situations where the data are statistically heterogeneous across the nodes and/or where communication is asynchronous, gossip SGD often gets trapped in local minimum since the model gradients are noticeably different. To overcome this issue, we solve a linearly constrained DNN cost minimization problem, which results in variable update rules that restrict differences among all node models. Our approach can be based on the Primal-Dual Method of Multipliers (PDMM) or the Alternating Direction Method of Multiplier (ADMM), but the cost function is linearized to be suitable for deep learning. It facilitates asynchronous communication. The results of our numerical experiments using CIFAR-10 indicate that the proposed algorithms converge to a global recognition model even though statistically heterogeneous data sets are placed on the nodes.
Kenta Niwa, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn
KDD2
2020 Multi-Delay Sparse Approach to Residual Crosstalk Reduction for Blind Source Separation
abstract
For reducing residual crosstalk in the output of blind source separation, we propose a frequency-domain post-filtering method that uses a multi-delay model of complex-valued residual crosstalk and sparsifies the estimates of the source signals. We formulate the reduction of residual crosstalk as an optimization problem using ℓ1norm and solve it using the alternating direction method of multiplier. The proposed method improved the source-to-interference ratio from 17.8 to 20.5 dB and the source-to-distortion ratio from 10.2 to 11.3 dB when it was combined with a brute-force solver of FastICA and reverberation time T60was 300 ms.
Satoru Emura, Hiroshi Sawada, Shoko Araki, Noboru Harada
IEEE Signal Process. Lett.4
2020 Microphone Array Wiener Post Filtering Using Monotone Operator Splitting
abstract
For array-based acoustic source enhancement, variants of multi-channel Wiener filters are commonly used. The approach includes a Wiener post-filter that requires the simultaneous estimation of the power spectral density (PSD) of the target source and of noise sources for each time-frame. Conventional methods generally do not exploit prior knowledge, such as sparsity of the source, in solving this simultaneous estimation problem. We show that, for common scenarios, the simultaneous PSD estimation with consideration of prior knowledge can be formulated as a convex optimization problem with linear constraints. We use monotone operator splitting (MOS) to solve the constrained optimization problem. Our experiments confirm that the proposed method improves the accuracy of the noise PSD estimation, and that the resulting enhanced target signal is of higher quality.
Kenta Niwa, Hironobu Chiba, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 A Two-class Hyper-spherical Autoencoder for Supervised Anomaly Detection
abstract
Supervised anomaly detection has been a tough problem due to its necessity of special handling of unseen anomalies. In this paper, we present a heuristic implementation of variational auto-encoder with von-Mises Fisher prior applied to a supervised anomaly detector. The closed latent space like sphere is suitable for detecting unseen anomalies because we have a possibility to "fill" the space with seen training samples. If it ideally works, the reconstruction error will be high for all unseen anomalies. Experiments show that our model can separate normal and anomaly samples in the spherical latent space. It is also shown that he proposed model improves the performance for seen anomalies without degrading the performance for unseen anomalies.
Yuta Kawachi, Yuma Koizumi, Shin Murata, Noboru Harada
ICASSP4
2019 Trainable Adaptive Window Switching for Speech Enhancement
abstract
This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask processing in the short-time Fourier transform (STFT)-domain is a typical speech enhancement method. To recover the target signal precisely, DNN-based short-time frequency transforms have recently been investigated and used instead of the STFT. However, since such a fixed-resolution short-time frequency transform method has a T-F resolution problem based on the uncertainty principle, not only the short-time frequency transform but also the length of the windowing function should be optimized. To overcome this problem, we incorporate AWS into the speech enhancement procedure, and the windowing function of each time-frame is manipulated using a DNN depending on the input signal. We confirmed that the proposed method achieved a higher signal-to-distortion ratio than conventional speech enhancement methods in fixed-resolution frequency domains.
Yuma Koizumi, Noboru Harada, Youichi Haneda
ICASSP2
2019 SNIPER: Few-shot Learning for Anomaly Detection to Minimize False-negative Rate with Ensured True-positive Rate
abstract
In anomaly detection systems, overlooking anomalies may result in serious incidents. Thus, when a system overlooks an anomaly, we need to update the system to never overlook the observed type of anomalies twice. There are roughly two possible approaches to solve this problem; re-training the whole system using all training data, or cascading a new specific detector for the overlooked anomaly. The first approach is the most effective solution; however, a huge computational cost and an amount of anomalous training data are required to re-train the system when it consists of a deep-learning-based anomaly detector. We focused on the latter approach and propose a training method for a cascaded specific anomaly detector using few-shot (just 1 to 3) samples. To suppress the false-negative rate of the overlooked anomaly, the proposed method works to decrease the false-positive rate under the constraint of true-positive rate equaling 1. Experimental results show that the proposed method outperformed conventional cross-entropy-based few-shot learning methods.
Yuma Koizumi, Shin Murata, Noboru Harada, Shoichiro Saito, Hisashi Uematsu
ICASSP3
2019 Deep Griffin-Lim Iteration
abstract
This paper presents a novel phase reconstruction method (only from a given amplitude spectrogram) by combining a signal-processing-based approach and a deep neural network (DNN). To retrieve a time-domain signal from its amplitude spectrogram, the corresponding phase is required. One of the popular phase reconstruction methods is the Griffin-Lim algorithm (GLA), which is based on the redundancy of the short-time Fourier transform. However, GLA often involves many iterations and produces low-quality signals owing to the lack of prior knowledge of the target signal. In order to address these issues, in this study, we propose an architecture which stacks a sub-block including two GLA-inspired fixed layers and a DNN. The number of stacked sub-blocks is adjustable, and we can trade the performance and computational load based on requirements of applications. The effectiveness of the proposed method is investigated by reconstructing phases from amplitude spectrograms of speeches.
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP5
2019 Non-negative Matrix Factorization Using Bregman Monotone Operator Splitting
abstract
A non-negative matrix factorization (NMF) algorithm based on Bregman monotone operator splitting (B-MOS) is proposed. Several commonly used NMF algorithms, such as the multiplicative update method, are often used in source separation for speech and image signals. To improve the convergence rate in the tail, applying the alternating direction method of multipliers (ADMM) is reported to be effective. However, a fixed step-size parameter has to be carefully chosen for fast and stable convergence. Our main idea to overcome this issue is to adaptively modify the variable space metric so that it matches the cost convexity. Besides this, selecting an appropriate MOS (e.g., Peaceman-Rachford splitting) instead of the Douglas-Rachford splitting used in the ADMM may effectively improve the convergence rate further. To realize these ideas w.r.t. adaptive metric modification and appropriate operator splitting selection, we apply B-MOS to the NMF problem and obtain a new NMF solver in this paper. Results of numerical experiments demonstrate that the proposed NMF solver with B-MOS improved the convergence rate in the tail.
Kenta Niwa, Noboru Harada
ICASSP2
2019 Function Designable Beamformer Based on Probabilistic Assumptions on Filter and Its Auxiliary Variables
abstract
We propose a novel beamformer design method that exploits probabilistic assumptions on auxiliary variables derived from filters and observed signals. Many conventional beamformer design methods can be understood in the context of optimization problems for some probabilistic cost functions. However, the class of cost functions used with these methods is quite limited to reflect multiple pieces of information and our demands on the filter, such as the sparsity assumption of the source signals and the low-latency constraint of the filter. We propose a method to design cost functions that incorporate multiple probabilistic assumptions. The assumptions are expressed as the sum of many convex terms, and every term has different auxiliary variables that are linearly constrained. Such cost functions can be optimized by iteratively optimizing with regard to each term alternately. This method enables us to more arbitrarily tune the beamformer. We conducted numerical simulations showing that our method effectively improves the performance from multiple perspectives.
Ryotaro Sato, Kenta Niwa, Noboru Harada
ICASSP3
2019 Data-driven Design of Perfect Reconstruction Filterbank for DNN-based Sound Source Enhancement
abstract
We propose a data-driven design method of perfect-reconstruction filterbank (PRFB) for sound-source enhancement (SSE) based on deep neural network (DNN). DNNs have been used to estimate a time-frequency (T-F) mask in the short-time Fourier transform (STFT) domain. Their training is more stable when a simple cost function as mean-squared error (MSE) is utilized comparing to some advanced cost such as objective sound quality assessments. However, such a simple cost function inherits strong assumptions on the statistics of the target and/or noise which is often not satisfied, and the mismatch of assumption results in degraded performance. In this paper, we propose to design the frequency scale of PRFB from training data so that the assumption on MSE is satisfied. For designing the frequency scale, the warped filterbank frame (WFBF) is considered as PRFB. The frequency characteristic of learned WFBF was in between STFT and the wavelet transform, and its effectiveness was confirmed by comparison with a standard STFT-based DNN whose input feature is compressed into the mel scale.
Daiki Takeuchi, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, Noboru Harada
ICASSP5
2019 AdaFlow: Domain-adaptive Density Estimator with Application to Anomaly Detection and Unpaired Cross-domain Translation
abstract
We tackle unsupervised anomaly detection (UAD), a problem of detecting data that significantly differ from normal data. UAD is typically solved by using density estimation. Recently, deep neural network (DNN)-based density estimators, such as Normalizing Flows, have been attracting attention. However, one of their drawbacks is the difficulty in adapting them to the change in the normal data's distribution. To address this difficulty, we propose AdaFlow, a new DNN-based density estimator that can be easily adapted to the change of the distribution. AdaFlow is a unified model of a Normalizing Flow and Adaptive Batch-Normalizations, a module that enables DNNs to adapt to new distributions. AdaFlow can be adapted to a new distribution by just conducting forward propagation once per sample; hence, it can be used on devices that have limited computational resources. We have confirmed the effectiveness of the proposed model through an anomaly detection in a sound task. We also propose a method of applying AdaFlow to the unpaired cross-domain translation problem, in which one has to train a cross-domain translation model with only unpaired samples. We have confirmed that our model can be used for the cross-domain translation problem through experiments on image datasets.
Masataka Yamaguchi, Yuma Koizumi, Noboru Harada
ICASSP3
2019 Unsupervised Detection of Anomalous Sound Based on Deep Learning and the Neyman-Pearson Lemma
abstract
This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of the unsupervised-ADS is to detect unknown anomalous sounds without training data of anomalous sounds. The use of an AE as a normal model is a state-of-the-art technique for the unsupervised-ADS. To decrease the false positive rate (FPR), the AE is trained to minimize the reconstruction error of normal sounds, and the anomaly score is calculated as the reconstruction error of the observed sound. Unfortunately, since this training procedure does not take into account the anomaly score for anomalous sounds, the true positive rate (TPR) does not necessarily increase. In this study, we define an objective function based on the Neyman-Pearson lemma by considering the ADS as a statistical hypothesis test. The proposed objective function trains the AE to maximize the TPR under an arbitrary low FPR condition. To calculate the TPR in the objective function, we consider that the set of anomalous sounds is the complementary set of normal sounds and simulate anomalous sounds by using a rejection sampling algorithm. Through experiments using synthetic data, we found that the proposed method improved the performance measures of the ADS under low FPR conditions. In addition, we confirmed that the proposed method could detect anomalous sounds in real environments.
Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, Yuta Kawachi, Noboru Harada
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 Sound Field Decomposition Using SPICE Decomposition
abstract
We propose a method to estimate a reverberant sound field by compressed sensing approach without using the hyper parameter for controlling sparsity. This method first applies sparse iterative covariance-based approach (SPICE) to the power spectral density matrix of the microphone signals and obtains the estimate of sensor noise power without using a hyper parameter. From this estimate, a convex optimization problem for a sparse solution is obtained that relates the microphone signals, plane-wave expansion coefficients, and the sensor noise power. The plane-wave expansion coefficients are obtained as the solution of this convex optimization problem. The sound field can be estimated from the plane-wave expansion coefficients.
Satoru Emura, Noboru Harada
ICASSP2
2018 Complementary Set Variational Autoencoder for Supervised Anomaly Detection
abstract
Anomalies have broad patterns corresponding to their causes. In industry, anomalies are typically observed as equipment failures. Anomaly detection aims to detect such failures as anomalies. Although this is usually a binary classification task, the potential existence of unseen (unknown) failures makes this task difficult. Conventional supervised approaches are suitable for detecting seen anomalies but not for unseen anomalies. Although, unsupervised neural networks for anomaly detection now detect unseen anomalies well, they cannot utilize anomalous data for detecting seen anomalies even if some data have been made available. Thus, providing an anomaly detector that finds both seen and unseen anomalies well is still a tough problem. In this paper, we introduce a novel probabilistic representation of anomalies to solve this problem. The proposed model defines the normal and anomaly distributions using the analogy between a set and the complementary set. We applied these distributions to an unsupervised variational autoencoder (VAE)-based method and turned it into a supervised VAE-based method. We tested the proposed method with well-known data and real industrial data to show that the proposed method detects seen anomalies better than the conventional unsupervised method without degrading the detection performance for unseen anomalies.
Yuta Kawachi, Yuma Koizumi, Noboru Harada
ICASSP3
2018 End-to-End Sound Source Enhancement Using Deep Neural Network in the Modified Discrete Cosine Transform Domain
abstract
This paper presents an end-to-end deep neural network (DNN)-based source enhancement on the basis of a time-frequency (T-F) mask processing in the modified discrete cosine transform (MDCT)-domain. To retrieve the target signal perfectly in the discrete Fourier transform (DFT)-domain, both amplitude and phase of the spectrum need to be manipulated. However, since it is difficult to deal with complex values by neural network straightforward way, a real-valued T-F mask is commonly estimated and only amplitude spectrum is manipulated. In this study, we use the MDCT instead of the DFT and estimate real-valued T-F masks in the MDCT-domain. The perfect retrieval can be achieved by manipulating only the real-valued MDCT-spectra. To reduce time-domain aliasing arises from manipulating the MDCT spectrum, we build an end-to-end DNN-based source enhancement using T-F mask and train the DNN to minimize an objective function defined in the time-domain. In experiments using several kinds of objective sound quality scores, we observed that the scores were significantly improved.
Yuma Koizumi, Noboru Harada, Youichi Haneda, Yusuke Hioka, Kazunori Kobayashi
ICASSP2
2018 Distortionless Beamforming Optimized With ℓ1-Norm Minimization
abstract
We propose beamforming method that minimizes the ℓ1norm of a beamformer output vector under the same distortionless constraint as that of the conventional minimum power distortionless response (MPDR) beamformer. Using the ℓ1norm makes the beamformer output sparse. This leads to reducing the residual elements of the interference signal. In addition, the sensitivity of the proposed beamformer can be controlled by adding a norm constraint as in the MPDR beamformer. The proposed method improved the signal-to-interference-noise ratio by 7 dB from that of the MPDR beamformer for reverberation time T60= 300 ms in a simulation.
Satoru Emura, Shoko Araki, Tomohiro Nakatani, Noboru Harada
IEEE Signal Process. Lett.4
2018 Optimal Golomb-Rice Code Extension for Lossless Coding of Low-Entropy Exponentially Distributed Sources
abstract
This paper presents an extension of GolombRice (GR) code for coding low-entropy sources, which the gap between their entropy and the conventional GR code length gets larger. We mention here the following four facts related to the proposed code, extended-domain GR (XDGR) code: it is represented by multiple code trees, based on the idea of almost instantaneous fixed-to-variable length codes, with its algorithm being a generalization of unary coding; its structure naturally contains run-length coding; the gap between the entropy and its average code length is theoretically guaranteed to be asymptotically negligible as the entropy of the exponentially distributed sources tends to zero; and its coding parameter, corresponding to the negative-domain Rice parameter of GR code, can be estimated from the input source-symbol sequence. Experimental evaluations are also presented supporting the theorems. The proposed XDGR code, having simple algorithm and high compression performance, is expected to be used for many coding applications, which deals with exponentially distributed sources at low bit rates.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
IEEE Trans. Inf. Theory3
2015 Standardization of the new 3GPP EVS codec
abstract
A new codec for Enhanced Voice Services (EVS), the successor of the current mobile HD voice codec AMR-WB, was standardized by the 3rd Generation Partnership Project (3GPP) in September 2014. The EVS codec addresses 3GPP's needs for cutting-edge technology enabling operation of 3GPP mobile communication systems in the most competitive means in terms of communication quality and efficiency. This paper provides an in-depth insight into 3GPP's rigorous and transparent processes that made it possible for the mobile industry, with its many competing players, to successfully develop and standardize a codec in an open, fair and constructive process. This paper also enables an understanding of this achievement by providing an overview of the EVS codec technology, the standard specifications, and the performance of the codec that will elevate HD voice services to the next quality level.
Stefan Bruhn, Harald Pobloth, Markus Schnell, Bernhard Grill, Jon Gibbs, Lei Miao 0004, Kari Järvinen, Lasse Laaksonen, Noboru Harada, Nobuhiko Naka, Stéphane Ragot, Stéphane Proust, Takako Sanda, Imre Varga, Craig Greer, Milan Jelinek, Minjie Xie, Paolo Usai
ICASSP9
2015 Resolution Warped Spectral Representation for Low-Delay and Low-Bit-Rate Audio Coder
abstract
We have devised a high-quality frequency-domain audio coder based on the state-of-the-art monaural wide-band coder aiming at its use in low-delay and low-bit-rate conditions. The coder efficiently represents frequency spectral envelopes of the target signals with low computational complexity using optimally prepared non-negative sparse matrices. The experimental results reveal that this representation has positive effects on the objective and subjective quality of the coder resulting in the comparable quality to the same bit rate of 3GPP Extended Adaptive Multi-Rate WideBand ( AMR-WB+), a coder which permits more than four times longer delay compared with the proposed coder. Consequently, this coder is suitable for applications in mobile communications, which require low delay and low complexity.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Optimal Coding of Generalized-Gaussian-Distributed Frequency Spectra for Low-Delay Audio Coder With Powered All-Pole Spectrum Estimation
abstract
We present an optimal coding scheme that parameterizes the maximum-likelihood estimate of variance for frequency spectra belonging to the generalized Gaussian distribution, the distribution covering the Laplacian and the Gaussian. By slightly modifying the all-pole model of the conventional linear prediction (LP), we can estimate the variance with the same method as in LP, which has low computational costs. Experimental results show that incorporating the coding scheme in a state-of-the-art wide-band audio coder enhances its objective and subjective quality in a low-bit-rate and low-delay situation by increasing the compression efficiency. Thus, this coding scheme will be useful in applications like mobile communications, which requires highly efficient compression.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.3
2010 Lossless Compression of Mapped Domain Linear Prediction Residual for ITU-T Recommendation G.711.0
abstract
Summary form only given. The ITU-T Rec. G.711 is widely used for narrowband telephony applications, including PSTN/GSTN and packet-based network applications such as VoIP, and has been used for many decades because of its proven voice quality, ubiquity, and utility. ITU has just established a lossless coding technology for G.711 encoded payloads, ITU-T Rec. G.711.0-Lossless compression of G.711 pulse code modulation. This paper introduces some coding technologies proposed and applied to the G.711.0 codec, such as Plus-Minus zero mapping for the mapped domain linear predictive coding and escaped-Huffman coding combined with adaptive recursive Rice coding for lossless compression of the prediction residual. The proposed and conventional algorithms were implemented and the coding performances of the codec were evaluated using large speech corpus in various conditions in terms of the computational complexity/compression performance trade-off. Since the distribution of the prediction residual signal sometimes does not follow the expected Laplacian PDF model of Rice coding, E-Huffman coding can improve the compression performance for relatively smaller Rice quotients of residual and Recursive Rice coding can improve the compression performance for relatively larger Rice quotients of residual. It is shown that the PM zero mapping improves the compression performance by 0.2% for ?-law input. The E-Huffman coding combined with adaptive recursive Rice coding improves the compression by 0.16% while the complexity was increased by 0.046 WMOPS for the encoder/decoder pair, averaged for all test conditions, compare to the conventional Rice coding scheme. The worst-case complexity was increased by 0.079 WMOPS. Average computational complexity is 1.071 WMOPS for the encoder/decoder pair and the worst-case complexity is 1.667 WMOPS in total.
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya
DCC1
2010 Low-Complexity PARCOR Coefficient Quantizer and Prediction Order Estimator for G.711.0 (Lossless Speech Coding)
abstract
This paper presents two low-complexity tools used for the new ITU-T recommendation G.711.0, which is the standard for lossless compression of G.711 (A-law/Mu-law logarithmic PCM) speech data. One is an algorithm for quantizing the PARCOR/reflection coefficients and the other is an estimation method for the optimal prediction order. Both tools are based on a criterion that minimizes the entropy of the prediction residual signals and can be implemented in a fixed-point low-complexity algorithm. G.711.0 with the developed practical tools will be widely used everywhere because it can losslessly reduce the data rate of G.711, the prevailing speech-coding technology.
Yutaka Kamamoto, Takehiro Moriya, Noboru Harada
DCC3
2010 Enhanced Lossless Coding Tools of LPC Residual for ITU-T G.711.0
abstract
Motivated by the rapid increase of VoIP services with G.711 for telephone speech, a new ITU-T recommendation, G.711.0 (frame-wise stateless lossless compression scheme for G.711 log PCM symbols), has been standardized. The standard scheme has several coding parts, each of which is adaptively selected depending on the characteristics of the input. Among them, the mapped domain prediction part is the one most frequently activated for normal speech signals. This part consists of linear prediction in the mapped domain and variable length coding of the prediction residual. It is useful for log-compressed/expanded signal, such as ITU-T G.711. This paper describes three newly devised enhancement tools for the coding of prediction residual signals: progressive order prediction, quantized prediction order, and adaptive and sub-frame base coding for separation parameters. The design criterion is the maximization of the averaged FoM (figure of merit) over frame lengths of 40, 80, 160, 240, and 320 samples. The first tool, progressive order prediction associated with the adaptive modification of the separation parameter for the first and second samples, enhances the compression ratio by 0.5 % with a negligible increase of the complexity. The second tool, quantized prediction order, improves the compression ratio by 0.2 % with even reduced complexity. The third tool, sub-frame base adaptive coding of separation parameters, gives a 0.2 % improvement in the compression ratio with comparable complexity. All three schemes are consistently and independently effective for improving the compression ratio, although the amount of improvement with each tool is small. At the same time, none of the tools have any significant impact on computational complexity. Therefore, all the devised tools improve the FoM and have been adopted in the mapped domain prediction part of the ITU-T G.711.0 standard.
Takehiro Moriya, Yutaka Kamamoto, Noboru Harada
DCC3
2010 Escaped-Huffman and adaptive recursive rice coding for lossless compression of the mapped domain linear prediction residual
abstract
ITU-T Recommendation G.711.0 has just been established. It defines a lossless and stateless compression for G.711 packet payloads (for both A-law and μ-law). This paper introduces some coding technologies proposed and applied to the G.711.0 codec, such as Plus-Minus zero mapping for the mapped domain linear predictive coding and escaped-Huffman coding combined with adaptive recursive Rice coding for lossless compression of the prediction residual. Performance test results for those coding tools are shown in comparison with the results for the conventional technology. The performance is measured based on the figure of merit (FoM), which is a function of the trade-off between compression performance and computational complexity. The proposed tools improve the compression performance by 0.16% in total while keeping the computational complexity of encoder/decoder pair low (about 1.0 WMOPS in average and 1.667 WMOPS in the worst-case).
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya
ICASSP1
2010 Emerging ITU-T standard G.711.0 - lossless compression of G.711 pulse code modulation
abstract
The ITU-T Recommendation G.711 is the benchmark standard for narrowband telephony. It has been successful for many decades because of its proven voice quality, ubiquity and utility. A new ITU-T recommendation, denoted G.711.0, has been recently established defining a lossless compression for G.711 packet payloads typically found in IP networks. This paper presents a brief overview of technologies employed within the G.711.0 standard and summarizes the compression and complexity results. It is shown that G.711.0 provides greater than 50% average compression in typical service provider environments while keeping low computational complexity for the encoder/decoder pair (1.0 WMOPS average, <;1.7 WMOPS worst case) and low memory footprint (about 5k octets RAM, 5.7k octets ROM, and 3.6k program memory measured in number of basic operators).
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya, Yusuke Hiwasaki, Michael A. Ramalho, Lorin Netsch, Jacek Stachurski, Lei Miao 0004, Hervé Taddei, Fengyan Qi
ICASSP1
2010 Low-complexity PARCOR coefficient quantizer and prediction order estimator for lossless speech coding
abstract
This paper describes two low-complexity tools used for the new ITU-T recommendation G.711.0, the lossless coding of G.711 (A-law/μ-law logarithmic PCM) speech data. One is an algorithm for quantizing the PARCOR/reflection coefficients and the other is an estimation method for the optimal prediction order. Both tools are based on a criterion that minimizes the entropy of the prediction residual signals and can be implemented in a fixed-point low-complexity algorithm. G.711.0 with the developed practical tools will be widely used everywhere because it can losslessly reduce the data rate of G.711, the prevailing speech-coding technology.
Yutaka Kamamoto, Takehiro Moriya, Noboru Harada
ICASSP3
2010 Enhanced lossless coding tools for prediction residual
abstract
Three elementary coding tools - a progressive order prediction tool, quantized order prediction tool, and adaptive and sub-frame base coding tool for separation parameters - have been devised to enhance the compression performance of the prediction residual. These are intended for the lossless coding of G.711 log PCM symbols used in packet-based network application such as VoIP. All tools are shown to be effective for reducing the average code length without any significant increase of computational complexity. As a result, all have been adopted in the mapped domain predictive coding part of the ITU-T G.711.0 standard.
Takehiro Moriya, Yutaka Kamamoto, Noboru Harada
ICASSP3
2008 Interchannel dependency analysis of biomedical signals for efficient lossless compression by MPEG-4 ALS
abstract
This paper describes a new search algorithm that quickly finds interchannel relationships between a coding channel and a reference channel in the multichannel coding tool of the MPEG- 4 Audio Lossless Coding (ALS) international standard. The algorithm has tree structure and can reduce data size with significantly smaller computation load than that of the conventional one. The devised method is based on a restricted greedy algorithm. It chooses the most efficient branch which does not make any loops in the existing path. The results of comprehensive evaluations show that this method maintains the compression performance (compression to around 1/3) and performs 1000 times as fast as the conventional method for the 512-channel magnetoencephalography signals. This algorithm enables practical lossless compression of biomedical data by the ALS, and at the same time, opens the way to a new multichannel analysis tool that may be used for purposes other than compression. The continual maintenance of this standard will make it possible to perfectly reconstruct encoded files even 100 years from now.
Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
ICASSP2
2008 Lossless compression of biomedical signals by MPEG-4 ALS with enhanced encoding tools
abstract
Enhanced encoding tools for MPEG-4 audio lossless coding (ALS) international standard were developed, with the goal of improving compression performance of time-series biomedical data. The multichannel coding (MCC) tool of this standard exploits interchannel redundancies to reduce the bit rate. To improve compression performance with the MCC, we have devised a multichannel linear prediction tool, which achieves around a 0.1% better compression ratio than that of the conventional method. We have also developed an interchannel dependency analysis tool, which performs about 1000 times faster than the conventional one. By combining these tools, biomedical signals are losslessly compressed to about 1/3 in a practical computational load. Compressed biomedical data will be decoded even 100 years from now, because the bitstream still remains compliant with the MPEG standard.
Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
MMSP2