EDBT 2026 Demo / reviewers in the wild / expert
Ryosuke Sawata
dblp:168/3055
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-3230-4335ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hearing Anything AnywhereabstractRecent years have seen immense progress in 3D computer vision and computer graphics, with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However, alongside immersive visual experiences, immersive auditory experiences are equally vital to our holistic perception of an environment. In this paper, we aim to reconstruct the spatial acoustic characteristics of an arbitrary environment given only a sparse set of (roughly 12) room impulse response (RIR) recordings and a planar reconstruction of the scene, a setup that is easily achievable by ordinary users. To this end, we introduce DIFFRIR, a differentiable RIR rendering framework with interpretable parametric models of salient acoustic features of the scene, including sound source directivity and surface reflectivity. This allows us to synthesize novel auditory experiences through the space with any source audio. To evaluate our method, we collect a dataset of RIR recordings and music in four diverse, real environments. We show that our model outperforms state-of-the-art baselines on rendering monaural and binaural RIRs and music at unseen locations, and learns physically interpretable parameters characterizing acoustic properties of the sound source and surfaces in the scene. Mason L. Wang, Ryosuke Sawata, Samuel Clarke, Ruohan Gao, Shangzhe Wu, Jiajun Wu 0001 |
CVPR | 2 |
| 2023 | Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining CapabilityabstractIn this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our model to generate realistic looking piano rolls from pure Gaussian noise conditioned on spectrograms. This new AMT formulation enables DiffRoll to transcribe, generate and even inpaint music. Due to the classifier-free nature, DiffRoll is also able to be trained on unpaired datasets where only piano rolls are available. Our experiments show that DiffRoll outperforms its discriminative counterpart by 19 percentage points (ppt.) and our ablation studies also indicate that it outperforms similar existing methods by 4.8 ppt.Source code and demonstration are available at https://sony.github.io/DiffRoll/. Kin Wai Cheuk, Ryosuke Sawata, Toshimitsu Uesaka, Naoki Murata, Naoya Takahashi, Shusuke Takahashi, Dorien Herremans, Yuki Mitsufuji |
ICASSP | 2 |
| 2023 | Class-Aware Shared Gaussian Process Dynamic ModelabstractA new method of Gaussian process dynamic model (GPDM), named class-aware shared GPDM (CSGPDM), is presented in this paper. One of the most difference between our CSGPDM and existing GPDM is considering class information which helps to build the class label-based latent space being effective for the following class-related tasks. In terms of representation learning, CSGPDM is optimized by considering not only a non-linear relationship but also time-series relation and discriminative information of each class label. Then CSGPDM can reflect the following three points to the estimated latent space: i) the relationship between heterogeneous input sets, ii) time-series relations lurked in each input data, and iii) class information. Therefore, when input heterogeneous sets of features have time-series relations and class information, the above CSGPDM-based latent space can be beneficial for the obtaining the new CSGPDM-based feature sets for the post classification and estimating one side of the lacking samples by bridging the input heterogeneous feature sets via the latent space. Experimental results show that the estimated CSGPDM-based latent space outperformed those of GPDM and shared GPDM (SGPDM). Ryosuke Sawata, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2023 | Diffiner: A Versatile Diffusion-based Generative Refiner for Speech EnhancementabstractAlthough deep neural network (DNN)-based speech enhancement (SE) methods outperform the previous non-DNN-based ones, they often degrade the perceptual quality of generated outputs.To tackle this problem, we introduce a DNN-based generative refiner, Diffiner, aiming to improve perceptual speech quality pre-processed by an SE method.We train a diffusionbased generative model by utilizing a dataset consisting of clean speech only.Then, our refiner effectively mixes clean parts newly generated via denoising diffusion restoration into the degraded and distorted parts caused by a preceding SE method, resulting in refined speech.Once our refiner is trained on a set of clean speech, it can be applied to various SE methods without additional training specialized for each SE module.Therefore, our refiner can be a versatile post-processing module w.r.t.SE methods and has high potential in terms of modularity.Experimental results show that our method improved perceptual speech quality regardless of the preceding SE methods used.Our code is available at https://github.com/sony/diffiner. Ryosuke Sawata, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Takashi Shibuya 0001, Shusuke Takahashi, Yuki Mitsufuji |
INTERSPEECH | 1 |
| 2022 | Improving Character Error Rate is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-Box Acoustic ModelsabstractA deep neural network (DNN)-based speech enhancement (SE) aiming to maximize the performance of an automatic speech recognition (ASR) system is proposed in this paper. In order to optimize the DNN-based SE model in terms of the character error rate (CER), which is one of the metric to evaluate the ASR system and generally non-differentiable, our method uses two DNNs: one for speech processing and one for mimicking the output CERs derived through an acoustic model (AM). Then both of DNNs are alternately optimized in the training phase. Even if the AM is a black-box, e.g., like one provided by a third-party, the proposed method enables the DNN-based SE model to be optimized in terms of the CER since the DNN mimicking the AM is differentiable. Consequently, it becomes feasible to build CER-centric SE model that has no negative effect, e.g., additional calculation cost and changing network architecture, on the inference phase since our method is merely a training scheme for the existing DNN-based methods. Experimental results show that our method improved CER by 8.8% relative derived through a black-box AM although certain noise levels are kept. Ryosuke Sawata, Yosuke Kashiwagi, Shusuke Takahashi |
ICASSP | 1 |
| 2021 | Human-Centered Favorite Music Classification Using EEG-Based Individual Music Preference Via Deep Time-Series CCAabstractA method to classify a user’s like or dislike musical pieces based on the extraction of his or her music preference is proposed in this paper. New scheme of Canonical Correlation Analysis (CCA), called Deep Time-series CCA (DTCCA), which can consider the correlation between two sets of input features with considering the timeseries relation lurked in each input data is exploited to realize the aforementioned classification. One of the most difference between DTCCA and existing other CCAs is enabling to consider the above time-series relation, and thus DTCCA make the individual electroencephalogram (EEG)-based favorite music classification more effective than the methods using one of other CCAs instead of DTCCA since EEG and audio signals are respectively time-series data. Experimental results show that DTCCA-based favorite music classification outperformed not only method using original features without CCA but also methods using other existing CCAs including even state-of-the-art CCA. Ryosuke Sawata, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |
| 2021 | All For One And One For All: Improving Music Separation By Bridging NetworksabstractThis paper proposes several improvements for music separation with deep neural networks (DNNs), namely a multi-domain loss (MDL) and two combination schemes. First, by using MDL we take advantage of the frequency and time domain representation of audio signals. Next, we utilize the relationship among instruments by jointly considering them. We do this on the one hand by modifying the network architecture and introducing a CrossNet structure. On the other hand, we consider combinations of instrument estimates by using a new combination loss (CL). MDL and CL can easily be applied to many existing DNN-based separation methods as they are merely loss functions which are only used during training and do not affect the inference step. Experimental results show that the performance of Open-Unmix (UMX), a well-known and state-of-the-art open-source library for music separation, can be improved by utilizing our above schemes. Our modifications of UMX are open-sourced together with this paper. Ryosuke Sawata, Stefan Uhlich, Shusuke Takahashi, Yuki Mitsufuji |
ICASSP | 1 |
| 2021 | Manifold-Aware Deep Clustering: Maximizing Angles Between Embedding Vectors Based on Regular SimplexabstractThis paper presents a new deep clustering (DC) method called manifold-aware DC (M-DC) that can enhance hyperspace utilization more effectively than the original DC.The original DC has a limitation in that a pair of two speakers has to be embedded having an orthogonal relationship due to its use of the one-hot vector-based loss function, while our method derives a unique loss function aimed at maximizing the target angle in the hyperspace based on the nature of a regular simplex.Our proposed loss imposes a higher penalty than the original DC when the speaker is assigned incorrectly.The change from DC to M-DC can be easily achieved by rewriting just one term in the loss function of DC, without any other modifications to the network architecture or model parameters.As such, our method has high practicability because it does not affect the original inference part.The experimental results show that the proposed method improves the performances of the original DC and its expansion method. Keitaro Tanaka, Ryosuke Sawata, Shusuke Takahashi |
Interspeech | 2 |
| 2019 | Novel Audio Feature Projection Using KDLPCCA-Based Correlation with EEG Features for Favorite Music ClassificationabstractA novel audio feature projection using Kernel Discriminative Locality Preserving Canonical Correlation Analysis (KDLPCCA)-based correlation with electroencephalogram (EEG) features for favorite music classification is presented in this paper. The projected audio features reflect individual music preference adaptively since they are calculated by considering correlations with the user's EEG signals during listening to musical pieces that the user likes/dislikes via a novel CCA proposed in this paper. The novel CCA, called KDLPCCA, can consider not only a non-linear correlation but also local properties and discriminative information of each class sample, namely, music likes/dislikes. Specifically, local properties reflect intrinsic data structures of the original audio features, and discriminative information enhances the power of the final classification. Hence, the projected audio features have an optimal correlation with individual music preference reflected in the user's EEG signals, adaptively. If the KDLPCCA-based projection that can transform original audio features into novel audio features is calculated once, our method can extract projected audio features from a new musical piece without newly observing individual EEG signals. Our method therefore has a high level of practicability. Consequently, effective classification of user's favorite musical pieces via a Support Vector Machine (SVM) classifier using the new projected audio features becomes feasible. Experimental results show that our method for favorite music classification using projected audio features via the novel CCA outperforms methods using original audio features, EEG features and even audio features projected by other state-of-the-art CCAs. Ryosuke Sawata, Takahiro Ogawa 0001, Miki Haseyama |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | Novel favorite music classification using EEG-based optimal audio features selected via KDLPCCAabstractThis paper presents a novel method of favorite music classification using EEG-based optimal audio features. To select audio features related to user's music preference, our method utilizes a relationship between EEG features obtained from the user's EEG signals during listening to music and their corresponding audio features since EEG signals of human reflect his/her music preference. Specifically, cross-loadings, whose components denote the degree of the relationship, are calculated based on Kernel Discriminative Locality Preserving Canonical Correlation Analysis (KDLPCCA) which is newly derived in the proposed method. In contrast with standard CCA, KDLPCCA can consider (1) non-linear correlation, (2) class information and (3) local structures of input EEG and audio features, simultaneously. Therefore, KDLPCCA-based cross-loadings can reflect best correlation between the user's EEG and corresponding audio signals. Then an optimal set of audio features related to his/her music preference can be obtained by employing the cross-loadings as novel criteria for feature selection. Consequently, our method realizes favorite music classification successfully by using the EEG-based optimal audio features. Ryosuke Sawata, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |