EDBT 2026 Demo / reviewers in the wild / expert
Wei-Hsiang Liao 0001
dblp:18/4227
· DBLP profile ↗
17ranked-venue papers
0as first author
16since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music EditingabstractMusic editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality. Xinlei Niu, Kin Wai Cheuk, Jing Zhang 0052, Naoki Murata, Chieh-Hsin Lai, Michele Mancusi, Woosung Choi, Giorgio Fabbro, Wei-Hsiang Liao 0001, Charles P. Martin, Yuki Mitsufuji |
AAAI | 9 |
| 2025 | Variable Bitrate Residual Vector Quantization for Audio CodingabstractRecent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple input audio, such as silence. To address this limitation, we propose variable bitrate RVQ (VRVQ) for audio codecs, which allows for more efficient coding by adapting the number of codebooks used per frame. Furthermore, we propose a gradient estimation method for the non-differentiable masking operation that transforms from the importance map to the binary importance mask, improving model training via a straight-through estimator. We demonstrate that the proposed training framework achieves superior results compared to the baseline method and shows further improvement when applied to the current state-of-the-art codec. Audio samples are available at: https://yoongi43.github.io/VBRRVQ.github.io/ Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Kin Wai Cheuk, Marco A. Martínez Ramírez, Kyogu Lee, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
ICASSP | 10 |
| 2025 | Latent Diffusion Bridges for Unsupervised Musical Audio Timbre TransferabstractMusic timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired monophonic single-instrument audio data. Each diffusion model is trained on a specific instrument with a Gaussian prior. During inference, a model is designated as the source model to map the input audio to its corresponding Gaussian prior, and another model is designated as the target model to reconstruct the target audio from this Gaussian prior, thereby facilitating timbre transfer. We compare our approach against existing unsupervised timbre transfer models such as VAEGAN and Gaussian Flow Bridges (GFB). Experimental results demonstrate that our method achieves both better Fréchet Audio Distance (FAD) and melody preservation, as reflected by lower pitch distances (DPD) compared to VAEGAN and GFB. Additionally, we discover that the noise level from the Gaussian prior, σ, can be adjusted to control the degree of melody preservation and amount of timbre transferred. Michele Mancusi, Yurii Halychanskyi, Kin Wai Cheuk, Eloi Moliner, Chieh-Hsin Lai, Stefan Uhlich, Junghyun Koo, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Giorgio Fabbro, Yuki Mitsufuji |
ICASSP | 9 |
| 2025 | HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model FinetuningabstractControllable generation through Stable Diffusion (SD) fine-tuning aims to improve fidelity, safety, and alignment with human guidance. Existing reinforcement learning from human feedback methods usually rely on predefined heuristic reward functions or pretrained reward models built on large-scale datasets, limiting their applicability to scenarios where collecting such data is costly or difficult. To effectively and efficiently utilize human feedback, we develop a framework, HERO, which leverages online human feedback collected on the fly during model learning. Specifically, HERO features two key mechanisms: (1) Feedback-Aligned Representation Learning, an online training method that captures human feedback and provides informative learning signals for fine-tuning, and (2) Feedback-Guided Image Generation, which involves generating images from SD's refined initialization samples, enabling faster convergence towards the evaluator's intent. We demonstrate that HERO is 4x more efficient in online feedback for body part anomaly correction compared to the best existing method. Additionally, experiments show that HERO can effectively handle tasks like reasoning, counting, personalization, and reducing NSFW content with only 0.5K online feedback. The code and project page are available at [https://hero-dm.github.io/](https://hero-dm.github.io/). Ayano Hiranaka, Shang-Fu Chen, Chieh-Hsin Lai, Naoki Murata, Takashi Shibuya 0001, Wei-Hsiang Liao 0001, Shao-Hua Sun, Yuki Mitsufuji |
ICLR | 7 |
| 2025 | A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
Yigitcan Özer, Woosung Choi, Joan Serrà, Mayank Kumar Singh, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
INTERSPEECH | 5 |
| 2024 | Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music TranscriptionabstractIn recent years, research on music transcription has focused mainly on architecture design and instrument-specific data acquisition. With the lack of availability of diverse datasets, progress is often limited to solo-instrument tasks such as piano transcription. Several works have explored multi-instrument transcription as a means to bolster the performance of models on low-resource tasks, but these methods face the same data availability issues. We propose Timbre-Trap, a novel framework which unifies music transcription and audio reconstruction by exploiting the strong separability between pitch and timbre. We train a single autoencoder to simultaneously estimate pitch salience and reconstruct complex spectral coefficients, selecting between either output during the decoding stage via a simple switch mechanism. In this way, the model learns to produce coefficients corresponding to timbre-less audio, which can be interpreted as pitch salience. We demonstrate that the framework leads to performance comparable to state-of-the-art instrument-agnostic transcription methods, while only requiring a small amount of annotated data. Frank Cwitkowitz, Kin Wai Cheuk, Woosung Choi, Marco A. Martínez Ramírez, Keisuke Toyama 0002, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
ICASSP | 6 |
| 2024 | VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple GuidanceabstractRestoring degraded music signals is essential to enhance audio quality for downstream music manipulation. Recent diffusion-based music restoration methods have demonstrated impressive performance, and among them, diffusion posterior sampling (DPS) stands out given its intrinsic properties, making it versatile across various restoration tasks. In this paper, we identify that there are potential issues which will degrade current DPS-based methods’ performance and introduce the way to mitigate the issues inspired by diverse diffusion guidance techniques including the RePaint (RP) strategy and the Pseudoinverse-Guided Diffusion Models (ΠGDM). We demonstrate our methods for the vocal declipping and bandwidth extension tasks under various levels of distortion and cutoff frequency, respectively. In both tasks, our methods outperform the current DPS-based music restoration benchmarks. We refer to http://carlosholivan.github.io/demos/audio-restoration-2023.html for examples of the restored audio samples. Carlos Hernandez-Olivan, Koichi Saito, Naoki Murata, Chieh-Hsin Lai, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
ICASSP | 6 |
| 2024 | Manifold Preserving Guided DiffusionabstractDespite the recent advancements, conditional image generation still faces challenges of cost, generalizability, and the need for task-specific training. In this paper, we propose Manifold Preserving Guided Diffusion (MPGD), a training-free conditional generation framework that leverages pretrained diffusion models and off-the-shelf neural networks with minimal additional inference cost for a broad range of tasks. Specifically, we leverage the manifold hypothesis to refine the guided diffusion steps and introduce a shortcut algorithm in the process. We then propose two methods for on-manifold training-free guidance using pre-trained autoencoders and demonstrate that our shortcut inherently preserves the manifolds when applied to latent diffusion models. Our experiments show that MPGD is efficient and effective for solving a variety of conditional generation applications in low-compute settings, and can consistently offer up to 3.8× speed-ups with the same number of diffusion steps while maintaining high sample quality compared to the baselines. Naoki Murata, Chieh-Hsin Lai, Yuhta Takida, Toshimitsu Uesaka, Wei-Hsiang Liao 0001, Yuki Mitsufuji, J. Zico Kolter, Ruslan Salakhutdinov, Stefano Ermon |
ICLR | 7 |
| 2024 | Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of DiffusionabstractConsistency Models (CM) (Song et al., 2023) accelerate score-based diffusion model sampling at the cost of sample quality but lack a natural way to trade-off quality for speed. To address this limitation, we propose Consistency Trajectory Model (CTM), a generalization encompassing CM and score-based models as special cases. CTM trains a single neural network that can -- in a single forward pass -- output scores (i.e., gradients of log-density) and enables unrestricted traversal between any initial and final time along the Probability Flow Ordinary Differential Equation (ODE) in a diffusion process. CTM enables the efficient combination of adversarial training and denoising score matching loss to enhance performance and achieves new state-of-the-art FIDs for single-step diffusion model sampling on CIFAR-10 (FID 1.73) and ImageNet at 64X64 resolution (FID 1.92). CTM also enables a new family of sampling schemes, both deterministic and stochastic, involving long jumps along the ODE solution trajectories. It consistently improves sample quality as computational budgets increase, avoiding the degradation seen in CM. Furthermore, unlike CM, CTM's access to the score function can streamline the adoption of established controllable/conditional generation methods from the diffusion community. This access also enables the computation of likelihood. The code is available at https://github.com/sony/ctm. Chieh-Hsin Lai, Wei-Hsiang Liao 0001, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon |
ICLR | 3 |
| 2024 | MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
Yixiao Zhang 0002, Yukara Ikemiya, Gus Xia, Naoki Murata, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Yuki Mitsufuji, Simon Dixon |
IJCAI | 6 |
| 2024 | SilentCipher: Deep Audio Watermarking
Mayank Kumar Singh, Naoya Takahashi, Wei-Hsiang Liao 0001, Yuki Mitsufuji |
INTERSPEECH | 3 |
| 2024 | PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion TeacherabstractThe diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the training costs through three stages: training diffusion on downsampled data, distilling the pretrained diffusion, and progressive super-resolution. With the proposed pipeline, PaGoDA achieves a $64\times$ reduced cost in training its diffusion model on $8\times$ downsampled data; while at the inference, with the single-step, it performs state-of-the-art on ImageNet across all resolutions from $64\times64$ to $512\times512$, and text-to-image. PaGoDA's pipeline can be applied directly in the latent space, adding compression alongside the pre-trained autoencoder in Latent Diffusion Models (e.g., Stable Diffusion). The code is available at https://github.com/sony/pagoda. Chieh-Hsin Lai, Wei-Hsiang Liao 0001, Yuhta Takida, Naoki Murata, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon |
NeurIPS | 3 |
| 2023 | Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio EffectsabstractWe propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording. All our models are trained in a self-supervised manner from an already-processed wet multitrack dataset with an effective data preprocessing method that alleviates the data scarcity of obtaining unprocessed dry data. We analyze the proposed encoder for the disentanglement capability of audio effects and also validate its performance for mixing style transfer through both objective and subjective evaluations. From the results, we show the proposed system not only converts the mixing style of multitrack audio close to a reference but is also robust with mixture-wise style transfer upon using a music source separation model. Junghyun Koo, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Stefan Uhlich, Kyogu Lee, Yuki Mitsufuji |
ICASSP | 3 |
| 2022 | Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial NetworksabstractA central task of a Disc Jockey (DJ) is to create a mixset of music with seamless transitions between adjacent tracks. In this paper, we explore a data-driven approach that uses a generative adversarial network to create the song transition by learning from real-world DJ mixes. The generator uses two differentiable digital signal processing components, an equalizer (EQ) and a fader, to mix two tracks selected by a data generation pipeline. The generator has to set the parameters of the EQs and fader in such a way that the resulting mix resembles real mixes created by human DJ, as judged by the discriminator counterpart. Result of a listening test shows that the model can achieve competitive results compared with a number of baselines. Bo-Yu Chen, Wei-Han Hsu, Wei-Hsiang Liao 0001, Marco A. Martínez Ramírez, Yuki Mitsufuji, Yi-Hsuan Yang |
ICASSP | 3 |
| 2022 | SQ-VAE: Variational Bayes on Discrete Representation with Self-annealed Stochastic QuantizationabstractOne noted issue of vector-quantized variational autoencoder (VQ-VAE) is that the learned discrete representation uses only a fraction of the full capacity of the codebook, also known as codebook collapse. We hypothesize that the training scheme of VQ-VAE, which involves some carefully designed heuristics, underlies this issue. In this paper, we propose a new training scheme that extends the standard VAE via novel stochastic dequantization and quantization, called stochastically quantized variational autoencoder (SQ-VAE). In SQ-VAE, we observe a trend that the quantization is stochastic at the initial stage of the training but gradually converges toward a deterministic quantization, which we call self-annealing. Our experiments show that SQ-VAE improves codebook utilization without using common heuristics. Furthermore, we empirically show that SQ-VAE is superior to VAE and VQ-VAE in vision- and speech-related tasks. Yuhta Takida, Takashi Shibuya 0001, Wei-Hsiang Liao 0001, Chieh-Hsin Lai, Junki Ohmura, Toshimitsu Uesaka, Naoki Murata, Shusuke Takahashi, Toshiyuki Kumakura, Yuki Mitsufuji |
ICML | 3 |
| 2022 | Preventing oversmoothing in VAE via generalized variance parameterizationabstractVariational autoencoders (VAEs) often suffer from posterior collapse, which is a phenomenon in which the learned latent space becomes uninformative. This is often related to the hyperparameter resembling the data variance. It can be shown that an inappropriate choice of this hyperparameter causes the oversmoothness in the linearly approximated case and can be empirically verified for the general cases. Moreover, determining such appropriate choice becomes infeasible if the data variance is non-uniform or conditional. Therefore, we propose VAE extensions with generalized parameterizations of the data variance and incorporate maximum likelihood estimation into the objective function to adaptively regularize the decoder smoothness. The images generated from proposed VAE extensions show improved Fréchet inception distance (FID) on MNIST and CelebA datasets. Yuhta Takida, Wei-Hsiang Liao 0001, Chieh-Hsin Lai, Toshimitsu Uesaka, Shusuke Takahashi, Yuki Mitsufuji |
Neurocomputing | 2 |
| 2010 | A SOT based digital audio coder using reference frame ordering methodabstractIn SOT based coding methods, ordering of the coefficients could affect coding efficiency. A small coefficient placed close to tree root usually reduces the coding efficiency significantly. Hence, a good reordering scheme for a SOT is critical. In digital audio coding, information of neighboring frames is usually highly correlated. Based on this property, a SOT based audio coding scheme with a coefficient reordering method is proposed. The proposed coder may have the coefficients of the current frame reordered depending on the decoded results of some previous frames. The coding qualities of the proposed coder and MPEG-4 AAC coder are compared using ODG (Objective Differential Grade) on a variety of sound sources. Yu-Lin Wang, Wei-Hsiang Liao 0001, Alvin Wen-Yu Su |
ISCAS | 2 |