Xinyang Wu 0005

dblp:397/7300 · DBLP profile ↗
← Back
3ranked-venue papers in the field
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3 (2 first)
YearPublicationVenuePosition
2025 Effects of Tempo and Tonality on Listener Enjoyment of Automated Pop Mashups
abstract
Pop mashup is a form of audio remixing prevalent among musicians and automatic mashup systems due to the inherent compatibility of songs within this genre. Recent advancements in machine learning models have facilitated tools for source separation, harmonic analysis, and matching, enabling the generation of high-quality mashups. However, selecting the optimal musical features, specifically tempo and tonality, to maximize listener enjoyment remains a challenge. To address this question, we conducted a series of listening experiments featuring mashups composed of two pop songs, systematically varying these musical features and evaluating the overall preferences of survey participants. Our findings contribute to an understanding of how adjustments to mashup components influence audience perception, offering valuable insights for the development of automated mashup algorithms aimed at enhancing output quality.
AnhDung Dinh, Xinyang Wu 0005, Andrew Brian Horner
IEEE Big Data2
2024 Graph Neural Network Guided Music Mashup Generation
abstract
Music mashups integrate elements from different songs to create surprising and engaging listening experiences. Typically, a mashup combines the vocal track of a base song with the instrumental tracks of complementary songs. Automating the production of mashups has been an area of research for decades. Traditional approaches utilize rule-based methods, such as matching tempo and harmonic similarity, to select optimal segments for mashup generation. More recent techniques leverage neural networks to classify segment compatibility. However, both approaches primarily focus on layering segments that are generally compatible, without ensuring their detailed integration and alignment with the vocal track. In contrast, we introduce a novel approach using Graph Neural Networks (GNNs) that learns to rearrange instrumental segments to better align with the base vocal track, resulting in more surprising and accurate music blends for mashup generation. Additionally, we conducted subjective listening tests to evaluate our generated mashups against a baseline model using the same song pairs and the original base songs, assessing our model’s performance. Generated mashups used for evaluation can be found in https://anonymousmus.github.io/bnr.github.io/.
Xinyang Wu 0005, Andrew Horner
IEEE Big Data1
2024 Diffusion Models for Automatic Music Mixing
abstract
Music mixing is a process that involves fine-tuning the levels, dynamics, and frequency content of musical elements to ensure clarity and harmony in the final music production. In this paper, we present an automatic mixing system based on diffusion models to correct imbalances in music mixes. We manipulate the well mixed stems’ short-time Fourier transform randomly to simulate the frequency and level imbalances commonly encountered in real-world scenarios. The difference between the imbalanced mix and the original mix is treated as noise for the diffusion model to predict, enabling the reverse denoising process to generate an automated mix. We evaluate our model’s performance by calculating the signal-to-distortion ratio between the original and predicted mixes. These results are compared with baseline automatic mixing models, demonstrating significant improvements. Test set results in audio: https://aimg2025submission.github.io/diffmusicmix/
Xinyang Wu 0005, Andrew Horner
IEEE Big Data1