Gus Xia

dblp:217/2241 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
19since 2021 · last 2025
0000-0003-3629-906XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Improvised Performance Following in Real Time for Automatic Accompaniment
abstract
Real-time performance tracking is the core component of automatic accompaniment systems. Previous methods typically require full information on the performance score (the human part) and also assume very limited improvisations off the score, otherwise, the tracker will get lost. In this paper, we aim to track a fully improvised performance for automatic accompaniment, in which the only reference is the accompaniment score (the machine part). The key idea is to incorporate semantic-level alignment to model the correspondence between the user performance and the accompaniment. Our system contains two real-time components: The first component is an accompaniment-conditioned quantization model using a Long Short-Term Memory (LSTM) layer. For each performance note onset, we first compute its rough projected score position using previously estimated performance-to-score mapping and then refine the mapping using the quantization results. The second component is a playback control system, which updates the performance-to-score mapping via the performance-quantized time pair (similar to the performance-score alignment pair in traditional automatic accompaniment systems). Experiments show the system achieves better tracking results on expressive piano performance than baselines, allowing score-free tracking on both texture and tempo variations. Video demos are accessible on our demo page1.
Junyan Jiang, Akira Maezawa, Gus Xia
ICASSP3
2025 MuPT: A Generative Symbolic Music Pretrained Transformer
abstract
In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition. To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks. Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions.
Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009
ICLR21
2025 Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints
abstract
We contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms that rely on domain-specific labels or knowledge, our method is based on the insight of domain-general statistical differences between content and style --- content varies more among different fragments within a sample but maintains an invariant vocabulary across data samples, whereas style remains relatively invariant within a sample but exhibits more significant variation across different samples. We integrate such inductive bias into an encoder-decoder architecture and name our method after V3 (variance-versus-invariance). Experimental results show that V3 generalizes across multiple domains and modalities, successfully learning disentangled content and style representations, such as pitch and timbre from music audio, digit and color from images of hand-written digits, and action and character appearance from simple animations. V3 demonstrates strong disentanglement performance compared to existing unsupervised methods, along with superior out-of-distribution generalization and few-shot learning capabilities compared to supervised counterparts. Lastly, symbolic-level interpretability emerges in the learned content codebook, forging a near one-to-one alignment between machine representation and human knowledge.
Ziyu Wang 0008, Bhiksha Raj, Gus Xia
ICLR4
2025 Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization
abstract
We present a unified framework for automatic multitrack music arrangement that enables a single pre-trained symbolic music model to handle diverse arrangement scenarios, including reinterpretation, simplification, and additive generation. At its core is a segment-level reconstruction objective operating on token-level disentangled content and style, allowing for flexible any-to-any instrumentation transformations at inference time. To support track-wise modeling, we introduce REMI-z, a structured tokenization scheme for multitrack symbolic music that enhances modeling efficiency and effectiveness for both arrangement tasks and unconditional generation. Our method outperforms task-specific state-of-the-art models on representative tasks in different arrangement scenarios---band arrangement, piano reduction, and drum arrangement, in both objective metrics and perceptual evaluations. Taken together, our framework demonstrates strong generality and suggests broader applicability in symbolic music-to-music transformation.
Longshen Ou, Ziyu Wang 0008, Gus Xia, Qihao Liang, Torin Hopkins, Ye Wang 0007
NeurIPS4
2024 Musical Scene Detection in Comics: Comparing Perception of Humans and GPT-4
abstract
This paper aims to detect musical scenes for improving background music generation for comics. Musical scenes can be defined as scenes where background music is enhancing the media experience. A significant barrier in detecting such scenes for comics is the absence of a ground truth to compare against. This makes our work one of the first in this field to address this task. Hence, we first analyse musical scenes through the dialogues in anime films adapted from Japanese comics (manga). Through our analysis we discover that, apart from the existing literature on dimensions of change in scene segmentation, musical scenes are also triggered by emotions. We then build an internal dataset to engineer prompts for GPT-4 to recognise musical scenes. The results of the generated outcomes are evaluated against the internal dataset as well as human evaluation. We are able to prove that prompt engineering can objectively improve the musical scene detection. Our results indicate that GPT-4 has 62.5% agreement rate with human evaluators. The implications of this research on background music generation are relevant for various media. This paper invites further investigation on this topic.
Muhammad Taimoor Haseeb, Gus Xia, Yoshimasa Tsuruoka
IEEE Big Data3
2024 GPT-4 Driven Cinematic Music Generation Through Text Processing
abstract
This paper presents Herrmann-11, a multimodal framework to generate background music tailored to movie scenes, by integrating state-of-the-art vision, language, music, and speech processing models. Our pipeline begins by extracting visual and speech information from a movie scene, performing emotional analysis on it, and converting these into descriptive texts. Then, GPT-4 translates these high-level descriptions into low-level music conditions. Finally, these text-based music conditions guide a text-to-music model to generate music that resonates with input movie scenes. Comprehensive objective and subjective evaluations attest to the high synthesis quality, congruence, and superiority of our pipeline.
Muhammad Taimoor Haseeb, Ahmad Hammoudeh, Gus Xia
ICASSP3
2024 MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
abstract
Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music. To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training. In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance. This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT). Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters. Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores.
Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001
ICLR15
2024 Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
abstract
Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured **whole-song** generation. In this paper, we make the first attempt to model a full music piece under the realization of *compositional hierarchy*. With a focus on symbolic representations of pop songs, we define a hierarchical language, in which each level of hierarchy focuses on the semantics and context dependency at a certain music scope. The high-level languages reveal whole-song form, phrase, and cadence, whereas the low-level languages focus on notes, chords, and their local patterns. A cascaded diffusion model is trained to model the hierarchical language, where each level is conditioned on its upper levels. Experiments and analysis show that our model is capable of generating full-piece music with recognizable global verse-chorus structure and cadences, and the music quality is higher than the baselines. Additionally, we show that the proposed model is *controllable* in a flexible way. By sampling from the interpretable hierarchical languages or adjusting pre-trained external representations, users can control the music flow via various features such as phrase harmonic structures, rhythmic patterns, and accompaniment texture.
Ziyu Wang 0008, Lejun Min, Gus Xia
ICLR3
2024 MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
Yixiao Zhang 0002, Yukara Ikemiya, Gus Xia, Naoki Murata, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Yuki Mitsufuji, Simon Dixon
IJCAI3
2024 Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls
Gus Xia, Yixiao Zhang 0002, Junyan Jiang
IJCAI2
2024 Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling
abstract
In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing computational efficiency. In this paper, we introduce a novel system that leverages prior modelling over disentangled style factors to address these challenges. Our method presents a two-stage process: initially, a piano arrangement is derived from the lead sheet by retrieving piano texture styles; subsequently, a multi-track orchestration is generated by infusing orchestral function styles into the piano arrangement. Our key design is the use of vector quantization and a unique multi-stream Transformer to model the long-term flow of the orchestration style, which enables flexible, controllable, and structured music generation. Experiments show that by factorizing the arrangement task into interpretable sub-stages, our approach enhances generative capacity while improving efficiency. Additionally, our system supports a variety of music genres and provides style control at different composition hierarchies. We further show that our system achieves superior coherence, structure, and overall arrangement quality compared to existing baselines.
Gus Xia, Ziyu Wang 0008, Ye Wang 0007
NeurIPS2
2023 Self-Supervised Hierarchical Metrical Structure Modeling
abstract
We propose a novel method to model hierarchical metrical structures for both symbolic music and audio signals in a self-supervised manner with minimal domain knowledge. The model trains and inferences on beat-aligned music signals and predicts an 8-layer hierarchical metrical tree from beat, measure to the section level. The training procedure does not require any hierarchical metrical labeling except for beats, purely relying on the nature of metrical regularity and inter-voice consistency as inductive biases. We show in experiments that the method achieves comparable performance with supervised baselines on multiple metrical structure analysis tasks on both symbolic music and audio signals. All demos, source code and pre-trained models are publicly available on GitHub1.
Junyan Jiang, Gus Xia
ICASSP2
2023 Controllable Music Inpainting with Mixed-Level and Disentangled Representation
abstract
Music inpainting, which is to complete the missing part of a piece given some context, is an important task of automated music generation. In this study, we contribute a controllable inpainting model by combining the high expressivity of mixed-level, disentangled music representations and the strong predictive power of masked language modeling. The model enables flexible user controls over both time scope (inpainted length and location) and semantic features that composers often consider during composition, say rhythm pattern and chords. The key model design is to simultaneously predict disentangled representations of different time ranges. Such design aims to mirror the thought process of a professional composer who can take into account of the music flow of various semantic features at different hierarchies in parallel. Results show that our model produces much higher quality music compared to the baseline, and the subjective evaluation shows that our model generates much better results than the baseline and can generate melodies that are similar to human composition.
Shiqi Wei, Ziyu Wang 0008, Weiguo Gao, Gus Xia
ICASSP4
2023 Calliffusion: Chinese Calligraphy Generation and Style Transfer with Diffusion Modeling
Qisheng Liao, Gus Xia, Zhinuo Wang
ICCC2
2023 Q&A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement
abstract
Music rearrangement is a common music practice of reconstructing and reconceptualizing a piece using new composition or instrumentation styles, which is also an important task of automatic music generation. Existing studies typically model the mapping from a source piece to a target piece via supervised learning. In this paper, we tackle rearrangement problems via self-supervised learning, in which the mapping styles can be regarded as conditions and controlled in a flexible way. Specifically, we are inspired by the representation disentanglement idea and propose Q&A, a query-based algorithm for multi-track music rearrangement under an encoder-decoder framework. Q&A learns both a content representation from the mixture and function (style) representations from each individual track, while the latter queries the former in order to rearrange a new piece. Our current model focuses on popular music and provides a controllable pathway to four scenarios: 1) re-instrumentation, 2) piano cover generation, 3) orchestration, and 4) voice separation. Experiments show that our query system achieves high-quality rearrangement results with delicate multi-track structures, significantly outperforming the baselines.
Gus Xia, Ye Wang 0007
IJCAI2
2023 Learning Interpretable Low-dimensional Representation via Physical Symmetry
abstract
We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music domain knowledge. It remains an open question what general computational principles *give rise to* interpretable representations, especially low-dim factors that agree with human perception. In this study, we take inspiration from modern physics and use *physical symmetry* as a self-consistency constraint for the latent space. Specifically, it requires the prior model that characterises the dynamics of the latent states to be *equivariant* with respect to certain group transformations. We show that physical symmetry leads the model to learn a *linear* pitch factor from unlabelled monophonic music audio in a self-supervised fashion. In addition, the same methodology can be applied to computer vision, learning a 3D Cartesian space from videos of a simple moving object without labels. Furthermore, physical symmetry naturally leads to *counterfactual representation augmentation*, a new technique which improves sample efficiency.
Xuanjie Liu, Daniel Chin, Gus Xia
NeurIPS4
2023 MARBLE: Music Audio Representation Benchmark for Universal Evaluation
abstract
In the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research.
Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001
NeurIPS19
2022 Audio-To-Symbolic Arrangement Via Cross-Modal Music Representation Learning
abstract
Could we automatically derive the score of a piano accompaniment based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio content but also have prior knowledge of piano composition (so that the generation "sounds like" the audio and meanwhile maintains musicality). To this end, we contribute a cross-modal representation-learning model, which 1) extracts chord and melodic information from the audio, and 2) learns texture representation from both audio and a corrupted ground truth arrangement. We further introduce a tailored training strategy that gradually shifts the source of texture information from corrupted score to audio. In the end, the score-based texture posterior is reduced to a standard normal distribution, and only audio is needed for inference. Experiments show that our model captures major audio information and outperforms baselines in generation quality.1
Ziyu Wang 0008, Dejing Xu, Gus Xia, Ying Shan
ICASSP3
2022 Music Phrase Inpainting Using Long-Term Representation and Contrastive Loss
abstract
Deep generative modeling has already become the leading technique for music automation. However, long-term generation remains a challenging task as most methods fall short in preserving a natural structure and the overall musicality when the generation scope exceeds several beats. In this study, we tackle the problem of long-term, phrase-level symbolic melody inpainting by equipping a sequence prediction model with phrase-level representation (as an extra condition) and contrastive loss (as an extra optimization term). The underlying ideas are twofold. First, to predict phrase-level music, we need phrase-level representations as a better context. Second, we should predict notes and their high-level representations simultaneously, while contrastive loss serves as a better target for abstract representations. Experimental results show that our method significantly outperforms the baselines. In particular, contrastive loss plays a critical role in the generation quality, and the phase-level representation further enhances the structure of long-term generation.1
Shiqi Wei, Gus Xia, Yixiao Zhang 0002, Weiguo Gao
ICASSP2
2020 Transformer VAE: A Hierarchical Model for Structure-Aware and Interpretable Music Representation Learning
abstract
Structure awareness and interpretability are two of the most desired properties of music generation algorithms. Structure-aware models generate more natural and coherent music with long-term dependencies, while interpretable models are more friendly for human-computer interaction and co-creation. To achieve these two goals simultaneously, we designed the Transformer Variational AutoEncoder, a hierarchical model that unifies the efforts of two recent breakthroughs in deep music generation: 1) the Music Transformer and 2) Deep Music Analogy. The former learns long-term dependencies using attention mechanism, and the latter learns interpretable latent representations using a disentangled conditional-VAE. We showed that Transformer VAE is essentially capable of learning a context-sensitive hierarchical representation, regarding local representations as the context and the dependencies among the local representations as the global structure. By interacting with the model, we can achieve context transfer, realizing the imaginary situation of "what if" a piece is developed following the music flow of another piece.
Junyan Jiang, Gus Xia, Dave B. Carlton, Chris N. Anderson, Ryan H. Miyakawa
ICASSP2
2019 Transferring Piano Performance Control across Environments
abstract
Player pianos driven by computers are able to record and reproduce various performance control parameters, including pitch, timing, velocity and pedaling. However, the resulting sound of performance is not 100% reproducible in a new environment due to the difference in room acoustics and physical properties of the piano. Inspired by the Psychoacoustic studies which showed that human pianists adjust their controls in new environments for better performances, we have developed a system that automatically transfers performance control across environments in order to make the reproduced sound as similar as the original one. In specific, our work includes (1) a systematic measurement of the control-sound relationship of player pianos under different environments, and (2) a novel algorithm to adjust the control parameters through interpolating the measured control-sound functions. We evaluated the effectiveness of our method by conducting a listening test. Experimental results show that our algorithm outperforms the baseline significantly.
Maoran Xu, Ziyu Wang 0008, Gus Xia
ICASSP3