Zhejing Hu

dblp:267/7383 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0002-9424-9071ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMs
abstract
Large Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studies validate that symbolic music can be modeled as a human language, yet composing structured music from partial symbolic inputs through natural language interaction remains underexplored. Even LLMs struggle to generate structurally coherent compositions in such hybrid input-output scenarios, highlighting a fundamental gap that calls for a domain-specific learning paradigm. To this end, we propose Inspiration-to-Structure (IoS), a cognitively inspired framework that enables LLMs to generate structured musical sections from melodic ideas. IoS employs a three-phase process—semantic, structural, and collaborative cognition—and is supported by two key components: (1) a new dataset and construction protocol called Structured Triplet Data (STD), and (2) a training method, Dual-Instance Structural Contrastive Optimization (DiSCO), designed to enhance structural awareness. Experiments show that IoS improves structural coherence by 47.8% and artistic creativity by 21.8% compared to conventional language modeling paradigm, supervised fine-tuning, and even enables smaller LLMs to surpass larger LLMs. These results suggest that symbolic music, while language-like, demands specialized modeling beyond standard language modeling paradigms. IoS enables LLMs to transform music theory knowledge into structured composition, empowering users to compose music interactively via language and advancing toward general creative AI.
Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Aiwei Zhang, Shenghua Zhong, Bruce X. B. Yu, Gong Chen 0006
AAAI1
2026 MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
abstract
Adolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers technical barriers and delivers efficient outputs, its role in fostering adolescents’ expressive growth has been overlooked. We propose MusicScaffold, an adolescent-centered framework that transforms classical AI roles from broad conceptualizations into stage-specific, actionable developmental scaffolds designed to make expressive strategies transparent and learnable and to support adolescents in mastering creative expression. In a four-week study with middle school students (ages 12–14), MusicScaffold enhanced cognitive specificity, behavioral regulation, and affective autonomy in music creation. By reframing generative AI as a scaffold rather than a generator, this work bridges the machine efficiency of generative systems with human growth in adolescent creativity education.
Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Gong Chen 0006, Bruce X. B. Yu, Jiannong Cao 0001
CHI1
2025 Compose with Me: Collaborative Music Inpainter for Symbolic Music Infilling
abstract
The field of music generation has seen a surge of interest from both academia and industry, with innovative platforms such as Suno, Udio, and SkyMusic earning widespread recognition. However, the challenge of music infilling—modifying specific music segments without reconstructing the entire piece—remains a significant hurdle for both audio-based and symbolic-based models, limiting their adaptability and practicality. In this paper, we address symbolic music infilling by introducing the Collaborative Music Inpainter (CMI), an advanced human-in-the-loop (HITL) model for music infilling. The CMI features the Joint Embedding Predictive Autoregressive Generative Architecture (JEP-AGA), which learns the high-level predictive representations of the masked part that needs to be infilled during the autoregressive generative process, akin to how humans perceive and interpret music. The newly developed Dynamic Interaction Learner (DIL) achieves HITL by iteratively refining the infilled output based on user interactions alone, significantly reducing the interaction cost without requiring further input. Experimental results confirm CMI’s superior performance in music infilling, demonstrating its efficiency in producing high-quality music.
Zhejing Hu, Bruce X. B. Yu
AAAI1
2025 CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
abstract
Generative artificial intelligence in music has made significant strides, yet it still falls short of the substantial achievements seen in natural language processing, primarily due to the limited availability of music data. Knowledge-informed approaches have been shown to enhance the performance of music generation models, even when only a few pieces of musical knowledge are integrated. This paper seeks to leverage comprehensive music theory in AI-driven music generation tasks, such as algorithmic composition and style transfer, which traditionally require significant manual effort with existing techniques. We introduce a novel automatic music lexicon construction model that generates a lexicon, named CompLex, comprising 37,432 items derived from just 9 manually input category keywords and 5 sentence prompt templates. A new multi-agent algorithm is proposed to automatically detect and mitigate hallucinations. CompLex demonstrates impressive performance improvements across three state-of-the-art text-to-music generation models, encompassing both symbolic and audio-based methods. Furthermore, we evaluate CompLex in terms of completeness, accuracy, non-redundancy, and executability, confirming that it possesses the key characteristics of an effective lexicon.
Zhejing Hu, Bruce X. B. Yu
IJCAI1
2024 Responding to the Call: Exploring Automatic Music Composition Using a Knowledge-Enhanced Model
abstract
Call-and-response is a musical technique that enriches the creativity of music, crafting coherent musical ideas that mirror the back-and-forth nature of human dialogue with distinct musical characteristics. Although this technique is integral to numerous musical compositions, it remains largely uncharted in automatic music composition. To enhance the creativity of machine-composed music, we first introduce the Call-Response Dataset (CRD) containing 19,155 annotated musical pairs and crafted comprehensive objective evaluation metrics for musical assessment. Then, we design a knowledge-enhanced learning-based method to bridge the gap between human and machine creativity. Specifically, we train the composition module using the call-response pairs, supplementing it with musical knowledge in terms of rhythm, melody, and harmony. Our experimental results underscore that our proposed model adeptly produces a wide variety of creative responses for various musical calls.
Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Xiao Ma 0023, Shenghua Zhong, Qianwen Luo
AAAI1
2024 Make a song curative: A spatio-temporal therapeutic music transfer model for anxiety reduction
Zhejing Hu, Gong Chen 0006, Yan Liu 0004, Xiao Ma 0023, Nianhong Guan, Xiaoying Wang 0007
Expert Syst. Appl.1
2024 The Beauty of Repetition: An Algorithmic Composition Model With Motif-Level Repetition Generator and Outline-to-Music Generator in Symbolic Music Generation
abstract
Most musical compositions utilize repetition as a fundamental element to create captivating aesthetic experiences. However, the potential of repetition in machine-learning-based algorithmic composition has not been thoroughly investigated. This article aims to make an initial attempt at repetition modeling by generating motif-level repetitions and integrating them into music through a combination of example-based and domain knowledge–based learning techniques. The article presents a new Motif-to-music Generation Model (MGM) that combines a motif-level repetition generator (MRG) and an outline-to-music generator (O2MG). To train this model, a new music repetition dataset (MRD) has been created, which includes 584,329 samples from various categories of motif repetition and 3,545 outline-music sequences from pop piano music. The MRG uses a Transformer encoder to learn the representation of music notes from MRD, while the repetition-aware learner in MRG takes advantage of the unique characteristics of repetitions based on music theory. The O2MG applies a novel outline-to-music learning strategy to learn the relationships among motif-level repetitions in the music and generate music based on these repetitions. The experiments show that MGM can generate a variety of beautiful repetitions with any given motif, improving the music quality and structure of machine-composed music.
Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003, Roger B. Dannenberg
IEEE Trans. Multim.1
2023 Noise-robust oversampling for imbalanced data classification
Yongxu Liu 0003, Yan Liu 0004, Bruce X. B. Yu, Shenghua Zhong, Zhejing Hu
Pattern Recognit.5
2023 Can Machines Generate Personalized Music? A Hybrid Favorite-Aware Method for User Preference Music Transfer
abstract
User preference music transfer (UPMT) is a new problem in music style transfer that can be applied to many scenarios but remains understudied. Transferring an arbitrary song to fit a user’s preferences increases musical diversity and improves user engagement, which can greatly benefit individuals’ mental health. Most music style transfer approaches rely on data-driven methods. In general, however, constructing a large training dataset is challenging because users can rarely provide enough of their favorite songs. To address this problem, this paper proposes a novel hybrid method called User Preference Transformer (UP-Transformer) which uses prior knowledge of only one piece of a user’s favorite music. Based on the distribution of music events in the provided music, we propose a new favorite-aware loss function to fine-tune the Transformer-based model. Two steps are proposed in the transfer phase to achieve UPMT based on the extracted music pattern in a user’s favorite music. Additionally, to alleviate the problem of evaluating melodic similarity in music style transfer, we propose a new concept called pattern similarity (PS) to measure the similarity between two pieces of music. Statistical tests indicate that the results of PS are consistent with the similarity score in a qualitative experiment. Furthermore, experimental results on subjects show that the transferred music achieves better performance in musicality, similarity, and user preferences.
Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003
IEEE Trans. Multim.1
2022 The Beauty of Repetition in Machine Composition Scenarios
abstract
Repetition, a basic form of artistic creation, appears in most musical works and delivers enthralling aesthetic experiences. However, repetition remains underexplored in terms of automatic music composition. As an initial effort in repetition modelling, this paper focuses on generating motif-level repetitions via domain knowledge-based and example-based learning techniques. A novel repetition transformer (R-Transformer) that combines a Transformer encoder and a repetition-aware learner is trained on a new repetition dataset with 584,329 samples from different categories of motif repetition. The Transformer encoder learns the representation among music notes from the repetition dataset; the novel repetition-aware learner exploits repetitions' unique characteristics based on music theory. Experiments show that, with any given motif, R-Transformer can generate a large number of variable and beautiful repetitions. With ingenious fusion of these high-quality pieces, the musicality and appeal of machine-composed music have been greatly improved.
Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003
ACM Multimedia1
2021 Finding therapeutic music for anxiety using scoring model
abstract
A large number of people suffer from anxiety in modern society. As an effective treatment with few side effects, music therapy has been used to reduce anxiety for decades in clinical practice. Yet therapists continue to perform music selection, a key step in music therapy, manually. Considering the growing need for music therapy services and social distancing amid public emergencies, an automatic method for music selection would be of great practical utility. This paper marks the first effort to identify music with therapeutic effects on anxiety reduction via a novel music scoring model. We formulate the calculation of a therapeutic score as a quadratic programming problem, which minimizes score variance among known therapeutic songs while maintaining their superiority over other songs. The proposed model can uncover common features that contribute to anxiety reduction by learning from small and unbalanced data. Using a music therapy experiment, we find that the proposed model outperforms existing techniques in predicting therapeutic songs. Feature analysis is also conducted, revealing that high-frequency spectrums are important in therapeutic scoring.
Gong Chen 0006, Zhejing Hu, Nianhong Guan, Xiaoying Wang 0007
Int. J. Intell. Syst.2
2020 Make Your Favorite Music Curative: Music Style Transfer for Anxiety Reduction
abstract
Anxiety is the most common mental problem that affects nearly 300 million individuals worldwide. The situation is even worse recently. In clinical practice, music therapy has been used for more than forty years because of its effectiveness and few side effects in emotion regulation. This paper proposes a novel style transfer model to generate the therapeutic music according to user's preference. It is widely recognized that the favorite music greatly increases the engagement of the user, hence results in much better curative effects. But in general, users can provide only one or several favorite songs, which are insufficient for the customization of therapeutic music. To address this difficulty, a new domain adaption algorithm that transfers the learning result for music genre classification to the music personalization, is designed. Targeting the joint minimization of the loss functions, three convolutional neural networks are utilized to generate the therapeutic music with only one labelled data of favorite song. The experiment on the anxiety suffers shows that the customized therapeutic music has achieved better and stable performance in anxiety reduction.
Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Shenghua Zhong, Aiwei Zhang
ACM Multimedia1
2020 Continuous blood pressure measurement from one-channel electrocardiogram signal using deep-learning techniques
Fen Miao, Zhejing Hu, Giancarlo Fortino, Xi-Ping Wang, Zeng-Ding Liu, Ye Li 0002
Artif. Intell. Medicine3
2020 A novel hybrid network of fusing rhythmic and morphological features for atrial fibrillation detection on mobile ECG signals
Xiaomao Fan, Zhejing Hu, Ruxin Wang 0001, Liyan Yin, Ye Li 0002, Yunpeng Cai
Neural Comput. Appl.2