Gong Chen 0006

dblp:49/4553-6 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-2098-1291ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMs
abstract
Large Language Models (LLMs) have demonstrated remarkable proficiency in diverse tasks. This success raises a fundamental question in machine composition: Can symbolic music be considered a special form of language that can be jointly modeled with natural language for composition tasks? Recent studies validate that symbolic music can be modeled as a human language, yet composing structured music from partial symbolic inputs through natural language interaction remains underexplored. Even LLMs struggle to generate structurally coherent compositions in such hybrid input-output scenarios, highlighting a fundamental gap that calls for a domain-specific learning paradigm. To this end, we propose Inspiration-to-Structure (IoS), a cognitively inspired framework that enables LLMs to generate structured musical sections from melodic ideas. IoS employs a three-phase process—semantic, structural, and collaborative cognition—and is supported by two key components: (1) a new dataset and construction protocol called Structured Triplet Data (STD), and (2) a training method, Dual-Instance Structural Contrastive Optimization (DiSCO), designed to enhance structural awareness. Experiments show that IoS improves structural coherence by 47.8% and artistic creativity by 21.8% compared to conventional language modeling paradigm, supervised fine-tuning, and even enables smaller LLMs to surpass larger LLMs. These results suggest that symbolic music, while language-like, demands specialized modeling beyond standard language modeling paradigms. IoS enables LLMs to transform music theory knowledge into structured composition, empowering users to compose music interactively via language and advancing toward general creative AI.
Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Aiwei Zhang, Shenghua Zhong, Bruce X. B. Yu, Gong Chen 0006
AAAI7
2026 MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
abstract
Adolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers technical barriers and delivers efficient outputs, its role in fostering adolescents’ expressive growth has been overlooked. We propose MusicScaffold, an adolescent-centered framework that transforms classical AI roles from broad conceptualizations into stage-specific, actionable developmental scaffolds designed to make expressive strategies transparent and learnable and to support adolescents in mastering creative expression. In a four-week study with middle school students (ages 12–14), MusicScaffold enhanced cognitive specificity, behavioral regulation, and affective autonomy in music creation. By reframing generative AI as a scaffold rather than a generator, this work bridges the machine efficiency of generative systems with human growth in adolescent creativity education.
Zhejing Hu, Yan Liu 0004, Zhi Zhang 0004, Gong Chen 0006, Bruce X. B. Yu, Jiannong Cao 0001
CHI4
2025 Mixture of Knowledge Minigraph Agents for Literature Review Generation
abstract
Literature reviews play a crucial role in scientific research for understanding the current state of research, identifying gaps, and guiding future studies on specific topics. However, the process of conducting a comprehensive literature review is yet time-consuming. This paper proposes a novel framework, collaborative knowledge minigraph agents (CKMAs), to automate scholarly literature reviews. A novel prompt-based algorithm, the knowledge minigraph construction agent (KMCA), is designed to identify relations between concepts from academic literature and automatically constructs knowledge minigraphs. By leveraging the capabilities of large language models on constructed knowledge minigraphs, the multiple path summarization agent (MPSA) efficiently organizes concepts and relations from different viewpoints to generate literature review paragraphs. We evaluate CKMAs on three benchmark datasets. Experimental results show the effectiveness of the proposed method, further revealing promising applications of LLMs in scientific research.
Zhi Zhang 0004, Yan Liu 0004, Shenghua Zhong, Gong Chen 0006, Yu Yang 0012, Jiannong Cao 0001
AAAI4
2024 Responding to the Call: Exploring Automatic Music Composition Using a Knowledge-Enhanced Model
abstract
Call-and-response is a musical technique that enriches the creativity of music, crafting coherent musical ideas that mirror the back-and-forth nature of human dialogue with distinct musical characteristics. Although this technique is integral to numerous musical compositions, it remains largely uncharted in automatic music composition. To enhance the creativity of machine-composed music, we first introduce the Call-Response Dataset (CRD) containing 19,155 annotated musical pairs and crafted comprehensive objective evaluation metrics for musical assessment. Then, we design a knowledge-enhanced learning-based method to bridge the gap between human and machine creativity. Specifically, we train the composition module using the call-response pairs, supplementing it with musical knowledge in terms of rhythm, melody, and harmony. Our experimental results underscore that our proposed model adeptly produces a wide variety of creative responses for various musical calls.
Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Xiao Ma 0023, Shenghua Zhong, Qianwen Luo
AAAI3
2024 Sal-Guide Diffusion: Saliency Maps Guide Emotional Image Generation through Adapter
abstract
The existing text-to-image generation methods based on stable diffusion yield better results in low-semantic prompt but often neglect the generation quality of high-semantic prompt such as emotional vocabulary, resulting in poor emotional image generation. In order to address this issue, we propose a novel approach called Sal-Guide Diffusion, which leverages saliency maps to guide emotional image generation with the goal of producing superior emotionally expressive images. In order to let the saliency maps guide the diffusion process, we introduce a lightweight adapter to extract emotional information from saliency maps and incorporate it into the diffusion process. Experimental results demonstrate that our proposed method generates higher-quality images across eight emotional dimensions, excelling in both generalization, emotional congruence, and subjective preference compared to stable diffusion or similar methods.
Xiangru Lin, Shenghua Zhong, Yan Liu 0004, Gong Chen 0006
ICME4
2024 Make a song curative: A spatio-temporal therapeutic music transfer model for anxiety reduction
Zhejing Hu, Gong Chen 0006, Yan Liu 0004, Xiao Ma 0023, Nianhong Guan, Xiaoying Wang 0007
Expert Syst. Appl.2
2024 The Beauty of Repetition: An Algorithmic Composition Model With Motif-Level Repetition Generator and Outline-to-Music Generator in Symbolic Music Generation
abstract
Most musical compositions utilize repetition as a fundamental element to create captivating aesthetic experiences. However, the potential of repetition in machine-learning-based algorithmic composition has not been thoroughly investigated. This article aims to make an initial attempt at repetition modeling by generating motif-level repetitions and integrating them into music through a combination of example-based and domain knowledge–based learning techniques. The article presents a new Motif-to-music Generation Model (MGM) that combines a motif-level repetition generator (MRG) and an outline-to-music generator (O2MG). To train this model, a new music repetition dataset (MRD) has been created, which includes 584,329 samples from various categories of motif repetition and 3,545 outline-music sequences from pop piano music. The MRG uses a Transformer encoder to learn the representation of music notes from MRD, while the repetition-aware learner in MRG takes advantage of the unique characteristics of repetitions based on music theory. The O2MG applies a novel outline-to-music learning strategy to learn the relationships among motif-level repetitions in the music and generate music based on these repetitions. The experiments show that MGM can generate a variety of beautiful repetitions with any given motif, improving the music quality and structure of machine-composed music.
Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003, Roger B. Dannenberg
IEEE Trans. Multim.4
2023 Can Machines Generate Personalized Music? A Hybrid Favorite-Aware Method for User Preference Music Transfer
abstract
User preference music transfer (UPMT) is a new problem in music style transfer that can be applied to many scenarios but remains understudied. Transferring an arbitrary song to fit a user’s preferences increases musical diversity and improves user engagement, which can greatly benefit individuals’ mental health. Most music style transfer approaches rely on data-driven methods. In general, however, constructing a large training dataset is challenging because users can rarely provide enough of their favorite songs. To address this problem, this paper proposes a novel hybrid method called User Preference Transformer (UP-Transformer) which uses prior knowledge of only one piece of a user’s favorite music. Based on the distribution of music events in the provided music, we propose a new favorite-aware loss function to fine-tune the Transformer-based model. Two steps are proposed in the transfer phase to achieve UPMT based on the extracted music pattern in a user’s favorite music. Additionally, to alleviate the problem of evaluating melodic similarity in music style transfer, we propose a new concept called pattern similarity (PS) to measure the similarity between two pieces of music. Statistical tests indicate that the results of PS are consistent with the similarity score in a qualitative experiment. Furthermore, experimental results on subjects show that the transferred music achieves better performance in musicality, similarity, and user preferences.
Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003
IEEE Trans. Multim.3
2022 EGCN: An Ensemble-based Learning Framework for Exploring Effective Skeleton-based Rehabilitation Exercise Assessment
abstract
Recently, some skeleton-based physical therapy systems have been attempted to automatically evaluate the correctness or quality of an exercise performed by rehabilitation subjects. However, in terms of algorithms and evaluation criteria, the task remains not fully explored regarding making full use of different skeleton features. To advance the prior work, we propose a learning framework called Ensemble-based Graph Convolutional Network (EGCN) for skeleton-based rehabilitation exercise assessment. As far as we know, this is the first attempt that utilizes both two skeleton feature groups and investigates different ensemble strategies for the task. We also examine the properness of existing evaluation criteria and focus on evaluating the prediction ability of our proposed method. We then conduct extensive cross-validation experiments on two latest public datasets: UI-PRMD and KIMORE. Results indicate that the model-level ensemble scheme of our EGCN achieves better performance than existing methods. Code is available: https://github.com/bruceyo/EGCN.
Bruce X. B. Yu, Yan Liu 0004, Gong Chen 0006, Keith C. C. Chan
IJCAI4
2022 The Beauty of Repetition in Machine Composition Scenarios
abstract
Repetition, a basic form of artistic creation, appears in most musical works and delivers enthralling aesthetic experiences. However, repetition remains underexplored in terms of automatic music composition. As an initial effort in repetition modelling, this paper focuses on generating motif-level repetitions via domain knowledge-based and example-based learning techniques. A novel repetition transformer (R-Transformer) that combines a Transformer encoder and a repetition-aware learner is trained on a new repetition dataset with 584,329 samples from different categories of motif repetition. The Transformer encoder learns the representation among music notes from the repetition dataset; the novel repetition-aware learner exploits repetitions' unique characteristics based on music theory. Experiments show that, with any given motif, R-Transformer can generate a large number of variable and beautiful repetitions. With ingenious fusion of these high-quality pieces, the musicality and appeal of machine-composed music have been greatly improved.
Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003
ACM Multimedia4
2021 Finding therapeutic music for anxiety using scoring model
abstract
A large number of people suffer from anxiety in modern society. As an effective treatment with few side effects, music therapy has been used to reduce anxiety for decades in clinical practice. Yet therapists continue to perform music selection, a key step in music therapy, manually. Considering the growing need for music therapy services and social distancing amid public emergencies, an automatic method for music selection would be of great practical utility. This paper marks the first effort to identify music with therapeutic effects on anxiety reduction via a novel music scoring model. We formulate the calculation of a therapeutic score as a quadratic programming problem, which minimizes score variance among known therapeutic songs while maintaining their superiority over other songs. The proposed model can uncover common features that contribute to anxiety reduction by learning from small and unbalanced data. Using a music therapy experiment, we find that the proposed model outperforms existing techniques in predicting therapeutic songs. Feature analysis is also conducted, revealing that high-frequency spectrums are important in therapeutic scoring.
Gong Chen 0006, Zhejing Hu, Nianhong Guan, Xiaoying Wang 0007
Int. J. Intell. Syst.1
2020 Make Your Favorite Music Curative: Music Style Transfer for Anxiety Reduction
abstract
Anxiety is the most common mental problem that affects nearly 300 million individuals worldwide. The situation is even worse recently. In clinical practice, music therapy has been used for more than forty years because of its effectiveness and few side effects in emotion regulation. This paper proposes a novel style transfer model to generate the therapeutic music according to user's preference. It is widely recognized that the favorite music greatly increases the engagement of the user, hence results in much better curative effects. But in general, users can provide only one or several favorite songs, which are insufficient for the customization of therapeutic music. To address this difficulty, a new domain adaption algorithm that transfers the learning result for music genre classification to the music personalization, is designed. Targeting the joint minimization of the loss functions, three convolutional neural networks are utilized to generate the therapeutic music with only one labelled data of favorite song. The experiment on the anxiety suffers shows that the customized therapeutic music has achieved better and stable performance in anxiety reduction.
Zhejing Hu, Yan Liu 0004, Gong Chen 0006, Shenghua Zhong, Aiwei Zhang
ACM Multimedia3
2018 Nurturing the Companion ChatBot
abstract
Although recent technical progress of artificial intelligence is impressive, the affective effect of intelligent agents has not been investigated sufficiently. This research focuses on the affective effect of ChatBot. We observe an interesting phenomenon that not only the affective response from ChatBot influences user's experience, but also the manner that user interacts with ChatBot affects the development of the ChatBot's intelligence. So this work proposes an ethical issue that it is necessary to regularize the user's behavior. We validate this argument by designing a novel paradigm, which enables the users to nurture companion ChatBots via developmental artificial intelligence techniques. With only twenty days nurturing, the users build affective bonding with the ChatBots and the ChatBots show significant progress in communication skills.
Gong Chen 0006
AIES1
2018 Musicality-Novelty Generative Adversarial Nets for Algorithmic Composition
abstract
Algorithmic composition, which enables computer to generate music like human composers, has lasting charm because it intends to approximate artistic creation, most mysterious part of human intelligence. To deliver both melodious and refreshing music, this paper proposes the Musicality-Novelty Generative Adversarial Nets for algorithmic composition. With the same generator, two adversarial nets alternately optimize the musicality and novelty of the machine-composed music. A new model called novelty game is presented to maximize the minimal distance between the machine-composed music sample and any human-composed music sample in the novelty space, where all well-known human composed music products are far from each other. We implement the proposed framework using three supervised CNNs with one for generator, one for musicality critic and one for novelty critic on the time-pitch feature space. Specifically, the novelty critic is implemented by Siamese neural networks with temporal alignment using dynamic time warping. We provide empirical validations by generating the music samples under various scenarios.
Gong Chen 0006, Yan Liu 0004, Shenghua Zhong
ACM Multimedia1
2017 Who Composes the Music?: Musicality Evaluation for Algorithmic Composition via Electroencephalography
abstract
It is very challenging to evaluate the creative work of artificial intelligence, such as algorithmic composition. Due to the nature of creativity, most existing criteria of music analysis, for example, similarity of the data, cannot be used directly to measure the quality of a new piece of music composed by computer. Subjective evaluation based on questionnaire lacks quantitative evaluation with solid evidence. To address these difficulties, this paper proposes a novel computational model combined with a novel psychological paradigm. Utilizing brain imaging techniques, the proposed evaluation method can provide reliable musicality score for machine-composed music.
Gong Chen 0006
ACM Multimedia1
2016 Learning Music Emotion Primitives via Supervised Dynamic Clustering
abstract
This paper explores a fundamental problem in music emotion analysis, i.e., how to segment the music sequence into a set of basic emotive units, which are named as emotion primitives. Current works on music emotion analysis are mainly based on the fixed-length music segments, which often leads to the difficulty of accurate emotion recognition. Short music segment, such as an individual music frame, may fail to evoke emotion response. Long music segment, such as an entire song, may convey various emotions over time. Moreover, the minimum length of music segment varies depending on the types of the emotions. To address these problems, we propose a novel method dubbed supervised dynamic clustering (SDC) to automatically decompose the music sequence into meaningful segments with various lengths. First, the music sequence is represented by a set of music frames. Then, the music frames are clustered according to the valence-arousal values in the emotion space. The clustering results are used to initialize the music segmentation. After that, a dynamic programming scheme is employed to jointly optimize the subsequent segmentation and grouping in the music feature space. Experimental results on standard dataset show both the effectiveness and the rationality of the proposed method.
Yang Liu 0007, Yan Liu 0004, Gong Chen 0006
ACM Multimedia4