Ganghui Ru

dblp:327/8380 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 BeatKAN: An Efficient and Drum-Attuned Beat Tracking Method Using Kolmogorov-Arnold Networks
abstract
In this paper, we propose an efficient and drum-attuned beat tracking method based on Kolmogorov-Arnold networks (KAN). Traditional MLP-based frameworks struggle with complex musical signals due to limited capacity in modeling intricate patterns. Inspired by KAN’s efficient ability to capture complex time-frequency relationships, we leverage it to enhance our model by employing learnable nonlinear activation functions on convolutional kernels. Additionally, we utilize music source separation techniques to extract drum tracks from the original audio, thereby expanding the existing beat datasets and simulating the human perception of beats through drum sounds. Experimental results demonstrate that our approach significantly reduces the parameter count while maintaining high accuracy in beat tracking.
Ganghui Ru, Wei Li 0012
ICASSP2
2025 KCE-Unet: A novel music denoising method with KANConv ECA Unet
abstract
During concerts, people often spontaneously record memorable moments with their phones. However, these recordings are frequently accompanied by noise, such as cheering and applause, which diminishes the playback experience. In this paper, we introduce a novel task specifically designed for denoising music in concert environments, a challenge that has been largely overlooked in previous research. To support this task, we created a new concert denoising dataset that includes songs performed in various major languages at concerts, with noise segments like cheering and applause. Building on this, we propose KANConv ECA Unet (KCE-Unet), a method that combines the U-Net network, efficient channel attention (ECA), and the recently proposed KAN network to flexibly remove noise in the mid-to-high frequency range of spectrograms. Extensive experiments demonstrate that our method outperforms previous models in denoising performance and effectively restore disrupted musical structures.
Yulun Wu 0002, Ganghui Ru, Yi Yu 0001, Wei Li 0012
ICASSP3
2025 HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking
abstract
Fine-tuning pre-trained foundation models has made significant progress in music information retrieval. However, applying these models to beat tracking tasks remains unexplored as the limited annotated data renders conventional fine-tuning methods ineffective. To address this challenge, we propose HingeNet, a novel and general parameter-efficient fine-tuning method specifically designed for beat tracking tasks. HingeNet is a lightweight and separable network, visually resembling a hinge, designed to tightly interface with pre-trained foundation models by using their intermediate feature representations as input. This unique architecture grants HingeNet broad generalizability, enabling effective integration with various pre-trained foundation models. Furthermore, considering the significance of harmonics in beat tracking, we introduce harmonic-aware mechanism during the fine-tuning process to better capture and emphasize the harmonic structures in musical signals. Experiments on benchmark datasets demonstrate that HingeNet achieves state-of-the-art performance in beat and downbeat tracking.
Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012
ICME1
2025 BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
abstract
Beat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles and accurately capture complex rhythmic structures. To overcome these challenges, we propose a novel beat tracking paradigm BeatFM, which introduces a pre-trained music foundation model and leverages its rich semantic knowledge to improve beat tracking performance. Pre-training on diverse music datasets endows music foundation models with a robust understanding of music, thereby effectively addressing these challenges. To further adapt it for beat tracking, we design a plug-and-play multi-dimensional semantic aggregation module, which is composed of three parallel sub-modules, each focusing on semantic aggregation in the temporal, frequency, and channel domains, respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance in beat and downbeat tracking across multiple benchmark datasets.
Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012
ICME1
2023 Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations Perspective
abstract
Music genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use audio content and lyrics content inefficiently. In addition, as genres normally co-occur in a music track, it is desirable to capture and model the genre correlations to improve the performance of multi-label music genre classification. To solve these issues, we present a novel multi-modal method leveraging audio-lyrics contrastive loss and two symmetric cross-modal attention, to align and fuse features from audio and lyrics. Furthermore, based on the nature of the multi-label classification, a genre correlations extraction module is presented to capture and model potential genre correlations. Extensive experiments demonstrate that our proposed method significantly surpasses other multi-label music genre classification methods and achieves state-of-the-art result on Music4All dataset.
Ganghui Ru, Xulong Zhang 0001, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006
ICASSP1
2023 Rethinking and Improving Few-Shot Segmentation From a Contour-Aware Perspective
abstract
Existing few-shot segmentation approaches basically adopt the idea of comparing the semantic prototype vector of the query image and support images, and then obtaining the segmentation result. However, recent studies have shown that a single feature vector in feature map cannot accurately represent pixel-level categories, thus leading to poor segmentation of object boundary and semantic ambiguity. To address this common problem, we propose a novel contour-aware network (CTANet) for few-shot segmentation in this paper. Unlike the usual practice of classifying each pixel separately, CTANet regards all pixels within the same contour as a whole, which can take advantage of the internal consistency of objects to obtain a more accurate representation of category information. To obtain the accurate object contour, our network consists of a contour generation module and a contour refinement module, where the former exploits multiple levels of features to generate a primary contour map and the latter learns to refine the primary contour map. Furthermore, a novel contour-aware mixed loss is proposed to fuse the common BCE loss and our contour-aware loss to supervise the training process on two levels, pixel-level and contour-level. Extensive experiments demonstrate that our CTANet achieves a new state-of-the-art performance on$ \text{PASCAL-5}^{i}$and$ \text{COCO-20}^{i}$. Hopefully, our new perspective could provide more clues for future research on few-shot segmentation. Our code is freely available at:https://github.com/hardtogetA/CTANet.
Weimin Tan, Ganghui Ru, Yueming Jiang, Bo Yan 0001
IEEE Trans. Multim.2
2022 Multimodal Music Emotion Recognition with Hierarchical Cross-Modal Attention Network
abstract
Computational music emotion recognition is to recognize the emotional content in music tracks. In computational music emotion recognition studies, researchers have paid close attention to the audio content of the music tracks. Although lyrics content and music context contribute greatly to the perceived emotion, these kinds of emotional information are usually ignored. Based on this finding, we propose a multimodal music emotion recognition method jointly predicting the valence and arousal values by combining the audio, lyrics, track name, and artist of a given track. Audio features, lyrics features and context features are extracted separately and fused by a cross-modal attention mechanism, forming a hierarchical structure. Our proposed model outperforms two baselines by a large margin and achieves state-of-the-art performance on two public datasets.
Ganghui Ru, Yi Yu 0001, Yulun Wu 0002, Dichucheng Li, Wei Li 0012
ICME2