Simon Lui

dblp:88/7595 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-0829-2867ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
abstract
Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. To overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs’ comprehensive music knowledge to generate contextually rich descriptions. Extensive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform.
Tristan Tsoi, Jiajun Deng, Yaolong Ju, Benno Weck, Holger Kirchhoff, Simon Lui
ICME6
2024 Cycle Frequency-Harmonic-Time Transformer for Note-Level Singing Voice Transcription
abstract
Singing voice transcription (SVT) is the task of converting singing voice music into symbolic note series. Although most SVT models utilized the time-frequency information from the input spectrogram, the useful harmonic information in singing voices has not been utilized enough. In this paper, we propose a novel 3D Cycle Frequency-Harmonic-Time Transformer (CFT) to explicitly capture the harmonic series of singing voices, where we first define a tokenization scheme that captures harmonics across multiple octaves, then the harmonic features are aggregated into the frequency-harmonic-time representations via a cyclic architecture. Results show that our method achieves state-of-the-art performances on several public datasets, including note-wise accuracy increases of 5.76% for MIR-ST500 and 13.56% for Cmedia.
Yaolong Ju, Simon Lui, Xuhao Du
ICME3
2023 Improving Automatic Singing Skill Evaluation with Timbral Features, Attention, and Singing Voice Separation
abstract
Most automatic singing skill evaluation (ASSE) models focus only on solo singing, resulting in a limited application scope since singing is usually mixed with instrumental accompaniment in music. In this paper, we propose a more general ASSE model which applies to both solo singing and singing with accompaniment. For this purpose, we employ an existing singing voice separation tool for accompaniment removal and compare ASSE models trained with and without accompaniment. Results show that accompaniment removal achieves better performances. Furthermore, we explore different features and model architectures, concluding that the additions of timbral features, attention mechanism, and dense layer further improve the performance. Finally, we show that our proposed model achieves a Pearson correlation coefficient of 0.562, a 62.4% relative improvement compared to 0.346 for the baseline model.
Yaolong Ju, Chunyang Xu, Jinhu Li, Simon Lui
ICME5
2022 MDAN: Multi-level Dependent Attention Network for Visual Emotion Analysis
abstract
Visual Emotion Analysis (VEA) is attracting increasing attention. One of the biggest challenges of VEA is to bridge the affective gap between visual clues in a picture and the emotion expressed by the picture. As the granularity of emotions increases, the affective gap increases as well. Existing deep approaches try to bridge the gap by directly learning discrimination among emotions globally in one shot. They ignore the hierarchical relationship among emotions at different affective levels, and the variation in the affective level of emotions to be classified. In this paper, we present the multi-level dependent attention network (MDAN) with two branches to leverage the emotion hierarchy and the correlation between different affective levels and semantic levels. The bottom-up branch directly learns emotions at the highest affective level and largely prevents hierarchy violation by explicitly following the emotion hierarchy while predicting emotions at lower affective levels. In contrast, the top-down branch aims to disentangle the affective gap by one-to-one mapping between semantic levels and affective levels, namely, Affective Semantic Mapping. A local classifier is appended at each semantic level to learn discrimination among emotions at the corresponding affective level. Then, we integrate global learning and local learning into a unified deep framework and optimize it simultaneously. Moreover, to properly model channel dependencies and spatial attention while disentangling the affective gap, we carefully designed two attention modules: the Multi-head Cross Channel Attention module and the Level-dependent Class Activation Map module. Finally, the proposed deep framework obtains new state-of-the-art performance on six VEA benchmarks, where it outperforms existing state-of-the-art methods by a large margin, e.g., +3.85% on the WEBEmo dataset at 25 classes classification accuracy.
Simon Lui
CVPR4
2021 Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis
abstract
LiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from the musical score, and generates acoustic features from these conditions. The full conditions in this paper consist of dynamic spectrogram energy, voiced/unvoiced (V/UV) decision and dynamic pitch curve, which are proven related to the expressiveness. We predict the pitch and the timbre features separately, avoiding the interdependence between these two features. Instead of neural network vocoders, a parametric WORLD vocoder is employed for the pitch curve consistency. Experiment results show that LiteSing outperforms the baseline model using feed-forward Transformer by 1.386 times faster on inference speed, 15 times smaller on training parameters number, and achieves a similar MOS on sound quality. Through an A/B test, LiteSing achieves 67.3% preference rate over baseline in pitch curve and dynamic spectrogram energy prediction. which demonstrates the advantage of LiteSing over the other compared models.
Xiaobin Zhuang, Szu-Yu Chou, Simon Lui
ICASSP6
2021 Large-scale singer recognition using deep metric learning: an experimental study
abstract
Singer recognition aims to automatically recognize the singer of a given recording. Compared to spoken voices, singing voice is characterized by a much higher degree of vocal style. The task becomes more challenging when it operates on numerous singers. This paper explores different strategies in a deep metric learning framework, with special focus on their performance in a large-scale dataset consisting of audio samples from 5057 singers. We conduct thorough experiments to compare loss functions, including triplet loss, generalized end-to-end (GE2E) loss, and prototypical network (PN) loss. Effects of vocal source separation is also investigated. Using audio inputs with separated vocals, our model trained with PN loss outperforms other evaluated methods in the identification task. While in the verification task with one-on-one comparison of two single embeddings, triplet loss achieves the best results. However, verification using PN loss shows superior performance to methods with triplet loss when using the centroid of 5 embed dings to represent the singer embedding. Using longer segments for a singer representation consistently improves the performance for all evaluated tasks.
Shichao Hu, Beici Liang, Zhouxuan Chen, Ethan Zhao 0002, Simon Lui
IJCNN6
2021 Real-time video super-resolution using lightweight depthwise separable group convolutions with channel shuffling
Zhijiao Xiao, Kwok-Wai Hung, Simon Lui
J. Vis. Commun. Image Represent.4
2021 Real-time video super resolution network using recurrent multi-branch dilated convolutions
Yubin Zeng, Zhijiao Xiao, Kwok-Wai Hung, Simon Lui
Signal Process. Image Commun.4
2020 Phase-Aware Music Super-Resolution Using Generative Adversarial Networks
abstract
Audio super-resolution is a challenging task of recovering the missing high-resolution features from a low-resolution signal.To address this, generative adversarial networks (GAN) have been used to achieve promising results by training the mappings between magnitudes of the low and high-frequency components.However, phase information is not well-considered for waveform reconstruction in conventional methods.In this paper, we tackle the problem of music super-resolution and conduct a thorough investigation on the importance of phase for this task.We use GAN to predict the magnitudes of the highfrequency components.The corresponding phase information can be extracted using either a GAN-based waveform synthesis system or a modified Griffin-Lim algorithm.Experimental results show that phase information plays an important role in the improvement of the reconstructed music quality.Moreover, our proposed method significantly outperforms other state-ofthe-art methods in terms of objective evaluations.
Shichao Hu, Beici Liang, Ethan Zhao 0002, Simon Lui
INTERSPEECH5
2020 Singing voice separation using a deep convolutional neural network trained by ideal binary mask and cross entropy
Kin Wah Edward Lin, Balamurali B. T., Enyan Koh, Simon Lui, Dorien Herremans
Neural Comput. Appl.4
2019 A novel music-based game with motion capture to support cognitive and motor function in the elderly
abstract
This paper presents a novel game prototype that uses music and motion detection as preventive medicine for the elderly. Given the aging populations around the globe, and the limited resources and staff able to care for these populations, eHealth solutions are becoming increasingly important, if not crucial, additions to modern healthcare and preventive medicine. Furthermore, because compliance rates for performing physical exercises are often quite low in the elderly, systems able to motivate and engage this population are a necessity. Our prototype uses music not only to engage listeners, but also to leverage the efficacy of music to improve mental and physical wellness. The game is based on a memory task to stimulate cognitive function, and requires users to perform physical gestures to mimic the playing of different musical instruments. To this end, the Microsoft Kinect sensor is used together with a newly developed gesture detection module in order to process users’ gestures. The resulting prototype system supports both cognitive functioning and physical strengthening in the elderly.
Kat Agres, Simon Lui, Dorien Herremans
CoG2
2017 Sinusoidal Partials Tracking for Singing Analysis Using the Heuristic of the Minimal Frequency and Magnitude Difference
Kin Wah Edward Lin, Hans Anderson, Clifford So, Simon Lui
INTERSPEECH4
2014 Visualising Singing Style under Common Musical Events Using Pitch-Dynamics Trajectories and Modified TRACLUS Clustering
abstract
We present a novel method for visualising the singing style of vocalists. To illustrate our method, we take 26 audio recordings of A cappella solo vocal music from two different professional singers and we visualise the performance style of each vocalist in a two-dimensional space of pitch and dynamics. We use our own novel modification of a trajectory clustering algorithm called TRACLUS to generate four representative paths, called trajectories, in that two dimensional space. Each trajectory represents the characteristic style of a vocalist during one of four common musical events: (1) Crescendo, (2) Diminuendo, (3) Ascending Pitches and (4) Descending Pitches. The unique shapes of these trajectories characterize the singing style of each vocalist with respect to each of these events. We present the details of our modified version of the TRACULUS algorithm and demonstrate graphically how the plots produced indicate distinct stylistic differences between singers. Potential applications for this method include: (a) automatic identification of singers and automatic classification of singing styles and (b) automatic retargeting of performance style to add human expression to computer generated vocal performances and allow singing synthesisers to imitate the styles of specific famous professional vocalists.
Kin Wah Edward Lin, Hans Anderson, Natalie Agus, Clifford So, Simon Lui
ICMLA5
2014 Modelling Mutual Information between Voiceprint and Optimal Number of Mel-Frequency Cepstral Coefficients in Voice Discrimination
abstract
In this paper, we study the relationship between the voiceprint and the optimal number of Mel-frequency Cepstral Coefficients (MFCCs). The voiceprint is modelled as sub-MFCCs matrix with the first d number of MFCCs. We model the relationship through information theory and formulate it as the mutual information maximization problem subject to the probabilities constraint. The solution of this optimization problem provides the optimal number of MFCCs, D among these d, which yields the highest classification accuracy of the voice discrimination, together with a confidence level. This study is dictated by the need to understand the use of MFCCs, which have proliferated since its invention to discriminate voice. We evaluate our model by comparing the leave-one-out cross validation (LOOCV) results of usual multi-class classifier, the Supervised Learning Gaussian Mixture Model (SLGMM), with a set of spoken words and A capella solo vocal performances. The experimental results show that our model is a more comprehensive feature selection criteria for the MFCCs than the de-facto technique, LOOCV.
Kin Wah Edward Lin, Tian Feng 0001, Natalie Agus, Clifford So, Simon Lui
ICMLA5
2012 Sensor-enabled yo-yos as new musical instruments
abstract
An interactive and programmable musical yo-yo system is designed. The aim of it is to demonstrate the feasibility of converting any sensor-enabled objects into potential musical instruments. This involves three design phases. First, the physical yo-yo is designed to house Iris sensors. The software is developed to sense the movement of yo-yo and transmit its measurements to Max/MSP for corresponding music generation. Finally, aurally pleasing and real-time musical sounds are designed and generated in effect of yo-yo by the computer music composer.
Jung-Hyun Jun, Sunardi, Joel W. Matthys, Simon Lui, Yu Gu 0001
IPSN5