VLDB 2026 Research / reviewers in the wild / expert
Simon Lui
dblp:88/7595
· DBLP profile ↗
15ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-0829-2867ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and MiningabstractMusic similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contrastive learning framework that leverages the open-ended nature of text descriptions to guide music similarity modeling, addressing the limitations of traditional uni-modal approaches in capturing complex musical relationships. To overcome the scarcity of high-quality text-music paired data, this paper introduces a dual-source data acquisition approach combining online scraping and LLM-based prompting, where carefully designed prompts leverage LLMs’ comprehensive music knowledge to generate contextually rich descriptions. Extensive experiments demonstrate that the proposed framework achieves significant performance improvements over existing benchmarks through objective metrics, subjective evaluations, and real-world A/B testing on the Huawei Music streaming platform. Tristan Tsoi, Jiajun Deng, Yaolong Ju, Benno Weck, Holger Kirchhoff, Simon Lui |
ICME | 6 |
| 2024 | Cycle Frequency-Harmonic-Time Transformer for Note-Level Singing Voice TranscriptionabstractSinging voice transcription (SVT) is the task of converting singing voice music into symbolic note series. Although most SVT models utilized the time-frequency information from the input spectrogram, the useful harmonic information in singing voices has not been utilized enough. In this paper, we propose a novel 3D Cycle Frequency-Harmonic-Time Transformer (CFT) to explicitly capture the harmonic series of singing voices, where we first define a tokenization scheme that captures harmonics across multiple octaves, then the harmonic features are aggregated into the frequency-harmonic-time representations via a cyclic architecture. Results show that our method achieves state-of-the-art performances on several public datasets, including note-wise accuracy increases of 5.76% for MIR-ST500 and 13.56% for Cmedia. Yaolong Ju, Simon Lui, Xuhao Du |
ICME | 3 |
| 2023 | Improving Automatic Singing Skill Evaluation with Timbral Features, Attention, and Singing Voice SeparationabstractMost automatic singing skill evaluation (ASSE) models focus only on solo singing, resulting in a limited application scope since singing is usually mixed with instrumental accompaniment in music. In this paper, we propose a more general ASSE model which applies to both solo singing and singing with accompaniment. For this purpose, we employ an existing singing voice separation tool for accompaniment removal and compare ASSE models trained with and without accompaniment. Results show that accompaniment removal achieves better performances. Furthermore, we explore different features and model architectures, concluding that the additions of timbral features, attention mechanism, and dense layer further improve the performance. Finally, we show that our proposed model achieves a Pearson correlation coefficient of 0.562, a 62.4% relative improvement compared to 0.346 for the baseline model. Yaolong Ju, Chunyang Xu, Jinhu Li, Simon Lui |
ICME | 5 |
| 2022 | MDAN: Multi-level Dependent Attention Network for Visual Emotion AnalysisabstractVisual Emotion Analysis (VEA) is attracting increasing attention. One of the biggest challenges of VEA is to bridge the affective gap between visual clues in a picture and the emotion expressed by the picture. As the granularity of emotions increases, the affective gap increases as well. Existing deep approaches try to bridge the gap by directly learning discrimination among emotions globally in one shot. They ignore the hierarchical relationship among emotions at different affective levels, and the variation in the affective level of emotions to be classified. In this paper, we present the multi-level dependent attention network (MDAN) with two branches to leverage the emotion hierarchy and the correlation between different affective levels and semantic levels. The bottom-up branch directly learns emotions at the highest affective level and largely prevents hierarchy violation by explicitly following the emotion hierarchy while predicting emotions at lower affective levels. In contrast, the top-down branch aims to disentangle the affective gap by one-to-one mapping between semantic levels and affective levels, namely, Affective Semantic Mapping. A local classifier is appended at each semantic level to learn discrimination among emotions at the corresponding affective level. Then, we integrate global learning and local learning into a unified deep framework and optimize it simultaneously. Moreover, to properly model channel dependencies and spatial attention while disentangling the affective gap, we carefully designed two attention modules: the Multi-head Cross Channel Attention module and the Level-dependent Class Activation Map module. Finally, the proposed deep framework obtains new state-of-the-art performance on six VEA benchmarks, where it outperforms existing state-of-the-art methods by a large margin, e.g., +3.85% on the WEBEmo dataset at 25 classes classification accuracy. Simon Lui |
CVPR | 4 |
| 2021 | Litesing: Towards Fast, Lightweight and Expressive Singing Voice SynthesisabstractLiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from the musical score, and generates acoustic features from these conditions. The full conditions in this paper consist of dynamic spectrogram energy, voiced/unvoiced (V/UV) decision and dynamic pitch curve, which are proven related to the expressiveness. We predict the pitch and the timbre features separately, avoiding the interdependence between these two features. Instead of neural network vocoders, a parametric WORLD vocoder is employed for the pitch curve consistency. Experiment results show that LiteSing outperforms the baseline model using feed-forward Transformer by 1.386 times faster on inference speed, 15 times smaller on training parameters number, and achieves a similar MOS on sound quality. Through an A/B test, LiteSing achieves 67.3% preference rate over baseline in pitch curve and dynamic spectrogram energy prediction. which demonstrates the advantage of LiteSing over the other compared models. Xiaobin Zhuang, Szu-Yu Chou, Simon Lui |
ICASSP | 6 |
| 2021 | Large-scale singer recognition using deep metric learning: an experimental studyabstractSinger recognition aims to automatically recognize the singer of a given recording. Compared to spoken voices, singing voice is characterized by a much higher degree of vocal style. The task becomes more challenging when it operates on numerous singers. This paper explores different strategies in a deep metric learning framework, with special focus on their performance in a large-scale dataset consisting of audio samples from 5057 singers. We conduct thorough experiments to compare loss functions, including triplet loss, generalized end-to-end (GE2E) loss, and prototypical network (PN) loss. Effects of vocal source separation is also investigated. Using audio inputs with separated vocals, our model trained with PN loss outperforms other evaluated methods in the identification task. While in the verification task with one-on-one comparison of two single embeddings, triplet loss achieves the best results. However, verification using PN loss shows superior performance to methods with triplet loss when using the centroid of 5 embed dings to represent the singer embedding. Using longer segments for a singer representation consistently improves the performance for all evaluated tasks. Shichao Hu, Beici Liang, Zhouxuan Chen, Ethan Zhao 0002, Simon Lui |
IJCNN | 6 |
| 2021 | Real-time video super-resolution using lightweight depthwise separable group convolutions with channel shuffling
Zhijiao Xiao, Kwok-Wai Hung, Simon Lui |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Real-time video super resolution network using recurrent multi-branch dilated convolutions
Yubin Zeng, Zhijiao Xiao, Kwok-Wai Hung, Simon Lui |
Signal Process. Image Commun. | 4 |
| 2020 | Phase-Aware Music Super-Resolution Using Generative Adversarial NetworksabstractAudio super-resolution is a challenging task of recovering the missing high-resolution features from a low-resolution signal.To address this, generative adversarial networks (GAN) have been used to achieve promising results by training the mappings between magnitudes of the low and high-frequency components.However, phase information is not well-considered for waveform reconstruction in conventional methods.In this paper, we tackle the problem of music super-resolution and conduct a thorough investigation on the importance of phase for this task.We use GAN to predict the magnitudes of the highfrequency components.The corresponding phase information can be extracted using either a GAN-based waveform synthesis system or a modified Griffin-Lim algorithm.Experimental results show that phase information plays an important role in the improvement of the reconstructed music quality.Moreover, our proposed method significantly outperforms other state-ofthe-art methods in terms of objective evaluations. Shichao Hu, Beici Liang, Ethan Zhao 0002, Simon Lui |
INTERSPEECH | 5 |
| 2020 | Singing voice separation using a deep convolutional neural network trained by ideal binary mask and cross entropy
Kin Wah Edward Lin, Balamurali B. T., Enyan Koh, Simon Lui, Dorien Herremans |
Neural Comput. Appl. | 4 |
| 2019 | A novel music-based game with motion capture to support cognitive and motor function in the elderlyabstractThis paper presents a novel game prototype that uses music and motion detection as preventive medicine for the elderly. Given the aging populations around the globe, and the limited resources and staff able to care for these populations, eHealth solutions are becoming increasingly important, if not crucial, additions to modern healthcare and preventive medicine. Furthermore, because compliance rates for performing physical exercises are often quite low in the elderly, systems able to motivate and engage this population are a necessity. Our prototype uses music not only to engage listeners, but also to leverage the efficacy of music to improve mental and physical wellness. The game is based on a memory task to stimulate cognitive function, and requires users to perform physical gestures to mimic the playing of different musical instruments. To this end, the Microsoft Kinect sensor is used together with a newly developed gesture detection module in order to process users’ gestures. The resulting prototype system supports both cognitive functioning and physical strengthening in the elderly. Kat Agres, Simon Lui, Dorien Herremans |
CoG | 2 |
| 2017 | Sinusoidal Partials Tracking for Singing Analysis Using the Heuristic of the Minimal Frequency and Magnitude Difference
Kin Wah Edward Lin, Hans Anderson, Clifford So, Simon Lui |
INTERSPEECH | 4 |
| 2014 | Visualising Singing Style under Common Musical Events Using Pitch-Dynamics Trajectories and Modified TRACLUS ClusteringabstractWe present a novel method for visualising the singing style of vocalists. To illustrate our method, we take 26 audio recordings of A cappella solo vocal music from two different professional singers and we visualise the performance style of each vocalist in a two-dimensional space of pitch and dynamics. We use our own novel modification of a trajectory clustering algorithm called TRACLUS to generate four representative paths, called trajectories, in that two dimensional space. Each trajectory represents the characteristic style of a vocalist during one of four common musical events: (1) Crescendo, (2) Diminuendo, (3) Ascending Pitches and (4) Descending Pitches. The unique shapes of these trajectories characterize the singing style of each vocalist with respect to each of these events. We present the details of our modified version of the TRACULUS algorithm and demonstrate graphically how the plots produced indicate distinct stylistic differences between singers. Potential applications for this method include: (a) automatic identification of singers and automatic classification of singing styles and (b) automatic retargeting of performance style to add human expression to computer generated vocal performances and allow singing synthesisers to imitate the styles of specific famous professional vocalists. Kin Wah Edward Lin, Hans Anderson, Natalie Agus, Clifford So, Simon Lui |
ICMLA | 5 |
| 2014 | Modelling Mutual Information between Voiceprint and Optimal Number of Mel-Frequency Cepstral Coefficients in Voice DiscriminationabstractIn this paper, we study the relationship between the voiceprint and the optimal number of Mel-frequency Cepstral Coefficients (MFCCs). The voiceprint is modelled as sub-MFCCs matrix with the first d number of MFCCs. We model the relationship through information theory and formulate it as the mutual information maximization problem subject to the probabilities constraint. The solution of this optimization problem provides the optimal number of MFCCs, D among these d, which yields the highest classification accuracy of the voice discrimination, together with a confidence level. This study is dictated by the need to understand the use of MFCCs, which have proliferated since its invention to discriminate voice. We evaluate our model by comparing the leave-one-out cross validation (LOOCV) results of usual multi-class classifier, the Supervised Learning Gaussian Mixture Model (SLGMM), with a set of spoken words and A capella solo vocal performances. The experimental results show that our model is a more comprehensive feature selection criteria for the MFCCs than the de-facto technique, LOOCV. Kin Wah Edward Lin, Tian Feng 0001, Natalie Agus, Clifford So, Simon Lui |
ICMLA | 5 |
| 2012 | Sensor-enabled yo-yos as new musical instrumentsabstractAn interactive and programmable musical yo-yo system is designed. The aim of it is to demonstrate the feasibility of converting any sensor-enabled objects into potential musical instruments. This involves three design phases. First, the physical yo-yo is designed to house Iris sensors. The software is developed to sense the movement of yo-yo and transmit its measurements to Max/MSP for corresponding music generation. Finally, aurally pleasing and real-time musical sounds are designed and generated in effect of yo-yo by the computer music composer. Jung-Hyun Jun, Sunardi, Joel W. Matthys, Simon Lui, Yu Gu 0001 |
IPSN | 5 |