VLDB 2026 Research / reviewers in the wild / expert
Wei Li 0012
dblp:64/6025-12
· DBLP profile ↗
56ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0002-4486-8341ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 29 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Computer networks · 5 · 2 first-author · 5 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Capacity of Cooperative Networks with Local Traffic Patterns
Wei Li 0012, Min Sheng, Junyu Liu, Yang Zheng 0003, Jiandong Li 0001 |
ICC | 1 |
| 2026 | Efficient Music Denoising with Channel Attention and Multi-Scale Sequence EncodingabstractMusic denoising aims to suppress unwanted noise while preserving the fidelity of musical signals. In real-world recordings such as pub performances, street busking, and concerts, smartphone audio is often degraded by complex and non-stationary noise sources (e.g., cheering, chatter, and environmental sounds), making robust and efficient denoising particularly challenging. While recent neural architectures developed for audio separation and enhancement have shown promise, their direct application to music denoising remains underexplored, especially under realistic recording conditions. To address this gap, we propose MusicTLE-U, a lightweight U-shaped denoising architecture built on a band-split backbone and augmented with two complementary modules: TLE (Temporal-LSTM-ECA) and MSSE (Multi-Scale Sequence Encoding). The TLE module improves temporal modeling by combining fully connected layers, an LSTM, and Efficient Channel Attention in a compact manner, while MSSE captures hierarchical global context through downsampling, multi-head attention, and upsampling operations. Experiments conducted on real-world noisy music mixtures demonstrate that MusicTLE-U achieves competitive or improved performance in objective evaluation metrics compared to strong baseline models, while maintaining computational efficiency. To support reproducibility, we release the training and evaluation protocol along with code and audio examples. Seungmin Ha, Wei Li 0012, Yulun Wu 0002 |
ICMR | 2 |
| 2026 | Teaching Audio-Language Models to Reason over TimeabstractLarge Audio-Language Models (LALMs) have achieved remarkable progress in general audio understanding. Nevertheless, they still exhibit significant limitations when reasoning about the complex temporal relationships between sound events. This bottleneck arises from two key problems: the lack of a systematic benchmark for evaluation and the absence of a reasoning-oriented training paradigm for audio. To resolve this, we introduce TASC-Bench (Temporal Audio Scene Comprehension Benchmark), a large-scale question-answering (QA) benchmark focused on the temporal understanding of audio scenes, built upon AudioSet. TASC-Bench comprises over 320 hours of audio and 9.5 million QA pairs, covering seven dimensions, including temporal boundaries, event durations, and inter-event relations. To teach LALMs to reason over time and overcome the limitations of Supervised Fine-Tuning (SFT) in tackling complex temporal reasoning, we propose a cross-modal Chain-of-Thought (CoT) distillation strategy. This progressive approach begins by bootstrapping the model’s reasoning capabilities using a large language model (LLM) as a teacher to generate CoT, then transitions to self-driven reasoning, and culminates in aligning its behavior with preference data. For this final alignment stage, we introduce our improved Equilibrated Chain Direct Preference Optimization (EC-DPO) method alongside SFT to mitigate hallucinations and enhance stability. Our experiments demonstrate that our dataset reveals the current shortcomings of LALMs in temporal reasoning, while our method significantly enhances these capabilities. We will make the TASC-Bench dataset publicly available to foster further research in audio temporal reasoning. Yunjia Li, Wei Li 0012 |
ICMR | 4 |
| 2025 | A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary RegularizationabstractSinging melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequency-supervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method1. Xiaoliang He, Kangjie Dong, Jingkai Cao, Shuai Yu 0002, Wei Li 0012, Yi Yu 0001 |
ICASSP | 5 |
| 2025 | Ultra Lightweight Singing Melody Extraction via Combination of Convolution and MLPabstractSinging melody extraction serves as an important foundation in the realm of music information retrieval (MIR). Although fully convolutional neural networks (CNNs) are commonly employed for singing melody extraction, they are constrained by inductive biases and face challenges in establishing long range dependency. Transformer-based networks have better performance, but the computational load is high. Recently, many multi-layer perceptron (MLP) architectures have been applied for a variety of computer vision tasks, demonstrating competitive performance. However, its potential ability in the task of singing melody extraction remains to be further explored. In this paper, we propose the lightweight convolutional MLP (LcMLP), an ultra lightweight model without sacrificing the performance. Firstly, we improve the original MLP-Mixer. We change the sequential MLPs to parallel ones and add some skip connections. Secondly, we propose a multi-level convolution fusion module that facilitates the interaction of features at various depths in MLP-Mixer. We conducted extensive experiments on several well-known public datasets, and our model demonstrates significant advantages in inference speed and computational load, while also achieving competitive performance. Kangjie Dong, Qiubo Huang, Shuai Yu 0002, Wei Li 0012 |
ICASSP | 5 |
| 2025 | BeatKAN: An Efficient and Drum-Attuned Beat Tracking Method Using Kolmogorov-Arnold NetworksabstractIn this paper, we propose an efficient and drum-attuned beat tracking method based on Kolmogorov-Arnold networks (KAN). Traditional MLP-based frameworks struggle with complex musical signals due to limited capacity in modeling intricate patterns. Inspired by KAN’s efficient ability to capture complex time-frequency relationships, we leverage it to enhance our model by employing learnable nonlinear activation functions on convolutional kernels. Additionally, we utilize music source separation techniques to extract drum tracks from the original audio, thereby expanding the existing beat datasets and simulating the human perception of beats through drum sounds. Experimental results demonstrate that our approach significantly reduces the parameter count while maintaining high accuracy in beat tracking. Ganghui Ru, Wei Li 0012 |
ICASSP | 3 |
| 2025 | KCE-Unet: A novel music denoising method with KANConv ECA UnetabstractDuring concerts, people often spontaneously record memorable moments with their phones. However, these recordings are frequently accompanied by noise, such as cheering and applause, which diminishes the playback experience. In this paper, we introduce a novel task specifically designed for denoising music in concert environments, a challenge that has been largely overlooked in previous research. To support this task, we created a new concert denoising dataset that includes songs performed in various major languages at concerts, with noise segments like cheering and applause. Building on this, we propose KANConv ECA Unet (KCE-Unet), a method that combines the U-Net network, efficient channel attention (ECA), and the recently proposed KAN network to flexibly remove noise in the mid-to-high frequency range of spectrograms. Extensive experiments demonstrate that our method outperforms previous models in denoising performance and effectively restore disrupted musical structures. Yulun Wu 0002, Ganghui Ru, Yi Yu 0001, Wei Li 0012 |
ICASSP | 5 |
| 2025 | HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat TrackingabstractFine-tuning pre-trained foundation models has made significant progress in music information retrieval. However, applying these models to beat tracking tasks remains unexplored as the limited annotated data renders conventional fine-tuning methods ineffective. To address this challenge, we propose HingeNet, a novel and general parameter-efficient fine-tuning method specifically designed for beat tracking tasks. HingeNet is a lightweight and separable network, visually resembling a hinge, designed to tightly interface with pre-trained foundation models by using their intermediate feature representations as input. This unique architecture grants HingeNet broad generalizability, enabling effective integration with various pre-trained foundation models. Furthermore, considering the significance of harmonics in beat tracking, we introduce harmonic-aware mechanism during the fine-tuning process to better capture and emphasize the harmonic structures in musical signals. Experiments on benchmark datasets demonstrate that HingeNet achieves state-of-the-art performance in beat and downbeat tracking. Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012 |
ICME | 8 |
| 2025 | BeatFM: Improving Beat Tracking with Pre-trained Music Foundation ModelabstractBeat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles and accurately capture complex rhythmic structures. To overcome these challenges, we propose a novel beat tracking paradigm BeatFM, which introduces a pre-trained music foundation model and leverages its rich semantic knowledge to improve beat tracking performance. Pre-training on diverse music datasets endows music foundation models with a robust understanding of music, thereby effectively addressing these challenges. To further adapt it for beat tracking, we design a plug-and-play multi-dimensional semantic aggregation module, which is composed of three parallel sub-modules, each focusing on semantic aggregation in the temporal, frequency, and channel domains, respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance in beat and downbeat tracking across multiple benchmark datasets. Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012 |
ICME | 8 |
| 2025 | Robust Throughput Capacity of Multi-Connectivity Wireless NetworksabstractIn this paper, we study the robust throughput capacity of multi-connectivity wireless networks when the network encounters zone node failures. In order to reveal the inherent relationship between the robustness of the network structure and the capability of wireless networks to carry information, robust throughput capacity, which is the product of the fraction of served source and destination (S-D) pairs, the number of S-D pairs and feasible throughput, is defined. It is shown that the robust throughput capacity is$\Theta \left ({{\sqrt {\frac {n}{k\log n}}}}\right)$for$\beta \gt 2$and$\Theta \left ({{\frac {1}{k\log n}\sqrt {\frac {n}{k\log n}}}}\right)$for$1 {\lt }\beta \leq 2$, where n is the number of nodes,$\beta $is the failure exponent and$k\:(\geq 1)$is the connectivity parameter representing the number of disjoint data paths between any two nodes. To balance the tradeoff between the throughput capacity and the robustness of the network structure, the feasible regions of connectivity parameters, which are limited by the failure exponent, are given for$\beta \gt 2$and$1\lt \beta \leq 2$, respectively. Correspondingly, the robust throughput capacity is$\Theta \left ({{\sqrt {\frac {n^{1-\gamma }}{\log n}}}}\right)$and$\Theta \left ({{\sqrt {\frac {n^{1-3\gamma }}{\log n}}}}\right)$, respectively, where$\gamma \in [0,1$) is the robustness exponent. These results can provide guidance for designing network protocols with fault tolerance in large-scale wireless networks. Min Sheng, Wei Li 0012, Junyu Liu, Jiandong Li 0001 |
IEEE Trans. Commun. | 2 |
| 2024 | Mertech: Instrument Playing Technique Detection Using Self-Supervised Pretrained Model with Multi-Task FinetuningabstractInstrument playing techniques (IPTs) constitute a pivotal component of musical expression. However, the development of automatic IPT detection methods suffers from limited labeled data and inherent class imbalance issues. In this paper, we propose to apply a self-supervised learning model pre-trained on large-scale unlabeled music data and finetune it on IPT detection tasks. This approach addresses data scarcity and class imbalance challenges. Recognizing the significance of pitch in capturing the nuances of IPTs and the importance of onset in locating IPT events, we investigate multi-task finetuning with pitch and onset detection as auxiliary tasks. Additionally, we apply a post-processing approach for event-level prediction, where an IPT activation initiates an event only if the onset output confirms an onset in that frame. Our method outperforms prior approaches in both frame-level and event-level metrics across multiple IPT benchmark datasets. Further experiments demonstrate the efficacy of multi-task finetuning on each IPT class.1 Dichucheng Li, Yinghao Ma, Weixing Wei, Qiuqiang Kong, Yulun Wu 0002, Mingjin Che, Emmanouil Benetos, Wei Li 0012 |
ICASSP | 9 |
| 2024 | A Scalable Sparse Transformer Model for Singing Melody ExtractionabstractExtracting the melody of a singing voice is an essential task within the realm of music information retrieval (MIR). Recently, transformer based models have drawn great attention in the field of MIR. However, due to the expensive computation cost and extensive parameters, it is difficult to train and deploy a transformer-based model for practical singing melody extraction. In this paper, we propose a simple yet effective scalable sparse transformer for singing melody extraction. To be specific, we first propose to employ a sparse transformer to reduce computation cost and the amount of parameters. Then, we proposed to scale the self-attention region of the sparse transformer in the spectrogram to obtain more accurate performance. Moreover, we propose to combine a scalable sparse transformer (S2Former) with CNN-based model to extract global and local features in the spectrogram. The proposed scalable transformer model can achieve a better balance between a standard transformer and a sparse transformer. To better fuse the features from transformer and CNN, we further propose a transformer-CNN fusion (TCF) module to combine significant features from transformer and CNN. The proposed model obtains state-of-the-art results on several public datasets. The conducted experiments confirm the effectiveness of the model we proposed. Shuai Yu 0002, Yi Yu 0001, Wei Li 0012 |
ICASSP | 4 |
| 2024 | Improving Drum Source Separation with Temporal-Frequency Statistical DescriptorsabstractDrum Source Separation (DSS) aims to separate drum mixtures into individual drum sounds, such as kick and snare. Deep neural network methods have been successfully applied for source separation. However, due to the limited size of existing datasets and the strongly overlap of drums in frequency and time, these methods still have certain shortcomings. To address these challenges, we construct a large drum sound dataset and propose a novel training objective to improve performance of DSS task. The training objective leverages three temporal-frequency statistical descriptors (spectral centroid, spectral spread, and spectral flux) to separate drum sources. Our experimental results demonstrate that our method can make a SDR improvement of 0.98 dB on UNet and 1.07 dB on MERT. Furthermore, our method achieves consistent improvements in low-resource and cross-dataset scenarios. Our code and dataset are available at https://github.com/150042/Drum-Separation-TF. Dichucheng Li, Xinlu Liu, Yongwei Gao, Wei Li 0012 |
ICME | 7 |
| 2024 | Harmonic Frequency-Separable Transformer for Instrument-Agnostic Music TranscriptionabstractAutomatic Music Transcription (AMT) aims to convert music audio into symbolic representations. Recently, transformer-based methods have been successfully applied to instrument-agnostic music transcription. This allows transcription models can no longer focus on specific characteristics for an instrument class. However, these transformer-based methods designs for AMT were mainly motivated by other research fields and uses additional large-scale datasets, without considering the intrinsic features and patterns of the music signals. In this paper, we propose the Harmonic Frequency-Separable Transformer (HFSFormer), providing effective prior information based on music knowledge for instrument-agnostic transcription. The HFSFormer can capture the harmonic structure of music and separate time-frequency representations to decouple multiple pitches and different timbres, which can better explicitly model the note’s onset/offset and pitch. Experimental results show that our proposed method outperforms state-of-the-art peers on public datasets while having an order of magnitude fewer parameters. Yulun Wu 0002, Weixing Wei, Dichucheng Li, Mengbo Li, Yi Yu 0001, Yongwei Gao, Wei Li 0012 |
ICME | 7 |
| 2023 | Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention MechanismabstractInstrument playing technique (IPT) is a key element of musical presentation. However, most of the existing works for IPT detection only concern monophonic music signals, yet little has been done to detect IPTs in polyphonic instrumental solo pieces with overlapping IPTs or mixed IPTs. In this paper, we formulate it as a frame-level multi-label classification problem and apply it to Guzheng, a Chinese plucked string instrument. We create a new dataset, Guzheng Tech99, containing Guzheng recordings and onset, offset, pitch, IPT annotations of each note. Because different IPTs vary a lot in their lengths, we propose a new method to solve this problem using multi-scale network and self-attention. The multi-scale network extracts features from different scales, and the self-attention mechanism applied to the feature maps at the coarsest scale further enhances the long-range feature extraction. Our approach outperforms existing works by a large margin, indicating its effectiveness in IPT detection. Dichucheng Li, Mingjin Che, Wenwu Meng, Yulun Wu 0002, Yi Yu 0001, Wei Li 0012 |
ICASSP | 7 |
| 2023 | LC-Beating: An Online System for Beat and Downbeat Tracking using Latency-Controlled MechanismabstractBeat and downbeat tracking is to predict beat and downbeat time steps from a given music piece. Some deep learning models with a dilated structure such as Temporal Convolutional Network (TCN) and Dilated Self-Attention Network (DSAN) have achieved promising performance for this task. However, most of them have to see the whole music context during inference, which limits their deployment to online systems. In this paper, we propose LC-Beating, a novel latency-controlled (LC) mechanism for online beat and downbeat tracking, in which the model only looks ahead a few frames. By appending limited future information, the model can better capture the activity of relevant musical beats, which significantly boosts the performance of online algorithms with limited latency. Moreover, LC-Beating applies a novel real-time implementation of the LC mechanism to TCN and DSAN. The experimental results show that our proposed method outperforms the previous online models by a large margin and is close to the results of the offline models. Xinlu Liu, Jiale Qian, Qiqi He, Yi Yu 0001, Wei Li 0012 |
ICME | 5 |
| 2023 | MFAE: Masked frame-level autoencoder with hybrid-supervision for low-resource music transcriptionabstractAutomantic Music Transcription (AMT) is an essential topic in music information retrieval (MIR), and it aims to transcribe audio recordings into symbolic representations. Recently, large-scale piano datasets with high-quality notations have been proposed for high-resolution piano transcription, which resulted in domain-specific AMT models achieved state-of- the-art results. However, those methods are hardly generalized to other ’low-resource’ instruments (such as guitar, cello, clarinet, etc.) transcription. In this paper, we propose a hybrid-supervised framework, the masked frame-level autoencoder (MFAE), to solve this issue. The proposed MFAE reconstructs the frame-level features of low-resource data to understand generic representations of low-resource instruments and improves low-resource transcription performance. Experimental results on several low- resource datasets (MAPS, MusicNet, and Guitarset) show that our framework achieves state-of-the-art performance in note-wise scores (Note F1 83.4%\64.1%\86.7%, Note-with-offset F1 59.8%\41.4%\71.6%). Moreover, our framework can be well generalized to various genres of instrument transcription, both in data-plentiful and data-limited scenarios. Yulun Wu 0002, Yi Yu 0001, Wei Li 0012 |
ICME | 4 |
| 2023 | The capacity of k-connectivity d-dimensional wireless networks with node failure
Wei Li 0012, Junyu Liu, Min Sheng, Jiandong Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2023 | A neural harmonic-aware network with gated attentive fusion for singing melody extraction
Shuai Yu 0002, Yi Yu 0001, Xiaoheng Sun, Wei Li 0012 |
Neurocomputing | 4 |
| 2023 | Multi-scale network with shared cross-attention for audio-visual correlation learning
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Wei Li 0012 |
Neural Comput. Appl. | 4 |
| 2023 | Variational Autoencoder with CCA for Audio-Visual Cross-modal RetrievalabstractCross-modal retrieval is to utilize one modality as a query to retrieve data from another modality, which has become a popular topic in information retrieval, machine learning, and databases. Finding a method to effectively measure the similarity between different modality data is the major challenge of cross-modal retrieval. Although several research works have calculated the correlation between different modality data via learning a common subspace representation, the encoder’s ability to extract features from multi-modal information is not satisfactory. In this article, we present a novel variational autoencoder architecture for audio–visual cross-modal retrieval by learning paired audio–visual correlation embedding and category correlation embedding as constraints to reinforce the mutuality of audio–visual information. On the one hand, audio encoder and visual encoder separately encode audio data and visual data into two different latent spaces. Further, two mutual latent spaces are respectively constructed by canonical correlation analysis. On the other hand, probabilistic modeling methods are used to deal with possible noise and missing information in the data. Additionally, in this way, the cross-modal discrepancies from intra-modal and inter-modal information are simultaneously eliminated in the joint embedding subspace. We conduct extensive experiments over two benchmark datasets. The experimental results confirm that the proposed architecture is effective in learning audio–visual correlation and is appreciably better than the existing cross-modal retrieval methods. Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Wei Li 0012 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Melody Generation from Lyrics with Local InterpretabilityabstractMelody generation aims to learn the distribution of real melodies to generate new melodies conditioned on lyrics, which has been a very interesting topic in the area of artificial intelligence and music. However, a challenging issue still limits the quality and reliability of melody generation conditioned on lyrics: how to enhance the interpretability between the input lyrics and generated melodies so humans can understand their relationships. To solve this issue, in this article, we propose a model for melody generation from lyrics with local interpretability, which contains two significant contributions: (i) Mutual information between input lyrics and generated melody is exploited to instruct the training of the network, which avoids the loss of content consistency during the training stage. (ii) Transformer is explored to efficiently extract semantic features from lyrics sequences, which provides more interpretable correlations between different syllables in lyrics. Experiments on a large-scale dataset with paired lyrics-melodies demonstrate that the proposed approach can generate higher-quality melodies from lyrics compared with existing methods. Wei Duan 0004, Yi Yu 0001, Xulong Zhang 0001, Suhua Tang, Wei Li 0012, Keizo Oyama |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Robust Capacity of Wireless Networks Under Cascading FailuresabstractIn this paper, we study the impact of node cascading failures and network structure robustness on the capacity of wireless networks. Especially, in order to quantify how much information can be conveyed by wireless networks under node cascading failures, robust capacity is defined to capture the influence of the intensity of the initial failure nodes$m$and the connectivity parameter$k$on capacity. Note that increasing$k$could provide$k$disjoint data paths for any two nodes, thereby combating the cascading failure. Denoting the intensity of the nodes as$n$and$m=n^{\frac{1}{\beta}}$, it is shown that robust capacity$O\left(\sqrt{\frac{n}{k\log n}}\right)$can be achieved when the initial failure exponent$\beta > 2$. In contrast, robust capacity would converge to zero with increasing$n$when$1 < \beta\leq 2$since the data paths of most source-destination pairs are interrupted due to the cascading failures. Moreover, increasing the connectivity parameter$k$, although capable of enhancing the network structure robustness, is shown to degrade cascading failures and robust capacity. Wei Li 0012, Junyu Liu, Min Sheng, Jiandong Li 0001 |
GLOBECOM | 1 |
| 2022 | Tonet: Tone-Octave Network for Singing Melody Extraction from Polyphonic MusicabstractSinging melody extraction is an important problem in the field of music information retrieval. Existing methods typically rely on frequency-domain representations to estimate the sung frequencies. However, this design does not lead to human-level performance in the perception of melody information for both tone (pitch-class) and octave. In this paper, we propose TONet1, a plug-and-play model that improves both tone and octave perceptions by leveraging a novel input representation and a novel network architecture. First, we present an improved input representation, the Tone-CFP, that explicitly groups harmonics via a rearrangement of frequency-bins. Second, we introduce an encoder-decoder architecture that is designed to obtain a salience feature map, a tone feature map, and an octave feature map. Third, we propose a tone-octave fusion mechanism to improve the final salience feature map. Experiments are done to verify the capability of TONet with various baseline backbone models. Our results show that tone-octave fusion with Tone-CFP can significantly improve the singing voice extraction performance across various datasets – with substantial gains in octave and tone accuracy. Ke Chen 0021, Shuai Yu 0002, Cheng-i Wang, Wei Li 0012, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 4 |
| 2022 | Deepchorus: A Hybrid Model of Multi-Scale Convolution And Self-Attention for Chorus DetectionabstractChorus detection is a challenging problem in musical signal processing as the chorus often repeats more than once in popular songs, usually with rich instruments and complex rhythm forms. Most of the existing works focus on the receptiveness of chorus sections based on some explicit features such as loudness and occurrence frequency. These pre-assumptions for chorus limit the generalization capacity of these methods, causing misdetection on other repeated sections such as verse. To solve the problem, in this paper we propose an end-to-end chorus detection model DeepChorus, reducing the engineering effort and the need for prior knowledge. The proposed model includes two main structures: i) a Multi-Scale Network to derive preliminary representations of chorus segments, and ii) a Self-Attention Convolution Network to further process the features into probability curves representing chorus presence. To obtain the final results, we apply an adaptive threshold to binarize the original curve. The experimental results show that DeepChorus outperforms existing state-of-the-art methods in most cases. Qiqi He, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012 |
ICASSP | 4 |
| 2022 | Hierarchical Graph-Based Neural Network for Singing Melody ExtractionabstractSinging melody extraction from polyphonic music is a critical and challenging task in music information retrieval (MIR). However, due to the interfere of the accompaniment and the background noise, it is key and challenging to obtain a global semantic representation that discriminates the singing melody line. To address this issue, we consider the two aspects that regards to obtaining the global semantic representation: the global relationships in the spectrum and the relationships between channels. In this paper, we propose a novel hierarchical graph-based network for singing melody extraction. In particular, according to its characteristics of the spectrum, we first model the spectrum into graph structure, a two-layer graph convolution network is used to obtain the global semantic representation in the spectrum. Then to capture the relationships between channels, channel-wise graph convolution module is devised to capture and reasoning the relationship between channels. The conducted experiments demonstrate the effectiveness of the proposed network. Shuai Yu 0002, Wei Li 0012 |
ICASSP | 3 |
| 2022 | A Glance-and-Gaze Network for Respiratory Sound ClassificationabstractA plethora of great successes has been achieved by the existing convolutional neural networks (CNN) for respiratory sound classification. Nevertheless, simultaneously capturing both the local and global features can never be an easy task due to the limitation of a CNN’s structure. In this contribution, we propose a novel glance-and-gaze network to address the aforementioned issue. The glance block aims to learn global information, while the gaze block is responsible for learning local patterns and suppressing the noises that attenuates the final performance. In the proposed method, both the global and local information can be extracted. Moreover, the spectral and temporal representations can be learnt via a feature fusion module. Experimental results on the largest public respiratory sound database demonstrate that the proposed model outperforms the state-of-the-art methods. Shuai Yu 0002, Yiwei Ding, Kun Qian 0003, Bin Hu 0001, Wei Li 0012, Björn W. Schuller |
ICASSP | 5 |
| 2022 | HarmoF0: Logarithmic Scale Dilated Convolution for Pitch EstimationabstractSounds, especially music, contain various harmonic components scattered in the frequency dimension. It is difficult for normal convolutional neural networks to ob-serve these overtones. This paper introduces a multiple rates dilated causal convolution (MRDC-Conv) method to capture the harmonic structure in logarithmic scale spectrograms efficiently. The harmonic is helpful for pitch estimation, which is important for many sound processing applications. We propose HarmoF0, a fully convolutional network, to evaluate the MRDC-Conv and other dilated convolutions in pitch estimation. The re-sults show that this model outperforms the DeepF0, yields state-of-the-art performance in three datasets, and simultaneously reduces more than 90% parameters. We also find that it has stronger noise resistance and fewer octave errors. Weixing Wei, Yi Yu 0001, Wei Li 0012 |
ICME | 4 |
| 2022 | Multimodal Music Emotion Recognition with Hierarchical Cross-Modal Attention NetworkabstractComputational music emotion recognition is to recognize the emotional content in music tracks. In computational music emotion recognition studies, researchers have paid close attention to the audio content of the music tracks. Although lyrics content and music context contribute greatly to the perceived emotion, these kinds of emotional information are usually ignored. Based on this finding, we propose a multimodal music emotion recognition method jointly predicting the valence and arousal values by combining the audio, lyrics, track name, and artist of a given track. Audio features, lyrics features and context features are extracted separately and fused by a cross-modal attention mechanism, forming a hierarchical structure. Our proposed model outperforms two baselines by a large margin and achieves state-of-the-art performance on two public datasets. Ganghui Ru, Yi Yu 0001, Yulun Wu 0002, Dichucheng Li, Wei Li 0012 |
ICME | 6 |
| 2022 | Singing Voice Detection via Similarity-Based Semi-Supervised LearningabstractData-driven methods play an important role in Singing Voice Detection (SVD). However, datasets with precise annotations are scarce. In this paper, we propose an SVD method via similarity-based semi-supervised learning (SSSL_SVD). For one thing, we propose to enrich the diversity of training data using the self-training semi-supervised method (SSL). In SSL, pseudo labels of the unlabeled data are first generated by a pre-trained teacher model and are then used to train a student model. For another thing, we propose to measure the audio frame from a similarity-based perspective. Taking it into consideration, we could provide more appropriate learning targets. Finally, experiment results indicate that the proposed method achieved comparable results with state-of-the-art (SOTA) algorithms. Yongwei Gao, Wei Li 0012 |
MMAsia | 3 |
| 2022 | Melody Generation from Lyrics Using Three Branch Conditional LSTM-GAN
Wei Duan 0004, Rajiv Ratn Shah, Suhua Tang, Wei Li 0012, Yi Yu 0001 |
MMM (1) | 6 |
| 2022 | A personality-guided affective brain - computer interface for implementation of emotional intelligence in machinesabstractAffective brain—computer interfaces have become an increasingly important topic to achieve emotional intelligence in human—machine collaboration. However, due to the complexity of electroencephalogram (EEG) signals and the individual differences in emotional response, it is still a great challenge to design a reliable and effective model. Considering the influence of personality traits on emotional response, it would be helpful to integrate personality information and EEG signals for emotion recognition. This study proposes a personality-guided attention neural network that can use personality information to learn effective EEG representations for emotion recognition. Specifically, we first use a convolutional neural network to extract rich temporal and regional representations of EEG signals, and a special convolution kernel is designed to learn inter- and intra-regional correlations simultaneously. Second, inspired by the fact that electrodes within distinct brain scalp regions play different roles in emotion recognition, a personality-guided regional-attention mechanism is proposed to further explore the contributions of electrodes within a region and between regions. Finally, attention-based long short-term memory is designed to explore the temporal dynamics of EEG signals. Experiments on the AMIGOS dataset, which is a dataset for multimodal research for affect, personality traits, and mood on individuals and groups, show that the proposed method can significantly improve the performance of subject-independent emotion recognition and outperform state-of-the-art methods. Wei Li 0012, Zejian Xing, Wenjie Yuan 0001, Xiaowei Zhang 0001, Bin Hu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | An Hrnet-Blstm Model With Two-Stage Training For Singing Melody ExtractionabstractWell-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate a large number of fundamental frequency (F0) values for model training. Since the refined pitch values used for the first round of training contain errors, a small set of well-labeled data is used for a second round of training. A high-resolution network (HRNet), initially developed for human pose estimation, is introduced for melody extraction. It considers multi-resolution feature learning, making the resulting representation semantically richer. Subsequently, a bidirectional long short-term memory (BLSTM) layer is used to exploit the temporal information of melody. In addition, a new loss function where the unvoiced frames only contribute to voicing detection but not to pitch classification is also proposed to alleviate the class imbalance problem. Experiment results on three public datasets show that the proposed system outperforms four state-of-the-art algorithms in most cases. Yongwei Gao, Xingjian Du, Bilei Zhu, Xiaoheng Sun, Wei Li 0012, Zejun Ma 0001 |
ICASSP | 5 |
| 2021 | Frequency-Temporal Attention Network for Singing Melody ExtractionabstractMusical audio is generally composed of three physical properties: frequency, time and magnitude. Interestingly, human auditory periphery also provides neural codes for each of these dimensions to perceive music. Inspired by these intrinsic characteristics, a frequency-temporal attention network is proposed to mimic human auditory for singing melody extraction. In particular, the proposed model contains frequency-temporal attention modules and a selective fusion module corresponding to these three physical properties. The frequency attention module is used to select the same activation frequency bands as did in cochlear and the temporal attention module is responsible for analyzing temporal patterns. Finally, the selective fusion module is suggested to recalibrate magnitudes and fuse the raw information for prediction. In addition, we propose to use another branch to simultaneously predict the presence of singing voice melody. The experimental results show that the proposed model outperforms existing state-of-the-art methods1. Shuai Yu 0002, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012 |
ICASSP | 4 |
| 2021 | Singer Identification Using Deep Timbre Feature Learning with KNN-NETabstractIn this paper, we study the issue of automatic singer identification (SID) in popular music recordings, which aims to recognize who sang a given piece of song. The main challenge for this investigation lies in the fact that a singer’s singing voice changes and intertwines with the signal of background accompaniment in time domain. To handle this challenge, we propose the KNN-Net for SID, which is a deep neural network model with the goal of learning local timbre feature representation from the mixture of singer voice and background music. Unlike other deep neural networks using the softmax layer as the output layer, we instead utilize the KNN as a more interpretable layer to output target singer labels. Moreover, attention mechanism is first introduced to highlight crucial timbre features for SID. Experiments on the existing artist20 dataset show that the proposed approach outperforms the state-of-the-art method by 4%. We also create singer32 and singer60 datasets consisting of Chinese pop music to evaluate the reliability of the proposed method. The more extensive experiments additionally indicate that our proposed model achieves a significant performance improvement compared to the state-of-the-art methods. Xulong Zhang 0001, Jiale Qian, Yi Yu 0001, Yifu Sun, Wei Li 0012 |
ICASSP | 5 |
| 2021 | HANME: Hierarchical Attention Network for Singing Melody ExtractionabstractSinging melody extraction in polyphonic musical audio is a very critical and challenging task in music information retrieval (MIR). Contextual frame-level information has proven its effectiveness in this task. However, existing works assign equal weight to each contextual frame, which may hinder the further improvement of the performance. To this end, we propose a hierarchical attention network for singing melody extraction (HANME) to extract the discriminative attention-aware features and alleviate the workload of the convolutional recurrent neural network (CRNN) for extracting local spatial and temporal features. Specifically, the first attention layer learns the context vector based on local spatial features extracted by residual convolutional neural network (CNN), and the second attention layer learns the temporal context vector based on long-term features extracted by Bidirectional Gated Recurrent Units (BiGRU). Due to the scarcity of labeled training data, we further propose a partial parameter adaptation approach to address the imbalance distribution of the labels for this task. We use the RWC dataset and part of vocal tracks of the MedleyDB dataset for training the model and evaluate the performance on the ADC2004, MIREX 05 and MedleyDB datasets. The experimental study demonstrates the superiority of our method compared with other state-of-the-art ones. Shuai Yu 0002, Yi Yu 0001, Wei Li 0012 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Residual Attention Based Network for Automatic Classification of Phonation ModesabstractPhonation mode is an essential characteristic of singing style as well as an important expression of performance. It can be classified into four categories, called neutral, breathy, pressed and flow. Previous studies used voice quality features and feature engineering for classification. While deep learning has achieved significant progress in other fields of music information retrieval (MIR), there are few attempts in the classification of phonation modes. In this study, a Residual Attention based network is proposed for automatic classification of phonation modes. The network consists of a convolutional network performing feature processing and a soft mask branch enabling the network focus on a specific area. In comparison experiments, the models with proposed network achieve better results in three of the four datasets than previous works, among which the highest classification accuracy is 94.58%, 2.29% higher than the baseline. Xiaoheng Sun, Yiliang Jiang, Wei Li 0012 |
ICME | 3 |
| 2019 | Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch RefinementabstractData-driven methods for melody extraction from polyphonic music generally require large amounts of labeled data for model training. However, musical data with annotations of melody fundamental frequency (F0) are rare and hard to obtain. To overcome this limitation, in this paper we propose to use melody MIDI files, which are more massively available, as the sources of labels to train a deep neural network (DNN) model for melody extraction. For each testing audio, the pitch sequence estimated by DNN is comprised of note numbers quantized at semitone level, and their resolution is relatively low. Therefore, we further propose a salience-based method to refine the pitch estimate of DNN to a higher resolution of 10 cents. Experimental results on three public datasets indicate that our method outperforms four state-of-the-art melody extraction methods in most cases. Yongwei Gao, Bilei Zhu, Wei Li 0012, Ke Li 0015, Yongjian Wu 0001, Feiyue Huang |
ICASSP | 3 |
| 2019 | Automatic Audio Chord Recognition With MIDI-Trained Deep Feature and BLSTM-CRF Sequence Decoding ModelabstractWith the advances of machine learning technologies, data-driven feature extraction and sequence modeling approaches are being widely explored for automatic chord recognition tasks. Currently, there is a bottleneck in the amount of enough annotated data for training robust acoustic models, as hand-annotating time-synchronized chord labels requires professional musical skills and considerable labor. To cope with this limitation, in this paper, we propose a convolutional neural network (CNN) based deep feature extractor, which is trained on a large set of time, synchronized musical instrument digital interface audio data pairs and can robustly estimate pitch class activations of real-world music audio recordings. The CNN feature extractor plus a bidirectional long short-term memory conditional random field decoding model forms the proposed hybrid system for automatic chord recognition. Experiments show that the proposed model is compatible for both regular major/minor triad chord classification and larger vocabulary chord recognition, and outperforms other state-of-the-art chord recognition systems. Yiming Wu 0003, Wei Li 0012 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Music Chord Recognition Based on Midi-Trained Deep Feature and BLSTM-CRF Hybird DecodingabstractIn this paper, we design a novel deep learning based hybrid system for automatic chord recognition. Currently, there is a bottleneck in the amount of enough annotated data for training robust acoustic models, as hand annotating time-synchronized chord labels requires professional musical skills and considerable labor. As a solution to this problem, we construct a large set of time synchronized MIDI-audio pairs, and use these data to train a Deep Residual Network (DRN) feature extractor, which can then estimate pitch class activations of real-world music audio recordings. Sequence classification and decoding are then performed with a trained Bidirectional LSTM and Conditional Random Fields (CRF) network. Experiments show that the proposed model is compatible for both regular major/minor triad chord classification and larger vocabulary chord recognition, the performance is good and no less than other state-of-the-art systems. The proposed system also achieved good evaluation score in MIREX 2017 Automatic Chord Estimation task. Yiming Wu 0003, Wei Li 0012 |
ICASSP | 2 |
| 2015 | Latent time-frequency component analysis: A novel pitch-based approach for singing voice separationabstractMonaural singing voice separation has aroused considerable attention. Many pitch-based methods have been proposed to address this task, but generally have limited performance. The most crucial difficulties lie in the inaccurate judgment on voiced pitches and the failed recognition on unvoiced singing sounds. In this paper, we propose a novel algorithm based on the latent component analysis of time-frequency representation to overcome these difficulties. Specifically, the time-frequency (T-F) representations of the song are firstly decomposed into components, and each component approximately originates from a single sound source. We then construct non-overlapping T-F segments with these components, to complete the omitted useful singing voice information. Extensive experiments on the MIR-1K public dataset shows the effectiveness of the proposed algorithm. Wei Li 0012, Bilei Zhu |
ICASSP | 2 |
| 2015 | Towards Solving the Bottleneck of Pitch-based Singing Voice SeparationabstractSinging voice separation from accompaniment in monaural music recordings is a crucial technique in music information retrieval. A majority of existing algorithms are based on singing pitch detection, and take the detected pitch as the cue to identify and separate the harmonic structure of the singing voice. However, as a key yet undependable premise, vocal pitch detection makes the separation performance of these algorithms rather limited. To overcome the inherent weakness of pitch-based inference algorithms, two novel methods based on non-negative matrix factorization (NMF) are devised in this paper. The first one combines NMF with the distribution regularities of vocals under different time frequency resolutions, so that many vocal unrelated portions are eliminated and the singing voice is hence enhanced. In consequence, the accuracy of vocal pitch detection is significantly improved. The second method applies NMF to decompose the spectrogram into non-overlapping and indivisible segments, which can be used as another cue besides the pitch to help identify the vocal harmonic structure. The two proposed methods are integrated into the framework of pitch-based inference. Extensive testing on the MIR-1K public dataset shows that both of them are rather effective, and the overall performances outperform other state-of-the-art singing separation algorithms. Bilei Zhu, Wei Li 0012 |
ACM Multimedia | 2 |
| 2013 | Multi-Stage Non-Negative Matrix Factorization for Monaural Singing Voice SeparationabstractSeparating singing voice from music accompaniment can be of interest for many applications such as melody extraction, singer identification, lyrics alignment and recognition, and content-based music retrieval. In this paper, a novel algorithm for singing voice separation in monaural mixtures is proposed. The algorithm consists of two stages, where non-negative matrix factorization (NMF) is applied to decompose the mixture spectrograms with long and short windows respectively. A spectral discontinuity thresholding method is devised for the long-window NMF to select out NMF components originating from pitched instrumental sounds, and a temporal discontinuity thresholding method is designed for the short-window NMF to pick out NMF components that are from percussive sounds. By eliminating the selected components, most pitched and percussive elements of the music accompaniment are filtered out from the input sound mixture, with little effect on the singing voice. Extensive testing on the MIR-1K public dataset of 1000 short audio clips and the Beach-Boys dataset of 14 full-track real-world songs showed that the proposed algorithm is both effective and efficient. Bilei Zhu, Wei Li 0012, Ruijiang Li, Xiangyang Xue 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | On the music content authenticationabstractDigital audio has been ubiquitous over the past decade. Since it can be easily modified by editing tools, there has been a strong need to protect its content for secure multimedia applications. Existing audio authentication algorithms are mainly focused on either human speech or general audio with music as part of the test data, while special research on music authentication has been somewhat neglected. In this article, we propose a novel algorithm to protect the integrity and authenticity of music signals. Its main contributions include: (1) Music is segmented into beat-based frames, which not only endows the authentication units with more semantic meaning but also perfectly resolves the challenging synchronization problem; (2) Robust hashes are generated from Chroma-based mid-level audio feature which can appropriately characterize the music content, and integrated with an encryption procedure to ensure the security against malicious block-wise vector quantization attack; (3) Fuzzy logic is adopted to make the authentication decision in light of three measures defined on bit errors, coinciding with the inherent blurred nature of authentication. Experiments exhibit good discriminative ability between admissible and malicious operations. Wei Li 0012, Bilei Zhu, Zhurong Wang |
ACM Multimedia | 1 |
| 2012 | A Double-Ranking Strategy for Long-Tail Product RecommendationabstractIn this paper we attempt to retrieve the items in the long-tail for top-N recommendation. That is, to recommend products that the end-user likes, but that are not generally popular, which has been getting more and more notice lately. By analysing the existing issue of current recommendation algorithms, a strategy is proposed that succeeds in maintaining recommendation accuracy while reducing the concentration of the recommendation on popular items in the system. Evaluating on the publicly available Movie lens and Yahoo! datasets, the results show the recommendation algorithm proposed in this work retrieves items in the users' relatively unpopular tastes without losing the performance in their popular tastes, which ultimately results in a better overall accuracy for the system. Mi Zhang 0001, Neil J. Hurley, Wei Li 0012, Xiangyang Xue 0001 |
Web Intelligence | 3 |
| 2011 | Towards content-based audio fragment authenticationabstractAudio authentication is a technique to protect the integrity and originality of audio signals. Due to the long duration, it is often desirable to authenticate only a segment of audio signal. To the authors' knowledge, this important issue has not been seriously researched so far. In this paper, a novel authentication algorithm is proposed for the purpose of audio fragment authentication. SIFT descriptor originated from computer vision field is introduced and calculated on audio spectrogram to accomplish the tasks of fragment alignment, time-domain blocking, cropping and inserting identification etc. Experiments show reliable discrimination between admissible and malicious manipulations, precise tamper localization and classification. Xiangyang Xue 0001, Wei Li 0012 |
ACM Multimedia | 2 |
| 2010 | Robust hashing for music copyright protection by combining beat segmentation and chromaabstractTime-scale modification and pitching shifting are two recognized challenging attacks to music copyright protection. To resist them simultaneously, a novel robust hashing method is proposed by combining the strength of music beat segmentation and chroma-based music feature. These two measures are aimed at solving the problem of desynchronization and frequency shifting respectively. Moreover, two layers of scrambling are performed to ensure the security. Experiments exhibit remarkable robustness against various attacks including pitch [email protected]%, time-scale [email protected]%, and [email protected]/10 etc. Wei Li 0012, Zhurong Wang, Bilei Zhu, Xiangyang Xue 0001 |
ACM Multimedia | 1 |
| 2010 | A novel audio fingerprinting method robust to time scale modification and pitch shiftingabstractA novel audio fingerprinting method that is highly robust to Time Scale Modification (TSM) and pitch shifting is proposed. Instead of simply employing spectral or tempo-related features, our system is based on computer-vision techniques. We transform each 1-D audio signal into a 2-D image and treat TSM and pitch shifting of the audio signal as stretch and translation of the corresponding image. Robust local descriptors are extracted from the image and matched against those of the reference audio signals. Experimental results show that our system is highly robust to various audio distortions, including the challenging TSM and pitch shifting. Bilei Zhu, Wei Li 0012, Zhurong Wang, Xiangyang Xue 0001 |
ACM Multimedia | 2 |
| 2010 | Robust audio identification for MP3 popular musicabstractAudio identification via fingerprint has been an active research field with wide applications for years. Many technical papers were published and commercial software systems were also employed. However, most of these previously reported methods work on the raw audio format in spite of the fact that nowadays compressed format audio, especially MP3 music, has grown into the dominant way to store on personal computers and transmit on the Internet. It would be interesting if a compressed unknown audio fragment is able to be directly recognized from the database without the fussy and time-consuming decompression-identification-recompression procedure. So far, very few algorithms run directly in the compressed domain for music information retrieval, and most of them take advantage of MDCT coefficients or derived energy type of features. As a first attempt, we propose in this paper utilizing compressed-domain spectral entropy as the audio feature to implement a novel audio fingerprinting algorithm. The compressed songs stored in a music database and the possibly distorted compressed query excerpts are first partially decompressed to obtain the MDCT coefficients as the intermediate result. Then by grouping granules into longer blocks, remapping the MDCT coefficients into 192 new frequency lines to unify the frequency distribution of long and short windows, and defining 9 new subbands which cover the main frequency bandwidth of popular songs in accordance with the scale-factor bands of short windows, we calculate the spectral entropy of all consecutive blocks and come to the final fingerprint sequence by means of magnitude relationship modeling. Experiments show that such fingerprints exhibit strong robustness against various audio signal distortions like recompression, noise interference, echo addition, equalization, band-pass filtering, pitch shifting, and slight time-scale modification etc. For 5s-long query examples which might be severely degraded, an average top-five retrieval precision rate of more than 90% can be obtained in our test data set composed of 1822 popular songs. Wei Li 0012, Yaduo Liu, Xiangyang Xue 0001 |
SIGIR | 1 |
| 2010 | Robust music identification based on low-order zernike moment in the compressed domainabstractIn this paper, we devise a novel robust music identification algorithm utilizing compressed-domain audio Zernike moment adapted from image processing techniques as the pivotal feature. Audio fingerprint derived from this feature exhibits strong robustness against various audio signal distortions including the challenging pitch shifting and time-scale modification. Experiments show that in our test dataset composed of 1822 popular songs, a 5s music query example which might have been severely corrupted is still sufficient to identify its original near-duplicate copy, with more than 90% top five precision rate. Wei Li 0012, Yaduo Liu, Xiangyang Xue 0001 |
SIGIR | 1 |
| 2009 | A Robust Mesh Watermarking Scheme Based on PCAabstractThis paper proposed a novel robust oblivious watermarking scheme in the spatial domain suitable for 3-D mesh object, which combines the principal component analysis (PCA) method with the construction of cone bins. Firstly, PCA is used to calculate the eigenvectors of covariance matrix of vertex coordinates. Then the object is rotated and translated so that its center of mass and the three eigenvectors coincide with the origin and the three axes of the Cartesian coordinate system. Subsequently, many cone bins are constructed, each bin center according to two angle parameters of spherical coordinate produced by pseudorandom number generators and the vertices are classified into the appropriate cone bins. Each cone bin is divided into a number of sub-bins in order to embed the watermark bit, and the size of sub-bin can make tradeoff between invisible and robustness. Experiment results show the remarkable ability of the proposed mechanism to resist against various attacks such as adding noise, clipping, similarity transform and vertex re-ordering. Xiaoqiang Li 0002, Wei Li 0012 |
ICIG | 3 |
| 2006 | Localized audio watermarking technique robust against time-scale modificationabstractSynchronization attacks like random cropping and time-scale modification are very challenging problems to audio watermarking techniques. To combat these attacks, a novel content-dependent localized robust audio watermarking scheme is proposed. The basic idea is to first select steady high-energy local regions that represent music edges like note attacks, transitions or drum sounds by using different methods, then embed the watermark in these regions. Such regions are of great importance to the understanding of music and will not be changed much for maintaining high auditory quality. In this way, the embedded watermark has the potential to escape all kinds of distortions. Experimental results show strong robustness against common audio signal processing, time-domain synchronization attacks, and most distortions introduced in Stirmark for Audio. Wei Li 0012, Xiangyang Xue 0001, Peizhong Lu |
IEEE Trans. Multim. | 1 |
| 2003 | An Optimized Multi-bits Blind Watermarking Scheme
Xiaoqiang Li 0002, Xiangyang Xue 0001, Wei Li 0012 |
ICICS | 3 |
| 2003 | Audio Watermarking Based on Music Content Analysis: Robust against Time Scale Modification
Wei Li 0012, Xiangyang Xue 0001 |
IWDW | 1 |
| 2000 | Speech enhancement using the constrained-optimization techniqueabstractWe address a problem of speech enhancement: recovering a speech source from a mixture of its delayed versions and additive noise. By using the constrained-optimization technique, the second order statistics based algorithm is developed. The new proposed algorithm requires no strong limitations to the speech signal and the noise. Simulation results show that our algorithm achieves a better performance as compared to other algorithms. Wei Li 0012, Wan-Chi Siu |
IEEE Signal Process. Lett. | 1 |
| 1999 | Recovery of single source signal from noisy and reverberant environments using second-order statistics
Wei Li 0012, J. C. H. Poon, Wan-Chi Siu |
Signal Process. | 1 |