Ye Wang 0007

dblp:44/6292-7 · DBLP profile ↗
← Back
105ranked-venue papers
9as first author
29since 2021 · last 2025
0000-0002-0123-1260ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 83 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 22 · 17 since 2021Computer networks · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Lead Instrument Detection from Multitrack Music
abstract
Prior approaches to lead instrument detection primarily analyze mixture audio, limited to coarse classifications and lacking generalization ability. This paper presents a novel approach to lead instrument detection in multitrack music audio by crafting expertly annotated datasets and designing a novel framework that integrates a self-supervised learning model with a track-wise, frame-level attention-based classifier. This attention mechanism dynamically extracts and aggregates track-specific features based on their auditory importance, enabling precise detection across varied instrument types and combinations. Enhanced by track classification and permutation augmentation, our model substantially outperforms existing SVM and CRNN models, showing robustness on unseen instruments and out-of-domain testing. We believe our exploration provides valuable in-sights for future research on audio content analysis in multitrack music settings.
Longshen Ou, Yu Takahashi, Ye Wang 0007
ICASSP3
2025 SPSinger: Multi-Singer Singing Voice Synthesis with Short Reference Prompt
abstract
Current singing voice synthesis systems often struggle in multi-singer scenarios due to limited training data that only includes a few singers. Existing zero-shot multi-singer singing voice synthesis systems are criticized for their reliance on global timbre embeddings from single reference audio, which fail to capture sufficient timbre details. This paper introduces SPSinger, a multi-singer singing voice synthesizer that generates singer-specific voices from brief reference audio (around 5 seconds) without prior training on the singer’s voice. SPSinger builds on the StableDiffusion framework by adding a global encoder to capture consistent timbre features from short reference prompts and an attention-based local encoder to capture detailed variations from long prompts, used only during training. To overcome the challenge of requiring long audio prompts during inference, we introduce the Latent Prompt Adaptation Model (LPAM), a Transformer-based module that derives timbre features from global embeddings. This approach eliminates the need for long reference prompts. Additionally, we propose a novel pitch shift algorithm that uses LPAM to predict the pitch shift values. Our experiments show that SPSinger achieves high-quality singing voice synthesis that preserves the identity of the target singer, even when using only short reference audio inputs in zero-shot scenarios.
Junchuan Zhao, Chetwin Low, Ye Wang 0007
ICASSP3
2025 On Calibration of LLM-based Guard Models for Reliable Content Moderation
abstract
Large language models (LLMs) pose significant risks due to the potential for generating harmful content or users attempting to evade guardrails. Existing studies have developed LLM-based guard models designed to moderate the input and output of threat LLMs, ensuring adherence to safety policies by blocking content that violates these protocols upon deployment. However, limited attention has been given to the reliability and calibration of such guard models. In this work, we empirically conduct comprehensive investigations of confidence calibration for 9 existing LLM-based guard models on 12 benchmarks in both user input and model output classification. Our findings reveal that current LLM-based guard models tend to 1) produce overconfident predictions, 2) exhibit significant miscalibration when subjected to jailbreak attacks, and 3) demonstrate limited robustness to the outputs generated by different types of response models. Additionally, we assess the effectiveness of post-hoc calibration methods to mitigate miscalibration. We demonstrate the efficacy of temperature scaling and, for the first time, highlight the benefits of contextual calibration for confidence calibration of guard models, particularly in the absence of validation sets. Our analysis and experiments underscore the limitations of current LLM-based guard models and provide valuable insights for the future development of well-calibrated guard models toward more reliable content moderation. We also advocate for incorporating reliability evaluation of confidence calibration when releasing future LLM-based guard models.
Hongfu Liu 0002, Hengguan Huang, Xiangming Gu, Hao Wang 0014, Ye Wang 0007
ICLR5
2025 When Attention Sink Emerges in Language Models: An Empirical View
abstract
Auto-regressive language Models (LMs) assign significant attention to the first token, even if it is not semantically important, which is known as **attention sink**. This phenomenon has been widely adopted in applications such as streaming/long context generation, KV cache optimization, inference acceleration, model quantization, and others. Despite its widespread use, a deep understanding of attention sink in LMs is still lacking. In this work, we first demonstrate that attention sinks exist universally in auto-regressive LMs with various inputs, even in small models. Furthermore, attention sink is observed to emerge during the LM pre-training, motivating us to investigate how *optimization*, *data distribution*, *loss function*, and *model architecture* in LM pre-training influence its emergence. We highlight that attention sink emerges after effective optimization on sufficient training data. The sink position is highly correlated with the loss function and data distribution. Most importantly, we find that attention sink acts more like key biases, *storing extra attention scores*, which could be non-informative and not contribute to the value computation. We also observe that this phenomenon (at least partially) stems from tokens' inner dependence on attention scores as a result of softmax normalization. After relaxing such dependence by replacing softmax attention with other attention operations, such as sigmoid attention without normalization, attention sinks do not emerge in LMs up to 1B parameters. The code is available at https://github.com/sail-sg/Attention-Sink.
Xiangming Gu, Tianyu Pang, Qian Liu 0033, Fengzhuo Zhang, Cunxiao Du, Ye Wang 0007
ICLR7
2025 LivePoem: Improving the Learning Experience of Classical Chinese Poetry with AI-Generated Musical Storyboards
abstract
Textbook reading has long dominated classical poetry education in Chinese-speaking communities. However, research has shown that extensive text-based learning can lead to learner disengagement and a less pleasant experience. This paper aims to improve the experience of classical Chinese poetry learning by introducing LivePoem—a system that generates musical storyboards (storyboards with background music) as audiovisual aids to support poetry comprehension. We employ a pre-trained diffusion model for storyboard generation and train a prosody-based poem-to-melody generator using a Transformer model, both validated by standard objective metrics to ensure generation quality. Through a within-subjects study involving 25 non-native Chinese learners, we compared learning outcomes from textbook reading and musical storyboard viewing through standardised reading comprehension tests. Additionally, the learning experience was assessed by the Self-Assessment Manikin (SAM) and an inductive thematic analysis of learners' open-ended feedback. Experimental results show that musical storyboards retained the learning outcomes of textbooks, while more effectively engaging learners and providing a more pleasant learning experience.
Qihao Liang, Xichu Ma, Torin Hopkins, Ye Wang 0007
IJCAI4
2025 Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization
abstract
We present a unified framework for automatic multitrack music arrangement that enables a single pre-trained symbolic music model to handle diverse arrangement scenarios, including reinterpretation, simplification, and additive generation. At its core is a segment-level reconstruction objective operating on token-level disentangled content and style, allowing for flexible any-to-any instrumentation transformations at inference time. To support track-wise modeling, we introduce REMI-z, a structured tokenization scheme for multitrack symbolic music that enhances modeling efficiency and effectiveness for both arrangement tasks and unconditional generation. Our method outperforms task-specific state-of-the-art models on representative tasks in different arrangement scenarios---band arrangement, piano reduction, and drum arrangement, in both objective metrics and perceptual evaluations. Taken together, our framework demonstrates strong generality and suggests broader applicability in symbolic music-to-music transformation.
Longshen Ou, Ziyu Wang 0008, Gus Xia, Qihao Liang, Torin Hopkins, Ye Wang 0007
NeurIPS7
2025 KeYric: Unsupervised Keywords Extraction and Expansion from Music for Coherent Lyrics Generation
abstract
We address the challenge of enhancing coherence in generated lyrics from symbolic music, particularly for creating singing-based language learning materials. Coherence, defined as the quality of being logical and consistent, forming a unified whole, is crucial for lyrics at multiple levels—word, sentence, and full-text. Additionally, it involves lyrics’ musicality—matching of style and sentiment of the music. To tackle this, we introduce KeYric, a novel system that leverages keyword skeletons to strengthen both coherence and musicality in lyrics generation. KeYric employs an innovative approach with an unsupervised keyword skeleton extractor and a graph-based skeleton expander, designed to produce a style-appropriate keyword skeleton from input music. This framework integrates the skeleton with the input music via a three-layer coherence mechanism, significantly enhancing lyric coherence by 5% in objective evaluations. Subjective assessments confirm that KeYric-generated lyrics are perceived as 19% more coherent and suitable for language learning through singing compared to existing models. Our analyses indicate that integrating genre-relevant elements, such as pitch, into music encoding is crucial, as musical genres significantly affect lyric coherence.
Xichu Ma, Min-Yen Kan, Wee Sun Lee, Ye Wang 0007
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Advancing Test-Time Adaptation in Wild Acoustic Test Settings
abstract
Acoustic foundation models, fine-tuned for Automatic Speech Recognition (ASR), suffer from performance degradation in wild acoustic test settings when deployed in real-world scenarios.Stabilizing online Test-Time Adaptation (TTA) under these conditions remains an open and unexplored question.Existing wild vision TTA methods often fail to handle speech data effectively due to the unique characteristics of high-entropy speech frames, which are unreliably filtered out even when containing crucial semantic content.Furthermore, unlike static vision data, speech signals follow short-term consistency, requiring specialized adaptation strategies.In this work, we propose a novel wild acoustic TTA method tailored for ASR fine-tuned acoustic foundation models.Our method, Confidence-Enhanced Adaptation, performs frame-level adaptation using a confidence-aware weight scheme to avoid filtering out essential information in high-entropy frames.Additionally, we apply consistency regularization during test-time optimization to leverage the inherent short-term consistency of speech signals.Our experiments on both synthetic and real-world datasets demonstrate that our approach outperforms existing baselines under various wild acoustic test settings, including Gaussian noise, environmental sounds, accent variations, and sung speech 1 .
Hongfu Liu 0002, Hengguan Huang, Ye Wang 0007
EMNLP3
2024 Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
abstract
A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further intervention from the adversary, (almost) all agents will become infected exponentially fast and exhibit harmful behaviors. To validate the feasibility of infectious jailbreak, we simulate multi-agent environments containing up to one million LLaVA-1.5 agents, and employ randomized pair-wise chat as a proof-of-concept instantiation for multi-agent interaction. Our results show that feeding an (infectious) adversarial image into the memory of any randomly chosen agent is sufficient to achieve infectious jailbreak. Finally, we derive a simple principle for determining whether a defense mechanism can provably restrain the spread of infectious jailbreak, but how to design a practical defense that meets this principle remains an open question to investigate.
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Qian Liu 0033, Ye Wang 0007, Jing Jiang 0001
ICML6
2024 XAI-Lyricist: Improving the Singability of AI-Generated Lyrics with Prosody Explanations
Qihao Liang, Xichu Ma, Finale Doshi-Velez, Brian Lim, Ye Wang 0007
IJCAI5
2024 End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding
Wei Zeng 0018, Xian He, Ye Wang 0007
IJCAI3
2024 Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling
abstract
In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing computational efficiency. In this paper, we introduce a novel system that leverages prior modelling over disentangled style factors to address these challenges. Our method presents a two-stage process: initially, a piano arrangement is derived from the lead sheet by retrieving piano texture styles; subsequently, a multi-track orchestration is generated by infusing orchestral function styles into the piano arrangement. Our key design is the use of vector quantization and a unique multi-stream Transformer to model the long-term flow of the orchestration style, which enables flexible, controllable, and structured music generation. Experiments show that by factorizing the arrangement task into interpretable sub-stages, our approach enhances generative capacity while improving efficiency. Additionally, our system supports a variety of music genres and provides style control at different composition hierarchies. We further show that our system achieves superior coherence, structure, and overall arrangement quality compared to existing baselines.
Gus Xia, Ziyu Wang 0008, Ye Wang 0007
NeurIPS4
2024 SinTechSVS: A Singing Technique Controllable Singing Voice Synthesis System
abstract
The precise control of singing techniques is of utmost importance in achieving emotionally expressive vocal performances. To bridge the gap between current Singing Voice Synthesis (SVS) systems and human singers, our paper focuses on developing an SVS system that allows for control over singing techniques. In this paper, we introduce SinTechSVS, a singing technique controllable SVS system composed of a singing technique annotator, a singing technique controllable synthesizer, and a singing technique recommender. Our approach leverages transfer learning for efficient singing technique annotation and adapts the DiffSinger framework with additional style encoders and an attention-based singing technique local score (STLS) module to enhance singing technique controllability. We also propose a Seq2Seq singing technique recommender for the new task of Singing Technique Recommendation (STR). Experimental results demonstrate that SinTechSVS significantly improves the quality and expressiveness of synthesized vocal performances, with comparable general synthesis capabilities to state-of-the-art SVS systems and enhanced control over singing techniques, as evidenced by objective and subjective evaluations. To the best of our knowledge, SinTechSVS is the first SVS capable of controlling singing techniques.
Junchuan Zhao, Chetwin Low, Ye Wang 0007
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Drawlody: Sketch-Based Melody Creation With Enhanced Usability and Interpretability
abstract
Sketch-based melody creation systems enable people to compose melodies by converting human-sketched melody contours into coherent melodies that fit the depicted contours. This remains one of the most intuitive approaches to interactive music creation. However, previous studies are still stagnating in limitations regarding usability and interpretability, which hinders effective interactions between people and AI. For one thing, these studies entail additional complex musical conditions as auxiliary inputs (e.g. chord progressions, contextual melodies, and predetermined rhythms), supporting only fixed-length and rule-based melody generation. This makes existing systems less usable, with generated melodies lacking diversity and coherence. Moreover, users without enough musical expertise might find it difficult to define appropriate inputs and to interpret the role of these inputs in guiding melody generation. To address these limitations, we present Drawlody, a novel sketch-based melody creation system with enhanced usability and interpretability. Specifically, Drawlody simplifies user input requirements by excluding all complex musical conditions, using only a simplified melody contour representation named Generalised Melody Contour (GMC) as input. This simplification clarifies the role of user controls, making the system more usable for people without musical training. To guide coherent melody generation from GMC, we propose FlexMIDI music representation, which simulates the tonal structure of melodies and faithfully explains how human-sketched contours guide melody generation. We employ a CNN-Transformer-based architecture as the foundation model to achieve arbitrary-length melody generation. Drawlody is evaluated by both objective and subjective music quality studies, as well as a usability and interpretability study. The results support its enhanced usability, interpretability, and high-quality melody generation capabilities. Video demos of the system are presentedhere.
Qihao Liang, Ye Wang 0007
IEEE Trans. Multim.2
2024 Symbolic Music Generation From Graph-Learning-Based Preference Modeling and Textual Queries
abstract
This paper investigates the domain of automatic music generation (AMG) and its capacity to produce music that is aligned with user preferences. The incorporation of user music preference (UMP) awareness in AMG technology has the potential to reduce reliance on musicians and domain experts while encouraging users to engage in activities that promote human health and potential. Current research in AMG has been limited to the qualitative control of a constrained set of attributes in the generated music such as selecting a genre from a given list. This constraint makes it challenging to develop music that is both aligned with UMP and suitable for practical text query-based applications. To address this challenge, we propose to apply deep-graph-networks on music community data, jointly modeling UMP and music features. Moreover, users' textual descriptions of expected music can be transformed into graphs that are compatible with UMPs. Node embeddings representing user queries' connotation are extracted to condition the music generator. The results on objective and subjective metrics demonstrate a significant improvement in UMP accuracy by 31.3%, UMP-aware AMG by 63.5%, and text-to-music AMG effectiveness by 76.5%. Our detailed analysis indicates that the generated music aligns best with queries comprised of short sentences and commonly used words.
Xichu Ma, Ye Wang 0007
IEEE Trans. Multim.3
2024 Automatic Lyric Transcription and Automatic Music Transcription from Multimodal Singing
abstract
Automatic lyric transcription (ALT) refers to transcribing singing voices into lyrics, while automatic music transcription (AMT) refers to transcribing singing voices into note events, i.e., musical MIDI notes. Despite these two tasks having significant potential for practical application, they are still nascent. This is because the transcription of lyrics and note events solely from singing audio is notoriously difficult due to the presence of noise contamination, e.g., musical accompaniment, resulting in a degradation of both the intelligibility of sung lyrics and the recognizability of sung notes. To address this challenge, we propose a general framework for implementing multimodal ALT and AMT systems. Additionally, we curate the first multimodal singing dataset, comprising N20EMv1 and N20EMv2, which encompasses audio recordings and videos of lip movements, together with ground truth for lyrics and note events. For model construction, we propose adapting self-supervised learning models from the speech domain as acoustic encoders and visual encoders to alleviate the scarcity of labeled data. We also introduce a residual cross-attention mechanism to effectively integrate features from the audio and video modalities. Through extensive experiments, we demonstrate that our single-modal systems exhibit state-of-the-art performance on both ALT and AMT tasks. Subsequently, through single-modal experiments, we also explore the individual contributions of each modality to the multimodal system. Finally, we combine these and demonstrate the effectiveness of our proposed multimodal systems, particularly in terms of their noise robustness.
Xiangming Gu, Longshen Ou, Wei Zeng 0018, Nicholas Wong, Ye Wang 0007
ACM Trans. Multim. Comput. Commun. Appl.6
2023 FedNP: Towards Non-IID Federated Learning via Federated Neural Propagation
abstract
Traditional federated learning (FL) algorithms, such as FedAvg, fail to handle non-i.i.d data because they learn a global model by simply averaging biased local models that are trained on non-i.i.d local data, therefore failing to model the global data distribution. In this paper, we present a novel Bayesian FL algorithm that successfully handles such a non-i.i.d FL setting by enhancing the local training task with an auxiliary task that explicitly estimates the global data distribution. One key challenge in estimating the global data distribution is that the data are partitioned in FL, and therefore the ground-truth global data distribution is inaccessible. To address this challenge, we propose an expectation-propagation-inspired probabilistic neural network, dubbed federated neural propagation (FedNP), which efficiently estimates the global data distribution given non-i.i.d data partitions. Our algorithm is sampling-free and end-to-end differentiable, can be applied with any conventional FL frameworks and learns richer global data representation. Experiments on both image classification tasks with synthetic non-i.i.d image data partitions and real-world non-i.i.d speech recognition tasks demonstrate that our framework effectively alleviates the performance deterioration caused by non-i.i.d data.
Xueyang Wu 0001, Hengguan Huang, Youlong Ding, Hao Wang 0014, Ye Wang 0007, Qian Xu 0005
AAAI5
2023 Songs Across Borders: Singable and Controllable Neural Lyric Translation
abstract
The development of general-domain neural machine translation (NMT) methods has advanced significantly in recent years, but the lack of naturalness and musical constraints in the outputs makes them unable to produce singable lyric translations.This paper bridges the singability quality gap by formalizing lyric translation into a constrained translation problem, converting theoretical guidance and practical techniques from translatology literature to promptdriven NMT approaches, exploring better adaptation methods, and instantiating them to an English-Chinese lyric translation system.Our model achieves 99.85%, 99.00%, and 95.52% on length accuracy, rhyme accuracy, and word boundary recall.In our subjective evaluation, our model shows a 75% relative enhancement on overall quality, compared against naive finetuning 1 .
Longshen Ou, Xichu Ma, Min-Yen Kan, Ye Wang 0007
ACL (1)4
2023 Phonation Mode Detection in Singing: A Singer Adapted Model
abstract
Phonation modes play a vital role in voice quality evaluation and vocal health diagnosis. Existing studies on phonation modes cover feature analysis and classification of vowels, which does not apply to real-life scenarios. In this paper, we define the phonation mode detection (PMD) problem, which entails the prediction of phonation mode labels as well as their onset and offset timestamps. To address the PMD problem, we propose the first dataset PMSing, and an end-to-end PMD network (P-Net) to integrate phonation mode identification and boundary detection, which also prevents the oversegmentation of frame-level output. Furthermore, we introduce an adapted P-Net model (AP-Net) based on an adversarial discriminative training process using labeled data from one singer and unlabeled data from unseen singers. Experiments show that the P-Net outperforms the state-of-the-art methods with an F-score of 0.680, and the AP-Net also achieves an F-score of 0.658 for unseen singers.
Yixin Wang 0007, Wei Wei 0037, Ye Wang 0007
ICASSP3
2023 Q&A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement
abstract
Music rearrangement is a common music practice of reconstructing and reconceptualizing a piece using new composition or instrumentation styles, which is also an important task of automatic music generation. Existing studies typically model the mapping from a source piece to a target piece via supervised learning. In this paper, we tackle rearrangement problems via self-supervised learning, in which the mapping styles can be regarded as conditions and controlled in a flexible way. Specifically, we are inspired by the representation disentanglement idea and propose Q&A, a query-based algorithm for multi-track music rearrangement under an encoder-decoder framework. Q&A learns both a content representation from the mixture and function (style) representations from each individual track, while the latter queries the former in order to rearrange a new piece. Our current model focuses on popular music and provides a controllable pathway to four scenarios: 1) re-instrumentation, 2) piano cover generation, 3) orchestration, and 4) voice separation. Experiments show that our query system achieves high-quality rearrangement results with delicate multi-track structures, significantly outperforming the baselines.
Gus Xia, Ye Wang 0007
IJCAI3
2023 Zero-Shot Automatic Pronunciation Assessment
Hongfu Liu 0002, Mingqian Shi, Ye Wang 0007
INTERSPEECH3
2023 Elucidate Gender Fairness in Singing Voice Transcription
abstract
It is widely known that males and females typically possess different sound characteristics when singing, such as timbre and pitch, but it has never been explored whether these gender-based characteristics lead to a performance disparity in singing voice transcription (SVT), whose target includes pitch. Such a disparity could cause fairness issues and severely affect the user experience of downstream SVT applications. Motivated by this, we first demonstrate the female superiority of SVT systems, which is observed across different models and datasets. We find that different pitch distributions, rather than gender data imbalance, contribute to this disparity. To address this issue, we propose using an attribute predictor to predict gender labels and adversarially training the SVT system to enforce the gender-invariance of acoustic representations. Leveraging the prior knowledge that pitch distributions may contribute to the gender bias, we propose conditionally aligning acoustic representations between demographic groups by feeding note events to the attribute predictor. Empirical experiments on multiple benchmark SVT datasets show that our method significantly reduces gender bias (up to more than 50%) with negligible degradation of overall SVT performance, on both in-domain and out-of-domain singing data, thus offering a better fairness-utility trade-off.
Xiangming Gu, Wei Zeng 0018, Ye Wang 0007
ACM Multimedia3
2023 Disentangled Adversarial Domain Adaptation for Phonation Mode Detection in Singing and Speech
abstract
Phonation mode detection predicts phonation modes and their temporal boundaries in singing and speech, holding promise for characterizing voice quality and vocal health. However, it is very challenging due to the domain disparities between training data and unannotated real-world recordings. To tackle this problem, we develop a disentangled adversarial domain adaptation network, which adapts the phonation mode detection model with the structure of the convolutional recurrent neural network pre-trained on the source domain to the target domain without phonation mode labels. Based on our curated sung and spoken dataset for phonation mode detection, we demonstrate that the subject and the singing-speech mismatches cause performance decline. By disentangling domain-invariant phonation mode and domain-specific embeddings, our method greatly enhances the effectiveness and explainability of unsupervised adversarial domain adaptation. Experiments show that the performance drop caused by the subject mismatch is alleviated via adaptation, resulting in 44.7% and 6.8% improvement of the F-score for singing and speech, respectively. The singing and speech domain adaptation experiment indicates that a model trained on singing data can be adapted to speech, yielding an F-score of 0.56, commensurate with the F-score of 0.59 achieved using a model exclusively trained on speech data. By further investigating the disentangled embeddings, we find that the phonation mode feature shared by singing and speech is invariant to pitch. These results inspire reliable and versatile applications in voice quality evaluation and paralinguistic information retrieval.
Yixin Wang 0007, Wei Wei 0037, Xiangming Gu, Xiaohong Guan, Ye Wang 0007
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Exploring Transformer's Potential on Automatic Piano Transcription
abstract
Most recent research about automatic music transcription (AMT) uses convolutional neural networks and recurrent neural networks to model the mapping from music signals to symbolic notation. Based on a high-resolution piano transcription system, we explore the possibility of incorporating another powerful sequence transformation tool—the Transformer—to deal with the AMT problem. We argue that the properties of the Transformer make it more suitable for certain AMT subtasks. We confirm the Transformer’s superiority on the velocity detection task by experiments on the MAESTRO dataset and a cross-dataset evaluation on the MAPS dataset. We observe a performance improvement on both frame-level and note-level metrics after introducing the Transformer network.
Longshen Ou, Emmanouil Benetos, Jiqing Han 0001, Ye Wang 0007
ICASSP5
2022 MM-ALT: A Multimodal Automatic Lyric Transcription System
abstract
Automatic lyric transcription (ALT) is a nascent field of study attracting increasing interest from both the speech and music information retrieval communities, given its significant application potential. However, ALT with audio data alone is a notoriously difficult task due to instrumental accompaniment and musical constraints resulting in degradation of both the phonetic cues and the intelligibility of sung lyrics. To tackle this challenge, we propose the MultiModal Automatic Lyric Transcription system (MM-ALT), together with a new dataset, N20EM, which consists of audio recordings, videos of lip movements, and inertial measurement unit (IMU) data of an earbud worn by the performing singer. We first adapt the wav2vec 2.0 framework from automatic speech recognition (ASR) to the ALT task. We then propose a video-based ALT method and an IMU-based voice activity detection (VAD) method. In addition, we put forward the Residual Cross Attention (RCA) mechanism to fuse data from the three modalities (i.e., audio, video, and IMU). Experiments show the effectiveness of our proposed MM-ALT system, especially in terms of noise robustness. Project page is at https://n20em.github.io.
Xiangming Gu, Longshen Ou, Danielle Ong, Ye Wang 0007
ACM Multimedia4
2022 Content based User Preference Modeling in Music Generation
abstract
Automatic music generation (AMG) has been an emerging research topic in AI in recent years. However, generating user-preferred music remains an unsolved problem. To address this challenge, we propose a hierarchical convolutional recurrent neural network with self-attention (CRNN-SA) to extract user music preference (UMP) and map it into an embedding space where the common UMPs are in the center and uncommon UMPs are scattered towards the edge. We then propose an explainable music distance measure as a bridge between the UMP and AMG; this measure computes the distance between a seed song and the user's UMP. That distance is then employed to adjust the AMG's parameters which control the music generation process in an iterative manner, so that the generated song will be closer to the user's UMP in every iteration. Experiments demonstrate that the proposed UMP embedding model successfully captures individual UMPs and that our proposed system is capable of generating user-preferred songs.
Xichu Ma, Ye Wang 0007
ACM Multimedia3
2022 Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation
abstract
Human intelligence has shown remarkably lower latency and higher precision than most AI systems when processing non-stationary streaming data in real-time. Numerous neuroscience studies suggest that such abilities may be driven by internal predictive modeling. In this paper, we explore the possibility of introducing such a mechanism in unsupervised domain adaptation (UDA) for handling non-stationary streaming data for real-time streaming applications. We propose to formulate internal predictive modeling as a continuous-time Bayesian filtering problem within a stochastic dynamical system context. Such a dynamical system describes the dynamics of model parameters of a UDA model evolving with non-stationary streaming data. Building on such a dynamical system, we then develop extrapolative continuous-time Bayesian neural networks (ECBNN), which generalize existing Bayesian neural networks to represent temporal dynamics and allow us to extrapolate the distribution of model parameters before observing the incoming data, therefore effectively reducing the latency. Remarkably, our empirical results show that ECBNN is capable of continuously generating better distributions of model parameters along the time axis given historical data only, thereby achieving (1) training-free test-time adaptation with low latency, (2) gradually improved alignment between the source and target features and (3) gradually improved model performance over time during the real-time testing stage.
Hengguan Huang, Xiangming Gu, Hao Wang 0014, Chang Xiao 0004, Hongfu Liu 0002, Ye Wang 0007
NeurIPS6
2021 STRODE: Stochastic Boundary Ordinary Differential Equation
abstract
Perception of time from sequentially acquired sensory inputs is rooted in everyday behaviors of individual organisms. Yet, most algorithms for time-series modeling fail to learn dynamics of random event timings directly from visual or audio inputs, requiring timing annotations during training that are usually unavailable for real-world applications. For instance, neuroscience perspectives on postdiction imply that there exist variable temporal ranges within which the incoming sensory inputs can affect the earlier perception, but such temporal ranges are mostly unannotated for real applications such as automatic speech recognition (ASR). In this paper, we present a probabilistic ordinary differential equation (ODE), called STochastic boundaRy ODE (STRODE), that learns both the timings and the dynamics of time series data without requiring any timing annotations during training. STRODE allows the usage of differential equations to sample from the posterior point processes, efficiently and analytically. We further provide theoretical guarantees on the learning of STRODE. Our empirical results show that our approach successfully infers event timings of time series data. Our method achieves competitive or superior performances compared to existing state-of-the-art methods for both synthetic and real-world datasets.
Hengguan Huang, Hongfu Liu 0002, Hao Wang 0014, Chang Xiao 0004, Ye Wang 0007
ICML5
2021 AI-Lyricist: Generating Music and Vocabulary Constrained Lyrics
abstract
We propose AI-Lyricist: a system to generate novel yet meaningful lyrics given a required vocabulary and a MIDI file as inputs. This task involves multiple challenges, including automatically identifying the melody and extracting a syllable template from multi-channel music, generating creative lyrics that match the input music's style and syllable alignment, and satisfying vocabulary constraints. To address these challenges, we propose an automatic lyrics generation system consisting of four modules: (1) A music structure analyzer to derive the musical structure and syllable template from a given MIDI file, utilizing the concept of expected syllable number to better identify the melody, (2) a SeqGAN-based lyrics generator optimized by multi-adversarial training through policy gradients with twin discriminators for text quality and syllable alignment, (3) a deep coupled music-lyrics embedding model to project music and lyrics into a joint space to allow fair comparison of both melody and lyric constraints, and a module called (4) Polisher, to satisfy vocabulary constraints by applying a mask to the generator and substituting the words to be learned. We trained our model on a dataset of over 7,000 music-lyrics pairs, enhanced with manually annotated labels in terms of theme, sentiment and genre. Both objective and subjective evaluations show AI-Lyricist's superior performance against the state-of-the-art for the proposed tasks.
Xichu Ma, Ye Wang 0007, Min-Yen Kan, Wee Sun Lee
ACM Multimedia2
2020 A-CRNN: A Domain Adaptation Model for Sound Event Detection
abstract
This paper presents a domain adaptation model for sound event detection. A common challenge for sound event detection is how to deal with the mismatch among different datasets. Typically, the performance of a model will decrease if it is tested on a dataset which is different from the one that the model is trained on. To address this problem, based on convolutional recurrent neural networks (CRNNs), we propose an adapted CRNN (A-CRNN) as an unsupervised adversarial domain adaptation model for sound event detection. We have collected and annotated a dataset in Singapore with two types of recording devices to complement existing datasets in the research community, especially with respect to domain adaptation. We perform experiments on recordings from different datasets and from different recordings devices. Our experimental results show that the proposed A-CRNN model can achieve a better performance on an unseen dataset in comparison with the baseline non-adapted CRNN model.
Wei Wei 0037, Hongning Zhu, Emmanouil Benetos, Ye Wang 0007
ICASSP4
2020 Deep Graph Random Process for Relational-Thinking-Based Speech Recognition
abstract
Lying at the core of human intelligence, relational thinking is characterized by initially relying on innumerable unconscious percepts pertaining to relations between new sensory signals and prior knowledge, consequently becoming a recognizable concept or object through coupling and transformation of these percepts. Such mental processes are difficult to model in real-world problems such as in conversational automatic speech recognition (ASR), as the percepts (if they are modelled as graphs indicating relationships among utterances) are supposed to be innumerable and not directly observable. In this paper, we present a Bayesian nonparametric deep learning method called deep graph random process (DGP) that can generate an infinite number of probabilistic graphs representing percepts. We further provide a closed-form solution for coupling and transformation of these percept graphs for acoustic modeling. Our approach is able to successfully infer relations among utterances without using any relational data during training. Experimental evaluations on ASR tasks including CHiME-2 and CHiME-5 demonstrate the effectiveness and benefits of our method.
Hengguan Huang, Fuzhao Xue, Hao Wang 0014, Ye Wang 0007
ICML4
2020 Automatic Leaderboard: Evaluation of Singing Quality Without a Standard Reference
abstract
Automatic evaluation of singing quality can be done with the help of a reference singing or the digital sheet music of the song. However, such a standard reference is not always available. In this article, we propose a framework to rank a large pool of singers according to their singing quality without any standard reference. We define musically motivated absolute measures based on pitch histogram, and relative measures based on inter-singer statistics to evaluate the quality of singing attributes such as intonation, and rhythm. The absolute measures evaluate the goodness of pitch histogram specific to a singer, while the relative measures use the similarity between singers in terms of pitch, rhythm, and timbre as an indicator of singing quality. With the relative measures, we formulate the concept of veracity or truth-finding for the ranking of singing quality. We successfully validate a self-organizing approach to rank-ordering a large pool of singers. The fusion of absolute and relative measures results in an average Spearman's rank correlation of 0.71 with human judgments in a 10-fold cross-validation experiment, which is close to the inter-judge correlation.
Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Automatic Evaluation of Song Intelligibility Using Singing Adapted STOI and Vocal-Specific Features
abstract
An objective machine-driven measure of song intelligibility would be of great utility for various music information retrieval tasks. Song intelligibility mostly depends on two factors, the amount of interference caused by background accompaniment, and the quality of singing vocal. We leverage these two factors to determine the intelligibility of a song. For the first factor, we adapt a well known method for intelligibility prediction of noisy speech, short term objective intelligibility (STOI), to singing. The singing-adapted STOI considers the polyphonic song as a time-frequency weighted noisy version of the extracted singing vocal. We use U-net based audio source separation method to extract singing vocal from a polyphonic song. The singing vocal shares the same underlying physiological mechanism for production as that of speech, with some differences in the pronunciation and prosody of the phonemes. Therefore, for the second factor, we have introduced vocal-specific features to measure the intelligibility of the singing vocal, which are excitation source, spectral, and prosodic singing characteristics. We perform detailed analysis on each of these features to establish their efficacy for quantifying song intelligibility. We train a regression model to derive the intelligibility scores using a combination of the vocal-specific features and singing adapted STOI, obtaining a significant improvement in performance. The correlation between the intelligibility score obtained using proposed framework and human-rated intelligibility score is 0.81, which shows the efficacy of the proposed approach.
Bidisha Sharma, Ye Wang 0007
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 SubSpectralNet - Using Sub-spectrogram Based Convolutional Neural Networks for Acoustic Scene Classification
abstract
Acoustic Scene Classification (ASC) is one of the core research problems in the field of Computational Sound Scene Analysis. In this work, we present SubSpectralNet, a novel model which captures discriminative features by incorporating frequency band-level differences to model soundscapes. Using mel-spectrograms, we propose the idea of using band-wise crops of the input time-frequency representations and train a convolutional neural network (CNN) on the same. We also propose a modification in the training method for more efficient learning of the CNN models. We first give a motivation for using sub-spectrograms by giving intuitive and statistical analyses and finally we develop a sub-spectrogram based CNN architecture for ASC. The system is evaluated on the public ASC development dataset provided for the "Detection and Classification of Acoustic Scenes and Events" (DCASE) 2018 Challenge. Our best model achieves an improvement of +14% in terms of classification accuracy with respect to the DCASE 2018 baseline system. Code and figures are available at https://github.com/ssrp/SubSpectralNet.
Sai Samarth R. Phaye, Emmanouil Benetos, Ye Wang 0007
ICASSP3
2019 Automatic Lyrics-to-audio Alignment on Polyphonic Music Using Singing-adapted Acoustic Models
abstract
Lyrics-to-audio alignment is to automatically align the lyrical words with the mixed singing audio (singing voice+musical accompaniment). Such alignment can be achieved with an automatic speech recognition (ASR) system. We propose to adapt the acoustic model of a speech recognizer towards solo singing voice. This avoids the hurdles of annotating a large polyphonic music training dataset. Moreover, a lexicon-modification based duration modelling has been incorporated to account for the long duration vowels in singing. As practical application demand the alignment on polyphonic music, we study the effect of different singing vocal separation methods in the task of lyrics-to-audio alignment in polyphonic music. The extracted vocals are forced-aligned with the singing-adapted models. We demonstrate that the use of audio source separation method and effective end-pointing of the songs has a high impact on the alignment performance through the experiments. We report a mean average absolute error of 3.87 seconds, which is comparable with the state-of-the-art lyrics-to-audio alignment system that is trained on a large polyphonic music database.
Bidisha Sharma, Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007
ICASSP4
2018 MANA: Designing and Validating a User-Centered Mobility Analysis System
abstract
In this paper, we demonstrate a new IMU-based wearable system (dubbed MANA or Mobility ANAlytics) for measuring gait in a clinical setting. The design process and choices that were made to ensure that the technology was invisible and accessible are described. We collect a rich and diverse dataset of walking data from sixty participants, including forty people with Parkinson's Disease (PD). The system is then validated in a clinical setting with this dataset. We present novel and innovative algorithms to measure common gait parameters. The system is able to estimate these gait parameters with high accuracy, with a mean absolute error of 4.0 cm for stride length and 2.6 cm for step length, outperforming all state-of-the-art methods that included data from people with PD.
Boyd Anderson, Shenggao Zhu, Hugh Anderson, Chao Xu Tay, Vincent Y. F. Tan, Ye Wang 0007
ASSETS8
2018 Automatic Pronunciation Evaluation of Singing
Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007
INTERSPEECH3
2018 SLIONS: A Karaoke Application to Enhance Foreign Language Learning
abstract
Singing songs can be an engaging and effective activity when learning a foreign language. In this paper, we describe a multi-language karaoke application called SLIONS: Singing and Listening to Improve Our Natural Speaking. When developing this application, we followed a user-centered design process which was informed by conducting interviews with domain experts, extensive usability testing, and reviewing existing gamified karaoke and language learning applications. The key feature of SLIONS is that we used automatic speech recognition (ASR) to provide students with personalized, granular feedback based on their singing pronunciation. We also provided multi-modal instruction: audio of music and singing tracks, video of a professional singer and translated text of lyrics to help students learn and master each song in the foreign language. To test the efficacy of SLIONS, we conducted a one-week pilot study with English and Chinese language learning students (N=15). The initial quantitative results show that our application can improve pronunciation and may improve vocabulary. In addition, the qualitative feedback from the students suggests that SLIONS is both fun to use and motivates students to practice speaking and singing in a foreign language.
Dania Murad, Riwu Wang, Douglas Turnbull, Ye Wang 0007
ACM Multimedia4
2016 A Computer Vision-Based System for Stride Length Estimation using a Mobile Phone Camera
abstract
Conditions such as Parkinson's disease (PD), a chronic neurodegenerative disorder which severely affects the motor system, will be an increasingly common problem for our growing and aging population. Gait analysis is widely used as a noninvasive method for PD diagnosis and assessment. However, current clinical systems for gait analysis usually require highly specialized cameras and lab settings, which are expensive and not scalable. This paper presents a computer vision-based gait analysis system using a camera on a common mobile phone. A simple PVC mat was designed with markers printed on it, on which a subject can walk whilst being recorded by a mobile phone camera. A set of video analysis methods were developed to segment the walking video, detect the mat and feet locations, and calculate gait parameters such as stride length. Experiments showed that stride length measurement has a mean absolute error of 0.62 cm, which is comparable with the "gold standard" walking mat system GAITRite. We also tested our system on Parkinson's disease patients in a real clinical environment. Our system is affordable, portable, and scalable, indicating a potential clinical gait measurement tool for use in both hospitals and the homes of patients.
Boyd Anderson, Shenggao Zhu, Ye Wang 0007
ASSETS4
2014 Improving Content-based and Hybrid Music Recommendation using Deep Learning
abstract
Existing content-based music recommendation systems typically employ a \textit{two-stage} approach. They first extract traditional audio content features such as Mel-frequency cepstral coefficients and then predict user preferences. However, these traditional features, originally not created for music recommendation, cannot capture all relevant information in the audio and thus put a cap on recommendation performance. Using a novel model based on deep belief network and probabilistic graphical model, we unify the two stages into an automated process that simultaneously learns features from audio content and makes personalized recommendations. Compared with existing deep learning based models, our model outperforms them in both the warm-start and cold-start stages without relying on collaborative filtering (CF). We then present an efficient hybrid method to seamlessly integrate the automatically learnt features and CF. Our hybrid method not only significantly improves the performance of CF but also outperforms the traditional feature mbased hybrid method.
Xinxi Wang, Ye Wang 0007
ACM Multimedia2
2014 Validating an iOS-based Rhythmic Auditory Cueing Evaluation (iRACE) for Parkinson's Disease
abstract
Movement disorders such as Parkinson's disease (PD) will affect a rapidly growing segment of the population as society continues to age. Rhythmic Auditory Cueing (RAC) is a well-supported evidence-based intervention for the treatment of gait impairments in PD. RAC interventions have not been widely adopted, however, due to limitations in access to personnel, technological, and financial resources. To help "scale up" RAC for wider distribution, we have developed an iOS-based Rhythmic Auditory Cueing Evaluation (iRACE) mobile application to deliver RAC and assess motor performance in PD patients. The touchscreen of the mobile device is used to assess motor timing during index finger tapping, and the device's built-in tri-axial accelerometer and gyroscope to assess step time and step length during walking. Novel machine learning-based gait analysis algorithms have been developed for iRACE, including heel strike detection, step length quantification, and left-versus-right foot identification. The concurrent validity of iRACE was assessed using a clinic-standard instrumented walking mat and a pair of force-sensing resistor sensors. Results from 10 PD patients reveal that iRACE has low error rates (<±1.0%) across a set of four clinically relevant outcome measures, indicating a potentially useful clinical tool.
Shenggao Zhu, Robert J. Ellis, Gottfried Schlaug, Yee Sien Ng, Ye Wang 0007
ACM Multimedia5
2014 Exploration in Interactive Personalized Music Recommendation: A Reinforcement Learning Approach
abstract
Current music recommender systems typically act in a greedy manner by recommending songs with the highest user ratings. Greedy recommendation, however, is suboptimal over the long term: it does not actively gather information on user preferences and fails to recommend novel songs that are potentially interesting. A successful recommender system must balance the needs to explore user preferences and to exploit this information for recommendation. This article presents a new approach to music recommendation by formulating this exploration-exploitation trade-off as a reinforcement learning task. To learn user preferences, it uses a Bayesian model that accounts for both audio content and the novelty of recommendations. A piecewise-linear approximation to the model and a variational inference algorithm help to speed up Bayesian inference. One additional benefit of our approach is a single unified model for both music recommendation and playlist generation. We demonstrate the strong potential of the proposed approach with simulation results and a user study.
Xinxi Wang, Yi Wang 0006, David Hsu, Ye Wang 0007
ACM Trans. Multim. Comput. Commun. Appl.4
2013 Non-reference audio quality assessment for online live music recordings
abstract
Immensely popular video sharing websites such as YouTube have become the most important sources of music information for Internet users and the most prominent platform for sharing live music. The audio quality of this huge amount of live music recordings, however, varies significantly due to factors such as environmental noise, location, and recording device. However, most video search engines do not take audio quality into consideration when retrieving and ranking results. Given the fact that most users prefer live music videos with better audio quality, we propose the first automatic, non-reference audio quality assessment framework for live music video search online. We first construct two annotated datasets of live music recordings. The first dataset contains 500 human-annotated pieces, and the second contains 2,400 synthetic pieces systematically generated by adding noise effects to clean recordings. Then, we formulate the assessment task as a ranking problem and try to solve it using a learning-based scheme. To validate the effectiveness of our framework, we perform both objective and subjective evaluations. Results show that our framework significantly improves the ranking performance of live music recording retrieval and can prove useful for various real-world music applications.
Ju-Chiang Wang, Jingli Cai, Zhiyan Duan, Hsin-Min Wang, Ye Wang 0007
ACM Multimedia6
2013 Query-Document-Dependent Fusion: A Case Study of Multimodal Music Retrieval
abstract
In recent years, multimodal fusion has emerged as a promising technology for effective multimedia retrieval. Developing the optimal fusion strategy for different modalities (e.g., content, metadata) has been the subject of intensive research. Given a query, existing methods derive a unified fusion strategy for all documents with the underlying assumption that the relative significance of a modality remains the same across all documents. However, this assumption is often invalid. We thus propose a general multimodal fusion framework, query-document-dependent fusion (QDDF), which derives the optimal fusion strategy for each query-document pair via intelligent content analysis of both queries and documents. By investigating multimodal fusion strategies adaptive to both queries and documents, we demonstrate that existing multimodal fusion approaches are special cases of QDDF and propose two QDDF approaches to derive fusion strategies. The dual-phase QDDF explicitly derives and fuses query- and document-dependent weights, and the regression-based QDDF determines the fusion weight for a query-document pair via a regression model derived from training data. To evaluate the proposed approaches, comprehensive experiments have been conducted using a multimedia data set with around 17 K full songs and over 236 K social queries. Results indicate that the regression-based QDDF is superior in handling single-dimension queries. In comparison, the dual-phase QDDF outperforms existing approaches for most query types. We found that document-dependent weights are instrumental in enhancing multimedia fusion performance. In addition, efficiency analysis demonstrates the scalability of QDDF over large data sets.
Bingjun Zhang, Yi Yu 0001, Jialie Shen 0001, Ye Wang 0007
IEEE Trans. Multim.5
2013 Scalable Content-Based Music Retrieval Using Chord Progression Histogram and Tree-Structure LSH
abstract
With more and more multimedia content made available on the Internet, music information retrieval is becoming a critical but challenging research topic, especially for real-time online search of similar songs from websites. In this paper we study how to quickly and reliably retrieve relevant songs from a large-scale dataset of music audio tracks according to melody similarity. Our contributions are two-fold: (i) Compact and accurate representation of audio tracks by exploiting music semantics. Chord progressions are recognized from audio signals based on trained music rules, and the recognition accuracy is improved by multi-probing. A concise chord progression histogram (CPH) is computed from each audio track as a mid-level feature, which retains the discriminative capability in describing audio content. (ii) Efficient organization of audio tracks according to their CPHs by using only one locality sensitive hash table with a tree-structure. A set of dominant chord progressions of each song is used as the hash key. Average degradation of ranks is further defined to estimate the similarity of two songs in terms of their dominant chord progressions, and used to control the number of probing in the retrieval stage. Experimental results on a large dataset with 74,055 music audio tracks confirm the scalability of the proposed retrieval algorithm. Compared to state-of-the-art methods, our algorithm improves the accuracy of summarization and indexing, and makes a further step towards the optimal performance determined by an exhaustive sequence comparison.
Yi Yu 0001, Roger Zimmermann, Ye Wang 0007, Vincent Oria
IEEE Trans. Multim.3
2012 Recognition and Summarization of Chord Progressions and Their Application to Music Information Retrieval
abstract
Accurate and compact representation of music signals is a key component of large-scale content-based music applications such as music content management and near duplicate audio detection. This problem is not well solved yet despite many research efforts in this field. In this paper, we suggest mid-level summarization of music signals based on chord progressions. More specially, in our proposed algorithm, chord progressions are recognized from music signals based on a supervised learning model, and recognition accuracy is improved by locally probing n-best candidates. By investigating the properties of chord progressions, we further calculate a histogram from the probed chord progressions as a summary of the music signal. We show that the chord progression-based summarization is a powerful feature descriptor for representing harmonic progressions and tonal structures of music signals. The proposed algorithm is evaluated with content-based music retrieval as a typical application. The experimental results on a dataset with more than 70,000 songs confirm that our algorithm can effectively improve summarization accuracy of musical audio contents and retrieval performance, and enhance music retrieval applications on large-scale audio databases.
Yi Yu 0001, Roger Zimmermann, Ye Wang 0007, Vincent Oria
ISM3
2012 A domain-specific music search engine for gait training
abstract
This paper demonstrates a domain-specific music retrieval system to help music therapists find appropriate music for Parkinson's disease patients in their gait training. Different from existing music search engines, this system incorporates multiple music dimensions (i.e., tempo, cultural style, and beat strength) required in gait training, and facilitates the searching process by allowing music retrieval directly on these dimensions. To support music search by tempo, a user-perception based method is also proposed to improve state-of-the-art tempo estimation algorithms. We conducted a user study and evaluated the system efficacy in searching music using different types of queries based on these music dimensions. Experimental results demonstrate the effectiveness and usability of our system in therapeutic gait training.
Ye Wang 0007
ACM Multimedia2
2012 Context-aware mobile music recommendation for daily activities
abstract
Existing music recommendation systems rely on collaborative filtering or content-based technologies to satisfy users' long-term music playing needs. Given the popularity of mobile music devices with rich sensing and wireless communication capabilities, we present in this paper a novel approach to employ contextual information collected with mobile devices for satisfying users' short-term music playing needs. We present a probabilistic model to integrate contextual information with music content analysis to offer music recommendation for daily activities, and we present a prototype implementation of the model. Finally, we present evaluation results demonstrating good accuracy and usability of the model and prototype.
Xinxi Wang, David S. Rosenblum, Ye Wang 0007
ACM Multimedia3
2012 A daily, activity-aware, mobile music recommender system
abstract
Existing music recommender systems rely on collaborative filtering or content-based technologies to satisfy users' long-term music playing needs. Given the popularity of mobile music devices with rich sensing and wireless communication capabilities, we demonstrate in this demo a novel system to employ contextual information collected with mobile devices for satisfying users' short-term music playing needs. In our system, contextual information is integrated with music content analysis to offer recommendation for daily activities.
Xinxi Wang, Ye Wang 0007, David S. Rosenblum
ACM Multimedia2
2012 MOGAT: a cloud-based mobile game system with auditory training for children with cochlear implants
abstract
Musical auditory habilitation is an essential process in adapting cochlear implant recipients to the musical hearing context provided by cochlear implants. However, due to the cost and time limitation, it is impossible for hearing healthcare professionals to provide intensive and extensive musical auditory habilitation for every cochlear implant recipient. In order to provide an efficient and cost-effective musical auditory training for children with cochlear implants, we designed and developed MObile Games with Auditory Training (MOGAT) on off-the-shelf mobile devices. MOGAT includes three intuitive and interesting mobile games for training pitch perception and production, and a cloud-based web service for music therapists to support and evaluate individual habilitation. We demonstrate MOGAT for enhancing musical habilitation for children with cochlear implants.
Yinsheng Zhou, Toni-Jan Keith Palma Monserrat, Ye Wang 0007
ACM Multimedia3
2012 MOGAT: mobile games with auditory training for children with cochlear implants
abstract
Cochlear implants have improved the lives of tens of thousands of the hearing impaired by providing sufficient auditory perception for speech, but these devices are far from satisfactory for music perception. Many cochlear implant recipients, especially pre-lingually deafened children, have difficulty recognizing and producing specific pitches. To improve musical auditory habilitation for children post cochlear implantation, we developed MOGAT: MObile Games with Auditory Training. The system includes three musical games built with off-the-shelf mobile devices to train their pitch perception and intonation skills respectively, and a cloud-based web service which allows music therapists to monitor and design individual training for children. The design of the games and web service was informed by a pilot survey (N=60 children). To ensure widespread use with low-cost mobile devices, we minimized the computation load while retaining highly accurate audio analysis. A 6-week user study (N=15 children) showed that the music habilitation with MOGAT was intuitive, enjoyable and motivating. It has improved most children's pitch discrimination and production, and several children's improvement was statistically significant (p<0.05).
Yinsheng Zhou, Khe Chai Sim, Patsy Tan, Ye Wang 0007
ACM Multimedia4
2011 MOGCLASS: evaluation of a collaborative system of mobile devices for classroom music education of young children
abstract
Composition, listening, and performance are essential activities in classroom music education, yet conventional music classes impose unnecessary limitations on students' ability to develop these skills. Based on in-depth fieldwork and a user-centered design approach, we created MOGCLASS, a multimodal collaborative music environment that enhances students' musical experience and improves teachers' management of the classroom.
Yinsheng Zhou, Graham Percival, Xinxi Wang, Ye Wang 0007, Shengdong Zhao 0001
CHI4
2011 Document dependent fusion in multimodal music retrieval
abstract
In this paper, we propose a novel multimodal fusion framework, document dependent fusion (DDF), which derives the optimal combination strategy for each individual document in the fusion process. For each document, we derive a document weight vector by estimating the descriptive abilities of its different modalities. The document weight vector also enables our framework to be easily integrated with existing multimodal fusion schemes, and achieve a better combination strategy for each document given a query. Experiments are conducted on a 17174-song music database to compare the retrieval accuracy of traditional query independent fusion and query dependent fusion approaches, and that obtained after integrating DDF with them. Experimental results indicate that DDF can significantly improve the retrieval performance of current fusion approaches.
Bingjun Zhang, Ye Wang 0007
ACM Multimedia3
2011 Sensor-Assisted Video Encoding for Mobile Devices in Real-World Environments
abstract
In this paper, we present a comprehensive study on sensor-assisted video encoding (SaVE) schemes for video capturing on mobile devices in real-world environments. Our purpose is to reduce the computational complexity of video encoding by leveraging sensors that are increasingly available on mobile devices, e.g., accelerometers and digital compasses. Motion estimation is a key component of video encoding. In this paper, SaVE calculates the rotational movement of a camera (on mobile devices) and then infers the global motion in the camera imager. SaVE subsequently employs the estimated global motion as predictors to simplify motion estimation algorithms for state-of-the-art H.264/AVC video coding. We have constructed a prototype of SaVE and evaluated its performance with a pair of accelerometers, a digital compass, and their combination. Our experimental results show that SaVE can significantly reduce the computations of motion estimation while achieving equal or better video quality. Our results also show that SaVE has a strong noise-resistant capability. Therefore, it can be practically employed in real-world environments.
Xiaoming Chen 0006, Zhendong Zhao, Ahmad Rahmati, Ye Wang 0007, Lin Zhong 0001
IEEE Trans. Circuits Syst. Video Technol.4
2010 A music search engine for therapeutic gait training
abstract
A music retrieval system is introduced that incorporate tempo, cultural, and beat strength features to help music therapists provide appropriate music for gait training for Parkinson's patients. Unlike current methods available to music therapists (e.g., personal CD/MP3 library search) we propose a domain-specific search engine that utilizes database of music found on YouTube. We independently evaluate the efficacy of our tempo, cultural, and beat strength features on a music database extracted from YouTube. Results from our user study demonstrate the effectiveness and usefulness of our search engine for this application.
Qiaoliang Xiang, Jason Hockman, Jianqing Yang, Yu Yi, Ichiro Fujinaga, Ye Wang 0007
ACM Multimedia7
2010 Automated sleep quality measurement using EEG signal: first step towards a domain specific music recommendation system
abstract
With the rapid pace of modern life, millions of people suffer from sleep problems. Music therapy, as a non-medication approach to mitigating sleep problems, has attracted increasing attention recently. However the adaptability of music therapy is limited by the time consuming task of choosing suitable music for users. Inspired by this observation, we discuss the concept of a domain specific music recommendation system, which automatically recommends music for users according to their sleep quality. The proposed system requires multidisciplinary efforts including automated sleep quality measurement and content-based music similarity measure. As a first step, we focus on the automated sleep quality measurement in this paper. An EEG-based approach is proposed to measure user's sleep quality. The advantages of our approach over standard Polysomnography (PSG) method are: 1) it measures sleep quality by recognizing three sleep categories rather than six sleep stages, thus higher accuracy can be expected; 2) three sleep categories are recognized by analyzing Electroencephalography (EEG) signal only, so the user experience is improved because he is attached with fewer sensors during sleep. We conduct experiments based on a standard data set. Our approach achieves high accuracy and shows promising potential for the music recommendation system.
Xinxi Wang, Ye Wang 0007
ACM Multimedia3
2010 Large-scale music tag recommendation with explicit multiple attributes
abstract
Social tagging can provide rich semantic information for large-scale retrieval in music discovery. Such collaborative intelligence, however, also generates a high degree of tags unhelpful to discovery, some of which obfuscate critical information. Towards addressing these shortcomings, tag recommendation for more robust music discovery is an emerging topic of significance for researchers. However, current methods do not consider diversity of music attributes, often using simple heuristics such as tag frequency for filtering out irrelevant tags. Music attributes encompass any number of perceived dimensions, for instance vocalness, genre, and instrumentation. Many of these are underrepresented by current tag recommenders. We propose a scheme for tag recommendation using Explicit Multiple Attributes based on tag semantic similarity and music content. In our approach, the attribute space is explicitly constrained at the outset to a set that minimizes semantic loss and tag noise, while ensuring attribute diversity. Once the user uploads or browses a song, the system recommends a list of relevant tags in each attribute independently. To the best of our knowledge, this is the first method to consider Explicit Multiple Attributes for tag recommendation. Our system is designed for large-scale deployment, on the order of millions of objects. For processing large-scale music data sets, we design parallel algorithms based on the MapReduce framework to perform large-scale music content and social tag analysis, train a model, and compute tag similarity. We evaluate our tag recommendation system on CAL-500 and a large-scale data set ($N = 77,448$ songs) generated by crawling Youtube and Last.fm. Our results indicate that our proposed method is both effective for recommending attribute-diverse relevant tags and efficient at scalable processing.
Zhendong Zhao, Xinxi Wang, Qiaoliang Xiang, Andy M. Sarroff, Ye Wang 0007
ACM Multimedia6
2010 MOGCLASS: a collaborative system of mobile devices forclassroom music education
abstract
We introduce MOGCLASS: a system of networked mobile devices to amplify and extend children's capabilities to perceive, perform and produce music collaboratively in classroom context. MOGCLASS includes various features for students to enhance their motivation, interest, and collaboration in music class. It provides a wide-ranging palette of easy-to-use musical instruments for students to choose from, and supports both collaborative silent practice with headphones, and collaborative performance with loudspeakers. To facilitate classroom management, the teacher's interface is used to control students' activities. Our evaluation results indicate that MOGCLASS is effective in increasing students' motivation in learning music and in supporting teachers' classroom management
Yinsheng Zhou, Graham Percival, Xinxi Wang, Ye Wang 0007, Shengdong Zhao 0001
ACM Multimedia4
2009 Cultural style based music classification of audio signals
abstract
Music classification based on cultural style is useful for music analysis and has potential applications in retrieval and recommendation systems. In this paper, we present the first attempt to classify audio signals automatically according to their cultural styles, which are characterized by timbre, rhythm, wavelet coefficients and musicology-based features. Machine learning algorithms are employed to investigate the effectiveness of various features on a data set of 1300 music pieces. Experimental results show that the proposed method can achieve an overall accuracy of 86% for six cultural styles, which shows the feasibility of integrating cultural style classification into music retrieval systems.
Qiaoliang Xiang, Ye Wang 0007, Lianhong Cai
ICASSP3
2009 SaVE: sensor-assisted motion estimation for efficient h.264/AVC video encoding
abstract
Motion estimation is a key component of modern video encoding and is very compute-intensive. We present a novel Sensor-assisted Video Encoding (SaVE) method to reduce the computational complexity of motion estimation in H.264/AVC encoders, leveraging accelerometers and digital compasses that are increasingly available on mobile devices. Using these sensors, SaVE calculates the rotational movement of a camera and then infers the global motion in the camera image sensor; it subsequently employs the estimated global motion to simplify the state-of-the-art motion estimation algorithms, UMHS and EPZS used in H.264/AVC encoders. We have constructed a prototype of SaVE and report extensive evaluation of it. Our experimental results show that SaVE can reduce the computations of UMHS and EPZS algorithms by up to 27% and 18%, respectively, while achieving the same or better video quality.
Xiaoming Chen 0006, Zhendong Zhao, Ahmad Rahmati, Ye Wang 0007, Lin Zhong 0001
ACM Multimedia4
2009 Comprehensive query-dependent fusion using regression-on-folksonomies: a case study of multimodal music search
abstract
The combination of heterogeneous knowledge sources has been widely regarded as an effective approach to boost retrieval accuracy in many information retrieval domains. While various technologies have been recently developed for information retrieval, multimodal music search has not kept pace with the enormous growth of data on the Internet. In this paper, we study the problem of integrating multiple online information sources to conduct effective query dependent fusion (QDF) of multiple search experts for music retrieval. We have developed a novel framework to construct a knowledge space of users' information need from online folksonomy data. With this innovation, a large number of comprehensive queries can be automatically constructed to train a better generalized QDF system against unseen user queries. In addition, our framework models QDF problem by regression of the optimal combination strategy on a query. Distinguished from the previous approaches, the regression model of QDF (RQDF) offers superior modeling capability with less constraints and more efficient computation. To validate our approach, a large scale test collection has been collected from different online sources, such as Last.fm, Wikipedia, and YouTube. All test data will be released to the public for better research synergy in multimodal music search. Our performance study indicates that the accuracy, efficiency, and robustness of the multimodal music search can be improved significantly by the proposed folksonomy-RQDF approach. In addition, since no human involvement is required to collect training examples, our approach offers great feasibility and practicality in system development.
Bingjun Zhang, Qiaoliang Xiang, Huanhuan Lu, Jialie Shen 0001, Ye Wang 0007
ACM Multimedia5
2009 CompositeMap: a novel music similarity measure for personalized multimodal music search
abstract
How to measure and model the similarity between different music items is one of the most fundamental yet challenging research problems in music information retrieval. This paper demonstrates a novel multimodal and adaptive music similarity measure (CompositeMap) with its application in a personalized multimodal music search system. CompositeMap can effectively combine music properties from different aspects into compact signatures via supervised learning, which lays the foundation for effective and efficient music search. In addition, an incremental Locality Sensitive Hashing algorithm is developed to support more efficient search processes. Experimental results based on two large music collections reveal various advantages in effectiveness, efficiency, adaptiveness, and scalability of the proposed music similarity measure and the music search system.
Bingjun Zhang, Qiaoliang Xiang, Ye Wang 0007, Jialie Shen 0001
ACM Multimedia3
2009 MOGFUN: musical mObile group for FUN
abstract
The computational power and sensory capabilities of mobile devices are increasing dramatically these days, rendering them suitable for real-time sound synthesis and various musical expressions. In this paper, we demonstrate a novel mobile music making system which leverages the ubiquity, ultra-mobility, and multi-modality of mobile devices (iPod touch) for people to create and compose music collaboratively. Unlike the conventional music making applications which generate the music on a single mobile device with a preset sound and interface, our system allows several players in a group to be connected together through wireless LAN network, creating music with different sounds and interfaces. Finally, the performance can be recorded as a single music file and played back in the future. The paper also shows some application scenarios for this collaborative music making system in future research.
Yinsheng Zhou, Dillion Tan, Graham Percival, Ye Wang 0007
ACM Multimedia5
2009 CompositeMap: a novel framework for music similarity measure
abstract
With the continuing advances in data storage and communication technology, there has been an explosive growth of music information from different application domains. As an effective technique for organizing, browsing, and searching large data collections, music information retrieval is attracting more and more attention. How to measure and model the similarity between different music items is one of the most fundamental yet challenging research problems. In this paper, we introduce a novel framework based on a multimodal and adaptive similarity measure for various applications. Distinguished from previous approaches, our system can effectively combine music properties from different aspects into a compact signature via supervised learning. In addition, an incremental Locality Sensitive Hashing algorithm has been developed to support efficient retrieval processes with different kinds of queries. Experimental results based on two large music collections reveal various advantages of the proposed framework including effectiveness, efficiency, adaptiveness, and scalability.
Bingjun Zhang, Jialie Shen 0001, Qiaoliang Xiang, Ye Wang 0007
SIGIR4
2009 A joint encoder-decoder framework for supporting energy efficient audio decoding
Wendong Huang, Ye Wang 0007
Multim. Syst.2
2009 An optimal speed control scheme supported by media servers for low-power multimedia applications
Wendong Huang, Ye Wang 0007
Multim. Syst.2
2008 Onset detection in pitched non-percussive music using warping-compensated correlation
abstract
Automatically extracting temporal information from musical recordings is inarguably one of the most critical subtasks of many music information retrieval systems. In this paper we present a system for automatic note onset detection in pitched non-percussive (PNP) musical sounds, which is the most challenging audio signal group for this task. We propose a new approach based on stable pitch cues and signal energy. A computationally inexpensive method for feature extraction, which efficiently suppresses vibrato, is combined with information derived from the signal energy in the feature space. Onsets are localized by a median filter based peak picking method. The proposed method is tested against a database of annotated violin recordings, covering a wide range of tempo and playing styles like vibrato and staccato. Our system outperforms prior state of the art systems with results for true positives of 91.2% and false positives of 9.2%.
Olaf Schleusing, Bingjun Zhang, Ye Wang 0007
ICASSP3
2008 Multimedia power management on a platter: from audio to video & games
abstract
Today, battery-life is a major design concern for all portable devices ranging from cell phones to PDAs and portable game consoles. The purpose of this tutorial will be to give an overview of power management techniques that are applicable to multimedia applications running on such battery-operated portable devices. In particular, we will discuss a host of techniques, some of which are applicable to audio processing applications, some to video processing, and the others to interactive 3D game applications. The tutorial will be helpful to students, researchers, application developers and engineers who have a background in traditional real-time multimedia applications and would like to get an overview of the important issues and solutions pertaining to using and developing power management techniques for the multimedia domain.
Samarjit Chakraborty, Ye Wang 0007
ACM Multimedia2
2008 SenseCoding: accelerometer-assisted motion estimation for efficient video encoding
abstract
Accelerometers have appeared on many camcorders, cameras and mobile phones. We present algorithms that estimate camera movement from accelerometer readings and apply the estimation to significantly improve the compute-intense motion estimation in video encoding. We have implemented a working prototype that simultaneously captures video and three-axis acceleration data. The video is then compressed with a reference MPEG-2 encoder, modified to incorporate the accelerometer readings to assist motion estimation. Our experimental data shows a two to three times speed improvement for the entire encoding process, in comparison with full search.
Guangming Hong, Ahmad Rahmati, Ye Wang 0007, Lin Zhong 0001
ACM Multimedia3
2008 iDVT: an interactive digital violin tutoring system based on audio-visual fusion
abstract
iDVT (interactive Digital Violin Tutor) is a violin learning system exploiting physical and virtual resources and interactivity. It aims at providing the user with new effective learning experience. This demonstration paper briefly describes the structure of the system and the underlying audio-visual processing techniques employed in the system.
Huanhuan Lu, Bingjun Zhang, Ye Wang 0007, Wee Kheng Leow
ACM Multimedia3
2008 Decoding-workload-aware video encoding
abstract
This paper presents a novel decoding-workload-aware video encoding scheme. It takes raw video data and decoding workload constraint of a mobile client as input and generates a video bitstream which matches such a constraint while striving to achieve the best video quality. For a given constraint, the best overall video quality of the encoded bitstream is selected with a tradeoff between spatial and temporal distortions. The main contributions of this paper include: 1) the proposal of an efficient scheme which selects the most suitable target frame rate before the actual encoding; 2) The design of a workload control (analogous to the rate control) scheme which ensures an accurate control of the decoding workload when the bitstream is generated using the proposed encoding scheme. Experimental results demonstrate the feasibility and performance of the proposed scheme.
Guangming Hong, An Vu Tran, Ye Wang 0007
NOSSDAV4
2008 LyricAlly: Automatic Synchronization of Textual Lyrics to Acoustic Music Signals
abstract
We present LyricAlly, a prototype that automatically aligns acoustic musical signals with their corresponding textual lyrics, in a manner similar to manually-aligned karaoke. We tackle this problem based on a multimodal approach, using an appropriate pairing of audio and text processing to create the resulting prototype. LyricAlly's acoustic signal processing uses standard audio features but constrained and informed by the musical nature of the signal. The resulting detected hierarchical rhythm structure is utilized in singing voice detection and chorus detection to produce results of higher accuracy and lower computational costs than their respective baselines. Text processing is employed to approximate the length of the sung passages from the lyrics. Results show an average error of less than one bar for per-line alignment of the lyrics on a test bed of 20 songs (sampled from CD audio and carefully selected for variety). We perform a comprehensive set of system-wide and per-component tests and discuss their results. We conclude by outlining steps for further development.
Min-Yen Kan, Ye Wang 0007, Denny Iskandar, Tin Lay Nwe, Arun Shenoy
IEEE Trans. Speech Audio Process.2
2007 Pop Music Beat Detection in the Huffman Coded Domain
abstract
This paper presents a novel beat detector that operates in the Huffman coded domain of a MP3 audio bitstream. We seek to answer two main questions. First, whether it is possible to extract beats without even partial decoding of a MP3 audio. Second, how to construct a low complexity algorithm to detect beats for multimedia applications in small devices such as mobile phones, if the answer to the first question is positive. Our investigation shows that beat detection in the Huffman coded domain can achieve fairly good results with a significantly reduced demand for computation. We also propose a graph-based algorithm to tackle the beat detection problem from an algorithmic perspective.
Jia Zhu 0004, Ye Wang 0007
ICME2
2007 A compressed domain distortion measure for fast video transcoding
abstract
Video applications on different mobile devices are becoming increasingly popular. It is an attractive alternative to transcode a high quality non-scalable video bitstream to match constraints (such as bandwidth or processing power) of different platforms with a similar functionality as a scalable video format. In principle, such a transcoder can reduce either the bit per frame (bpf) or the frame per second (fps) of the original bitstream to meet a particular constraint. In the case that multiple candidates with different combinations of bpf and fps satisfy the constraint, an objective video quality measure is needed for the transcoder to choose the candidate with the overall best quality considering both the spatial quality (reflected by bpf) and the temporal quality (reflected by fps). Conventional measures, such as PSNR and MSE operate in the pixel-domain, require full decoding of both the original and candidate video bitstreams and are computationally very expensive. This drawback renders them unsuitable for real-time transcoding applications. To solve this problem, we propose a Mean Compressed Domain Error (MCDE) to predict the quality of the transcoded video. Experimental results show that the proposed MCDE can predict video quality accurately with a negligible computational complexity in comparison with the conventional MSE/PSNR.
An Vu Tran, Ye Wang 0007
ACM Multimedia3
2007 A workload prediction model for decoding mpeg video and its application to workload-scalable transcoding
abstract
Multimedia playback is restricted by the processing power of mobile devices, and in particular, the playback quality can be degraded due to insufficient processing power. To address this problem, we propose a new workload-scalable transcoding scheme which converts a pre-recorded video bitstream into a new video bitstream that satisfies the device's workload constraint, while keeping the transcoding distortion minimal. The key of this proposed transcoding scheme lies on a new workload prediction model, which is fast, accurate and is generic enough to apply to different video formats, decoder implementations and target platforms. The main contributions of this paper include 1) a workload prediction model for decoding MPEG video based on an offline bitstream analysis method; 2) a transcoding scheme that uses the proposed model to control the decoding workload on the target device. To facilitate our transcoding scheme, we have proposed a compressed domain distortion measure (CDDM) that takes effects from both frames per second (fps) and bits per frame (bpf) into consideration. CDDM ensures the transcoded video bitstream to have the best playback quality given the device's workload constraint. Both the workload prediction model and the transcoding scheme are evaluated experimentally.
An Vu Tran, Ye Wang 0007
ACM Multimedia3
2007 Visual analysis of fingering for pedagogical violin transcription
abstract
Automatic music transcription, in spite of decades of research, remains a challenging research problem. The traditional audio-only approach has yet to achieve a satisfactory performance for any computer-aided pedagogical system. Inspired by the high correlation between violin playing techniques (fingering, bowing) and the played acoustic notes, this paper presents a first attempt in visual analysis of violin fingering to compensate for the difficulties in audio-only music transcription. This is achieved by a robust multiple finger tracking algorithm and a string detection method that extract press, release, and fingertip position from the fingering video and automatically translate the fingering information into the played acoustic note, i.e., onset, offset, and pitches. Experimental results reveal high correctness in multiple finger tracking and string detection, thus paving the way for an improved audio-visual violin transcription system.
Bingjun Zhang, Jia Zhu 0004, Ye Wang 0007, Wee Kheng Leow
ACM Multimedia3
2006 A Violin Music Transcriber for Personalized Learning
abstract
This paper presents a new version of our violin music transcriber [1] to support personalized learning. The proposed method is designed to detect duo-pitch (two strings being bowed at the same time) from real-world violin audio signals recorded in a home environment. Our method uses a semitone band spectrogram, a signal spectral representation with direct musical relevance. We exploit constraints of violin sound to improve the transcription performance and speed in comparison with existing methods. We have carried out rigorous evaluations using (a) single pitch notes and duo-phonic pitch samples within the violin's playing range (G3-B6), and (b) music excerpts. For pitch and duo-pitch samples our method can achieve a transcription precision score of 93.1% and recall score of 96.7% respectively. For music excerpts, an average of 95% of all notes could be found (recall), and 93% of notes transcribed correctly (precision).
Wei Jie Jonathan Boo, Ye Wang 0007, Alex Loscos
ICME2
2006 Efficient Partial Spectrum Reconstruction using an Asymmetric PQMF Algorithm for MPEG-Coded Stereo Audio
abstract
This paper presents a novel algorithm of a scalable and efficient pseudo-quadrature mirror filters (PQMF), which is employed for partial decoding a single-layer audio bitstream such as MP3, typically coded in joint/MS mode. The proposed algorithm is a new extension to our previous work on scalable audio decoding and is designed for asymmetric partial spectrum reconstruction (APSR), where perceptually irrelevant computations are removed. Furthermore, an efficient up-sampling operation is introduced for right channel output. The slight distortions introduced by our simple up-sampling method are inaudible according to a set of perceptual evaluations. Simulation results show that 64.6% energy savings can be achieved for a typical configuration in comparison to the standard PQMF algorithm employed by MPEG-1 audio
Wendong Huang, Ye Wang 0007
ICME2
2006 Syllabic level automatic synchronization of music signals and text lyrics
abstract
10.1145/1180639.1180777
Denny Iskandar, Ye Wang 0007, Min-Yen Kan, Haizhou Li 0001
ACM Multimedia2
2006 Generic forward error correction of short frames for IP streaming applications
Jari Korhonen, Ye Wang 0007
Multim. Tools Appl.3
2005 Music transcription using an instrument model
abstract
We introduce a method to transcribe music with the help of an instrument model. One of the most important and difficult problems in music transcription is polyphonic pitch estimation. Common pitch estimation algorithms reported in the literature tend to have errors in the following three situations: missing fundamental; missing harmonics; shared frequencies. We believe an instrument model can make pitch estimation more robust in these three situations and thus can help to improve music transcription accuracy. We devise a spectrum subtraction algorithm to transcribe single and multiple instrument polyphonic music.
Terence Sim, Ye Wang 0007, Arun Shenoy
ICASSP (3)3
2005 Optimization of source and channel coding for voice over IP
abstract
Voice over Internet protocol (VoIP) applications must typically choose a tradeoff between the bits allocated for forward error correcting (FEC) and that for the source coding to achieve the best speech quality at a given packet loss rate. In this paper, we present a new scheme to optimize the speech quality subject to the bandwidth constraints and the packet loss rate. The scheme adopts adaptive multi-rate (AMR) speech codec along with a FEC scheme based on exclusive OR (XOR) operations. Retransmission is also taken into account if the round trip time (RTT) is within a certain limit. We use a simplified E-model as objective metric. Subjective listening tests show that our scheme improves the perceptual speech quality significantly compared to the non-adaptive baseline speech transmission system.
Jari Korhonen, Ye Wang 0007
ICME3
2005 Using offline bitstream analysis for power-aware video decoding in portable devices
abstract
Dynamic voltage/frequency scheduling algorithms for multimedia applications have recently been a subject of intensive research. Many of these algorithms use control-theoretic feedback techniques to predict the future execution demand of an application based on the demand in the recent past. Such techniques suffer from two major disadvantages: (i) they are computationally expensive, and (ii) it is difficult to give performance or quality-of-service guarantees based on these techniques (since the predictions can occasionally turn out to be incorrect). To address these shortcomings, in this paper we propose a completely new approach for dynamic voltage and frequency scaling. Our technique is based on an offline bitstream analysis of multimedia files. Based on this analysis, we insert metadata information describing the computational demand that will be generated when decoding the file. Such bitstream analysis and metadata insertion can be done when the multimedia file is being downloaded into a portable device from a desktop computer. In this paper we illustrate this technique using the MPEG-2 decoder application. We show that the amount of metadata that needs to be inserted is a very small fraction of the total size of the video clip and it can lead to significant energy savings. The metadata inserted will typically consist of the frequency value at which the processor needs to be run at different points in time during the decoding process. Lastly, in contrast to runtime prediction-based techniques, our scheme can be used to provide performance and quality-of-service guarantees and at the same time avoids any runtime computation overhead.
Samarjit Chakraborty, Ye Wang 0007
ACM Multimedia3
2005 Power-aware bandwidth and stereo-image scalable audio decoding
abstract
We propose a new workload-scalable audio decoding scheme that would enable users to control the tradeoff between playback quality and power consumption in battery-powered portable audio players. Our objective is to give users a control at the decoder side, similar to the Long Play (LP) recording mode at the encoder side in many media recording devices. The main contribution of this paper is a proposal for a Bandwidth and Stereo-image Scalable (BSS) decoding scheme for single-layer audio formats such as MP3. The proposed scheme is based on an analysis of the perceptual relevance of different audio components in the compressed bitstream. The bandwidth and stereo-image scalability directly translates into scalability in terms of the computational workload generated by the decoder. This can be exploited by a voltage/frequency scalable processor to save energy and prolong the battery life.
Wendong Huang, Ye Wang 0007, Samarjit Chakraborty
ACM Multimedia2
2005 Digital violin tutor: an integrated system for beginning violin learners
abstract
Prompt feedback is essential for beginning violin learners; however, most amateur learners can only meet with teachers and receive feedback once or twice a week. To help such learners, we have attempted an initial design of Digital Violin Tutor (DVT), an integrated system that provides the much-needed feedback when human teachers are not available. DVT combines violin audio transcription with visualization. Our transcription method is fast, accurate, and robust again noise for violin audio recorded in home environments. The visualization is designed to be intuitive and easily understandable by people with little music knowledge. The different visualization modalities--video, 2D fingerboard animation, 3D avatar animation--help learners to practice and learn more effectively. The entire system has been implemented with off-the-shelf hardware and shown to be practical in home environments. In our user study, the system has received very positive evaluation.
Ye Wang 0007, David Hsu
ACM Multimedia2
2005 Power-efficient streaming for mobile terminals
abstract
Wireless Network Interface (WNI) is one of the most critical components for power efficiency in multimedia streaming to mobile devices. A common strategy to save power is to switch WNI to active mode only when network activity is expected. In streaming systems, this approach is problematic because data are typically received continuously. One solution is to transmit data packets as bursts, which leaves WNI more time between bursts in standby mode. However, that subjects bursty transmission in high peak rates, which leaves it prone to congestion. In this paper, we study theoretically and empirically the impact of burst length and peak transmission rate for observed packet loss and delay characteristics as well as potential energy savings in a Wireless Local Area Network (WLAN) environment. We outline and implement a test system with adaptive burst length to achieve improved trade-off between power efficiency and congestion tolerance.
Jari Korhonen, Ye Wang 0007
NOSSDAV2
2005 Effect of packet size on loss rate and delay in wireless links
abstract
Transmitting large packets over wireless networks helps to reduce header overhead, but may have an adverse effect on loss rate due to corruptions in a radio link. Packet loss in lower layers, however, is typically hidden from the upper protocol layers by link or MAC layer protocols. For this reason, errors in the physical layer are observed by the application as higher variance in end-to-end delay rather than increased packet loss rate. We study the effect of packet size on loss rate and delay characteristics in a wireless real-time application. We derive an analytical model for the dependency between packet length and delay characteristics. We validate our theoretical analysis through experiments in an ad hoc network using WLAN technologies. We show that careful design of packetization schemes in the application layer may significantly improve radio link resource utilization in delay sensitive media streaming under difficult wireless network conditions.
Jari Korhonen, Ye Wang 0007
WCNC2
2005 Toward bandwidth-efficient and error-robust audio streaming over lossy packet networks
Jari Korhonen, Ye Wang 0007, David Isherwood
Multim. Syst.2
2004 Automatic music summarization in compressed domain
abstract
A novel compressed domain automatic music summarization approach is presented in this paper. The proposed method works directly in the compressed domain. Only the encoded subband samples are extracted and processed for characterizing music content and discovering the music structure. The experimental results and the evaluation by a subjective study have shown that the summarization based on MPEG-1 Layer 3 (MP3) music is comparable to the summarization based on uncompressed PCM music samples.
Xi Shao, Changsheng Xu, Ye Wang 0007, Mohan Kankanhalli
ICASSP (4)3
2004 Singing voice detection using twice-iterated composite Fourier transform
abstract
We propose a twice-iterated composite Fourier transform (TICFT) technique to detect the singing voice boundaries from acoustical polyphonic music signals. We show that the cumulative TICFT energy in the lower coefficients is capable of differentiating the harmonic structures of vocal and instrumental music in higher octaves. The musical signal is first segmented into frames based on quarter-notes. Then TICFT is used to measure the harmonic structure of each frame. Finally, the vocal and instrumental frames are classified by applying music domain knowledge. Experimental results show over 80% frame level accuracy can be achieved.
Namunu Chinthaka Maddage, Kong-Wah Wan, Changsheng Xu, Ye Wang 0007
ICME4
2004 Key determination of acoustic musical signals
abstract
The work presents a novel rule-based approach for determining the key of acoustic musical signals. Knowledge of the key enables to be derived, from music knowledge, the pitch class elements that a piece of music uses. Our technique is a combination of chroma based frequency analysis and music knowledge of rhythm structure and chord change patterns followed by rule-based inference. Experimental results illustrate that 90% accuracy is achieved for key determination and we have given an explanation for the cases that have generated an incorrect key. Steps for further development are also outlined.
Arun Shenoy, Roshni Mohapatra, Ye Wang 0007
ICME3
2004 Singing voice detection in popular music
abstract
We propose a novel technique for the automatic classification of vocal and non-vocal regions in an acoustic musical signal. Our technique uses a combination of harmonic content attenuation using higher level musical knowledge of key followed by sub-band energy processing to obtain features from the musical audio signal. We employ a Multi-Model Hidden Markov Model (MM-HMM) classifier for vocal and non-vocal classification that utilizes song structure information to create multiple models as opposed to conventional HMM training methods that employ only one model for each class. A statistical hypothesis testing approach followed by an automatic bootstrapping process is employed to further improve the accuracy of classification. An experimental evaluation on a database of 20 popular songs shows the validity of the proposed approach with an average classification accuracy of 86.7%.
Tin Lay Nwe, Arun Shenoy, Ye Wang 0007
ACM Multimedia3
2004 A framework for robust and scalable audio streaming
abstract
We propose a framework to achieve bandwidth efficient, error robust and bitrate scalable audio streaming. Our approach is compatible with most audio compression format. The main contributions of this paper include: 1) the proposal of a Multi-Stage Interleaving (MSI) strategy which translates packet loss into loss of separate frequency components that are less perceptually significant; and 2) the design of a Layered Unequal-Sized Packetization (LUSP) scheme which enables bitrate scalability and prioritized packet transmission. The combination of the proposed MSI and LUSP allows the use of a set of simple yet effective methods of error concealment in the compressed domain. Our approach offers significant advantages over existing methods in terms of memory consumption (a savings of over 40 times in the sample MP3 implementation), and computational complexity, which are critical issues for battery-powered small devices.
Ye Wang 0007, Wendong Huang, Jari Korhonen
ACM Multimedia1
2004 LyricAlly: automatic synchronization of acoustic musical signals and textual lyrics
abstract
We present a prototype that automatically aligns acoustic musical signals with their corresponding textual lyrics, in a manner similar to manually-aligned karaoke. We tackle this problem using a multimodal approach, where the appropriate pairing of audio and text processing helps create a more accurate system. Our audio processing technique uses a combination of top-down and bottom-up approaches, combining the strength of low-level audio features and high-level musical knowledge to determine the hierarchical rhythm structure, singing voice and chorus sections in the musical audio. Text processing is also employed to approximate the length of the sung passages using the textual lyrics. Results show an average error of less than one bar for per-line alignment of the lyrics on a test bed of 20 songs (sampled from CD audio and carefully selected for variety). We perform holistic and per-component testing and analysis and outline steps for further development.
Ye Wang 0007, Min-Yen Kan, Tin Lay Nwe, Arun Shenoy
ACM Multimedia1
2004 The creation of a music-driven digital violinist
abstract
This paper describes an initial attempt on a music-driven digital violinist (MDV) system, which automatically generates animation of a violinist based on violin music. MDV first analyzes the input audio signal and transcribes it into music notes. Next it uses the notes to synthesize the animated video of a violinist. Tests on the prototype system show that it achieves adequate visual realism and near real-time performance.
Ankur Dhanik, David Hsu, Ye Wang 0007
ACM Multimedia4
2003 Schemes for error resilient streaming of perceptually coded audio
abstract
This paper presents novel extensions to our earlier system for streaming perceptually coded audio over error prone channels such as Mobile IP. To improve error robustness while maintaining bandwidth efficiency, the new extensions combine the strength of an error resilient coding scheme in the sender, a prioritized packet transport scheme in the network and a compressed domain error concealment strategy in the terminal. Different concealment methods are used for each part of the coded audio data according to their perceptual importance and statistical characteristics. In our current implementation, we employed MPEG-2 Advanced Audio Coding (AAC) encoded bitstreams and an RTP/UDP-based test system for performance evaluation. Simulation results have shown that our improved streaming system is more robust against packet losses in comparison with conventional methods.
Jari Korhonen, Ye Wang 0007
ICASSP (5)2
2003 Parametric vector quantization for coding percussive sounds in music
abstract
This paper presents a novel parametric vector quantization (PVQ) scheme as the secondary encoding to code perceptually salient percussive sounds such as drums in the time domain. It is deployed to improve the quality of service in the case of packet losses during percussive events. As a generalization and an improvement of our earlier system, the new scheme can achieve a better balance between bandwidth efficiency and error robustness. The proposed coding technique has been implemented with the MPEG-2 AAC (Advanced Audio Coding) frame structure. Experimental results with music samples have shown the effectiveness of the proposed scheme.
Ye Wang 0007, Ali Ahmaniemi, Markus Vaalgamaa
ICASSP (5)1
2003 Schemes for error resilient streaming of perceptually coded audio
abstract
This paper presents novel extensions to our earlier system for streaming perceptually coded audio over error prone channels such as mobile IP. To improve error robustness while maintaining bandwidth efficiency, the new extensions combine the strength of an error resilient coding scheme in the sender, prioritized packet transport scheme in the network and a compressed domain error concealment strategy in the terminal. Different concealment methods are used for each part of the coded audio data according to their perceptual importance and statistical characteristics. In our current implementation, we employed MPEG-2 advanced audio coding (AAC) encoded bitstreams and an RTP/UDP-based test system for performance evaluation. Simulation results have shown that our improved streaming system is more robust against packet losses in comparison with conventional methods.
Jari Korhonen, Ye Wang 0007
ICME2
2003 Parametric vector quantization for coding percussive sounds in music
abstract
This paper presents a novel parametric vector quantization (PVQ) scheme as the secondary encoding to code perceptually salient percussive sounds such as drums in the time domain. It is deployed to improve the quality of service in the case of packet losses during percussive events. As a generalization and an improvement of our earlier system, the new scheme can achieve a better balance between bandwidth efficiency and error robustness. The proposed coding technique has been implemented with the MPEG-2 AAC (advanced audio coding) frame structure. Experimental results with music samples have shown the effectiveness of the proposed scheme.
Ye Wang 0007, Ali Ahmaniemi, Markus Vaalgamaa
ICME1
2003 Content-based UEP: a new scheme for packet loss recovery in music streaming
abstract
Bandwidth efficiency and error robustness are two essential and conflicting requirements for streaming media content over error-prone channels, such as wireless channels. This paper describes a new scheme called content-based unequal error protection (C-UEP), which aims to improve the user-perceived QoS in the case of packet loss. We use music streaming as an example to show the effectiveness of the new concept. C-UEP requires only a small fraction of the redundancy used in existing forward error correction (FEC) methods. C-UEP classifies every audio segment (e.g. an encoding frame) into different classes to improve encoding efficiency. Salient transients such as drumbeats and note onsets are encoded with more redundancy in a secondary bitstream used to recover lost packets by the receiver. Formal perceptual evaluations show that our scheme improves audio quality significantly over simple muting and packet repetition baselines. This improvement is achieved with a negligible amount of redundancy, which is transmitted to the receiver ahead of playback.
Ye Wang 0007, Ali Ahmaniemi, David Isherwood, Wendong Huang
ACM Multimedia1
2003 Application of a content-based percussive sound synthesizer to packet loss recovery in music streaming
abstract
This paper presents a novel method to recover lost packets in music streaming using a synthesizer to generate percussive sounds. As an improvement of the state-of-the-art system that uses a content-based audio codebook, the new method can greatly reduce the redundant information needed to recover perceptually critical lost packets.
Lonce L. Wyse, Ye Wang 0007, Xinglei Zhu
ACM Multimedia2
2002 A drumbeat-pattern based error concealment method for music streaming applications
abstract
This paper presents a novel drumbeat-pattern based error concealment scheme, which detects the drumbeat-pattern of music signals on the encoder side and embeds the beat information as ancillary data in a preceding data unit in the compressed bitstream. The embedded beat information is then used to perform an error concealment task on the decoder side. The proposed method was implemented using an MPEG-4 AAC (Advanced Audio Coding) codec. Informal evaluations have shown that the proposed active error concealment method clearly improved the overall subjective sound quality in comparison with conventional methods if the packet losses include the drumbeats.
Ye Wang 0007, Sebastian Streich
ICASSP1
2001 A Beat-Pattern based Error Concealment Scheme for Music Delivery with Burst Packet Loss
abstract
Error concealment is an important method to mitigate the degradation of the audio quality when compressed audio packets are lost in error prone channels, such as mobile Internet and digital audio broadcasting. This paper presents a novel error concealment scheme, which exploits the beat and rhythmic pattern of music signals. Preliminary simulations show significantly improved subjective sound quality in comparison with conventional methods in the case of burst packet losses. The new scheme is proposed as a complement to prior arts. It can be adopted to essentially all existing perceptual audio decoders such as an MP3 decoder for streaming music.
Ye Wang 0007
ICME1
2001 A compressed domain beat detector using MP3 audio bitstreams
abstract
This paper presents a novel beat detector that processes MPEG-1 Layer III (known as MP3) encoded audio bitstreams directly in the compressed domain. Most previous beat detection or tracking systems dealing with MIDI or PCM signals are not directly applicable to compressed audio bitstreams, such as MP3 bitstreams. We have developed the beat detector as a part of a beat-pattern based error concealment scheme for streaming music over error prone channels. Special effort was used to obtain a tailored trade-off between performance, complexity and memory consumption for this specific application. A comparison between the machine-detected results to the human annotation has shown that the proposed method correctly tracked beats in 4 out of 6 popular music test signals. The results were analyzed.
Ye Wang 0007, Miikka Vilermo
ACM Multimedia1
1999 An excitation level based psychoacoustic model for audio compression
abstract
This paper describes an excitation level based psychoacoustic model to estimate the simultaneous masking threshold for audio coding. The system has the following stages: 1) a windowing function; 2) a time-to-frequency transformation; 3) an excitation level calculation block similar to that in Moore and Glasberg's loudness model; 4) a correction factor for estimating masking threshold; 5) the inclusion of the absolute masking threshold; 6) the output Signal-to-Masking ratio. We have evaluated the performance by integrating the proposed psychoacoustic model into an audio coder similar to MPEG-2 AAC, which contains only the basic coding tools. Our model performs better than or as well as the psychoacoustic model suggested in the MPEG-2 AAC audio coding standard for all the test signals. We can achieve almost transparent quality with bitrate below 64 kbps for most of the critical test signals. Significant improvements have been achieved with speech signals, which are always difficult for transform audio coders.
Ye Wang 0007, Miikka Vilermo
ACM Multimedia (1)1