Zhiyao Duan

dblp:04/6716 · DBLP profile ↗
← Back
68ranked-venue papers
9as first author
30since 2021 · last 2025
0000-0002-8334-9974ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 6 first-author · 28 since 2021Artificial intelligence and machine learning · 25 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
abstract
Zero-shot online voice conversion (VC) holds significant promise for real-time communications and entertainment. However, current VC models struggle to preserve semantic fidelity under real-time constraints, deliver natural-sounding conversions, and adapt effectively to unseen speaker characteristics. To address these challenges, we introduce Conan, a chunkwise online zero-shot voice conversion model that preserves the content of the source while matching the speaker representation of reference speech. Conan comprises three core components: 1) a Stream Content Extractor that leverages Emformer for low-latency streaming content encoding; 2) an Adaptive Style Encoder that extracts fine-grained stylistic features from reference speech for enhanced style adaptation; 3) a Causal Shuffle Vocoder that implements a fully causal HiFiGAN using a pixel-shuffle mechanism. Experimental evaluations demonstrate that Conan outperforms baseline models in subjective and objective metrics. Audio samples can be found at https://aaronz345.github.io/ConanDemo.
Yu Zhang 0126, Baotong Tian, Zhiyao Duan
ASRU3
2025 Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges
abstract
In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and Acoustic Signal Processing Technical Commitee. In this paper, we reflect the main research achievements of MIR along the three EDICS related to music analysis, processing and generation. We then review a set of successful practices that fuel the rapid development of MIR research. One practice is the annual research benchmark, the Music Information Retrieval Evaluation eXchange, where participants compete on a set of research tasks. Another practice is the pursuit of reproducible and open research. The active engagement with industry research and products is another key factor for achieving large societal impacts and motivating younger generations of students to join the field. Last but not the least, the commitment to diversity, equity and inclusion ensures MIR to be a vibrant and open community where various ideas, methodologies, and career pathways collide. We finish by providing future challenges MIR will have to face.
Geoffroy Peeters, Zafar Rafii, Magdalena Fuentes, Zhiyao Duan, Emmanouil Benetos, Juhan Nam, Yuki Mitsufuji
ICASSP4
2025 Audio Visual Segmentation through Text Embeddings
abstract
The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt to overcome the challenge of limited data by leveraging the vision foundation model, Segment Anything Model (SAM), prompting it with audio to enhance its ability to segment sounding source objects. While this approach alleviates the model’s burden on understanding visual modality by utilizing knowledge of pre-trained SAM, it does not address the fundamental challenge of learning audio-visual correspondence with limited data. To address this limitation, we propose AV2T-SAM, a novel framework that bridges audio features with the text embedding space of pre-trained text-prompted SAM. Our method leverages multimodal correspondence learned from rich text-image paired datasets to enhance audio-visual alignment. Furthermore, we introduce a novel feature, fCLIP⊙fCLAP, which emphasizes shared semantics of audio and visual modalities while filtering irrelevant noise. Our approach outperforms existing methods on the AVSBench dataset by effectively utilizing pre-trained segmentation models and cross-modal semantic alignment. The source code is released at https://github.com/bok-bok/AV2T-SAM.
Kyungbok Lee, You Zhang 0001, Zhiyao Duan
ICIP3
2025 PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
You Zhang 0001, Baotong Tian, Lin Zhang 0054, Zhiyao Duan
INTERSPEECH4
2024 SingFake: Singing Voice Deepfake Detection
abstract
The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage. Unlike synthesized speech, synthesized singing voices are typically released in songs containing strong background music that may hide synthesis artifacts. Additionally, singing voices present different acoustic and linguistic characteristics from speech utterances. These unique properties make singing voice deepfake detection a relevant but significantly different problem from synthetic speech detection. In this work, we propose the singing voice deepfake detection task. We first present SingFake, the first curated in-the-wild dataset consisting of 28.93 hours of bonafide and 29.40 hours of deepfake song clips in five languages from 40 singers. We provide a train/validation/test split where the test sets include various scenarios. We then use SingFake to evaluate four state-of-the-art speech countermeasure systems trained on speech utterances. We find these systems lag significantly behind their performance on speech test data. When trained on SingFake, either using separated vocal tracks or song mixtures, these systems show substantial improvement. However, our evaluations also identify challenges associated with unseen singers, communication codecs, languages, and musical contexts, calling for dedicated research into singing voice deepfake detection. The SingFake dataset and related resources are available1.
Yongyi Zang, You Zhang 0001, Mojtaba Heydari, Zhiyao Duan
ICASSP4
2024 SynthTab: Leveraging Synthesized Data for Guitar Tablature Transcription
abstract
Guitar tablature is a form of music notation widely used among guitarists. It captures not only the musical content of a piece, but also its implementation and ornamentation on the instrument. Guitar Tablature Transcription (GTT) is an important task with broad applications in music education, composition, and entertainment. Existing GTT datasets are quite limited in size and scope, rendering models trained on them prone to overfitting and incapable of generalizing to out-of-domain data. In order to address this issue, we present a methodology for synthesizing large-scale GTT audio using commercial acoustic and electric guitar plugins. We procure SynthTab, a dataset derived from DadaGP, which is a vast and diverse collection of richly annotated symbolic tablature. The proposed synthesis pipeline produces audio which faithfully adheres to the original fingerings and a subset of techniques specified in the tablature, and covers multiple guitars and styles for each track. Experiments show that pre-training a baseline GTT model on SynthTab can improve transcription performance when fine-tuning and testing on an individual dataset. More importantly, cross-dataset experiments show that pre-training significantly mitigates issues with overfitting.
Yongyi Zang, Frank Cwitkowitz, Zhiyao Duan
ICASSP4
2024 Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech
abstract
Dimensional representations of speech emotions such as the arousal-valence (AV) representation provide a continuous and fine-grained description and control than their categorical counterparts. They have wide applications in tasks such as dynamic emotion understanding and expressive text-to-speech synthesis. Existing methods that predict the dimensional emotion representation from speech cast it as a supervised regression task. These methods face data scarcity issues, as dimensional annotations are much harder to acquire than categorical labels. In this work, we propose to learn the AV representation from categorical emotion labels of speech. We start by learning a rich and emotion-relevant high-dimensional speech feature representation using self-supervised pre-training and emotion classification fine-tuning. This representation is then mapped to the 2D AV space according to psychological findings through anchored dimensionality reduction. Experiments show that our method achieves a Concordance Correlation Coefficient (CCC) performance comparable to state-of-the-art supervised regression methods on IEMO-CAP without leveraging ground-truth AV annotations during training. This validates our proposed approach on AV prediction. Furthermore, visualization of AV predictions on MEAD and EmoDB datasets shows the interpretability of the learned AV representations.
Enting Zhou, You Zhang 0001, Zhiyao Duan
ICASSP3
2024 GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
Zehua Kcriss Li, Meiying Melissa Chen, Pinxin Liu, Zhiyao Duan
INTERSPEECH5
2024 CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
Yongyi Zang, Jiatong Shi, You Zhang 0001, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Wen-Xiao Zhao, Tomoki Toda, Zhiyao Duan
INTERSPEECH11
2024 A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
abstract
This paper addresses the challenge of developing a robust audio-visual deepfake detection model. In practical use cases, new generation algorithms are continually emerging, and these algorithms are not encountered during the development of detection methods. This calls for the generalization ability of the method. Additionally, to ensure the credibility of detection methods, it is beneficial for the model to interpret which cues from the video indicate it is fake. Motivated by these considerations, we then propose a multi-stream fusion approach with one-class learning as a representation-level regularization technique. We study the generalization problem of audio-visual deepfake detection by creating a new benchmark by extending and re-splitting the existing FakeAVCeleb dataset. The benchmark con-tains four categories of fake videos (Real Audio-Fake Visual, Fake Audio-Fake Visual, Fake Audio-Real Visual, and Unsynchronized videos). The experimental results demonstrate that our approach surpasses the previous models by a large margin. Furthermore, our proposed framework offers interpretability, indicating which modality the model identifies as more likely to be fake. The source code is released at https://github.com/bok-bok/MSOC.
Kyungbok Lee, You Zhang 0001, Zhiyao Duan
MMSP3
2024 SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
abstract
With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-generated singing voices from authentic singers. This challenge features two tracks: a controlled setting track (CtrSVDD) and an in-the-wild scenario track (WildSVDD). The CtrSVDD track utilizes publicly available singing vocal data to generate deepfakes using state-of-the-art singing voice synthesis and conversion systems. Meanwhile, the WildSVDD track expands upon the existing SingFake dataset, which includes data sourced from popular user-generated content websites. For the CtrSVDD track, we received submissions from 47 teams, with 37 surpassing our baselines and the top team achieving a 1.65% equal error rate. For the WildSVDD track, we benchmarked the baselines. This paper reviews these results, discusses key findings, and outlines future directions for SVDD research.
You Zhang 0001, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto, Tomoki Toda, Zhiyao Duan
SLT6
2024 MusicHiFi: Fast High-Fidelity Stereo Vocoding
abstract
Diffusion-based audio and music generation models commonly perform generation by constructing an image representation of audio (e.g., a mel-spectrogram) and then convert it to waveform using a phase reconstruction model or vocoder. Typical vocoders, however, produce monophonic audio at lower resolutions (e.g., 16-24 kHz), which limits their usefulness. We propose MusicHiFi—an efficient high-fidelity stereophonic vocoder. Our method employs a cascade of three generative adversarial networks (GANs) that convert low-resolution mel-spectrograms to audio, upsamples to high-resolution audio via bandwidth extension, and upmixes to stereophonic audio. Compared to past work, we propose 1) a unified GAN-based generator and discriminator architecture and training procedure for each stage of our cascade, 2) a new fast, near downsampling-compatible bandwidth extension module, and 3) a new fast downmix-compatible mono-to-stereo upmixer that ensures the preservation of monophonic content in the output. We evaluate our approach using objective and subjective listening tests and find our approach yields comparable or better audio quality, better spatialization control, and significantly faster inference speed compared to past work.
Juan Pablo Cáceres, Zhiyao Duan, Nicholas J. Bryan
IEEE Signal Process. Lett.3
2024 Cacophony: An Improved Contrastive Audio-Text Model
abstract
Despite recent advancements, audio-text models still lag behind their image-text counterparts in scale and performance. In this paper, we propose to improve both the data scale and the training procedure of audio-text contrastive models. Specifically, we craft a large-scale audio-text dataset containing 13,000 hours of text-labeled audio, using pretrained language models to process noisy text descriptions and automatic captioning to obtain text descriptions for unlabeled audio samples. We first train on audio-only data with a masked autoencoder (MAE) objective, which allows us to benefit from the scalability of unlabeled audio datasets. We then train a contrastive model with an auxiliary captioning objective with the audio encoder initialized from the MAE model. Our final model, which we name Cacophony, achieves state-of-the-art performance on audio-text retrieval tasks, and exhibits competitive results on the HEAR benchmark and other downstream tasks such as zero-shot classification.
Jordan Darefsky, Zhiyao Duan
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 SAMO: Speaker Attractor Multi-Center One-Class Learning For Voice Anti-Spoofing
abstract
Voice anti-spoofing systems are crucial auxiliaries for automatic speaker verification (ASV) systems. A major challenge is caused by unseen attacks empowered by advanced speech synthesis technologies. Our previous research on one-class learning has improved the generalization ability to unseen attacks by compacting the bona fide speech in the embedding space. However, such compactness lacks consideration of the diversity of speakers. In this work, we propose speaker attractor multi-center one-class learning (SAMO), which clusters bona fide speech around a number of speaker attractors and pushes away spoofing attacks from all the attractors in a high-dimensional embedding space. For training, we propose an algorithm for the co-optimization of bona fide speech clustering and bona fide/spoof classification. For inference, we propose strategies to enable anti-spoofing for speakers without enrollment. Our proposed system outperforms existing state-of-the-art single systems with a relative improvement of 38% on equal error rate (EER) on the ASVspoof2019 LA evaluation set.
Siwen Ding, You Zhang 0001, Zhiyao Duan
ICASSP3
2023 SingNet: a real-time Singing Voice beat and Downbeat Tracking System
abstract
Singing voice beat and downbeat tracking posses several applications in automatic music production, analysis and manipulation. Among them, some require real-time processing, such as live performance processing and auto-accompaniment for singing inputs. This task is challenging owing to the non-trivial rhythmic and harmonic patterns in singing signals. For real-time processing, it introduces further constraints such as inaccessibility to future data and the impossibility to correct the previous results that are inconsistent with the latter ones. In this paper, we introduce the first system that tracks the beats and downbeats of singing voices in real-time. Specifically, we propose a novel dynamic particle filtering approach that incorporates offline historical data to correct the online inference by using a variable number of particles. We evaluate the performance on two datasets: GTZAN with the separated vocal tracks, and an in-house dataset with the original vocal stems. Experimental result demonstrates that our proposed approach outperforms the baseline by 3–5%.
Mojtaba Heydari, Ju-Chiang Wang, Zhiyao Duan
ICASSP3
2023 HRTF Field: Unifying Measured HRTF Magnitude Representation with Neural Fields
abstract
Head-related transfer functions (HRTFs) are a set of functions describing the spatial filtering effect of the outer ear (i.e., torso, head, and pinnae) onto sound sources at different azimuth and elevation angles. They are widely used in spatial audio rendering. While the azimuth and elevation angles are intrinsically continuous, measured HRTFs in existing datasets employ different spatial sampling schemes, making it difficult to model HRTFs across datasets. In this work, we propose to use neural fields, a differentiable representation of functions through neural networks, to model HRTFs with arbitrary spatial sampling schemes. Such representation is unified across datasets with different spatial sampling schemes. HRTFs for arbitrary azimuth and elevation angles can be derived from this representation. We further introduce a generative model named HRTF field to learn the latent space of the HRTF neural fields across subjects. We demonstrate promising performance on HRTF interpolation and generation tasks and point out potential future work.
You Zhang 0001, Zhiyao Duan
ICASSP3
2023 Transcription Free Filler Word Detection with Neural Semi-CRFs
abstract
Non-linguistic filler words, such as "uh" or "um", are prevalent in spontaneous speech and serve as indicators for expressing hesitation or uncertainty. Previous works for detecting certain non-linguistic filler words are highly dependent on transcriptions from a well-established commercial automatic speech recognition (ASR) system. However, certain ASR systems are not universally accessible from many aspects, e.g., budget, target languages, and computational power. In this work, we investigate filler word detection system1that does not depend on ASR systems. We show that, by using the structured state space sequence model (S4) and neural semi-Markov conditional random fields (semi-CRFs), we achieve an absolute F1 improvement of 6.4% (segment level) and 3.1% (event level) on the PodcastFillers dataset. We also conduct a qualitative analysis on the detected results to analyze the limitations of our proposed system.
Yujia Yan, Juan Pablo Cáceres, Zhiyao Duan
ICASSP4
2023 ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
abstract
Recent advancements in neural speech synthesis have renewed interest in voice conversion (VC) to go beyond timbre transfer.Achieving controllability of para-linguistic parameters like pitch and speed is crucial in various applications.However, existing studies either lack interpretability or only provide global control at the utterance level.This paper introduces ControlVC, the first neural voice conversion system to enable time-varying controls on pitch and speed.ControlVC uses pre-trained encoders to generate pitch and linguistic embeddings, combined and converted to speech using a vocoder.Speed control is achieved by TD-PSOLA pre-processing, while pitch control is achieved by manipulating the pitch contour before feeding it into the encoder.Systematic subjective and objective evaluations show that this work significantly outperforms selfconstructed baselines on speech quality and controllability for non-parallel zero-shot conversion while achieving time-varying control 1 .
Meiying Chen, Zhiyao Duan
INTERSPEECH2
2023 Phase perturbation improves channel robustness for speech spoofing countermeasures
abstract
In this paper, we aim to address the problem of channel robustness in speech countermeasure (CM) systems, which are used to distinguish synthetic speech from human natural speech.On the basis of two hypotheses, we suggest an approach for perturbing phase information during the training of time-domain CM systems.Communication networks often employ lossy compression codec that encodes only magnitude information, therefore heavily altering phase information.Also, state-of-the-art CM systems rely on phase information to identify spoofed speech.Thus, we believe the information loss in the phase domain induced by lossy compression codec degrades the performance of the unseen channel.We first establish the dependence of timedomain CM systems on phase information by perturbing phase in evaluation, showing strong degradation.Then, we demonstrated that perturbing phase during training leads to a significant performance improvement, whereas perturbing magnitude leads to further degradation.
Yongyi Zang, You Zhang 0001, Zhiyao Duan
INTERSPEECH3
2022 A Novel 1D State Space for Efficient Music Rhythmic Analysis
abstract
Inferring music time structures has a broad range of applications in music production, processing and analysis. Scholars have proposed various methods to analyze different aspects of time structures, such as beat, downbeat, tempo and meter. Many state-of-the-art (SOFA) methods, however, are computationally expensive. This makes them inapplicable in real-world industrial settings where the scale of the music collections can be millions. This paper proposes a new state space and a semi-Markov model for music time structure analysis. The proposed approach turns the commonly used 2D state spaces into a 1D model through a jump-back reward strategy. It reduces the state spaces size drastically. We then utilize the proposed method for causal, joint beat, downbeat, tempo, and meter tracking, and compare it against several previous methods. The proposed method delivers similar performance with the SOFA joint causal models with a much smaller state space and a more than 30 times speedup.
Mojtaba Heydari, Matthew C. McCallum, Andreas F. Ehmann, Zhiyao Duan
ICASSP4
2022 Progressive Teacher-Student Training Framework for Music Tagging
abstract
Music tagging is the task of predicting multiple tags of a music excerpt, and plays an important role in modern music recommendation systems. To obtain superior performance, recent approaches of music tagging focus on developing sophisticated models or exploiting additional multi-modal information. However, none of them deal with the problem of label noise during the training process despite of its ubiquitous presence. In this paper, we propose a progressive two-stage teacher-student training framework to prevent the music tagging model from overfitting label noise. Experimental results suggest that the proposed method surpasses conventional label-noise-robust methods and exhibits scalability across different tagging models. Moreover, detailed analyses demonstrate that the two teachers in the framework gradually improve student model’s generalization performance and effectively avoid the impairment from label noise.
Rui Lu 0003, Baigong Zheng, Jiarui Hai, Zhiyao Duan, Ji Liu 0002
ICASSP5
2022 A Study of The Robustness of Raw Waveform Based Speaker Embeddings Under Mismatched Conditions
abstract
In this paper, we conduct a cross-dataset study on parametric and non-parametric raw-waveform based speaker embeddings through speaker verification experiments. In general, we observe a more significant performance degradation of these raw-waveform systems compared to spectral based systems. We then propose two strategies to improve the performance of raw-waveform based systems on cross-dataset tests. The first strategy is to change the real-valued filters into analytic filters to ensure shift-invariance. The second strategy is to apply variational dropout to non-parametric filters to prevent them from overfitting irrelevant nuance features. By combining these strategies, we achieve results comparable to spectral based systems on both the VoxCeleb and VOiCEs datasets. Futhermore, we demonstrate that the learned filters carry little noise compared to existing non-parametric learnable front-ends.
Frank Cwitkowitz, Zhiyao Duan
ICASSP3
2022 DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization
abstract
Speaker diarization aims to determine “who spoke when” in multi-speaker scenarios. Audio-visual speaker diarization leverages visual information in addition to audio signals and has shown improved performance. Existing audio-visual methods extract speaker embeddings for each video clip using audio and facial features, and then perform clustering according to their similarity. However, this approach would not work well for noisy or overlapped speech where audio features are corrupted, nor for off-screen speakers where visual features are missing. In this work, we propose dynamic vision-guided speaker embedding (DyViSE), a novel method for leveraging visual information to extract speaker embeddings in a multi-stage system. DyViSE uses dynamic lip movement information to denoise audio in a latent space and integrates facial features to obtain an identity-discriminative embedding for each speaking segment. DyViSE is trained with a deep clustering loss along with an exemplary loss. DyViSE demonstrates remarkable performance on both real-world videos and artificially assembled videos. Our code is available at https://github.com/urkax/DyViSE.
Abudukelimu Wuerkaixi, Kunda Yan, You Zhang 0001, Zhiyao Duan, Changshui Zhang
MMSP4
2022 Music Source Separation With Generative Flow
abstract
Fully-supervised models for source separation are trained on parallel mixture-source data and are currently state-of-the-art. However, such parallel data is often difficult to obtain, and it is cumbersome to adapt trained models to mixtures with new sources. Source-only supervised models, in contrast, only require individual source data for training. In this paper, we first leverage flow-based generators to train individual music source priors and then use these models, along with likelihood-based objectives, to separate music mixtures. We show that in singing voice separation and music separation tasks, our proposed method is competitive with a fully-supervised approach. We also demonstrate that we can flexibly add new types of sources, whereas fully-supervised approaches would require retraining of the entire model.
Jordan Darefsky, Fei Jiang 0007, Anton Selitskiy, Zhiyao Duan
IEEE Signal Process. Lett.5
2022 Speech Driven Talking Face Generation From a Single Image and an Emotion Condition
abstract
Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an end-to-end talking face generation system that takes a speech utterance, a single face image, and a categorical emotion label as input to render a talking face video synchronized with the speech and expressing the conditioned emotion. Objective evaluation on image quality, audiovisual synchronization, and visual emotion expression shows that the proposed system outperforms a state-of-the-art baseline system. Subjective evaluation of visual emotion expression and video realness also demonstrates the superiority of the proposed system. Furthermore, we conduct a human emotion recognition pilot study using generated videos with mismatched emotions among the audio and visual modalities. Results show that humans respond to the visual modality more significantly than the audio modality on this task.
Sefik Emre Eskimez, You Zhang 0001, Zhiyao Duan
IEEE Trans. Multim.3
2021 Don't Look Back: An Online Beat Tracking Method Using RNN and Enhanced Particle Filtering
abstract
Online beat tracking (OBT) has always been a challenging task. Due to the inaccessibility of future data and the need to make inference in real-time. We propose Don’t Look back! (DLB), a novel approach optimized for efficiency when performing OBT. DLB feeds the activations of a unidirectional RNN into an enhanced Monte-Carlo localization model to infer beat positions. Most preexisting OBT methods either apply some offline approaches to a moving window containing past data to make predictions about future beat positions or must be primed with past data at startup to initialize. Meanwhile, our proposed method only uses activation of the current time frame to infer beat positions. As such, without waiting at the beginning to receive a chunk, it provides an immediate beat tracking response, which is critical for many OBT applications. DLB significantly improves beat tracking accuracy over state-of-the-art OBT methods, yielding a similar performance to offline methods.
Mojtaba Heydari, Zhiyao Duan
ICASSP2
2021 An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
abstract
Spoofing countermeasure (CM) systems are critical in speaker verification; they aim to discern spoofing attacks from bona fide speech trials.In practice, however, acoustic condition variability in speech utterances may significantly degrade the performance of CM systems.In this paper, we conduct a cross-dataset study on several state-of-the-art CM systems and observe significant performance degradation compared with their singledataset performance.Observing differences of average magnitude spectra of bona fide utterances across the datasets, we hypothesize that channel mismatch among these datasets is one important reason.We then verify it by demonstrating a similar degradation of CM systems trained on original but evaluated on channel-shifted data.Finally, we propose several channel robust strategies (data augmentation, multi-task learning, adversarial learning) for CM systems, and observe a significant performance improvement on cross-dataset experiments.
You Zhang 0001, Fei Jiang 0007, Zhiyao Duan
Interspeech4
2021 Y-Vector: Multiscale Waveform Encoder for Speaker Embedding
abstract
State-of-the-art text-independent speaker verification systems typically use cepstral features or filter bank energies as speech features.Recent studies attempted to extract speaker embeddings directly from raw waveforms and have shown competitive results.In this paper, we propose a novel multi-scale waveform encoder that uses three convolution branches with different time scales to compute speech features from the waveform.These features are then processed by squeeze-and-excitation blocks, a multi-level feature aggregator, and a time delayed neural network (TDNN) to compute speaker embedding.We show that the proposed embeddings outperforms existing raw-waveformbased speaker embeddings on speaker verification by a large margin.A further analysis of the learned filters shows that the multi-scale encoder attends to different frequency bands at its different scales while resulting in a more flat overall frequency response than any of the single-scale counterparts.
Fei Jiang 0007, Zhiyao Duan
Interspeech3
2021 Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFs
abstract
Piano transcription systems are typically optimized to estimate pitch activity at each frame of audio. They are often followed by carefully designed heuristics and post-processing algorithms to estimate note events from the frame-level predictions. Recent methods have also framed piano transcription as a multi-task learning problem, where the activation of different stages of a note event are estimated independently. These practices are not well aligned with the desired outcome of the task, which is the specification of note intervals as holistic events, rather than the aggregation of disjoint observations. In this work, we propose a novel formulation of piano transcription, which is optimized to directly predict note events. Our method is based on Semi-Markov Conditional Random Fields (semi-CRF), which produce scores for intervals rather than individual frames. When formulating piano transcription in this way, we eliminate the need to rely on disjoint frame-level estimates for different stages of a note event. We conduct experiments on the MAESTRO dataset and demonstrate that the proposed model surpasses the current state-of-the-art for piano transcription. Our results suggest that the semi-CRF output layer, while still quadratic in complexity, is a simple, fast and well-performing solution for event-based prediction, and may lead to similar success in other areas which currently rely on frame-level estimates.
Yujia Yan, Frank Cwitkowitz, Zhiyao Duan
NeurIPS3
2021 One-Class Learning Towards Synthetic Voice Spoofing Detection
abstract
Human voices can be used to authenticate the identity of the speaker, but the automatic speaker verification (ASV) systems are vulnerable to voice spoofing attacks, such as impersonation, replay, text-to-speech, and voice conversion. Recently, researchers developed anti-spoofing techniques to improve the reliability of ASV systems against spoofing attacks. However, most methods encounter difficulties in detecting unknown attacks in practical use, which often have different statistical distributions from known attacks. Especially, the fast development of synthetic voice spoofing algorithms is generating increasingly powerful attacks, putting the ASV systems at risk of unseen attacks. In this work, we propose an anti-spoofing system to detect unknown synthetic voice spoofing attacks (i.e., text-to-speech or voice conversion) using one-class learning. The key idea is to compact the bona fide speech representation and inject an angular margin to separate the spoofing attacks in the embedding space. Without resorting to any data augmentation methods, our proposed system achieves an equal error rate (EER) of 2.19% on the evaluation set of ASVspoof 2019 Challenge logical access scenario, outperforming all existing single systems (i.e., those without model ensemble).
You Zhang 0001, Fei Jiang 0007, Zhiyao Duan
IEEE Signal Process. Lett.3
2020 RL-Duet: Online Music Accompaniment Generation Using Deep Reinforcement Learning
abstract
This paper presents a deep reinforcement learning algorithm for online accompaniment generation, with potential for real-time interactive human-machine duet improvisation. Different from offline music generation and harmonization, online music accompaniment requires the algorithm to respond to human input and generate the machine counterpart in a sequential order. We cast this as a reinforcement learning problem, where the generation agent learns a policy to generate a musical note (action) based on previously generated context (state). The key of this algorithm is the well-functioning reward model. Instead of defining it using music composition rules, we learn this model from monophonic and polyphonic training data. This model considers the compatibility of the machine-generated note with both the machine-generated context and the human-generated context. Experiments show that this algorithm is able to respond to the human part and generate a melodic, harmonic and diverse machine part. Subjective evaluations on preferences show that the proposed algorithm generates music pieces of higher quality than the baseline method.
Nan Jiang 0023, Sheng Jin 0007, Zhiyao Duan, Changshui Zhang
AAAI3
2020 Vroom!: A Search Engine for Sounds by Vocal Imitation Queries
abstract
Traditional search through collections of audio recordings compares a text-based query to text metadata associated with each audio file and does not address the actual content of the audio. Text descriptions do not describe all aspects of the audio content in detail. Query by vocal imitation (QBV) is a kind of query by example that lets users imitate the content of the audio they seek, providing an alternative search method to traditional text search. Prior work proposed several neural networks, such as TL-IMINET, for QBV, however, previous systems have not been deployed in an actual search engine nor evaluated by real users. We have developed a state-of-the-art QBV system (Vroom!) and a baseline query-by-text search engine (TextSearch). We deployed both systems in an experimental framework to perform user experiments with Amazon Mechanical Turk (AMT) workers. Results showed that Vroom! received significantly higher search satisfaction ratings than TextSearch did for sound categories that were difficult for subjects to describe by text. Results also showed a better overall ease-of-use rating for Vroom! than TextSearch on the sound library used in our experiments. These findings suggest that QBV, as a complimentary search approach to existing text-based search, can improve both search results and user experience.
Yichi Zhang 0008, Junbo Hu, Bryan Pardo, Zhiyao Duan
CHIIR5
2020 End-To-End Generation of Talking Faces from Noisy Speech
abstract
Acoustic cues are not the only component in speech communication; if the visual counterpart is present, it is shown to benefit speech comprehension. In this work, we propose an end-to-end (no pre- or post-processing) system that can generate talking faces from arbitrarily long noisy speech. We propose a mouth region mask to encourage the network to focus on mouth movements rather than speech irrelevant movements. In addition, we use generative adversarial network (GAN) training to improve the image quality and mouth-speech synchronization. Furthermore, we employ noise-resilient training to make our network robust to unseen non-stationary noise. We evaluate our system with image quality and mouth shape (landmark) measures on noisy speech utterances with five types of unseen non-stationary noise between -10 dB and 30 dB signal-to-noise ratio (SNR) with increments of 1 dB SNR. Results show that our system outperforms a state-of-the-art baseline system significantly, and our noise-resilient training improves performance for noisy speech in a wide range of SNR.
Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, Zhiyao Duan
ICASSP4
2020 When Counterpoint Meets Chinese Folk Melodies
abstract
Counterpoint is an important concept in Western music theory. In the past century, there have been significant interests in incorporating counterpoint into Chinese folk music composition. In this paper, we propose a reinforcement learning-based system, named FolkDuet, towards the online countermelody generation for Chinese folk melodies. With no existing data of Chinese folk duets, FolkDuet employs two reward models based on out-of-domain data, i.e. Bach chorales, and monophonic Chinese folk melodies. An interaction reward model is trained on the duets formed from outer parts of Bach chorales to model counterpoint interaction, while a style reward model is trained on monophonic melodies of Chinese folk songs to model melodic patterns. With both rewards, the generator of FolkDuet is trained to generate countermelodies while maintaining the Chinese folk style. The entire generation process is performed in an online fashion, allowing real-time interactive human-machine duet improvisation. Experiments show that the proposed algorithm achieves better subjective and objective results than the baselines.
Nan Jiang 0023, Sheng Jin 0007, Zhiyao Duan, Changshui Zhang
NeurIPS3
2020 Speaker Attractor Network: Generalizing Speech Separation to Unseen Numbers of Sources
abstract
Most existing speech separation research focuses on improving the separation performance under consistent source number conditions between training and testing. In real-world applications, however, the source number may be different from that in training sets. In this letter, we address this problem by thoroughly improving the deep attractor network in terms of the network architecture and learning objectives so that it can well generalize to separating an unseen number of sources. Experimental results show that, compared with existing models, the proposed method significantly improves the separation performance when generalizing to an unseen number of speakers, and can separate up to five speakers even the model is only trained on two-speaker mixtures.
Fei Jiang 0007, Zhiyao Duan
IEEE Signal Process. Lett.2
2020 Noise-Resilient Training Method for Face Landmark Generation From Speech
abstract
Visual cues such as lip movements, when available, play an important role in speech communication. They are especially helpful for the hearing impaired population or in noisy environments. When not available, having a system to automatically generate talking faces in sync with input speech would enhance speech communication and enable many novel applications. In this article, we present a new system that can generate 3D talking face landmarks from speech in an online fashion. We employ a neural network that accepts the raw waveform as an input. The network contains convolutional layers with 1D kernels and outputs the active shape model (ASM) coefficients of face landmarks. To promote smoother transitions between video frames, we present a variant of the model that has the same architecture but also accepts the previous frame's ASM coefficients as an additional input. To cope with background noise, we propose a new training method to incorporate speech enhancement ideas at the feature level. Objective evaluations on landmark prediction show that the proposed system yields statistically significantly smaller errors than two state-of-the-art baseline methods on both a single-speaker dataset and a multi-speaker dataset. Experiments on noisy speech input with five types of non-stationary unseen noise show statistically significant improvements of the system performance thanks to the noise-resilient training method. Finally, subjective evaluations show that the generated talking faces have a significantly more convincing match with the input audio, achieving a similarly convincing level of realism as the ground-truth landmarks.
Sefik Emre Eskimez, Ross K. Maddox, Chenliang Xu, Zhiyao Duan
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Hierarchical Cross-Modal Talking Face Generation With Dynamic Pixel-Wise Loss
abstract
We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we propose first to transfer audio to high-level structure, i.e., the facial landmarks, and then to generate video frames conditioned on the landmarks. Compared to a direct audio-to-image approach, our cascade approach avoids fitting spurious correlations between audiovisual signals that are irrelevant to the speech content. We, humans, are sensitive to temporal discontinuities and subtle artifacts in video. To avoid those pixel jittering problems and to enforce the network to focus on audiovisual-correlated regions, we propose a novel dynamically adjustable pixel-wise loss with an attention mechanism. Furthermore, to generate a sharper image with well-synchronized facial movements, we propose a novel regression-based discriminator structure, which considers sequence-level information along with frame-level information. Thoughtful experiments on several datasets and real-world samples demonstrate significantly better results obtained by our method than the state-of-the-art methods in both quantitative and qualitative comparisons.
Ross K. Maddox, Zhiyao Duan, Chenliang Xu
CVPR3
2019 Audio-Visual Deep Clustering for Speech Separation
abstract
Speech separation aims to separate individual voices from an audio mixture of multiple simultaneous talkers. Audio-only approaches show unsatisfactory performance when the speakers are of the same gender or share similar voice characteristics. This is due to challenges on learning appropriate feature representations for separating voices in single frames and streaming voices across time. Visual signals of speech (e.g., lip movements), if available, can be leveraged to learn better feature representations for separation. In this paper, we propose a novel audio-visual deep clustering model (AVDC) to integrate visual information into the process of learning better feature representations (embeddings) for Time-Frequency (T-F) bin clustering. It employs a two-stage audio-visual fusion strategy where speaker-wise audio-visual T-F embeddings are first computed after the first-stage fusion to model the audio-visual correspondence for each speaker. In the second-stage fusion, audio-visual embeddings of all speakers and audio embeddings calculated by deep clustering from the audio mixture are concatenated to form the final T-F embedding for clustering. Through a series of experiments, the proposed AVDC model is shown to outperform the audio-only deep clustering and utterance-level permutation invariant training baselines and three other state-of-the-art audio-visual approaches. Further analyses show that the AVDC model learns a better T-F embedding for alleviating the source permutation problem across frames. Other experiments show that the AVDC model is able to generalize across different numbers of speakers between training and testing and shows some robustness when visual information is partially missing.
Rui Lu 0003, Zhiyao Duan, Changshui Zhang
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Siamese Style Convolutional Neural Networks for Sound Search by Vocal Imitation
abstract
Conventional methods for finding audio in databases typically search text labels, rather than the audio itself. This can be problematic as labels may be missing, irrelevant to the audio content, or not known by users. Query by vocal imitation lets users query using vocal imitations instead. To do so, appropriate audio feature representations and effective similarity measures of imitations and original sounds must be developed. In this paper, we build upon our preliminary work to propose Siamese style convolutional neural networks to learn feature representations and similarity measures in a unified end-to-end training framework. Our Siamese architecture uses two convolutional neural networks to extract features, one from vocal imitations and the other from original sounds. The encoded features are then concatenated and fed into a fully connected network to estimate their similarity. We propose two versions of the system: IMINET is symmetric where the two encoders have an identical structure and are trained from scratch, while TL-IMINET is asymmetric and adopts the transfer learning idea by pretraining the two encoders from other relevant tasks: spoken language recognition for the imitation encoder and environmental sound classification for the original sound encoder. Experimental results show that both versions of the proposed system outperform a state-of-the-art system for sound search by vocal imitation, and the performance can be further improved when they are fused with the state of the art system. Results also show that transfer learning significantly improves the retrieval performance. This paper also provides insights to the proposed networks by visualizing and sonifying input patterns that maximize the activation of certain neurons in different layers.
Yichi Zhang 0008, Bryan Pardo, Zhiyao Duan
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Creating a Multitrack Classical Music Performance Dataset for Multimodal Music Analysis: Challenges, Insights, and Applications
abstract
We introduce a dataset for facilitating audio-visual analysis of music performances. The dataset comprises 44 simple multi-instrument classical music pieces assembled from coordinated but separately recorded performances of individual tracks. For each piece, we provide the musical score in MIDI format, the audio recordings of the individual tracks, the audio and video recording of the assembled mixture, and ground-truth annotation files including frame-level and note-level transcriptions. We describe our methodology for the creation of the dataset, particularly highlighting our approaches to address the challenges involved in maintaining synchronization and expressiveness. We demonstrate the high quality of synchronization achieved with our proposed approach by comparing the dataset with existing widely used music audio datasets. We anticipate that the dataset will be useful for the development and evaluation of existing music information retrieval (MIR) tasks, as well as for novel multimodal tasks. We benchmark two existing MIR tasks (multipitch analysis and score-informed source separation) on the dataset and compare them with other existing music audio datasets. In addition, we consider two novel multimodal MIR tasks (visually informed multipitch analysis and polyphonic vibrato analysis) enabled by the dataset and provide evaluation measurements and baseline systems for future comparisons (from our recent work). Finally, we propose several emerging research directions that the dataset enables.
Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, Gaurav Sharma 0001
IEEE Trans. Multim.4
2018 Lip Movements Generation at a Glance
Zhiheng Li 0002, Ross K. Maddox, Zhiyao Duan, Chenliang Xu
ECCV (7)4
2018 Audio-Visual Event Localization in Unconstrained Videos
Yapeng Tian, Jing Shi 0005, Bochen Li, Zhiyao Duan, Chenliang Xu
ECCV (2)4
2018 Unsupervised Learning Approach to Feature Analysis for Automatic Speech Emotion Recognition
abstract
The scarcity of emotional speech data is a bottleneck of developing automatic speech emotion recognition (ASER) systems. One way to alleviate this issue is to use unsupervised feature learning techniques to learn features from the widely available general speech and use these features to train emotion classifiers. These unsupervised methods, such as denoising autoencoder (DAE), variational autoencoder (VAE), adversarial autoencoder (AAE) and adversarial variational Bayes (AVB), can capture the intrinsic structure of the data distribution in the learned feature representation. In this work, we systematically investigate four kinds of unsupervised feature learning methods for improving speaker-independent ASER. We show that all methods improve the performance regarding unweighted accuracy rating (UAR) and Fl-score over methods that use hand-crafted features or that do not perform feature learning on external datasets. We also show that VAE, AAE and AVB methods, which control the distribution of the latent representation, outperform DAE that does not control such distribution. This suggests the benefits of using variational inference methods to learn features from general speech for the speech tasks such as ASER that has very limited labeled data.
Sefik Emre Eskimez, Zhiyao Duan, Wendi B. Heinzelman
ICASSP2
2018 Multi-Scale Recurrent Neural Network for Sound Event Detection
abstract
Sound event detection (SED) in real life is an interesting but challenging task due to the polyphonic and long-term dependent nature of sound events. Recently, multi-label recurrent neural networks (RNNs) have shown promises. However, even equipped with long short-term memory (LSTM) or gated recurrent unit (GRU) cells, RNNs are still limited to model the long-term dependency. In this paper, we propose a multiscale RNN to address this issue. By integrating information from different time resolutions, we can better capture both the fine-grained and long-term dependencies of sound events. We experiment on the development sets of Task3 of DCASE2016 and DCASE2017. Compared to our previously proposed single-scale RNN that won the third place among the 13 teams in Task3 of DCASE2017, the proposed multiscale model achieves statistically significantly better performance on the development datasets of both DECASE2016 and DCASE2017.
Rui Lu 0003, Zhiyao Duan, Changshui Zhang
ICASSP2
2018 Score-Aligned Polyphonic Microtiming Estimation
abstract
Accurate estimation of note onset timing is important for music ensemble performance analysis and synthesis. In this study, we present a method for the detection of onsets from polyphonic mixtures, using score information. First, a MIDI score is aligned to the audio signal using dynamic time warping, and pitches of performed notes are refined using a multi-pitch estimation technique. Notes in a signal are then isolated using a spectral masking method, based on the average harmonic structure learned from each source. Onset timing is finally estimated by maximizing the time derivative of the energy curve of the note within an observation window. We show that this method significantly improves the onset timing estimation accuracy, measured by both the align rate and onset time deviation, and outperforms a state-of-art reference method.
Ryan Stables, Bochen Li, Zhiyao Duan
ICASSP4
2018 Visualization and Interpretation of Siamese Style Convolutional Neural Networks for Sound Search by Vocal Imitation
abstract
Designing systems that allow users to search sounds through vocal imitation augments the current text-based search engines and advances human-computer interaction. Previously we proposed a Siamese style convolutional network called IMINET for sound search by vocal imitation, which jointly addresses feature extraction by Convolutional Neural Network (CNN) and similarity calculation by Fully Connected Network (FCN), and is currently the state of the art. However, how such architecture works is still a mystery. In this paper, we try to answer this question. First, we visualize the input patterns that maximize the activation of different neurons in each CNN tower; this helps us understand what features are extracted from vocal imitations and sound candidates. Second, we visualize the imitation-sound input pairs that maximize the activation of different neurons in the FCN layers; this helps us understand what kind of input pattern pairs are recognized during the similarity calculation. Interesting patterns are found to reveal the local-to-global and simple-to-conceptual learning mechanism of TL-IMINET. Experiments also show how transfer learning helps to improve TL-IMINET performance from the visualization aspect.
Yichi Zhang 0008, Zhiyao Duan
ICASSP2
2018 Joint Speaker Diarization and Recognition Using Convolutional and Recurrent Neural Networks
abstract
Speaker diarization (detecting who-spoke-when using relative identity labels) and speaker recognition (detecting absolute identity labels without timing) are different but related tasks that often need to be completed simultaneously in many scenarios. Traditional methods, however, address them independently. In this paper, we propose a method to jointly diarize and recognize speakers from a collection of conversations. This method benefits from the sparsity and temporal smoothness of speakers within a conversation and the large-scale timbre modeling across recordings and speakers. Specifically, we employ one convolutional neural network (CNN) to perform segment-level speaker classification and another CNN to detect the probability of speaker change within a conversation. We then concatenate the output of both CNNs and feed it into a recurrent neural network (RNN) for joint speaker diarization and recognition. Experiments on different datasets show promising performance of our proposed approach.
Zhihan Zhou 0004, Yichi Zhang 0008, Zhiyao Duan
ICASSP3
2018 Front-end speech enhancement for commercial speaker verification systems
Sefik Emre Eskimez, Peter Soufleris, Zhiyao Duan, Wendi B. Heinzelman
Speech Commun.3
2018 Listen and Look: Audio-Visual Matching Assisted Speech Source Separation
abstract
Source permutation, i.e., assigning separated signal snippets to wrong sources over time, is a major issue in the state-of-the-art speaker-independent speech source separation methods. In addition to auditory cues, humans also leverage visual cues to solve this problem at cocktail parties: matching lip movements with voice fluctuations helps humans to better pay attention to the speaker of interest. In this letter, we propose an audio-visual matching network to learn the correspondence between voice fluctuations and lip movements. We then propose a framework to apply this network to address the source permutation problem and improve over audio-only speech separation methods. The modular design of this framework makes it easy to apply the matching network to any audio-only speech separation method. Experiments on two-talker mixtures show that the proposed approach significantly improves the separation quality over the state-of-the-art audio-only method. This improvement is especially pronounced on mixtures that the audio-only method fails, in which the speakers often have similar voice characteristics.
Rui Lu 0003, Zhiyao Duan, Changshui Zhang
IEEE Signal Process. Lett.2
2017 Visually informed multi-pitch analysis of string ensembles
abstract
Multi-pitch analysis of polyphonic music requires estimating concurrent pitches (estimation) and organizing them into temporal streams according to their sound sources (streaming). This is challenging for approaches based on audio alone due to the polyphonic nature of the audio signals. Video of the performance, when available, can be useful to alleviate some of the difficulties. In this paper, we propose to detect the play/non-play (P/NP) activities from musical performance videos using optical flow analysis to help with audio-based multi-pitch analysis. Specifically, the detected P/NP activity provides a more accurate estimate of the instantaneous polyphony (i.e., the number of pitches at a time instant), and also helps with assigning pitch estimates to only active sound sources. As the first attempt towards audio-visual multi-pitch analysis of multi-instrument musical performances, we demonstrate the concept on 11 string ensembles. Experiments show a high overall P/NP detection accuracy of 85.3%, and a statistically significant improvement on both the multi-pitch estimation and streaming accuracy, under paired t-tests at a significance level of 0:01 in most cases.
Karthik Dinesh, Bochen Li, Xinzhao Liu, Zhiyao Duan, Gaurav Sharma 0001
ICASSP4
2017 See and listen: Score-informed association of sound tracks to players in chamber music performance videos
abstract
Both audio and visual aspects of a musical performance, especially their association, are important for expressing players' ideas and for engaging the audience. In this paper, we present a framework for combining audio and video analyses of multi-instrument chamber music performances to associate players in the video to the individual separated instrument sources from the audio, in a score-informed fashion. The instrument sources are first separated using a score-informed source separation techniques. The individual sources are then associated with different players in the video by correlating the onset instants of their aligned score tracks with the players' motion detected using optical flow. Experiments on 19 musical pieces with varying polyphony show that the proposed method obtains the correct association for 17 pieces, and an accuracy of 89.2% of the association of all individual tracks. The approach enables novel music enjoyment experiences by allowing users to target an audio source by clicking on the player in the video to separate/enhance it.
Bochen Li, Karthik Dinesh, Zhiyao Duan, Gaurav Sharma 0001
ICASSP3
2017 Deep ranking: Triplet MatchNet for music metric learning
abstract
Metric learning for music is an important problem for many music information retrieval (MIR) applications such as music generation, analysis, retrieval, classification and recommendation. Traditional music metrics are mostly defined on linear transformations of handcrafted audio features, and may be improper in many situations given the large variety of music styles and instrumentations. In this paper, we propose a deep neural network named Triplet MatchNet to learn metrics directly from raw audio signals of triplets of music excerpts with human-annotated relative similarity in a supervised fashion. It has the advantage of learning highly nonlinear feature representations and metrics in this end-to-end architecture. Experiments on a widely used music similarity measure dataset show that our method significantly outperforms three state-of-the-art music metric learning methods. Experiments also show that the learned features better preserve the partial orders of the relative similarity than handcrafted features.
Rui Lu 0003, Kailun Wu, Zhiyao Duan, Changshui Zhang
ICASSP3
2017 Piano Transcription With Convolutional Sparse Lateral Inhibition
abstract
This letter extends our prior work on context-dependent piano transcription to estimate the length of the notes in addition to their pitch and onset. This approach employs convolutional sparse coding along with lateral inhibition constraints to approximate a musical signal as the sum of piano note waveforms (dictionary elements) convolved with their temporal activations. The waveforms are pre-recorded for the specific piano to be transcribed in the specific environment. A dictionary containing multiple waveforms per pitch is generated by truncating a long waveform for each pitch to different lengths. During transcription, the dictionary elements are fixed and their temporal activations are estimated and postprocessed to obtain the pitch, onset, and note length estimation. A sparsity penalty promotes globally sparse activations of the dictionary elements, and a lateral inhibition term penalizes concurrent activations of different waveforms corresponding to the same pitch within a temporal neighborhood, to achieve note length estimation. Experiments on the MIDI aligned piano sounds dataset show that the proposed approach significantly outperforms a state-of-the-art music transcription method trained in the same context-dependent setting in transcription accuracy.
Andrea Cogliati, Zhiyao Duan, Brendt Wohlberg
IEEE Signal Process. Lett.2
2016 Emotion classification: How does an automated system compare to Naive human coders?
abstract
The fact that emotions play a vital role in social interactions, along with the demand for novel human-computer interaction applications, have led to the development of a number of automatic emotion classification systems. However, it is still debatable whether the performance of such systems can compare with human coders. To address this issue, in this study, we present a comprehensive comparison in a speech-based emotion classification task between 138 Amazon Mechanical Turk workers (Turkers) and a state-of-the-art automatic computer system. The comparison includes classifying speech utterances into six emotions (happy, neutral, sad, anger, disgust and fear), into three arousal classes (active, passive, and neutral), and into three valence classes (positive, negative, and neutral). The results show that the computer system outperforms the naive Turkers in almost all cases. Furthermore, the computer system can increase the classification accuracy by rejecting to classify utterances for which it is not confident, while the Turkers do not show a significantly higher classification accuracy on their confident utterances versus unconfi-dent ones.
Sefik Emre Eskimez, Kenneth Imade, Melissa Sturge-Apple, Zhiyao Duan, Wendi B. Heinzelman
ICASSP5
2016 IMISOUND: An unsupervised system for sound query by vocal imitation
abstract
Vocal imitation is widely used in human interactions. In this paper, we propose a novel human-computer interaction system called IMISOUND that listens to a vocal imitation and retrieves similar sounds from a sound library. This system allows users to search sounds even if they do not remember their semantic labels or the sounds do not have these labels (e.g., synthesized sound effects). IMISOUND employs a Stacked Auto-Encoder (SAE) to extract features from both the vocal imitation (query) and sounds in the library (candidates). The SAE is pre-trained using training vocal imitations of sounds not in the library to automatically learn more suitable feature representations than human-engineered features such as MFCC's. It then measures the similarity between the query and each sound candidate, using the K-L divergence and Dynamic Time Warping distance between their feature representations, and finally retrieves the closest sounds. IMISOUND is an unsupervised system in the sense that no training is performed for the target sound, nonetheless, experiments show that it achieves comparable performance to a previously proposed supervised system which requires pre-training on sounds to be retrieved. Experiments also show that IMISOUND significantly outperforms an unsupervised MFCC-based baseline system, validating the advantage of the SAE feature representation.
Yichi Zhang 0008, Zhiyao Duan
ICASSP2
2016 Context-Dependent Piano Music Transcription With Convolutional Sparse Coding
abstract
This paper presents a novel approach to automatic transcription of piano music in a context-dependent setting. This approach employs convolutional sparse coding to approximate the music waveform as the summation of piano note waveforms (dictionary elements) convolved with their temporal activations (onset transcription). The piano note waveforms are pre-recorded for the specific piano to be transcribed in the specific environment. During transcription, the note waveforms are fixed and their temporal activations are estimated and post-processed to obtain the pitch and onset transcription. This approach works in the time domain, models temporal evolution of piano notes, and estimates pitches and onsets simultaneously in the same framework. Experiments show that it significantly outperforms a state-of-the-art music transcription method trained in the same context-dependent setting, in both transcription accuracy and time precision, in various scenarios including synthetic, anechoic, noisy, and reverberant environments.
Andrea Cogliati, Zhiyao Duan, Brendt Wohlberg
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 An Approach to Score Following for Piano Performances With the Sustained Effect
abstract
One challenge in score following for piano music is the sustained effect, i.e., the waveform of a note lasts longer than what is notated in the score. This can be caused by expressive performing styles such as the legato articulation and the usage of the sustain and the sostenuto pedals and can also be caused by the reverberation in the recording environment. This effect creates nonnotated overlappings between sustained notes and latter notes in the audio. It decreases the audio-score alignment accuracy and robustness of score following systems and makes them be prone to delay errors, i.e., aligning audio to a score position that is earlier than the correct position. In this paper, we propose to modify the feature representation of the audio to attenuate the sustained effect. We show that this idea can be applied to both the chromagram and the spectral-peak representations, which are commonly used in score following systems. Experiments on the MAPS dataset show that the proposed method significantly improves the alignment accuracy and robustness of score following systems for piano performances, in both anechoic and highly reverberant environments.
Bochen Li, Zhiyao Duan
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Piano music transcription modeling note temporal evolution
abstract
Automatic music transcription (AMT) is the process of converting an acoustic musical signal into a symbolic musical representation such as a MIDI piano roll, which contains the pitches, the onsets and offsets of the notes and, possibly, their dynamic and source (i.e., instrument). Existing algorithms for AMT commonly identify pitches and their saliences in each frame and then form notes in a post-processing stage, which applies a combination of thresholding, pruning and smoothing operations. Very few existing methods consider the note temporal evolution over multiple frames during the pitch identification stage. In this work we propose a note-based spectrogram factorization method that uses the entire temporal evolution of piano notes as a template dictionary. The method uses an artificial neural network to detect note onsets from the audio spectral flux. Next, it estimates the notes present in each audio segment between two successive onsets with a greedy search algorithm. Finally, the spectrogram of each segment is factorized using a discrete combination of note templates comprised of full note spectrograms of individual piano notes sampled at different dynamic levels. We also propose a new psychoacoustically informed measure for spectrogram similarity.
Andrea Cogliati, Zhiyao Duan
ICASSP2
2014 A novel cepstral representation for timbre modeling of sound sources in polyphonic mixtures
abstract
We propose a novel cepstral representation called the uniform discrete cepstrum (UDC) to represent the timbre of sound sources in a sound mixture. Different from ordinary cepstrum and MFCC which have to be calculated from the full magnitude spectrum of a source after source separation, UDC can be calculated directly from isolated spectral points that are likely to belong to the source in the mixture spectrum (e.g., non-overlapping harmonics of a harmonic source). Existing cepstral representations that have this property are discrete cepstrum and regularized discrete cepstrum, however, compared to the proposed UDC, they are not as effective and are more complex to compute. The key advantage of UDC is that it uses a more natural and locally adaptive regularizer to prevent it from overfitting the isolated spectral points. We derive the mathematical relations between these cepstral representations, and compare their timbre modeling performances in the task of instrument recognition in polyphonic audio mixtures. We show that UDC and its mel-scale variant MUDC significantly outperform all the other representations.
Zhiyao Duan, Bryan Pardo, Laurent Daudet
ICASSP1
2014 Multi-pitch Streaming of Harmonic Sound Mixtures
abstract
Multi-pitch analysis of concurrent sound sources is an important but challenging problem. It requires estimating pitch values of all harmonic sources in individual frames and streaming the pitch estimates into trajectories, each of which corresponds to a source. We address the streaming problem for monophonic sound sources. We take the original audio, plus frame-level pitch estimates from any multi-pitch estimation algorithm as inputs, and output a pitch trajectory for each source. Our approach does not require pre-training of source models from isolated recordings. Instead, it casts the problem as a constrained clustering problem, where each cluster corresponds to a source. The clustering objective is to minimize the timbre inconsistency within each cluster. We explore different timbre features for music and speech. For music, harmonic structure and a newly proposed feature called uniform discrete cepstrum (UDC) are found effective; while for speech, MFCC and UDC works well. We also show that timbre-consistency is insufficient for effective streaming. Constraints are imposed on pairs of pitch estimates according to their time-frequency relationships. We propose a new constrained clustering algorithm that satisfies as many constraints as possible while optimizing the clustering objective. We compare the proposed approach with other state-of-the-art supervised and unsupervised multi-pitch streaming approaches that are specifically designed for music or speech. Better or comparable results are shown.
Zhiyao Duan, Jinyu Han, Bryan Pardo
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Combining rhythm-based and pitch-based methods for background and melody separation
abstract
Musical works are often composed of two characteristic components: the background (typically the musical accompaniment), which generally exhibits a strong rhythmic structure with distinctive repeating time elements, and the melody (typically the singing voice or a solo instrument), which generally exhibits a strong harmonic structure with a distinctive predominant pitch contour. Drawing from findings in cognitive psychology, we propose to investigate the simple combination of two dedicated approaches for separating those two components: a rhythm-based method that focuses on extracting the background via a rhythmic mask derived from identifying the repeating time elements in the mixture and a pitch-based method that focuses on extracting the melody via a harmonic mask derived from identifying the predominant pitch contour in the mixture. Evaluation on a data set of song clips showed that combining such two contrasting yet complementary methods can help to improve separation performance-from the point of view of both components-compared with using only one of those methods, and also compared with two other state-of-the-art approaches.
Zafar Rafii, Zhiyao Duan, Bryan Pardo
IEEE ACM Trans. Audio Speech Lang. Process.2
2012 Speech Enhancement by Online Non-negative Spectrogram Decomposition in Non-stationary Noise Environments
abstract
Classical single-channel speech enhancement algorithms have two convenient properties: they require pre-learning the noise model but not the speech model, and they work online. However, they often have difficulties in dealing with non-stationary noise sources. Source separation algorithms based on nonnegative spectrogram decompositions are capable of dealing with non-stationary noise, but do not possess the aforementioned properties. In this paper we present a novel algorithm that combines the advantages of both classical algorithms and non-negative spectrogram decomposition algorithms. Experiments show that it significantly outperforms four categories of classical algorithms in non-stationary noise environments.
Zhiyao Duan, Gautham J. Mysore, Paris Smaragdis
INTERSPEECH1
2011 A state space model for online polyphonic audio-score alignment
abstract
We present a novel online audio-score alignment approach for multi-instrument polyphonic music. This approach uses a 2-dimensional state vector to model the underlying score position and tempo of each time frame of the audio performance. The process model is defined by dynamic equations to transition between states. Two representations of the observed audio frame are proposed, resulting in two observation models: a multi-pitch-based and a chroma-based. Particle filtering is used to infer the hidden states from observations. Experiments on 150 music pieces with polyphony from one to four show the proposed approach outperforms an existing offline global string alignment-based score alignment approach. Results also show that the multi-pitch-based observation model works better than the chroma-based one.
Zhiyao Duan, Bryan Pardo
ICASSP1
2010 Song-level multi-pitch tracking by heavily constrained clustering
abstract
Given a set of monophonic, harmonic sound sources (e.g. human voices or wind instruments), multi-pitch estimation (MPE) is the task of determining the instantaneous pitches of each source. Multi-pitch tracking (MPT) connects the instantaneous pitch estimates provided by MPE algorithms into pitch trajectories of sources. A trajectory can be short (within a musical note), or long (an entire piece of music). While note-level MPT methods usually utilize local time-frequency proximity of pitches to connect them into a note, song-level MPT is much more difficult and needs more information. This is because pitches evolve discontinuously from note to note, and pitch trajectories can even interweave. In this paper, we cast the song-level MPT problem as a constrained clustering problem. The constraints are time-frequency locality of pitches and the clustering objective is their timbre consistency. Due to this problem's unique properties, existing constrained clustering algorithms cannot be directly applied. We propose a new constrained clustering algorithm. Experiments show that our approach produces good results on real-world music recordings of 4 musical instruments.
Zhiyao Duan, Jinyu Han, Bryan Pardo
ICASSP1
2010 Multiple Fundamental Frequency Estimation by Modeling Spectral Peaks and Non-Peak Regions
abstract
This paper presents a maximum-likelihood approach to multiple fundamental frequency (F0) estimation for a mixture of harmonic sound sources, where the power spectrum of a time frame is the observation and the F0s are the parameters to be estimated. When defining the likelihood model, the proposed method models both spectral peaks and non-peak regions (frequencies further than a musical quarter tone from all observed peaks). It is shown that the peak likelihood and the non-peak region likelihood act as a complementary pair. The former helps find F0s that have harmonics that explain peaks, while the latter helps avoid F0s that have harmonics in non-peak regions. Parameters of these models are learned from monophonic and polyphonic training data. This paper proposes an iterative greedy search strategy to estimate F0s one by one, to avoid the combinatorial problem of concurrent F0 estimation. It also proposes a polyphony estimation method to terminate the iterative process. Finally, this paper proposes a postprocessing method to refine polyphony and F0 estimates using neighboring frames. This paper also analyzes the relative contributions of different components of the proposed method. It is shown that the refinement component eliminates many inconsistent estimation errors. Evaluations are done on ten recorded four-part J. S. Bach chorales. Results show that the proposed method shows superior F0 estimation and polyphony estimation compared to two state-of-the-art algorithms.
Zhiyao Duan, Bryan Pardo, Changshui Zhang
IEEE Trans. Speech Audio Process.1
2008 Audio tonality mode classification without tonic annotations
abstract
Traditional tonality mode (major or minor) classification or audio key finding algorithms often rely on tonic annotations (key names) of the training songs. However, unlike classical music whose keys are usually explicitly labeled in their titles, the keys of numerous popular music are hard to obtain. In contrast, it is much easier to only label the mode for each song. With only modes labeled, traditional approaches to key or mode classification cannot be directly applied, due to the lack of the reference point to transpose and align the chroma features with different keys. In this paper, we present an alignment approach to transpose chroma features within each mode to a reference (but unknown) tonic. Then several methods, including Single Profile Correlation, Multiple Profile Correlation and Support Vector Machine, are exploited to address mode learning and classification. Experimental results show the feasibility of the proposed approach.
Zhiyao Duan, Lie Lu, Changshui Zhang
ICME1
2008 Unsupervised Single-Channel Music Source Separation by Average Harmonic Structure Modeling
abstract
Source separation of musical signals is an appealing but difficult problem, especially in the single-channel case. In this paper, an unsupervised single-channel music source separation algorithm based on average harmonic structure modeling is proposed. Under the assumption of playing in narrow pitch ranges, different harmonic instrumental sources in a piece of music often have different but stable harmonic structures; thus, sources can be characterized uniquely by harmonic structure models. Given the number of instrumental sources, the proposed algorithm learns these models directly from the mixed signal by clustering the harmonic structures extracted from different frames. The corresponding sources are then extracted from the mixed signal using the models. Experiments on several mixed signals, including synthesized instrumental sources, real instrumental sources, and singing voices, show that this algorithm outperforms the general nonnegative matrix factorization (NMF)-based source separation algorithm, and yields good subjective listening quality. As a side effect, this algorithm estimates the pitches of the harmonic instrumental sources. The number of concurrent sounds in each frame is also computed, which is a difficult task for general multipitch estimation (MPE) algorithms.
Zhiyao Duan, Yungang Zhang, Changshui Zhang, Zhenwei Shi 0001
IEEE Trans. Speech Audio Process.1
2007 Multi-Pitch Estimation Based on Partial Event and Support Transfer
abstract
This paper proposes a method for the multi-pitch estimation of polyphonic music signals. Instead of on the frame level, the estimation is based on the partial event, which is defined like the note event in MIDI. All partial events in a piece of music are extracted dynamically in the process of the frame by frame short time Fourier transform (STFT). For each event, net support degree received from other events is calculated and the events with the highest support degrees are selected to be the fundamental frequency (FO) events. From another point of view, the support is transferred from higher frequency partial events to lower ones and finally concentrated on the FO events. This method can estimate the number of concurrent sounds, the onset and offset times of the notes. Experiments on both randomly mixed chord signals and synthesized ensemble music signals in "wav" format are conducted and the results are promising.
Zhiyao Duan, Dan Zhang 0007, Changshui Zhang, Zhenwei Shi 0001
ICME1