Yangyang Xia

dblp:226/1929 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
6since 2021 · last 2021
0009-0007-4303-0751ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2021 Incorporating Real-World Noisy Speech in Neural-Network-Based Speech Enhancement Systems
abstract
Supervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-world degraded speech data that may better represent the scenarios where such systems are used. In this paper, we explore methods that enable supervised speech enhancement systems to train on real-world degraded speech data. Specifically, we propose a semi-supervised approach for speech enhancement in which we first train a modified vector-quantized variational autoencoder that solves a source separation task. We then use this trained autoencoder to further train an enhancement network using real-world noisy speech data by computing a triplet-based unsupervised loss function. Experiments show promising results for incorporating real-world data in training speech enhancement systems.
Yangyang Xia, Buye Xu, Anurag Kumar 0003
ASRU1
2021 Domain-Specific Suppression for Adaptive Object Detection
abstract
Domain adaptation methods face performance degradation in object detection, as the complexity of tasks require more about the transferability of the model. We propose a new perspective on how CNN models gain the transferability, viewing the weights of a model as a series of motion patterns. The directions of weights, and the gradients, can be divided into domain-specific and domain-invariant parts, and the goal of domain adaptation is to concentrate on the domain-invariant direction while eliminating the disturbance from domain-specific one. Current UDA object detection methods view the two directions as a whole while optimizing, which will cause domain-invariant direction mismatch even if the output features are perfectly aligned. In this paper, we propose the domain-specific suppression, an exemplary and generalizable constraint to the original convolution gradients in backpropagation to detach the two parts of directions and suppress the domain-specific one. We further validate our theoretical analysis and methods on several domain adaptive object detection tasks, including weather, camera configuration, and synthetic to real-world adaptation. Our experiment results show significant advance over the state-of-the-art methods in the UDA object detection field, performing a promotion of 10.2 ∼ 12.2% mAP on all these domain adaptation scenarios.
Rui Zhang 0040, Yangyang Xia, Xishan Zhang, Shaoli Liu
CVPR5
2021 A Modulation-Domain Loss for Neural-Network-Based Real-Time Speech Enhancement
abstract
We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then used to calculate a weighted mean-squared error (MSE) in the modulation domain for training a speech enhancement system. Experiments showed that adding the modulation-domain MSE to the MSE in the spectro-temporal domain substantially improved the objective prediction of speech quality and intelligibility for real-time speech enhancement systems without incurring additional computation during inference.
Tyler Vuong, Yangyang Xia, Richard M. Stern
ICASSP2
2021 The Application of Learnable STRF Kernels to the 2021 Fearless Steps Phase-03 SAD Challenge
Tyler Vuong, Yangyang Xia, Richard M. Stern
Interspeech2
2021 Temporal Context in Speech Emotion Recognition
Yangyang Xia, Alexander I. Rudnicky, Richard M. Stern
Interspeech1
2021 Distributed Data Collection in Age-Aware Vehicular Participatory Sensing Networks
abstract
The advent of vehicle-to-everything communication facilitates the emergence of vehicular sensing networks, where vehicles equipped with advanced sensors continuously sample informative status updates of its surroundings and forward the sampled data to roadside infrastructure based on a certain routing strategy. The collected data is analyzed to obtain real-time situational awareness to impose certain behaviors on the vehicles. In such networked control systems, the timeliness of collected data is of critical importance to system performance, which can be quantified by the concept of Age of Information. Note that to obtain timely perception of its surroundings, each vehicle tends to sample status updates at the maximum frequency, which may congest the network due to limited communication resource. Moreover, the highly dynamic nature of vehicular network poses a great challenge in finding a reliable route for timely data forwarding. Therefore, the data collection scheme should be carefully designed to balance the timeliness of collected information and network stability. In this article, we study an age optimization problem by jointly considering the data sampling at source vehicles and the data forwarding process for multiple information flows across the network. We employ the Lyapunov optimization technique to develop a distributed age-aware data collection scheme consists of a threshold-based sampling strategy at source vehicles and a learning-based data forwarding strategy. Simulation results show that our proposed scheme outperforms existing strategies in collecting status updates in a timely manner.
Xiaoqi Qin, Yangyang Xia, Hang Li 0003, Zhiyong Feng 0001, Ping Zhang 0003
IEEE Internet Things J.2
2020 Weighted Speech Distortion Losses for Neural-Network-Based Real-Time Speech Enhancement
abstract
This paper investigates several aspects of training a RNN (recurrent neural network) that impact the objective and subjective quality of enhanced speech for real-time single-channel speech enhancement. Specifically, we focus on a RNN that enhances short-time speech spectra on a single-frame-in, single-frame-out basis, a framework adopted by most classical signal processing methods. We propose two novel mean-squared-error-based learning objectives that enable separate control over the importance of speech distortion versus noise reduction. The proposed loss functions are evaluated by widely accepted objective quality and intelligibility measures and compared to other competitive online methods. In addition, we study the impact of feature normalization and varying batch sequence lengths on the objective quality of enhanced speech. Finally, we show subjective ratings for the proposed approach and a state-of-the-art real-time RNN-based method.
Yangyang Xia, Sebastian Braun, Chandan K. A. Reddy, Harishchandra Dubey, Ross Cutler, Ivan Tashev
ICASSP1
2020 Learnable Spectro-Temporal Receptive Fields for Robust Voice Type Discrimination
abstract
Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio that were played back such as traffic noise and television broadcasts ("Distractor Audio"). In this work, we propose a deep-learning-based VTD system that features an initial layer of learnable spectro-temporal receptive fields (STRFs). Our approach is also shown to provide very strong performance on a similar spoofing detection task in the ASVspoof 2019 challenge. We evaluate our approach on a new standardized VTD database that was collected to support research in this area. In particular, we study the effect of using learnable STRFs compared to static STRFs or unconstrained kernels. We also show that our system consistently improves a competitive baseline system across a wide range of signal-to-noise ratios on spoofing detection in the presence of VTD distractor noise.
Tyler Vuong, Yangyang Xia, Richard M. Stern
INTERSPEECH2
2018 A Priori SNR Estimation Based on a Recurrent Neural Network for Robust Speech Enhancement
Yangyang Xia, Richard M. Stern
INTERSPEECH1