Yin Cao

dblp:94/7894 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ERBMA-Net: Enhanced Random Binary Multilevel Attention Network for Facial Depression Recognition
abstract
Depression is a major contributor to the global disease burden and is closely associated with variations in facial expressions, making them a critical biomarker for its recognition. Traditional local binary convolutional neural networks (LBCNN) rely on fixed binary weights, which fail to capture complex and nuanced features, limiting their adaptability and generalization. Existing methods often lack extracting spatial hierarchies and model long-range dependencies within facial expressions. To address these gaps, an enhanced random binary multilevel attention network (ERBMA-Net) is proposed for facial depression recognition, which employs a multibranch architecture to extract global (face) and local (eyes and mouth) features. Subsequently, the enhanced random binary convolutional neural network (ERBCNN) introduces random binary filters and a refinement layer to improve adaptability over LBCNN. Furthermore, multilevel attention mechanisms, including spatial attention are applied independently to global and local features and self-attention on concatenated features to capture critical spatial and relational patterns. Comprehensive evaluations on the Audio-Visual Emotion Challenge 2014 (AVEC2014) and the Changzhou No. 2 People’s Hospital (CZ2024) datasets demonstrate the superior performance of ERBMA-Net, achieving mean absolute error and root mean square error: 6.80/8.18 and 6.37/8.08, respectively. These results establish ERBMA-Net as a robust framework for artificial intelligence-assisted mental health diagnostics. The code is available at:https://github.com/DrTuryalai/ERBMA-Net.
Muhammad Turyalai Khan, Yin Cao, Faisal Shafait, Wu Jun
IEEE Trans. Comput. Soc. Syst.2
2025 Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
Fang Kang, Yin Cao
INTERSPEECH2
2025 DiffStereo: End-to-End Mono-to-Stereo Audio Generation with Diffusion Transformer
Suqi Zhang, Zheqi Dai, Yongyi Zang, Yin Cao, Qiuqiang Kong
INTERSPEECH4
2025 EDTC: enhanced depth of text comprehension in automated audio captioning
abstract
Abstract Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multimodal domains. Facilitating models in comprehending text information plays a pivotal role in establishing a seamless connection between the two modalities of text and audio. While recent research has focused on closing the gap between these two modalities through contrastive learning, it is challenging to bridge the difference between both modalities using only simple contrastive loss. This paper introduces enhanced depth of text comprehension, which enhances the model’s understanding of text information from three different perspectives. First, a combined Local-Global feature Fusion module is introduced to fuse heterogeneous audio features, enabling the extraction of high-level semantic information and the discovery of latent inter-sample relationships. Next, a novel representation module, TRANSLATOR, constructs a twin-branch structure based on the conventional dual-stream model, mapping features from both modalities into a shared high-dimensional audio-text space. Finally, contrastive learning is integrated with momentum-based weight updates, allowing the system to effectively capture shared high-level semantic representations across the audio and text modalities.
Liwen Tan, Yi Zhou 0014, Yin Cao
Comput. J.4
2025 Improved Digital Arctangent Demodulation Method With Doppler Signal at Special Sampling Rate for Laser Heterodyne Interferometer
abstract
The importance of vibration monitoring in the Internet of Things (IoT) is reflected in many aspects, especially in the fields of industry, infrastructure, health monitoring, etc. Laser heterodyne interferometer has been widely used in the measurement of vibration displacement and velocity. The phase demodulation method of the Doppler signal is crucial to realizing the real-time and high precision measurement with wide bandwidth. To reduce the requirements for the data acquisition system and the resource consumption, the improved digital arctangent demodulation method is proposed, which is more suitable to run on the digital signal processor. Utilizing the symmetry of the Doppler signal spectrum, the Doppler signal can be acquired by special Nyquist or bandpass sampling rates. Then the pair of orthogonal signals are generated by specific delays of the digitized Doppler signals. With the sine approximation method (SAM), the vibration displacement or velocity can be obtained from the unwrapped phase after the arctangent calculation. Compared with the classical arctangent demodulation method and the commercial decoder of the laser Doppler vibrometer (LDV), experiments are designed to demonstrate the feasibility and performance of the proposed method with the vibration frequency from 500 Hz to 1 MHz and the peak amplitude of the vibration velocity from 31.63 lm/s to 3.16 m/s. Additionally, simulation experiments further explore its applicability in demodulating Doppler signals with asymmetric spectrum. The improved digital arctangent demodulation method offers significant potential for developing Doppler signal demodulation systems for laser heterodyne interferometers, particularly in real-time, efficient, and high-precision vibration monitoring applications.
Xiujuan Feng, Longbiao He, Feng Niu, Ronghua Fan, Yin Cao, Lijing Li
IEEE Internet Things J.8
2024 FastFaceCLIP: A lightweight text-driven high-quality face image manipulation
abstract
Abstract Although many new methods have emerged in text‐driven images, the large computational power required for model training causes these methods to have a slow training process. Additionally, these methods consume a considerable amount of video random access memory (VRAM) resources during training. When generating high‐resolution images, the VRAM resources are often insufficient, which results in the inability to generate high‐resolution images. Nevertheless, recent Vision Transformers (ViTs) advancements have demonstrated their image classification and recognition capabilities. Unlike the traditional Convolutional Neural Networks based methods, ViTs have a Transformer‐based architecture, leverage attention mechanisms to capture comprehensive global information, moreover enabling enhanced global understanding of images through inherent long‐range dependencies, thus extracting more robust features and achieving comparable results with reduced computational load. The adaptability of ViTs to text‐driven image manipulation was investigated. Specifically, existing image generation methods were refined and the FastFaceCLIP method was proposed by combining the image‐text semantic alignment function of the pre‐trained CLIP model with the high‐resolution image generation function of the proposed FastFace. Additionally, the Multi‐Axis Nested Transformer module was incorporated for advanced feature extraction from the latent space, generating higher‐resolution images that are further enhanced using the Real‐ESRGAN algorithm. Eventually, extensive face manipulation‐related tests on the CelebA‐HQ dataset challenge the proposed method and other related schemes, demonstrating that FastFaceCLIP effectively generates semantically accurate, visually realistic, and clear images using fewer parameters and less time.
Junping Qin, Qianli Ma 0006, Yin Cao
IET Comput. Vis.4
2024 Selective-Memory Meta-Learning With Environment Representations for Sound Event Localization and Detection
abstract
Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular acoustic settings, often show restricted generalization capabilities for diverse acoustic environments. Furthermore, obtaining annotated samples for spatial sound events is notably costly. Deploying a SELD system in a new environment requires extensive time for re-training and fine-tuning. To overcome these challenges, we propose environment-adaptive Meta-SELD, designed for efficient adaptation to new environments using minimal data. Our method specifically utilizes computationally synthesized spatial data and employs Model-Agnostic Meta-Learning (MAML) on a pre-trained, environment-independent model. The method then utilizes fast adaptation to unseen real-world environments using limited samples from the respective environments. Inspired by the Learning-to-Forget approach, we introduce the concept of selective memory as a strategy for resolving conflicts across environments. This approach involves selectively memorizing target-environment-relevant information and adapting to the new environments through the selective attenuation of model parameters. In addition, we introduce environment representations to characterize different acoustic settings, enhancing the adaptability of our attenuation approach to various environments. We evaluate our proposed method on the development set of the Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset and computationally synthesized scenes. Experimental results demonstrate the superior performance of the proposed method compared to conventional supervised learning methods, particularly in localization.
Jinbo Hu, Yin Cao, Ming Wu 0005, Qiuqiang Kong, Feiran Yang 0001, Mark D. Plumbley, Jun Yang 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 A Track-Wise Ensemble Event Independent Network for Polyphonic Sound Event Localization and Detection
abstract
Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a trackwise ensemble event independent network with a novel data augmentation method is proposed. The proposed model is based on our previous proposed Event-Independent Network V2 and is extended by conformer blocks and dense blocks. The track-wise ensemble model with track-wise output format is proposed to solve an ensemble model problem for track-wise output format that track permutation may occur among different models. The data augmentation approach contains several data augmentation chains, which are composed of random combinations of several data augmentation operations. The method also utilizes log-mel spectrograms, intensity vectors, and Spatial Cues-Augmented Log-Spectrogram (SALSA) for different models. We evaluate our proposed method in the Task of the L3DAS22 challenge and obtain the top ranking solution with a location-dependent F-score to be 0.699. Source code is released1.
Jinbo Hu, Yin Cao, Ming Wu 0005, Qiuqiang Kong, Feiran Yang 0001, Mark D. Plumbley, Jun Yang 0004
ICASSP2
2022 Statistical analysis of multichannel FxLMS algorithm for narrowband active noise control
Ming Wu 0005, Jing Chen 0090, Zeqiang Zhang, Yin Cao, Jun Yang 0004
Signal Process.6
2021 An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
abstract
Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA angles simultaneously. We study the SELD task from a multi-task learning perspective. Two open problems are addressed in this paper. Firstly, to detect overlapping sound events of the same type but with different DoAs, we propose to use a trackwise output format and solve the accompanying track permutation problem with permutation-invariant training. Multi-head self-attention is further used to separate tracks. Secondly, a previous finding is that, by using hard parameter-sharing, SELD suffers from a performance loss compared with learning the subtasks separately. This is solved by a soft parameter-sharing scheme. We term the proposed method as Event Independent Network V2 (EINV2), which is an improved version of our previously-proposed method and an end-to-end network for SELD. We show that our proposed EINV2 for joint SED and DoA estimation outperforms previous methods by a large margin, and has comparable performance to state-of-the-art ensemble models.
Yin Cao, Turab Iqbal, Qiuqiang Kong, Fengyan An, Wenwu Wang 0001, Mark D. Plumbley
ICASSP1
2020 Learning With Out-of-Distribution Data for Audio Classification
abstract
In supervised machine learning, the assumption that training data is labelled correctly is not always satisfied. In this paper, we investigate an instance of labelling error for classification tasks in which the dataset is corrupted with out-of-distribution (OOD) instances: data that does not belong to any of the target classes, but is labelled as such. We show that detecting and relabelling certain OOD instances, rather than discarding them, can have a positive effect on learning. The proposed method uses an auxiliary classifier, trained on data that is known to be in-distribution, for detection and relabelling. The amount of data required for this is shown to be small. Experiments are carried out on the FSDnoisy18k audio dataset, where OOD instances are very prevalent. The proposed method is shown to improve the performance of convolutional neural networks by a significant margin. Comparisons with other noise-robust techniques are similarly encouraging.
Turab Iqbal, Yin Cao, Qiuqiang Kong, Mark D. Plumbley, Wenwu Wang 0001
ICASSP2
2020 Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis
abstract
Source separation is the task of separating an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular sound classes such as speech and music. Much previous work requires mixtures and clean source pairs for training. In this work, we propose a source separation framework trained with weakly labelled data. Weakly labelled data only contains the tags of an audio clip, without the occurrence time of sound events. We first train a sound event detection system with AudioSet. The trained sound event detection system is used to detect segments that are most likely to contain a target sound event. Then a regression is learnt from a mixture of two randomly selected segments to a target segment conditioned on the audio tagging prediction of the target segment. Our proposed system can separate 527 kinds of sound classes from AudioSet within a single system. A U-Net is adopted for the separation system and achieves an average SDR of 5.67 dB over 527 sound classes in AudioSet.
Qiuqiang Kong, Yuxuan Wang 0002, Xuchen Song, Yin Cao, Wenwu Wang 0001, Mark D. Plumbley
ICASSP4
2020 Module dividing for brain functional networks by employing betweenness efficiency
Min Cai, Xuelian Ming, Yin Cao, Ling Zou 0002, Shuihua Wang
Multim. Tools Appl.4
2020 Rich club characteristics of dynamic brain functional networks in resting state
Min Cai, Yin Cao, Ling Zou 0002, Shuihua Wang
Multim. Tools Appl.4
2020 PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
abstract
Audio pattern recognition is an important research topic in the machine learning area, and includes several tasks such as audio tagging, acoustic scene classification, music classification, speech emotion classification and sound event detection. Recently, neural networks have been applied to tackle audio pattern recognition problems. However, previous systems are built on specific datasets with limited durations. Recently, in computer vision and natural language processing, systems pretrained on large-scale datasets have generalized well to several tasks. However, there is limited research on pretraining systems on large-scale datasets for audio pattern recognition. In this paper, we propose pretrained audio neural networks (PANNs) trained on the large-scale AudioSet dataset. These PANNs are transferred to other audio related tasks. We investigate the performance and computational complexity of PANNs modeled by a variety of convolutional neural networks. We propose an architecture called Wavegram-Logmel-CNN using both log-mel spectrogram and waveform as input feature. Our best PANN system achieves a state-of-the-art mean average precision (mAP) of 0.439 on AudioSet tagging, outperforming the best previous system of 0.392. We transfer PANNs to six audio pattern recognition tasks, and demonstrate state-of-the-art performance in several of those tasks. We have released the source code and pretrained models of PANNs:https://github.com/qiuqiangkong/audioset_tagging_cnn.
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang 0002, Wenwu Wang 0001, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Acoustic Scene Generation with Conditional Samplernn
abstract
Acoustic scene generation (ASG) is a task to generate waveforms for acoustic scenes. ASG can be used to generate audio scenes for movies and computer games. Recently, neural networks such as SampleRNN have been used for speech and music generation. However, ASG is more challenging due to its wide variety. In addition, evaluating a generative model is also difficult. In this paper, we propose to use a conditional SampleRNN model to generate acoustic scenes conditioned on the input classes. We also propose objective criteria to evaluate the quality and diversity of the generated samples based on classification accuracy. The experiments on the DCASE 2016 Task 1 acoustic scene data show that with the generated audio samples, a classification accuracy of 65.5% can be achieved compared to samples generated by a random model of 6.7% and samples from real recording of 83.1%. The performance of a classifier trained only on generated samples achieves an accuracy of 51.3%, as opposed to an accuracy of 6.7% with samples generated by a random model.
Qiuqiang Kong, Yong Xu 0004, Turab Iqbal, Yin Cao, Wenwu Wang 0001, Mark D. Plumbley
ICASSP4
2019 Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge
abstract
Humans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the creation of a data set that is structured according to the top-level of a taxonomy derived from human judgements and the design of an associated machine learning challenge, in which strong generalisation abilities are required to be successful. We introduce a baseline classification system, a deep convolutional network, which showed strong performance with an average accuracy on the evaluation data of 80.8%. The result is discussed in the light of two alternative explanations: An unlikely accidental category bias in the sound recordings or a more plausible true acoustic grounding of the high-level categories.
Christian Kroos, Oliver Bones, Yin Cao, Lara Harris, Philip J. B. Jackson, William J. Davies, Wenwu Wang 0001, Trevor J. Cox, Mark D. Plumbley
ICASSP3
2014 Effective connectivity analysis of fMRI data based on network motifs
Ling Zou 0002, Yin Cao, Nong Qian, Zhenghua Ma
J. Supercomput.3
2012 Monotonic Regression: A New Way for Correlating Subjective and Objective Ratings in Image Quality Research
abstract
To assess the performance of image quality metrics (IQMs), some regressions, such as logistic regression and polynomial regression, are used to correlate objective ratings with subjective scores. However, some defects in optimality are shown in these regressions. In this correspondence, monotonic regression (MR) is found to be an effective correlation method in the performance assessment of IQMs. Both theoretical analysis and experimental results have proven that MR performs better than any other regression. We believe that MR could be an effective tool for performance assessment in the IQM research.
Yu Han 0013, Yunze Cai, Yin Cao, Xiaoming Xu 0001
IEEE Trans. Image Process.3