EDBT 2026 Demo / reviewers in the wild / expert
Xinyuan Qian 0001
dblp:119/4340-1
· DBLP profile ↗
43ranked-venue papers
11as first author
38since 2021 · last 2026
0000-0002-9511-6713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 10 first-author · 31 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AV-SSAN: Audio-Visual Selective DOA Estimation Through Explicit Multi-Band Semantic-Spatial AlignmentabstractAudio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we introduce Cross-Instance Audio-Visual Localization (CI-AVL), a novel task that localizes target sound sources using visual prompts from different instances of the same semantic class. CI-AVL enables selective localization without spatially paired data. To solve this task, we propose AV-SSAN, a semantic-spatial alignment framework centered on a Multi-Band Semantic-Spatial Alignment Network (MB-SSA Net). MB-SSA Net decomposes the audio spectrogram into multiple frequency bands, aligns each band with semantic visual prompts, and refines spatial cues to estimate the direction-of-arrival (DoA). To facilitate this research, we construct VGGSound-SSL, a large-scale dataset comprising 13,981 spatial audio clips across 296 categories, each paired with visual prompts. AV-SSAN achieves a mean absolute error of 16.59° and an accuracy of 71.29%, significantly outperforming existing AV-SSL methods. Hongxu Zhu, Kainan Chen, Xinyuan Qian 0001 |
AAAI | 5 |
| 2026 | TPEech: Target Speaker Extraction and Noise Suppression With Historical Dialogue Text CuesabstractIn complex multi-speaker scenarios with significant speaker overlap and background noise, extracting the target speaker's speech remains a major challenge. This capability is crucial for dialogue-based applications such as AI speech assistants, where downstream tasks such as speech recognition depend on clean speech. A potential solution to address these challenges is Target Speaker Extraction (TSE), which leverages auxiliary information to extract target speech from mixed and noisy speech, thus overcoming the limitations of Speech Separation (SS) and Speech Enhancement (SE). In particular, we propose a multi-modal TSE network, namely Text Prompt Extractor with echo cue block (TPEech), which uses historical dialogue text as cues for extraction and incorporates the echo cue block (ECB) to further exploit this cue and enhance TSE performance. The experiments show the excellent extraction and denoising capabilities of our proposed network. TPEech achieves an SI-SDRi of 9.632 dB, an SDR of 13.045 dB, a PESQ of 2.814, and a STOI of 0.885, outperforming competitive baselines. Additionally, we experimentally verify that TPEech is robust against semantically incomplete textual prompts. Dataset and source code will be publicly available. Ziyang Jiang, Shuai Wang 0016, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different StylesabstractHumans can perceive speakers’ characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech ( TTS ) scholars grounded their investigations on real-person faces, thereby restricting effective speech synthesis from applying to vast potential usage scenarios with diverse characters and image styles. To solve this issue, we introduce a novel FaceSpeak approach. It extracts salient identity characteristics and emotional representations from a wide variety of image styles. Meanwhile, it mitigates the extraneous information (e.g., background, clothing, and hair color, etc.), resulting in synthesized speech closely aligned with a character’s persona. Furthermore, to overcome the scarcity of multi-modal TTS data, we have devised an innovative dataset, namely Expressive Multi-Modal TTS ( EM2TTS), which is diligently curated and annotated to facilitate research in this domain. The experimental results demonstrate our proposed FaceSpeak can generate portrait-aligned voice with satisfactory naturalness and quality. Tian-Hao Zhang, Xinyuan Qian 0001, Xu-Cheng Yin |
AAAI | 4 |
| 2025 | M2PAIR: A High-Quality Acoustic Impulse Response Computation ModelabstractAcoustic Impulse Response (AIR) provides crucial spatial information about the environment, significantly enhancing audio immersion. However, achieving high perceptual quality while computing AIR in real-time for interactive audio-video media (IAVM) presents a challenging problem. This study proposes the Mesh to Parametric AIR (M2PAIR), a method for computing AIR designed for IAVM. M2PAIR integrates neural networks with psychoacoustics. It takes the 3D scene mesh, the listener positions, and the sound source positions as inputs, utilizes perceptual parameters as intermediaries, and computes the desired high-quality AIR signal based on these parameters. Experimental results demonstrate that M2PAIR improves the perceptual quality of AIR output compared to existing methods while reducing the model complexity. Additionally, it meets the requirements of IAVM, including real-time computation, high sampling rates, and flexible duration for the output AIR. Xinpei Zhao, Jing Wang 0037, Xinyuan Qian 0001 |
ICASSP | 4 |
| 2025 | Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech RecognitionabstractRecently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to integrate acoustic and language models, leveraging its capacity to implicitly fuse language models within static graphs, thereby ensuring robust recognition while also facilitating rapid error correction. However, WFST necessitates a frame-by-frame search of CTC posterior probabilities through autoregression, which significantly hampers inference speed. In this work, we thoroughly investigate the spike property of CTC outputs and further propose the conjecture that adjacent frames to non-blank spikes carry semantic information beneficial to the model. Building on this, we propose the Spike Window Decoding algorithm, which greatly improves the inference speed by making the number of frames decoded in WFST linearly related to the number of spiking frames in the CTC output, while guaranteeing the recognition performance. Our method achieves SOTA recognition accuracy with significantly accelerates decoding speed, proven across both AISHELL-1 and large-scale In-House datasets, establishing a pioneering approach for integrating CTC output with WFST. Tian-Hao Zhang, Xinyuan Qian 0001, Xu-Cheng Yin |
ICASSP | 6 |
| 2025 | Analytic Class Incremental Learning for Sound Source Localization With Privacy ProtectionabstractSound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based Sound Source Localization (SSL) methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle to adapt to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due tocatastrophic forgettingby incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods. Xinyuan Qian 0001, Xianghu Yue, Huiping Zhuang, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2025 | Improving Bird Vocalization Recognition in Open-Set Cross-Corpus Scenarios With Semantic Feature Reconstruction and Dual Strategy ScoringabstractAutomated recognition of bird vocalizations (BVs) is essential for biodiversity monitoring through passive acoustic monitoring (PAM), yet deep learning (DL) models encounter substantial challenges in open environments. These include difficulties in detecting unknown classes, extracting species-specific features, and achieving robust cross-corpus recognition. To address these challenges, this letter presents a DL-based open-set cross-corpus recognition method for BVs that combines feature construction with open-set recognition (OSR) techniques. We introduce a three-channel spectrogram that integrates both amplitude and phase information to enhance feature representation. To improve the recognition accuracy of known classes across corpora, we employ a class-specific semantic reconstruction model to extract deep features. For unknown class discrimination, we propose a Dual Strategy Coupling Scoring (DSCS) mechanism, which synthesizes the log-likelihood ratio score (LLRS) and reconstruction error score (RES). Our method achieves the highest weighted accuracy among existing approaches on a public dataset, demonstrating its effectiveness for open-set cross-corpus bird vocalization recognition. Jiangjian Xie, Xinyuan Qian 0001, Junguo Zhang, Björn W. Schuller |
IEEE Signal Process. Lett. | 3 |
| 2025 | SSDQ: Target Speaker Extraction via Semantic and Spatial Dual QueryingabstractTarget Speaker Extraction (TSE) in real-world multi-speaker environments is highly challenging. Previous works have largely relied on pre-enrollment speech to extract the target speaker's voice. However, such methods are limited in spontaneous scenarios where pre-enrollment speech or spatial information is unavailable. To address this, we propose Semantic and Spatial Dual Querying (SSDQ), a unified framework that integrates natural language descriptions and region-based spatial queries to guide TSE. SSDQ employs dual query encoders for semantic and spatial cues, fusing them into the audio stream via a FiLM-based interaction module. A novel Controllable Feature Wrapping (CFW) mechanism further enables a dynamic balance between speaker identity and acoustic clarity. We also introduce SS-Libri, a spatialized mixture dataset designed to benchmark dual-query systems. Extensive experiments demonstrate that SSDQ achieves superior extraction accuracy and robustness under challenging conditions, yielding the SI-SNRi of 19.63 dB, SNRi of 20.30 dB, PESQ of 1.83, and STOI of 0.26. Xinjia Zhu, Xinyuan Qian 0001, Dong Liang 0008 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Enhancing Real-World Active Speaker Detection With Multi-Modal Extraction Pre-TrainingabstractAudio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable gap in addressing the challenges from real-world AV-ASD scenarios. Due to the presence of low-quality noisy videos in such cases, AV-ASD systems without a selective listening ability are short of effectively filtering out disruptive voice components from mixed audio inputs. In this paper, we propose a Multi-modal Speech Extraction-to-Detection framework named ‘MuSED’, which is pre-trained with audio-visual target speech extraction to learn the denoising ability, then it is fine-tuned with the AV-ASD task. Meanwhile, to better capture the multi-modal information and deal with real-world problems such as missing modality, MuSED is modelled on the time domain directly and integrates the multi-modal plus-and-minus augmentation strategy. Our experiments demonstrate that MuSED substantially outperforms the state-of-the-art AV-ASD methods and achieves 95.6% mAP on the AVA-ActiveSpeaker dataset, 98.3% AP on the ASW dataset, and 97.9% F1 on the Columbia AV-ASD dataset, respectively. We will publicly release the code in due course. Ruijie Tao, Xinyuan Qian 0001, Rohan Kumar Das, Xiaoxue Gao, Haizhou Li 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing MechanismabstractThe prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization algorithm with a selective hearing mechanism. Given a reference speech of the target speaker, we first produce a speaker-dependent spectrogram mask to eliminate interfering speakers’ speech. Subsequently, a Long-Short-Term Memory (LSTM) network is employed to extract the target speaker’s location from the filtered spectrogram. Experiments validate the superiority of our proposed method over existing algorithms for different scale-invariant signal-to-noise ratios (SNR) conditions. Specifically, at SNR = -10 dB, our proposed network LocSelect achieves a mean absolute error (MAE) of 3.55° and an accuracy (ACC) of 87.40%. Xinyuan Qian 0001, Zexu Pan, Kainan Chen, Haizhou Li 0001 |
ICASSP | 2 |
| 2024 | Visually Guided Binaural Audio Generation with Cross-Modal ConsistencyabstractBinaural audio delivers an immersive spatial auditory experience to human listeners, but most existing videos lack binaural audio due to the expertise required for recording environments. Recent studies have been dedicated to converting monaural audio into binaural ones conditioned on the visual inputs. In this paper, we propose a novel audio-visual spatialization network with two added audio decoders, which rely on carefully designed visual features to generate audio outputs for the left and right channels, respectively. In addition, we propose an audio-visual matching loss to further explore the correlation between binaural audio and the scene visual input. Experiment results show that the proposed method outperforms several state-of-the-art binaural audio generation methods on two benchmark datasets FAIR-Play and MUSIC-Stereo. Qualitative results are also presented to demonstrate the effectiveness of the proposed method. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001 |
ICASSP | 3 |
| 2024 | GLMB 3D Speaker Tracking with Video-Assisted Multi-Channel Audio Optimization FunctionsabstractSpeaker tracking plays a significant role in numerous real-world human robot interaction (HRI) applications. In recent years, there has been a growing interest in utilizing multi-sensory information, such as complementary audio and visual signals, to address the challenges of speaker tracking. Despite the promising results, existing approaches still encounter difficulties in accurately determining the speaker’s true location, particularly in adverse conditions such as speech pauses, reverberation, or visual occlusions, leading to missed detections or spurious estimates. In this paper, we propose a novel speaker tracking method based on the Generalized Labelled Multi-Bernoulli (GLMB) filter. Our method operates in 3D space using audio information captured by a microphone array and video streams obtained from a monocular camera. The GLMB-based tracker effectively handles outliers in location estimates and maintains tracking during periods of missed detections. Experiments conducted on the publicly available AV16.3 dataset show that our proposal surpasses other competitive methods with improved results. Xinyuan Qian 0001, Zexu Pan, Qiquan Zhang, Kainan Chen, Shoufeng Lin |
ICASSP | 1 |
| 2024 | Transmitted and Aggregated Self-Attention for Automatic Speech Recognition
Tian-Hao Zhang, Xinyuan Qian 0001, Feng Chen 0040, Xu-Cheng Yin |
INTERSPEECH | 2 |
| 2024 | An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2024 | ListenFormer: Responsive Listening Head Generation with Non-autoregressive TransformersabstractAs one of the crucial elements in human-robot interaction, responsive listening head generation has attracted considerable attention from researchers. It aims to generate a listening head video based on speaker's audio and video as well as a reference listener image. However, existing methods exhibit two limitations: 1) the generation capability of their models is limited, resulting in generated videos that are far from real ones, and 2) they mostly employ autoregressive generative models, unable to mitigate the risk of error accumulation. To tackle these issues, we propose Listenformer that leverages the powerful temporal modeling capability of transformers for generation. It can perform non-autoregressive prediction with the proposed two-stage training method, simultaneously achieving temporal continuity and overall consistency in the outputs. To fully utilize the information from the speaker inputs, we designed an audio-motion attention fusion module, which improves the correlation of audio and motion features for accurate response. Additionally, a novel decoding method called sliding window with a large shift is proposed for Listenformer, demonstrating both excellent computational efficiency and effectiveness. Extensive experiments show that Listenformer outperforms the existing state-of-the-art methods on ViCo and L2L datasets. And a perceptual user study demonstrates the comprehensive performance of our method in generating diversity, identity preserving, speaker-listener synchronization, and attitude matching. Our code is available at https://liushenme.github.io/ListenFormer.github.io/. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 3 |
| 2024 | MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental TasksabstractClass-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy. Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 7 |
| 2024 | M3TTS: Multi-modal text-to-speech of multi-scale style control for dubbing
Li-Fang Wei, Xinyuan Qian 0001, Tian-Hao Zhang, Song-Lu Chen, Xu-Cheng Yin |
Pattern Recognit. Lett. | 3 |
| 2024 | Audio-Visual Temporal Forgery Detection Using Embedding-Level Fusion and Multi-Dimensional Contrastive LossabstractAudio-visual deepfake detection is the process of identifying and detecting deepfakes that have been generated using both audio and visual content with AI algorithms. Most existing methods primarily focus on the overall authenticity while neglecting the position of forgeries in time. This can be particularly problematic, as even a small alteration in a clip can significantly impact its meaning. Such brand new attacks are dangerous and how to tackle such attacks remains an open question. In this paper, we present a novel neural network-based model to tackle the temporal forgery detection (TFD) problem. It consists of new audio and visual encoders with cross-modal attention for embedding extraction, and an embedding-level fusion mechanism with self-attention for forgery localization. Besides, a multi-dimensional contrastive loss is proposed which helps the model not only to capture audio-visual inconsistency for deepfake detection but also to exploit temporal inconsistency by coherently constraining the extracted embeddings. Extensive experiments on the LAV-DF dataset show that the presented method outperforms several state-of-the-art temporal forgery localization methods by up to 23.4% on [email protected] and 13.8% on AR@100. In addition, we also show the effectiveness of the proposed model on deepfake detection. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Deep Cross-Modal Retrieval Between Spatial Image and Acoustic SpeechabstractCross-modal Retrieval (CMR) is formulated for the scenarios where the queries and retrieval results are of different modalities. Existing Cross-modal Retrieval (CMR) studies mainly focus on the common contextualized information between text transcripts and images, and the synchronized event information in audio-visual recordings. Unlike all previous works, in this article, we investigate the geometric correspondence between images and speech recordings captured in the same space and formulate a novel CMR task, called Spatial Image-Acoustic Retrieval (SIAR). To this end, we first design a novel speech encoder that consists of convolution neural networks and transformer layers, to learn space-aware speech representations. Then, to eliminate the cross-modal inherent discrepancy, we propose the Contrastive Speech Image Retrieval (CSIR) method which uses supervised contrastive learning to attract the same-space cross-modal features while repelling the ones from different spaces. Finally, image and speech features are directly compared and we predict the SIAR result with the maximum similarity. Extensive experiments demonstrate that our proposed speech encoder can recognize space from human speeches with superior performance over the other prevailing networks. It also sets our penultimate goal of speech-to-speech retrieval. Furthermore, our CSIR proposal can successfully perform bi-directional SIAR between spatial images and reverberant speeches with promising results. Code and data will be available. Xinyuan Qian 0001, Wei Xue 0002, Qiquan Zhang, Ruijie Tao, Haizhou Li 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertabstractTalking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip movements i.e., the visual intelligibility of the spoken words, which is an important aspect of generation quality. To address the problem, we propose using a lipreading expert to improve the intelligibility of the generated lip regions by penalizing the incorrect generation results. Moreover, to compensate for data scarcity, we train the lip-reading expert in an audio-visual self-supervised manner. With a lip-reading expert, we propose a novel contrastive learning to enhance lip-speech synchronization, and a transformer to encode audio synchronically with video, while considering global temporal dependency of audio. For evaluation, we propose a new strategy with two different lip-reading experts to measure intelligibility of the generated videos. Rigorous experiments show that our proposal is superior to other State-of-the-art (SOTA) methods, such as Wav2Lip, in reading intelligibility i.e., over 38% Word Error Rate (WER) on LRS2 dataset and 27.8% accuracy on LRW dataset. We also achieve the SOTA performance in lip-speech synchronization and comparable performances in visual quality. Xinyuan Qian 0001, Malu Zhang, Robby T. Tan, Haizhou Li 0001 |
CVPR | 2 |
| 2023 | Stream Attention Based U-Net for L3DAS23 ChallengeabstractMachine learning applications of 3D audio are gaining increasing interest in recent years. In this paper, we propose a stream attention based U-Net to remove background noise and reverberation based on ICASSP Signal Processing Grand Challenge 2023: L3DAS23 Challenge1Audio-only track task1 3D Speech Enhancement. Results show that proposed method achieves superior performance than the official baseline model. Honglong Wang, Yanjie Fu, Meng Ge, Longbiao Wang, Xinyuan Qian 0001 |
ICASSP | 6 |
| 2023 | Self-Convolution for Automatic Speech RecognitionabstractSelf-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effective and superior in learning local information. Whereas it fails in self-interaction and capturing long-range dependence among input tokens. Accordingly, we take their complementary advantages and propose a new module, namely self-convolution, to compensate for each individual limitations. Specifically, self-convolution generates convolution kernels at each token (to model local information) which are then used to convolve itself (for self-interaction). Moreover, we bring in global information during the generation of convolution kernel to enhance the learning of long-range dependencies. In this way, the advantages of self-attention and CNN are both utilized. We conduct rigorous experiments on LibriSpeech, Tedlium2, and AIShell1 datasets and demonstrate that our proposed self-convolution can achieve superior ASR performance than self-attention with less computational cost. Qi Liu 0041, Xinyuan Qian 0001, Song-Lu Chen, Feng Chen 0040, Xu-Cheng Yin |
ICASSP | 3 |
| 2023 | Ripple Sparse Self-Attention for Monaural Speech EnhancementabstractThe use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the strong local correlations of speech signals. This study presents a simple yet effective sparse self-attention for speech enhancement, called ripple attention, which simultaneously performs fine- and coarse-grained modeling for local and global dependencies, respectively. Specifically, we employ local band attention to enable each frame to attend to its closest neighbor frames in a window at fine granularity, while employing dilated attention outside the window to model the global dependencies at a coarse granularity. We evaluate the efficacy of our ripple attention for speech enhancement on two commonly used training objectives. Extensive experimental results consistently confirm the superior performance of the ripple attention design over standard full self-attention, blockwise attention, and dual-path attention (Sep-Former) in terms of speech quality and intelligibility. Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 4 |
| 2023 | A Miniaturised Camera-based Multi-Modal Tactile SensorabstractIn conjunction with huge recent progress in cam-era and computer vision technology, camera-based sensors have increasingly shown considerable promise in relation to tactile sensing. In comparison to competing technologies (be they resistive, capacitive or magnetic based), they offer super-high-resolution, while suffering from fewer wiring problems. The human tactile system is composed of various types of mechanoreceptors, each able to perceive and process distinct information such as force, pressure, texture, etc. Camera-based tactile sensors such as GelSight mainly focus on high-resolution geometric sensing on a flat surface, and their force measurement capabilities are limited by the hysteresis and non-linearity of the silicone material. In this paper, we present a miniaturised dome-shaped camera-based tactile sensor that allows accurate force and tactile sensing in a single coherent system. The key novelty of the sensor design is as follows. First, we demonstrate how to build a smooth silicone hemispheric sensing medium with uniform markers on its curved surface. Second, we enhance the illumination of the rounded silicone with diffused LEDs. Third, we construct a force-sensitive mechanical structure in a compact form factor with usage of springs to accurately perceive forces. Our multi-modal sensor is able to acquire tactile information from multi-axis forces, local force distribution, and contact geometry, all in real-time. We apply an end-to-end deep learning method to process all the information. Kaspar Althoefer, Yonggen Ling, Wanlin Li, Xinyuan Qian 0001, Wang Wei Lee, Peng Qi 0001 |
ICRA | 4 |
| 2023 | InterFormer: Interactive Local and Global Features Fusion for Automatic Speech Recognition
Zhi-Hao Lai, Tian-Hao Zhang, Qi Liu 0041, Xinyuan Qian 0001, Li-Fang Wei, Feng Chen 0040, Song-Lu Chen, Xu-Cheng Yin |
INTERSPEECH | 4 |
| 2023 | Rethinking Speech Recognition with A Multimodal Perspective via Acoustic and Semantic Cooperative Decoding
Tian-Hao Zhang, Haibo Qin, Zhi-Hao Lai, Song-Lu Chen, Qi Liu 0041, Feng Chen 0040, Xinyuan Qian 0001, Xu-Cheng Yin |
INTERSPEECH | 7 |
| 2023 | Audio-Visual Cross-Attention Network for Robotic Speaker TrackingabstractAudio-visual signals can be used jointly for robotic perception as they complement each other. Such multi-modal sensory fusion has a clear advantage, especially under noisy acoustic conditions. Speaker localization, as an essential robotic function, was traditionally solved as a signal processing problem that now increasingly finds deep learning solutions. The question is how to fuse audio-visual signals in an effective way. Speaker tracking is not only more desirable, but also potentially more accurate than speaker localization because it explores the speaker's temporal motion dynamics for smoothed trajectory estimation. However, due to the lack of large annotated dataset, speaker tracking is not well studied as speaker localization. In this paper, we study robotic speaker Direction of Arrival (DoA) estimation with a focus on audio-visual fusion and tracking methodology. We propose a Cross-Modal Attentive Fusion (CMAF) mechanism, which explores self-attention to learn intra-modal temporal dependencies, and cross-attention mechanism for inter-modal alignment. We also collect a realistic dataset on a robotic platform to support the study. The experimental results demonstrate that our proposed network outperforms the state-of-the-art audio-visual localization and tracking methods under noisy conditions, with an improved accuracy of 5.82% and 3.62% at SNR = −20 dB, respectively. Xinyuan Qian 0001, Zhengdong Wang, Guohui Guan 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Device Features Based on Linear Transformation With Parallel Training Data for Replay Speech DetectionabstractReplay speech poses a growing threat to speaker verification systems, thus the detection of replay speech becomes increasingly important. A critical factor differentiating replay speech and genuine speech is the representation of device information. Replay speech carries physical device information that originates from recording device, playback device, and environmental noise. In this work, a device-related linear transformation strategy is proposed to disentangle non-device information from replay speech. First, we conduct factor analysis by introducing a common vector for both replay utterance and the corresponding genuine speech utterance on parallel training data; then, we derive an expectation maximization formula to obtain the parameters of the device-related linear transformation; subsequently, three device feature extraction methods are developed based on the device-related linear transformation. The developed device features are evaluated on ASVspoof 2017 version 2.0 and ASVspoof 2021 physical access corpora. The experimental results demonstrate that our proposed linear transformation strategy is effective for replay spoofing detection, and the resultant device features outperform many typical features. Moreover, our spoofing detection systems display superior performance over several competitive state-of-the-art systems. Longting Xu, Chang Huai You, Xinyuan Qian 0001, Daiyu Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | A Time-Frequency Attention Module for Neural Speech EnhancementabstractSpeech enhancement plays an essential role in a wide range of speech processing applications. Recent studies on speech enhancement tend to investigate how to effectively capture the long-term contextual dependencies of speech signals to boost performance. However, these studies generally neglect the time-frequency (T-F) distribution information of speech spectral components, which is equally important for speech enhancement. In this paper, we propose a simple yet very effective network module, which we term the T-F attention (TFA) module, that uses two parallel attention branches, i.e., time-frame attention and frequency-channel attention, to explicitly exploit position information to generate a 2-D attention map to characterise the salient T-F speech distribution. We validate our TFA module as part of two widely used backbone networks (residual temporal convolution network and Transformer) and conduct speech enhancement with four most popular training objectives. Our extensive experiments demonstrate that our proposed TFA module consistently leads to substantial enhancement performance improvements in terms of the five most widely used objective metrics, with negligible parameter overheads. In addition, we further evaluate the efficacy of speech enhancement as a front-end for a downstream speech recognition task. Our evaluation results show that the TFA module significantly improves the robustness of the system to noisy conditions. Qiquan Zhang, Xinyuan Qian 0001, Zhaoheng Ni, Aaron Nicolson, Eliathamby Ambikairajah, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Speech-Oriented Sparse Attention Denoising for Voice User Interface Toward Industry 5.0abstractThe adoption of voice user interface (VUI) will promote network automation with enhanced efficiency with reduced simplicity and operating expense in Industry 5.0. Given the noisy environments, speech denoising is indispensable for the VUI in Internet of Things (IoT) or Industrial IoT (IIoT). Despite Transformer's recent success in speech denoising, the adopted full self-attention suffers from quadratic complexity, which challenges the computational power of the IoT/IIoT components. Considering the strong local correlations of speech signals, a speech-oriented sparse attention denoising scheme is developed to keep the meaningful local and global dependencies while mitigating the redundant attentions, resulting in a significant reduction in computational complexity. With the full self-attention as the baseline, experimental results revealed that the proposed scheme achieves a better denoising performance and yields a lower computational cost, indicating the strong potential for various VUI application scenarios in IoT and IIoT toward Industry 5.0. Hongxu Zhu, Qiquan Zhang, Peng Gao 0005, Xinyuan Qian 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Iterative Sound Source Localization for Unknown Number of SourcesabstractSound source localization aims to seek the direction of arrival (DOA) of all sound sources from the observed multichannel audio.For the practical problem of unknown number of sources, existing localization algorithms attempt to predict a likelihood-based coding (i.e., spatial spectrum) and employ a pre-determined threshold to detect the source number and corresponding DOA value.However, these threshold-based algorithms are not stable since they are limited by the careful choice of threshold.To address this problem, we propose an iterative sound source localization approach called ISSL, which can iteratively extract each source's DOA without threshold until the termination criterion is met.Unlike threshold-based algorithms, ISSL designs an active source detector network based on binary classifier to accept residual spatial spectrum and decide whether to stop the iteration.By doing so, our ISSL can deal with an arbitrary number of sources, even more than the number of sources seen during the training stage.The experimental results show that our ISSL achieves significant performance improvements in both DOA estimation and source number detection compared with the existing threshold-based algorithms. Yanjie Fu, Meng Ge, Xinyuan Qian 0001, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2022 | Speaker Extraction With Co-Speech Gestures CueabstractSpeaker extraction seeks to extract the clean speech of a target speaker from a multi-talker mixture speech. There have been studies to use a pre-recorded speech sample or face image of the target speaker as the speaker cue. In human communication, co-speech gestures that are naturally timed with speech also contribute to speech perception. In this work, we explore the use of co-speech gestures sequence, e.g. hand and body movements, as the speaker cue for speaker extraction, which could be easily obtained from low-resolution video recordings, thus more available than face recordings. We propose two networks using the co-speech gestures cue to perform attentive listening on the target speaker, one that implicitly fuses the co-speech gestures cue in the speaker extraction process, the other performs speech separation first, followed by explicitly using the co-speech gestures cue to associate a separated speech to the target speaker. The experimental results show that the co-speech gestures cue is informative in associating with the target speaker. Zexu Pan, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Deep Audio-Visual Beamforming for Speaker LocalizationabstractGeneralized Cross Correlation (GCC) is the most popular localization technique over the past decades and can be extended with the beamforming method e.g. Steered Response Power (SRP) when multiple microphone pairs exist. Considering the promising results of Deep Learning (DL) strategies over classical approaches, in this work, instead of directly using Generalized Cross Correlation (GCC), SRP is derived with the DL-learnt ideal correlation functions for each pair of a microphone array. To deploy visual information, we explore the Conditional Variational Auto-Encoder (CVAE) framework in which the audio generative process is conditioned on the visual features encoded by face detections. The vision-derived auxiliary correlation function eventually contributes to the back-end beamformer for improved localization performance. To the best of our knowledge, this is the first deep-generative audiovisual method for speaker localization. Experimental results demonstrate our superior performance over other competitive methods, especially when the speech signal is corrupted by noise. Xinyuan Qian 0001, Qiquan Zhang, Guohui Guan 0001, Wei Xue 0002 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Audio-Visual Tracking of Concurrent SpeakersabstractAudio-visual tracking of an unknown number of concurrent speakers in 3D is a challenging task, especially when sound and video are collected with a compact sensing platform. In this paper, we propose a tracker that builds on generative and discriminative audio-visual likelihood models formulated in a particle filtering framework. We localize multiple concurrent speakers with a de-emphasized acoustic map assisted by the image detection-derived 3D video observations. The 3D multi-modal observations are either assigned to existing tracks for discriminative likelihood computation or used to initialize new tracks. The generative likelihoods rely on color distribution of the target and the de-emphasized acoustic map value. Experiments on AV16.3 and CAV3D datasets show that the proposed tracker outperforms the uni-modal trackers and the state-of-the-art approaches both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 1 |
| 2021 | Multi-Target DoA Estimation with an Audio-Visual Fusion MechanismabstractMost of the prior studies in the spatial Direction of Arrival (DoA) domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio and visual signals for multi-speaker localization. The use of heterogeneous sensors can provide complementary information to overcome uni-modal challenges, such as noise, reverberation, illumination variations, and occlusions. We attempt to address these issues by introducing an adaptive weighting mechanism for audio-visual fusion. We also propose a novel video simulation method that generates visual features from noisy target 3D annotations that are synchronized with acoustic features. Experimental results confirm that audio-visual fusion consistently improves the performance of speaker DoA estimation, while the adaptive weighting mechanism shows clear benefits. Xinyuan Qian 0001, Maulik C. Madhavi, Zexu Pan, Haizhou Li 0001 |
ICASSP | 1 |
| 2021 | GCC-PHAT with Speech-oriented Attention for Robotic Sound Source LocalizationabstractRobotic audition is a basic sense that helps robots perceive the surroundings and interact with humans. Sound Source Localization (SSL) is an essential module for a robotic system. However, the performance of most sound source localization techniques degrades in noisy and reverberant environments due to inaccurate Time Difference of Arrival (TDoA) estimation. In robotic sound source localization, we are more interested in detecting the arrival of human speech than other sound sources. Ideally, we expect an effective TDoA estimation to respond only to speech signals, while masking off other interferences. In this paper, we propose a novel technique that learns to attend to speech fundamental frequency and harmonics while suppressing noise interference and reverberation. The novel TDoA feature is referred to as Generalized Cross Correlation with Phase Transform and Speech Mask (GCC-PHAT-SM). We perform sound source localization experiments on real-world data captured from a robotic platform. Experiments show that GCC-PHAT-SM feature significantly outperforms traditional Generalized Cross Correlation (GCC) feature in noisy and reverberant acoustic environments. Xinyuan Qian 0001, Zihan Pan, Malu Zhang, Haizhou Li 0001 |
ICRA | 2 |
| 2021 | Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionabstractActive speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as well as audio-visual interaction. Unlike the prior work where systems make decision instantaneously using short-term features, we propose a novel framework, named TalkNet, that makes decision by taking both short-term and long-term features into consideration. TalkNet consists of audio and visual temporal encoders for feature representation, audio-visual cross-attention mechanism for inter-modality interaction, and a self-attention mechanism to capture long-term speaking evidence. The experiments demonstrate that TalkNet achieves 3.5% and 2.2% improvement over the state-of-the-art systems on the AVA-ActiveSpeaker dataset and Columbia ASD dataset, respectively. Code has been made available at: https://github.com/TaoRuijie/TalkNet_ASD. Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 0001, Zheng Shou 0001, Haizhou Li 0001 |
ACM Multimedia | 4 |
| 2021 | Three-Dimensional Speaker Localization: Audio-Refined Visual Scaling Factor EstimationabstractNeither a monocular RGB camera nor a small-size microphone array is capable of accurate three-dimensional (3D) speaker localization. By taking advantage of accurate visual object detection, and audio-visual complementary sensor fusion, we formulate the three-dimensional (3D) speaker localization problem as a visual scaling factor estimation problem. As a result, we effectively reduce the traditional audio-only 3D speaker localization from an exhaustive grid search to a one-dimensional (1D) optimization problem. We propose a multi-modal perception system with two optimization approaches. We show that the proposed methods are effective, accurate, and robust against interference and, as corroborated by indicative empirical results on real dataset, competitive to the conventional uni-modal and the state-of-the-art audio-visual speaker localization approaches. Xinyuan Qian 0001, Qi Liu 0005, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Audio-Visual Multi-Speaker Tracking Based on the GLMB Framework
Shoufeng Lin, Xinyuan Qian 0001 |
INTERSPEECH | 2 |
| 2019 | Accurate Target Annotation in 3D from Multimodal StreamsabstractAccurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverages annotations from reference streams (e.g. individual camera views) and measurements from unannotated additional streams (e.g. audio) to infer 3D trajectories through an optimization. The core of our approach is a multi-modal extension of Bundle Adjustment with a cross-modal correspondence detection that selectively uses measurements in the optimization. We apply the proposed approach to fully annotate a new multi-modal and multi-view dataset for multi-speaker 3D tracking. Oswald Lanz, Alessio Brutti, Alessio Xompero, Xinyuan Qian 0001, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 4 |
| 2019 | Multi-Speaker Tracking From an Audio-Visual Sensing DeviceabstractCompact multi-sensor platforms are portable and thus desirable for robotics and personal-assistance tasks. However, compared to physically distributed sensors, the size of these platforms makes person tracking more difficult. To address this challenge, we propose a novel 3-D audio-visual people tracker that exploits visual observations (object detections) to guide the acoustic processing by constraining the acoustic likelihood on the horizontal plane defined by the predicted height of a speaker. This solution allows the tracker to estimate, with a small microphone array, the distance of a sound. Moreover, we apply a color-based visual likelihood on the image plane to compensate for misdetections. Finally, we use a 3-D particle filter and greedy data association to combine visual observations, color-based, and acoustic likelihoods to track the position of multiple simultaneous speakers. We compare the proposed multimodal 3-D tracker against two state-of-the-art methods on the AV16.3 dataset and on a newly collected dataset with co-located sensors, which we make available to the research community. Experimental results show that our multimodal approach outperforms the other methods both in 3-D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 1 |
| 2018 | 3D Mouth Tracking from a Compact Microphone Array Co-Located with a cameraabstractWe address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood computation is assisted by video, which relies on a GCC-PHAT based acoustic map. By combining audio and video inputs, the proposed approach can cope with a reverberant and noisy environment, and can deal with situations when the person is occluded, outside the Field of View (FoV), or not facing the sensors. Experimental results show that the proposed tracker is accurate both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Xompero, Andrea Cavallaro, Alessio Brutti, Oswald Lanz, Maurizio Omologo |
ICASSP | 1 |
| 2017 | 3D audio-visual speaker tracking with an adaptive particle filterabstractWe propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter framework. The reliability of the audio signal is measured based on the maximum Global Coherence Field (GCF) peak value at each frame. The visual reliability is based on colour-histogram matching with detection results compared with a reference image in the RGB space. Experiments on the AV16.3 dataset show that the proposed adaptive audio-visual tracker outperforms both the individual modalities and a classical approach with fixed parameters in terms of tracking accuracy. Xinyuan Qian 0001, Alessio Brutti, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 1 |