VLDB 2026 Research / reviewers in the wild / expert
Longting Xu
dblp:173/6633
· DBLP profile ↗
14ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-2329-895XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-modal Speech Enhancement with Limited Electromyography ChannelsabstractSpeech enhancement (SE) aims to improve the clarity, intelligibility, and quality of speech signals for various speech enabled applications. However, air-conducted (AC) speech is highly susceptible to ambient noise, particularly in low signal-to-noise ratio (SNR) and non-stationary noise environments. Incorporating multi-modal information has shown promise in enhancing speech in such challenging scenarios. Electromyography (EMG) signals, which capture muscle activity during speech production, offer noise-resistant properties beneficial for SE in adverse conditions. Most previous EMG-based SE methods required 35 EMG channels, limiting their practicality. To address this, we propose a novel method that considers only 8-channel EMG signals with acoustic signals using a modified SEMamba network with added cross-modality modules. Our experiments demonstrate substantial improvements in speech quality and intelligibility over traditional approaches, especially in extremely low SNR settings. Notably, compared to the SE (AC) approach, our method achieves a significant PESQ gain of 0.235 under matched low SNR conditions and 0.527 under mismatched conditions, highlighting its robustness. Fuyuan Feng, Longting Xu, Rohan Kumar Das |
ICASSP | 2 |
| 2025 | A Lightweight Vision-Based Deep Learning Model for Fall DetectionabstractFall detection is crucial to ensure the safety and health of elderly people at home. Vision-based deep learning methods, known for their ease of deployment and non-intrusiveness, have gradually replaced traditional wearable device-based approaches. However, challenges such as high model complexity and significant computational demands remain unresolved. To address these issues, we propose a lightweight fall detection model centered around the LSE Block and the Local Feature Pyramid Network (LFPN). The LSE Block, which forms the Backbone's core, consists of the LightweightConv module, SS2D, and ECA module, enabling efficient capture of both local and global features while reducing complexity. The LFPN design further simplifies the network architecture, achieving high detection accuracy with lower computational requirements. Experiments on the Le2i and FDD datasets demonstrate that this lightweight model reduces the parameters and Floating Point Operations (FLOPs) by 62% and 63%, with a slight improvement in detection accuracy. Code available at: https://github.com/DHUspeech/Lightweight. Fuyuan Feng, Zhangcheng Yang, Longting Xu |
IJCNN | 6 |
| 2025 | Fall-Mamba: A Multimodal Fusion and Masked Mamba-Based Approach for Fall DetectionabstractFalls are a leading cause of injury and death among the elderly, making fall detection critically important. Traditional wearable sensors and environmental devices have limitations in terms of comfort, convenience, and accuracy. With the advancement of artificial intelligence and the Internet of Things (IoT), camera-based fall detection has become a research focus, but challenges such as occlusion and poor lighting conditions remain. To address these issues, this study introduces an innovative model named Fall-Mamba. Compared to previous approaches, Fall-Mamba utilizes a Cross-Attention mechanism to fuse video and audio data, significantly enhancing the comprehensive understanding and detection performance of fall events. Additionally, the model incorporates a Multi-Head Temporal Attention mechanism and Frame Masking strategy, improving its ability to capture key frames and increasing its robustness. Extensive experiments conducted on multi-view, multi-scene datasets, including Le2i, URFD, and Multicam, demonstrate the superior performance of Fall-Mamba, achieving an accuracy of 99.63% and exhibiting high robustness. This technology provides strong protection for the safety of the elderly in IoT-enabled smart homes. The code has been published at https://github.com/DHUspeech/fall-mamba. Qicheng Xu, Fuyuan Feng, Xiaochen Lu, Longting Xu |
IEEE Internet Things J. | 5 |
| 2025 | Robust and Imperceptible Watermarking Framework for Generative Audio ModelsabstractThe rapid development of generative audio models has raised concerns about copyright protection and traceability. To tackle these challenges, we first propose a robust and imperceptible watermarking framework embedded directly into the generative process. Our method embeds watermarks into the convolutional layers of the generative model, allowing the synthesis of watermarked audio without compromising quality. A trainable binary mask selectively modulates kernel weights, ensuring precise and efficient embedding. The framework incorporates a dedicated encoder-decoder architecture for accurate watermark embedding and extraction. A normalization step aligns the modified kernel weights with the original statistical properties to preserve the model's performance. Experimental results across multiple datasets demonstrate the method's robustness against common audio attacks. Additionally, the approach achieves high audio fidelity and near-perfect watermark recovery, offering a practical solution for traceable and secure audio synthesis. Fuyuan Feng, Guanglin Zhang, Longting Xu |
IEEE Signal Process. Lett. | 5 |
| 2025 | A Robust Coverless Audio Steganography Based on Differential Privacy ClusteringabstractConventional audio steganography methods typically require embedding secret information into the carrier, making them vulnerable to steganalysis. To address this issue, we propose a novel coverless audio steganography method that hides information by generating carriers and establishing mapping rules rather than embedding data directly. Our approach leverages a differential privacy clustering algorithm to cluster audio data and select representative audio files, thereby enhancing the security of the steganography. Additionally, we introduce an improved audio feature extraction method that combines traditional Mel-frequency cepstral coefficients (MFCC) with global statistical information, significantly boosting the robustness of the secret information against common audio attacks, particularly time-stretching attacks. Experimental results show that our method achieves a robustness rate of up to 95% against time-stretching and maintains an average security accuracy rate exceeding 97% across various attack scenarios. The proposed method ensures that the audio carrier remains unaltered, thus effectively resisting detection by steganalysis tools. This innovative approach provides a practical and efficient solution for the secure transmission of information in the digital era. Longting Xu, Xiaochen Lu, Guanglin Zhang, Wei Rao 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | Device Features Based on Linear Transformation With Parallel Training Data for Replay Speech DetectionabstractReplay speech poses a growing threat to speaker verification systems, thus the detection of replay speech becomes increasingly important. A critical factor differentiating replay speech and genuine speech is the representation of device information. Replay speech carries physical device information that originates from recording device, playback device, and environmental noise. In this work, a device-related linear transformation strategy is proposed to disentangle non-device information from replay speech. First, we conduct factor analysis by introducing a common vector for both replay utterance and the corresponding genuine speech utterance on parallel training data; then, we derive an expectation maximization formula to obtain the parameters of the device-related linear transformation; subsequently, three device feature extraction methods are developed based on the device-related linear transformation. The developed device features are evaluated on ASVspoof 2017 version 2.0 and ASVspoof 2021 physical access corpora. The experimental results demonstrate that our proposed linear transformation strategy is effective for replay spoofing detection, and the resultant device features outperform many typical features. Moreover, our spoofing detection systems display superior performance over several competitive state-of-the-art systems. Longting Xu, Chang Huai You, Xinyuan Qian 0001, Daiyu Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Graph Fourier Transform Based Audio Zero-WatermarkingabstractThe frequent exchange of multimedia information in the present era projects an increasing demand for copyright protection. In this work, we propose a novel audio zerowatermarking technology based on graph Fourier transform for enhancing the robustness with respect to copyright protection. In this approach, the combined shift operator is used to construct the graph signal, upon which the graph Fourier analysis is performed. The selected maximum absolute graph Fourier coefficients representing the characteristics of the audio segment are then encoded into a feature binary sequence using K-means algorithm. Finally, the resultant feature binary sequence is XORed with the watermark binary sequence to realize the embedding of the zero-watermark. The experimental studies show that the proposed approach performs more effectively in resisting common or synchronization attacks than the existing state ofthe-art methods. Longting Xu, Daiyu Huang, Syed Faham Ali Zaidi, Abdul Rauf 0004, Rohan Kumar Das |
IEEE Signal Process. Lett. | 1 |
| 2021 | Cascaded Convolutional Neural Network-Based Hyperspectral Image Resolution Enhancement via an Auxiliary Panchromatic ImageabstractOwing to the limits of incident energy and hardware system, hyperspectral (HS) images always suffer from low spatial resolution, compared with multispectral (MS) or panchromatic (PAN) images. Therefore, image fusion has emerged as a useful technology that is able to combine the characteristics of high spectral and spatial resolutions of HS and PAN/MS images. In this paper, a novel HS and PAN image fusion method based on convolutional neural network (CNN) is proposed. The proposed method incorporates the ideas of both hyper-sharpening and MS pan-sharpening techniques, thereby employing a two-stage cascaded CNN to reconstruct the anticipated high-resolution HS image. Technically, the proposed CNN architecture consists of two sub-networks, the detail injection sub-network and unmixing sub-network. The former aims at producing a latent high-resolution MS image, whereas the latter estimates the desired high-resolution abundance maps by exploring the spatial and spectral information of both HS and MS images. Moreover, two model-training fashions are presented in this paper for the sake of effectively training our network. Experiments on simulated and real remote sensing data demonstrate that the proposed method can improve the spatial resolution and spectral fidelity of HS image, and achieve better performance than some state-of-the-art HS pan-sharpening algorithms. Xiaochen Lu, Junping Zhang, Dezheng Yang, Longting Xu, Fengde Jia |
IEEE Trans. Image Process. | 4 |
| 2020 | Adversarial Dictionary Learning for Monaural Speech Enhancement
Yunyun Ji, Longting Xu, Wei-Ping Zhu 0001 |
INTERSPEECH | 2 |
| 2018 | Co-whitening of I-vectors for Short and Long Duration Speaker Verification
Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
INTERSPEECH | 1 |
| 2018 | Generative X-Vectors for Text-Independent Speaker VerificationabstractSpeaker verification (SV) systems using deep neural network embeddings, so-called the x-vector systems, are becoming popular due to its good performance superior to the i-vector systems. The fusion of these systems provides improved performance benefiting both from the discriminatively trained x-vectors and generative i-vectors capturing distinct speaker characteristics. In this paper, we propose a novel method to include the complementary information of i-vector and x-vector, that is called generative x-vector. The generative x-vector utilizes a transformation model learned from the i-vector and x-vector representations of the background data. Canonical correlation analysis is applied to derive this transformation model, which is later used to transform the standard x-vectors of the enrollment and test segments to the corresponding generative x-vectors. The SV experiments performed on the NIST SRE 2010 dataset demonstrate that the system using generative x-vectors provides considerably better performance than the baseline i-vector and x-vector systems. Furthermore, the generative x-vectors outperform the fusion of i-vector and x-vector systems for long-duration utterances, while yielding comparable results for short-duration utterances. Longting Xu, Rohan Kumar Das, Emre Yilmaz 0001, Haizhou Li 0001 |
SLT | 1 |
| 2018 | Generalizing I-Vector Estimation for Rapid Speaker RecognitionabstractAn i-vector is a compact representation that captures both the speaker and session variabilities rendered in a spoken utterance. Over the past years, it has prevailed over other techniques and is now the de facto representation for text-independent speaker recognition. Standard i-vector extraction requires intense computation at run-time. Reducing the computation will allow effective use of i-vector in more applications. Such intense computation arises from the posterior covariance matrix, when estimating the i-vector. There have been studies on how to simplify the computation of posterior covariance matrix with modest success. In this paper, we propose a novel approach to i-vector extraction without the need to evaluate the full posterior covariance thereby speeding up the run-time extraction process. This is achieved by generalizing the i-vector estimation in two ways. First, we introduce the use of occupancy reweighting in conjunction with whitening over the Baum-Welch statistics as part of the preprocessing step. Second, we introduce the so-called subspace-orthogonalizing prior (SOP) to replace the standard Gaussian prior in i-vector formulation. Experiments conducted on the extended-core task of NIST SRE'10 show that the proposed rapid SOP approach achieves considerable speed-up over the standard i-vector with comparable equal error rates. Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Sparse coding of total variability matrix
Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
INTERSPEECH | 1 |
| 2013 | Speaker identification based on sparse subspace modelabstractMel Frequency Cepstrum Coefficient(MFCC) has been proven extremely successful for text-independent speaker identification. We address the speaker identification problem by presenting a novel Sparse Representation-Subspace algorithm. We propose to develop an overcomplete dictionary of each speaker using the Mel filterbank log energies for all the training utterances. We therefore propose to represent learned dictionary as a linear combination of all the log energies, thereby generating a naturally sparse representation, which is the novel subspace of the speaker. Besides, DCT step of MFCC is a fixed matrix, learned dictionary for different speaker is more adaptive. In the identification process, the unknown vectors of Mel filterbank log energies coefficients are projected into each subspace to decide the matching speaker. Experiments have been conducted on the speech database in our anechoic chamber, and a comparison with MFCC based speaker identification algorithms yields a favorable performance index for the proposed algorithm. Different sparsity and dictionary size shows different results. Longting Xu, Zhen Yang 0001 |
APCC | 1 |