VLDB 2026 Research / reviewers in the wild / expert
Xu Xiang
dblp:41/3059
· DBLP profile ↗
16ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KeBugFix: Automated Program Repair Framework Based on Code Retrieval Enhancement and LLM Agent
Chaopeng Wang, Yi Sun 0006, Chao Wang 0061, Zan Zhou 0001, Yujiao Yuan, Xu Xiang, Fei Xiao 0005 |
IEEE Big Data | 9 |
| 2025 | Attributed graph clustering with multi-scale weight-based pairwise coarsening and contrastive learning
Binxiong Li, Yuefei Wang, Binyu Zhao 0002, Heyang Gao, Benhan Yang, Quanzhou Luo, Xu Xiang, Huijie Tang |
Neurocomputing | 8 |
| 2024 | Advancing speaker embedding learning: Wespeaker toolkit for research and production
Shuai Wang 0016, Zhengyang Chen, Bing Han 0008, Chengdong Liang, Xu Xiang, Wen Ding 0005, Johan Rohdin, Anna Silnova, Yanmin Qian, Haizhou Li 0001 |
Speech Commun. | 7 |
| 2023 | Wespeaker: A Research and Production Oriented Speaker Embedding Learning ToolkitabstractSpeaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit, Wespeaker. Wespeaker contains the implementation of scalable data management, state-of-the-art speaker embedding models, loss functions, and scoring back-ends, with highly competitive results achieved by structured recipes which were adopted in the winning systems in several speaker verification challenges. The application to other downstream tasks such as speaker diarization is also exhibited in the related recipe. Moreover, CPU- and GPU-compatible deployment codes are integrated for production-oriented development. The toolkit is publicly available at https://github.com/wenet-e2e/wespeaker. Chengdong Liang, Shuai Wang 0016, Zhengyang Chen, Xu Xiang, Yanlei Deng, Yanmin Qian |
ICASSP | 6 |
| 2023 | Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022
Zhengyang Chen, Bing Han 0008, Xu Xiang, Houjun Huang, Bei Liu 0003, Yanmin Qian |
INTERSPEECH | 3 |
| 2023 | Stable local interpretable model-agnostic explanations based on a variational autoencoder
Xu Xiang, Hong Yu 0007, Ye Wang 0006, Guoyin Wang 0001 |
Appl. Intell. | 1 |
| 2022 | MSDWild: Multi-modal Speaker Diarization Dataset in the Wild
Tao Liu 0068, Shuai Fan 0005, Xu Xiang, Shaoxiong Lin, Tianyuan Han, Binwei Yao, Yanmin Qian, Kai Yu 0004 |
INTERSPEECH | 3 |
| 2021 | AISpeech-SJTU Accent Identification System for the Accented English Speech Recognition ChallengeabstractThis paper describes the AISpeech-SJTU system for the accent identification track of the Interspeech-2020 Accented English Speech Recognition Challenge. In this challenge track, only 160-hour accented English data collected from 8 countries and the auxiliary Librispeech dataset are provided for training. To build an accurate and robust accent identification system, we explore the whole system pipeline in detail. First, we introduce the ASR based phone posteriorgram (PPG) feature to accent identification and verify its efficacy. Then, a novel TTS based approach is carefully designed to augment the very limited accent training data for the first time. Finally, we propose the test time augmentation and embedding fusion schemes to further improve the system performance. Our final system is ranked first in the challenge and outperforms all the other participants by a large margin. The submitted system achieves 83.63% average accuracy on the challenge evaluation data, ahead of the others by more than 10% in absolute terms. Houjun Huang, Xu Xiang, Yexin Yang, Rao Ma, Yanmin Qian |
ICASSP | 2 |
| 2021 | Unit Selection Synthesis Based Data Augmentation for Fixed Phrase Speaker VerificationabstractData augmentation is commonly used to help build a robust speaker verification system, especially in limited-resource case. However, conventional data augmentation methods usually focus on the diversity of acoustic environment, leaving the lexicon variation neglected. For text dependent speaker verification tasks, it’s well-known that preparing training data with the target transcript is the most effectual approach to build a well-performing system, however collecting such data is time-consuming and expensive. In this work, we propose a unit selection synthesis based data augmentation method to leverage the abundant text-independent data resources. In this approach text-independent speeches of each speaker are firstly broke up to speech segments each contains one phone unit. Then segments that contain phonetics in the target transcript are selected to produce a speech with the target transcript by concatenating them in turn. Experiments are carried out on the AISHELL Speaker Verification Challenge 2019 database, the results and analysis shows that our proposed method can boost the system performance significantly. Houjun Huang, Xu Xiang, Shuai Wang 0016, Yanmin Qian |
ICASSP | 2 |
| 2019 | Binary neural networks for speech recognitionabstractRecently, deep neural networks (DNNs) significantly outperform Gaussian mixture models in acoustic modeling for speech recognition. However, the substantial increase in computational load during the inference stage makes deep models difficult to directly deploy on low-power embedded devices. To alleviate this issue, structure sparseness and low precision fixed-point quantization have been applied widely. In this work, binary neural networks for speech recognition are developed to reduce the computational cost during the inference stage. A fast implementation of binary matrix multiplication is introduced. On modern central processing unit (CPU) and graphics processing unit (GPU) architectures, a 5–7 times speedup compared with full precision floatingpoint matrix multiplication can be achieved in real applications. Several kinds of binary neural networks and related model optimization algorithms are developed for large vocabulary continuous speech recognition acoustic modeling. In addition, to improve the accuracy of binary models, knowledge distillation from the normal full precision floating-point model to the compressed binary model is explored. Experiments on the standard Switchboard speech recognition task show that the proposed binary neural networks can deliver 3–4 times speedup over the normal full precision deep models. With the knowledge distillation from the normal floating-point models, the binary DNNs or binary convolutional neural networks (CNNs) can restrict the word error rate (WER) degradation to within 15.0%, compared to the normal full precision floating-point DNNs or CNNs, respectively. Particularly for the binary CNN with binarization only on the convolutional layers, the WER degradation is very small and is almost negligible with the proposed approach. Yanmin Qian, Xu Xiang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | Binary Deep Neural Networks for Speech Recognition
Xu Xiang, Yanmin Qian, Kai Yu 0004 |
INTERSPEECH | 1 |
| 2015 | Recurrent neural network language model with structured word embeddings for speech recognitionabstractDue to effective word context encoding and long-term context preserving, recurrent neural network language model (RNNLM) has attracted great interest by showing better performance over back-off n-gram models and feed-forward neural network language models (FNNLM). However, it still has the difficulty of modelling words of very low frequency in training data. To address this issue, a new framework of structured word embedding is introduced to RNNLM, where both input and target word embeddings are factorized into weighted sum of the corresponding sub-word embeddings. The framework is instantiated for Chinese, where characters can be naturally used as the sub-word units. Experiments on a Chinese twitter LVCSR task showed that the proposed approach effectively outperformed the standard RNNLM, yielding a relative PPL improvement of 8:8% and an absolute 0:59% CER improvement in N-Best re-scoring. Tianxing He, Xu Xiang, Yanmin Qian, Kai Yu 0004 |
ICASSP | 2 |
| 2011 | Behaviour of SFM algorithms with erroneous calibration
Loong Fah Cheong, Xu Xiang |
Comput. Vis. Image Underst. | 2 |
| 2008 | On the Minimum Distance Conjecture for Schubert CodesabstractIn 2000, S. R. Ghorpade and G. Lachaud presented a conjecture on the minimum distance of algebraic geometric codes associated to Schubert varieties. In this correspondence, we use the structure and properties of the exterior algebra to give an elementary proof of this conjecture, and provide a construction of codewords of minimum Hamming weight. Xu Xiang |
IEEE Trans. Inf. Theory | 1 |
| 2006 | Error Characteristics of SFM with Erroneous Focal Length
Loong Fah Cheong, Xu Xiang |
ACCV (1) | 2 |
| 2006 | How Do Movie Viewers Perceive Scene Structure from Dynamic CuesabstractCinema viewed from a location other than a Canonical Viewing Point (CVP) presents distortions to the viewer in both its static and dynamic aspects. Past works have investigated mainly the static aspect of this problem and attempted to explain why viewers still seem to perceive the scene very well. The dynamic aspect of depth perception, which is known as structure from motion, and its possible distortion, have not been well investigated. In this paper, we derive the dynamic depth cues perceived by the viewer and use the so-called isodistortion framework to understand its distortion. The result is that viewers seated at a reasonably central position experience a shift in the intrinsic parameters of their visual systems. Despite this shift, the key properties of the perceived depths remain largely the same, being determined in the main by the accuracy to which extrinsic motion parameters can be recovered. For a viewer seated at a non-central position and watching the movie screen with a slant angle, the view is related to the view at the CVP by a homography, resulting in various aberrations such as non-central projection. Loong Fah Cheong, Xu Xiang |
CVPR (1) | 2 |