VLDB 2026 Research / reviewers in the wild / expert
Xuesong Yang
dblp:94/10333
· DBLP profile ↗
23ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic PyramidabstractVision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic levels. To address this, we present Hierarchical window (Hiwin) transformer as a plug-and-play solution for MLLMs, centered around our inverse semantic pyramid (ISP). Hiwin transformer comprises two key modules: (i) a visual detail injection module, which progressively injects low-level visual details into high-level language-aligned semantics features, thereby constructing an ISP, and (ii) a hierarchical window attention module, which leverages cross-scale windows to condense multi-level semantics from the ISP. Notably, our design achieves an average boost of 3.7% across 14 benchmarks compared with the baseline method, 9.3% on DocVQA for instance. Zonghao Guo, Xuesong Yang, Chi Chen 0005, Yuan Yao 0013, Tat-Seng Chua, Maosong Sun 0001 |
AAAI | 5 |
| 2026 | Multi-reservoir computing with ordered aggregation for time series analysis
Xuesong Yang, Meiming You, Baoxiang Du |
Knowl. Based Syst. | 1 |
| 2025 | Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free GuidanceabstractShehzeen Samarah Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shehzeen Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li 0007 |
EMNLP | 3 |
| 2025 | NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
Edresson Casanova, Paarth Neekhara, Ryan Langman, Shehzeen Hussain, Subhankar Ghosh, Xuesong Yang, Ante Jukic, Jason Li 0007, Boris Ginsburg |
INTERSPEECH | 6 |
| 2025 | VoiceNoNG: Robust High-Quality Speech Editing Model without Hallucinations
Sung-Feng Huang, Heng-Cheng Kuo, Zhehuai Chen, Xuesong Yang, Pin-Jui Ku, Ante Jukic, Chao-Han Huck Yang, Yu Tsao 0001, Yu-Chiang Frank Wang, Hung-yi Lee, Szu-Wei Fu |
INTERSPEECH | 4 |
| 2025 | HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
Ryan Langman, Xuesong Yang, Paarth Neekhara, Shehzeen Hussain, Edresson Casanova, Evelina Bakhturina, Jason Li 0007 |
INTERSPEECH | 2 |
| 2024 | Pudica: Toward Near-Zero Queuing Delay in Congestion Control for Cloud Gaming
Shibo Wang 0002, Shusen Yang, Chenglei Wu, Longwei Jiang, Chenren Xu, Cong Zhao 0001, Xuesong Yang, Jianjun Xiao 0003, Changxi Zheng, Jing Wang 0077 |
NSDI | 8 |
| 2024 | Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech EditsabstractNeural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like $\mathrm{A}^{3} \mathrm{~T}$ and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made available at: https://jasonswfu.github.io/SINE_dataset/index.html Sung-Feng Huang, Heng-Cheng Kuo, Zhehuai Chen, Xuesong Yang, Chao-Han Huck Yang, Yu Tsao 0001, Yu-Chiang Frank Wang, Hung-yi Lee, Szu-Wei Fu |
SLT | 4 |
| 2024 | Depth asynchronous time delay reservoir for nonlinear time series forecasting task
Meiming You, Xuesong Yang |
Inf. Sci. | 4 |
| 2024 | Reservoir Computing Based on Memristor Arrays in Random StatesabstractReservoir computing is a machine learning paradigm with lower training costs that replaces traditional recurrent neural networks in some time series processing areas to simplify complexity. The compact network structure and low-complexity training method of this approach make it more suitable for hardware implementation, and reservoir computing exhibits unique advantages over other deep learning models. Memristor is a single device that can change its resistance state by memorizing the applied voltage or history current. The non-linear and time-memory characteristics of memristors are highly compatible with the dynamic properties required for reservoir computing. Consequently, memristors can be harnessed to construct nonlinear nodes within reservoirs, forming intricate dynamical units. This study introduces a novel type of reservoir building unit, termed as a memristor array in a random state, and proposes a reservoir computing hardware system based on memristor arrays in random states (MARS-RC). The randomness and nonlinear properties of this memristor array give the MARS-RC system the unique ability to more effectively capture the dynamic characteristics of the data. Simultaneously, we’ve designed an array random initialization circuit unit (ARI) to facilitate control over the memristor’s state, thus enhancing the system’s controllability. The MARS-RC system exhibits comparable predictive performance to conventional software reservoirs in predictive experiments involving chaotic time series. This is supported by comparisons with five distinct reservoir computing platforms. Furthermore, we’ve applied the MARS-RC system to the task of multi-classifying ECG signals. In these experiments, by employing multiple MARS-RC arrays in parallel, the system exhibits substantial robustness when dealing with varying input data sizes, attaining a remarkable classification accuracy of 99.375%. This study proposes a novel construction approach to advance reservoir computing hardware systems and provides new insights. Xuesong Yang, Meiming You, Liai Pang, Baoxiang Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | LibFewShot: A Comprehensive Library for Few-Shot LearningabstractFew-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning. Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | SLN-RED: Regularization by Simultaneous Local and Nonlocal Denoising for Image RestorationabstractRegularization by denoising (RED) framework has shown impressive performance for many imaging inverse problems, by leveraging the denoising method in defining an explicit regularization. In this letter, we propose a novel SLN-RED scheme for image restoration by exploiting the local and nonlocal denoisers simultaneously. Theoretically, we proves that forboundeddenoisers, the SLN-RED under ADMM scheme with a continuation strategy converges to a fixed-point. Numerical experiments on deblurring and super-resolution tasks demonstrate promising performance of the proposed algorithm. Liangtian He, Xuesong Yang, Yilun Wang 0004, Chao Wang 0091 |
IEEE Signal Process. Lett. | 3 |
| 2021 | REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with RelabelingabstractAccents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DAT). We unveil the magic behind DAT and provide, for the first time, a theoretical guarantee that DAT learns accent-invariant representations. We also prove that performing the gradient reversal in DAT is equivalent to minimizing the Jensen-Shannon divergence between domain output distributions. Motivated by the proof of equivalence, we introduce reDAT, a novel technique based on DAT, which relabels data using either unsupervised clustering or soft labels. Experiments on 23K hours of multi-accent data show that DAT achieves competitive results over accent-specific baselines on both native and non-native English accents but up to 13% relative WER reduction on unseen accents; our reDAT yields further improvements over DAT by 3% and 8% relatively on non-native accents of American and British English. Hu Hu, Xuesong Yang, Zeynab Raeesy, Jinxi Guo, Gökçe Keskin, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Roland Maas |
ICASSP | 2 |
| 2019 | When CTC Training Meets Acoustic LandmarksabstractConnectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in resource-constrained scenarios. In this paper, the convergence properties of CTC are improved by incorporating acoustic landmarks. We tailored a new set of acoustic landmarks to help CTC training converge more rapidly and smoothly while also reducing recognition error rates. We leveraged new target label sequences mixed with both phone and manner changes to guide CTC training. Experiments on TIMIT demonstrated that CTC based acoustic models converge significantly faster and smoother when they are augmented by acoustic landmarks. The models pretrained with mixed target labels can be further finetuned, resulting in phone error rates 8.72% below baseline on TIMIT. Consistent performance gain is also observed on WSJ (a larger corpus) and reduced TIMIT (smaller). With WSJ, we are the first to succeed in verifying the effectiveness of acoustic landmark theory on a mid-sized ASR task. Di He 0004, Xuesong Yang, Boon Pang Lim, Mark Hasegawa-Johnson, Deming Chen |
ICASSP | 2 |
| 2019 | AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder LossabstractDespite the progress in voice conversion, many-to-many voice conversion trained on non-parallel data, as well as zero-shot voice conversion, remains under-explored. Deep style transfer algorithms, generative adversarial networks (GAN) in particular, are being applied as new solutions in this field. However, GAN training is very sophisticated and difficult, and there is no strong evidence that its generated speech is of good perceptual quality. In this paper, we propose a new style transfer scheme that involves only an autoencoder with a carefully designed bottleneck. We formally show that this scheme can achieve distribution-matching style transfer by training only on self-reconstruction loss. Based on this scheme, we proposed AutoVC, which achieves state-of-the-art results in many-to-many voice conversion with non-parallel data, and which is the first to perform zero-shot voice conversion. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Mark Hasegawa-Johnson |
ICML | 4 |
| 2018 | Deep Learning Based Speech BeamformingabstractMulti-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would otherwise be too complicated. On the other hand, deep learning based enhancement approaches are able to learn complicated speech distributions and perform efficient inference, but they are unable to deal with variable number of input channels. Also, deep learning approaches introduce a lot of errors, particularly in the presence of unseen noise types and settings. We have therefore proposed an enhancement framework called DEEPBEAM, which combines the two complementary classes of algorithms. DEEPBEAM introduces a beamforming filter to produce natural sounding speech, but the filter coefficients are determined with the help of a monaural speech enhancement neural network. Experiments on synthetic and real-world data show that DEEPBEAM is able to produce clean, dry and natural sounding speech, and is robust against unseen noise. Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
ICASSP | 4 |
| 2018 | Joint Modeling of Accents and Acoustics for Multi-Accent Speech RecognitionabstractThe performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several accents during training and building a single model in multi-task fashion, where tasks correspond to individual accents. In this paper, we explore an alternate model where we jointly learn an accent classifier and a multi-task acoustic model. Experiments on the American English Wall Street Journal and British English Cambridge corpora demonstrate that our joint model outperforms the strong multi-task acoustic model baseline. We obtain a 5.94% relative improvement in word error rate on British English, and 9.47% relative improvement on American English. This illustrates that jointly modeling with accent information improves acoustic model performance. Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
ICASSP | 1 |
| 2018 | Improved ASR for Under-resourced Languages through Multi-task Learning with Acoustic LandmarksabstractFurui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady-state regions are secondary or supplemental. Acoustic landmarks are perceptually salient, even in a language one doesn't speak, and it has been demonstrated that non-speakers of the language can identify features such as the primary articulator of the landmark. These factors suggest a strategy for developing language-independent automatic speech recognition: landmarks can potentially be learned once from a suitably labeled corpus and rapidly applied to many other languages. This paper proposes enhancing the cross-lingual portability of a neural network by using landmarks as the secondary task in multi-task learning (MTL). The network is trained in a well-resourced source language with both phone and landmark labels (English), then adapted to an under-resourced target language with only word labels (Iban). Landmark-tasked MTL reduces source-language phone error rate by 2.9% relative, and reduces target-language word error rate by 1.9%-5.9% depending on the amount of target-language training data. These results suggest that landmark-tasked MTL causes the DNN to learn hidden-node features that are useful for cross-lingual adaptation. Di He 0004, Boon Pang Lim, Xuesong Yang, Mark Hasegawa-Johnson, Deming Chen |
INTERSPEECH | 3 |
| 2017 | End-to-end joint learning of natural language understanding and dialogue managerabstractNatural language understanding and dialogue policy learning are both essential in conversational systems that predict the next system actions in response to a current user utterance. Conventional approaches aggregate separate models of natural language understanding (NLU) and system action prediction (SAP) as a pipeline that is sensitive to noisy outputs of error-prone NLU. To address the issues, we propose an end-to-end deep recurrent neural network with limited contextual dialogue memory by jointly training NLU and SAP on DSTC4 multi-domain human-human dialogues. Experiments show that our proposed model significantly outperforms the state-of-the-art pipeline models for both NLU and SAP, which indicates that our joint model is capable of mitigating the affects of noisy NLU outputs, and NLU model can be refined by error flows backpropagating from the extra supervised signals of system actions. Xuesong Yang, Yun-Nung Chen, Dilek Hakkani-Tür, Paul A. Crook, Xiujun Li, Jianfeng Gao 0001, Li Deng 0001 |
ICASSP | 1 |
| 2017 | Speech Enhancement Using Bayesian Wavenet
Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Xuesong Yang, Dinei A. F. Florêncio, Mark Hasegawa-Johnson |
INTERSPEECH | 4 |
| 2014 | Machine learning approaches to improving pronunciation error detection on an imbalanced corpusabstractIn this paper, we investigate the task of phone-level pronunciation error detection as a binary classification problem, the performance of which is heavily affected by the imbalanced distribution of the classes in a manually annotated data set of non-native English. In order to address problems caused by this extreme class imbalance, methods for cost-sensitive learning (weighting inversely proportional to class frequencies) and over-sampling of synthetic instances (SMOTE) are investigated in order to improve classification performance. Experiments using classifiers consisting of features based on acoustic phonetics and word identity demonstrate that these machine learning approaches lead to performance improvements over the baseline system based on the extremely imbalanced data. In addition, several different types of classifiers were compared. Finally, the paper analyzes the robustness of classifier performance across different phones. Xuesong Yang, Anastassia Loukina, Keelan Evanini |
SLT | 1 |
| 2011 | Improvement of Segmental Mispronunciation Detection with Prior Knowledge Extracted from Large L2 Speech Corpus
Dean Luo, Xuesong Yang |
INTERSPEECH | 2 |
| 2011 | Sound source localization for mobile robot based on time difference feature and space grid matchingabstractAuditory is a convenient and efficient way for Human-Robot Interaction, however implementing a sound source localization system based on TDOA method encounters many problems, such as noise of real environments, and resolution of nonlinear equations, switch between far field and near field and lack of microphones for geometric positioning localization method. In this paper, a new spectral weighting GCC-PHAT method is proposed to deal with noise. Furthermore, the time difference feature of sound source and its spatial distribution are analyzed. Based on prosperities of the distribution, a space grid matching (SGM) algorithm is proposed for localization step, which handles those problems that geometric positioning method faces effectively. Decision tree and valid feature detection algorithm are also proposed to reduce computational complexity and improve performance. Experiments are achieved in real environments on a mobile robot platform, in which 2016 sets of speech data are tested using four microphones in 3D space. More than 95% azimuth localization rate with error less than 5 degrees and approximate 90% horizontal distance localization rate are obtained. Xiaofei Li 0001, Hong Liu 0008, Xuesong Yang |
IROS | 3 |