Jian Kang 0006

dblp:56/6072-6 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation
abstract
The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.
Longhao Li, Zhao Guo, Hongjie Chen 0001, Yuhang Dai, Hongfei Xue, Tianlun Zuo, Chengyou Wang, Shuiyuan Wang, Hui Bu, Jie Li 0001, Jian Kang 0006, Ruibin Yuan, Ziya Zhou, Wei Xue 0002, Lei Xie 0001
AAAI13
2026 DIFFA: Large Language Diffusion Models Can Listen and Understand
abstract
Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce DIFFA, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of large language diffusion models for efficient and scalable audio understanding, opening a new direction for speech-driven AI.
Jiaming Zhou 0001, Hongjie Chen 0001, Shiwan Zhao, Jian Kang 0006, Jie Li 0001, Enzhi Wang, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xuelong Li 0001
AAAI4
2025 Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
Hongjie Chen 0001, Qing Wang 0039, Hang Lv 0006, Jian Kang 0006, Jie Li 0001, Zhennan Lin, Lei Xie 0001
INTERSPEECH5
2022 Adversarial Sample Detection for Speaker Verification by Neural Vocoders
abstract
Automatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is seriously vulnerable to recently emerged adversarial attacks, yet effective counter-measures against them are limited. In this paper, we adopt neural vocoders to spot adversarial samples for ASV. We use the neural vocoder to re-synthesize audio and find that the difference between the ASV scores for the original and re-synthesized audio is a good indicator for discrimination between genuine and adversarial samples. This effort is, to the best of our knowledge, among the first to pursue such a technical direction for detecting time-domain adversarial samples for ASV, and hence there is a lack of established baselines for comparison. Consequently, we implement the Griffin-Lim algorithm as the detection baseline. The proposed approach achieves effective detection performance that outperforms the baselines in all the settings. We also show that the neural vocoder adopted in the detection framework is dataset-independent. Our codes will be made open-source for future works to do fair comparison1.
Po-Chun Hsu, Ji Gao, Shen Huang, Jian Kang 0006, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee
ICASSP6
2022 ICPR 2022 Challenge on Multi-Modal Subtitle Recognition
abstract
Video subtitle recognition, as one of the basic elements of video editing, has received increasing attention recently. However, the misaligned between audio and subtitles as well as the costly manual annotation remain a demanding issue toward subsequent intelligent processing. In this paper, we introduce a multi-modal subtitle recognition challenge for ICPR 2022, in which we present a large-scale video dataset (215 hours in total for visual and audio annotations), and 3 tracks including: 1) extracting subtitles in visual modality with audio annotation (ESV); 2) extracting subtitles in audio modality with visual annotations (ESA); and 3) extracting subtitles with both visual and audio annotation (ESVA). The challenge attracts 376 participants, among which the methods of top 3 teams on each track have been elaborated.
Shen Huang, Pengfei Hu 0004, Jian Kang 0006, Weida Liang, Yaqiang Wu, Yong Liu 0027
ICPR7
2022 Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition
abstract
General accent recognition (AR) models tend to directly extract low-level information from spectrums, which always significantly overfit on speakers or channels. Considering accent can be regarded as a series of shifts relative to native pronunciation, distinguishing accents will be an easier task with accent shift as input. But due to the lack of native utterance as an anchor, estimating the accent shift is difficult. In this paper, we propose linguistic-acoustic similarity based accent shift (LASAS) for AR tasks. For an accent speech utterance, after mapping the corresponding text vector to multiple accent-associated spaces as anchors, its accent shift could be estimated by the similarities between the acoustic embedding and those anchors. Then, we concatenate the accent shift with a dimension-reduced text vector to obtain a linguistic-acoustic bimodal representation. Compared with pure acoustic embedding, the bimodal representation is richer and more clear by taking full advantage of both linguistic and acoustic information, which can effectively improve AR performance. Experiments on Accented English Speech Recognition Challenge (AESRC) dataset show that our method achieves 77.42% accuracy on Test set, obtaining a 6.94% relative improvement over a competitive system in the challenge.
Qijie Shao, Jinghao Yan, Jian Kang 0006, Xian Shi, Pengfei Hu 0004, Lei Xie 0001
INTERSPEECH3
2021 Leveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech Recognition
abstract
In Uyghur speech, consonant and vowel reduction are often encountered, especially in spontaneous speech with high speech rate, which will cause a degradation of speech recognition performance. To solve this problem, we propose an effective phone mask training method for Conformer-based Uyghur end-to-end (E2E) speech recognition. The idea is to randomly mask off a certain percentage features of phones during model training, which simulates the above verbal phenomena and facilitates E2E model to learn more contextual information. According to experiments, the above issues can be greatly alleviated. In addition, deep investigations are carried out into different units in masking, which shows the effectiveness of our proposed masking unit. We also further study the masking method and optimize filling strategy of phone mask. Finally, compared with Conformer-based E2E baseline without mask training, our model demonstrates about 5.51% relative Word Error Rate (WER) reduction on reading speech and 12.92% on spontaneous speech, respectively. The above approach has also been verified on test-set of open-source data THUYG-20, which shows 20% relative improvements.
Pengfei Hu 0004, Jian Kang 0006, Shen Huang
Interspeech3
2021 The TNT Team System Descriptions of Cantonese and Mongolian for IARPA OpenASR20
Zhiqiang Lv, Ambyer Han, Guan-Bo Wang, Gui-Xin Shi, Jian Kang 0006, Jinghao Yan, Pengfei Hu 0004, Shen Huang, Weiqiang Zhang 0001
Interspeech6
2019 Multimedia Simultaneous Translation System for Minority Language Communication with Mandarin
Shen Huang, Bojie Hu, Pengfei Hu 0004, Jian Kang 0006, Zhiqiang Lv, Jinghao Yan, Qi Ju 0002, Shiyin Kang, Deyi Tuo, Guangzhi Li, Nurmemet Yolwas
INTERSPEECH5
2017 Gated convolutional networks based hybrid acoustic models for low resource speech recognition
abstract
In acoustic modeling for large vocabulary speech recognition, recurrent neural networks (RNN) have shown great abilities to model temporal dependencies. However, the performance of RNN is not prominent in resource limited tasks, even worse than the traditional feedforward neural networks (FNN). Furthermore, training time for RNN is much more than that for FNN. In recent years, some novel models are provided. They use non-recurrent architectures to model long term dependencies. In these architectures, they show that using gate mechanism is an effective method to construct acoustic models. On the other hand, it has been proved that using convolution operation is a good method to learn acoustic features. We hope to take advantages of both these two methods. In this paper we present a gated convolutional approach to low resource speech recognition tasks. The gated convolutional networks use convolutional architectures to learn input features and a gate to control information. Experiments are conducted on the OpenKWS, a series of low resource keyword search evaluations. From the results, the gated convolutional networks relatively decrease the WER about 6% over the baseline LSTM models, 5% over the DNN models and 3% over the BLSTM models. In addition, the new models accelerate the learning speed by more than 1.8 and 3.2 times compared to that of the baseline LSTM and BLSTM models.
Jian Kang 0006, Weiqiang Zhang 0001, Jia Liu 0001
ASRU1
2017 An LSTM-CTC based verification system for proxy-word based OOV keyword search
abstract
Proxy-word based out of vocabulary (OOV) keyword search has been proven to be quite effective in keyword search. In proxy-word based OOV keyword search, each OOV keyword is assigned several proxies and detections of the proxies are regarded as detections of the OOV keywords. However, the confidence scores of these detections are still those of the proxies from lattices. To obtain a better confidence measure, we employ an LSTM-CTC verification method in this work and the confidence scores are regenerated. OOV keyword search results on the evalpart1 dataset of the OpenKWS16 Evaluation have shown consistent improvement and the maximum relative improvement can reach 21.06% for the MWTW metric.
Zhiqiang Lv, Jian Kang 0006, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP2
2015 High-performance Swahili keyword search with very limited language pack: The THUEE system for the OpenKWS15 evaluation
abstract
This paper presents the Swahili keyword search system developed by the THUEE team for the OpenKWS15 evaluation, which is conducted by NIST under the IARPA Babel program. There are several highlights in the development of the system, including automatic generation of the pronunciation lexicon, aggressive data augmentation, the multilingual bottleneck feature extractor trained from 6 languages, text selection from web data for language model training, semi-supervised training for acoustic models and language models, out-of-vocabulary keyword detection using morphemes and a rich diversity of the systems for combination. A wide variety of acoustic modeling techniques are explored and compared. Up to 12 different individual systems are used for combination. The system achieves the state-of-the-art performance in the required condition of the evaluation.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Jia Liu 0001
ASRU4
2015 Improved system fusion for keyword search
abstract
It has been demonstrated that system fusion can significantly improve the performance of keyword search. In this paper, we compare the performance of several widely-used arithmetic-based fusion methods using different normalization pipeline and try to find the best pipeline. A novel arithmetic-based fusion method is proposed in this work. The method supplies a more effective way to incorporate the number of systems which have non-zero scores for a detection. When tested on the development test dataset of the OpenKWS15 Evaluation, the proposed method achieves the highest maximum term-weighted value (MTWV) and actual term-weighted value (ATWV) among all other arithmetic-based fusion methods. Usually, discriminative fusion methods employing classifiers can outperform arithmetic-based fusion methods. A DNN-based fusion method is explored in this work. After word-burst information is added, the DNN-based fusion method outperforms all other methods. In addition, it is notable that our arithmetic-based method achieves the same MTWV as the DNN-based method.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Weiqiang Zhang 0001, Jia Liu 0001
ASRU4
2015 Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognition
abstract
Exploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold, and set the outputs of inactive nodes to zero. Thus, only a few nonzero outputs are fed to the next layer. Using this method, the output vector of each hidden layer becomes very sparse, so that the computational cost of the feed-forward algorithm can be reduced by adopting sparse matrix operations. The proposed method is evaluated in both small and large vocabulary speech recognition tasks, and results demonstrate that we can reduce the nonzero outputs to fewer than 20% of the total number of hidden nodes, without sacrificing speech recognition performance.
Jian Kang 0006, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP1