EDBT 2026 Demo / reviewers in the wild / expert
Yuekai Zhang
dblp:191/5931
· DBLP profile ↗
15ranked-venue papers
3as first author
12since 2021 · last 2025
0009-0000-7121-8822ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Deployment of Large Speech Recognition Models on GPUabstractLarge automatic speech recognition (ASR) models have achieved remarkable progress in recent years, but their deployment in production faces significant challenges due to large model sizes and autoregressive decoding methods. This paper presents comprehensive solutions for efficiently deploying large ASR models on GPUs using NVIDIA Triton Inference Server and TensorRT-LLM. Our deployment framework supports both encoder-decoder architectures and speech LLMs. With a modular design based on NVIDIA Triton, the framework can be easily extended to other model architectures. We implement optimized TensorRT-LLM engines for Whisper models and speech LLMs. Compared to existing implementations, our Whisper TensorRTLLM solution achieves more than $\mathbf{5 0 \%}$ throughput improvement. The complete deployment solutions are open-sourced and provide one-click deployment through docker-compose, facilitating rapid adoption in production environments.121https://github.com/k2-fsa/sherpa/tree/master/triton/whisper2https://github.com/k2-fsa/sherpa/tree/master/triton/speech_llm Yuekai Zhang, Junjie Lai |
ASRU | 1 |
| 2025 | PINNsAgent: Automated PDE Surrogation with Large Language ModelsabstractSolving partial differential equations (PDEs) using neural methods has been a long-standing scientific and engineering research pursuit. Physics-Informed Neural Networks (PINNs) have emerged as a promising alternative to traditional numerical methods for solving PDEs. However, the gap between domain-specific knowledge and deep learning expertise often limits the practical application of PINNs. Previous works typically involve manually conducting extensive PINNs experiments and summarizing heuristic rules for hyperparameter tuning. In this work, we introduce PINNsAgent, a novel surrogation framework that leverages large language models (LLMs) to bridge the gap between domain-specific knowledge and deep learning. PINNsAgent integrates Physics-Guided Knowledge Replay (PGKR) for efficient knowledge transfer from solved PDEs to similar problems, and Memory Tree Reasoning for exploring the search space of optimal PINNs architectures. We evaluate PINNsAgent on 14 benchmark PDEs, demonstrating its effectiveness in automating the surrogation process and significantly improving the accuracy of PINNs-based solutions. Qingpo Wuwu, Chonghan Gao, Tianyu Chen 0017, Yuekai Zhang, Jianxin Li 0002, Haoyi Zhou, Shanghang Zhang |
ICML | 5 |
| 2025 | DACFusion: Dual Asymmetric Cross-Attention guided feature fusion for multispectral object detection
Jingchen Qian, Baiyou Qiao, Yuekai Zhang, Tongyan Liu, Gang Wu 0007, Donghong Han |
Neurocomputing | 3 |
| 2025 | Multi-perspective empathy modeling for empathetic dialogue generation
Baiyou Qiao, Yuekai Zhang, Xinchi Li, Donghong Han |
Knowl. Based Syst. | 2 |
| 2023 | TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for PytorchabstractTorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and performant PyTorch components. Its contributors routinely engage with users to understand their needs and fulfill them by developing impactful features. Here, we survey TorchAudio’s development principles and contents and highlight key features we include in its latest version (2.1): self-supervised learning pre-trained pipelines and training recipes, high-performance CTC decoders, speech recognition models and training recipes, advanced media I/O capabilities, and tools for performing forced alignment, multi-channel speech enhancement, and reference-less speech assessment. For a selection of these features, through empirical studies, we demonstrate their efficacy and show that they achieve competitive or state-of-the-art performance. Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang 0007, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma 0001, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Kumar 0003, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Shinji Watanabe 0001, Yangyang Shi, Yumeng Tao |
ASRU | 10 |
| 2023 | Lightvessel: Exploring Lightweight Coronary Artery Vessel Segmentation Via Similarity Knowledge DistillationabstractIn recent years, deep convolution neural networks (DCNNs) have achieved great prospects in coronary artery vessel segmentation. However, it is difficult to deploy complicated models in clinical scenarios since high-performance approaches have excessive parameters and high computation costs. To tackle this problem, we propose LightVessel, a Similarity Knowledge Distillation Framework, for lightweight coronary artery vessel segmentation. Primarily, we propose a Feature-wise Similarity Distillation (FSD) module for semantic-shift modeling. Specifically, we calculate the feature similarity between the symmetric layers from the encoder and decoder. Then the similarity is transferred as knowledge from a cumbersome teacher network to a non-trained lightweight student network. Meanwhile, for encouraging the student model to learn more pixel-wise semantic information, we introduce the Adversarial Similarity Distillation (ASD) module. Concretely, the ASD module aims to construct the spatial adversarial correlation between the annotation and prediction from the teacher and student models, respectively. Through the ASD module, the student model obtains fined-grained subtle edge segmented results of the coronary artery vessel. Extensive experiments conducted on Clinical Coronary Artery Vessel Dataset demonstrate that LightVessel outperforms various knowledge distillation counterparts. Hao Dang, Yuekai Zhang, Xingqun Qi, Muyi Sun |
ICASSP | 2 |
| 2023 | TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length PenaltyabstractIn this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which does not require any alignment. We demonstrate that TrimTail is computationally cheap and can be applied online and optimized with any training loss or any model architecture on any dataset without any extra effort by applying it on various end-to-end streaming ASR networks either trained with CTC loss [1] or Transducer loss [2]. We achieve 100 ~ 200ms latency reduction with equal or even better accuracy on both Aishell-1 and Librispeech. Moreover, by using TrimTail, we can achieve a 400ms algorithmic improvement of User Sensitive Delay (USD) with an accuracy loss of less than 0.2. Xingchen Song, Di Wu 0061, Zhiyong Wu 0001, Yuekai Zhang, Zhendong Peng, Wenpeng Li, Fuping Pan, Changbao Zhu |
ICASSP | 5 |
| 2022 | ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnetabstractAs Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolkits that can be used to generate reproducible results on different Spoken Language Understanding (SLU) benchmarks. Hence, there is a need to build an open source standard that can be used to have a faster start into SLU research. We present ESPnet-SLU, which is designed for quick development of spoken language understanding in a single framework. ESPnet-SLU is a project inside end-to-end speech processing toolkit, ESPnet, which is a widely used open-source standard for various speech processing tasks like ASR, Text to Speech (TTS) and Speech Translation (ST). We enhance the toolkit to provide implementations for various SLU benchmarks that enable researchers to seamlessly mix-and-match different ASR and NLU models. We also provide pretrained models with intensively tuned hyper-parameters that can match or even outperform the current state-of-the-art performances. The toolkit is publicly available at https://github.com/espnet/espnet. Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng 0003, Yuekai Zhang, Sujay Kumar, Karthik Ganesan 0003, Brian Yan, Ngoc Thang Vu, Alan W. Black, Shinji Watanabe 0001 |
ICASSP | 7 |
| 2021 | Recent Developments on Espnet Toolkit Boosted By ConformerabstractIn this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, such as automatic speech recognition (ASR), speech translations (ST), speech separation (SS) and text-to-speech (TTS). Our experiments reveal various training tips and significant performance benefits obtained with the Conformer on different tasks. These results are competitive or even outperform the current state-of-art Transformer models. We are preparing to release all-in-one recipes using open source and publicly available corpora for all the above tasks with pre-trained models. Our aim for this work is to contribute to our research community by reducing the burden of preparing state-of-the-art research environments usually requiring high resources. Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, Jing Shi 0003, Shinji Watanabe 0001, Wangyou Zhang, Yuekai Zhang |
ICASSP | 15 |
| 2021 | Sequence-To-Sequence Singing Voice Synthesis With Perceptual Entropy LossabstractThe neural network (NN) based singing voice synthesis (SVS) systems require sufficient data to train well and are are prone to over-fitting due to data scarcity. However, we often encounter data limitation problem in building SVS systems because of high data acquisition and annotation cost,. In this work, we propose a Perceptual Entropy (PE) loss derived from a psycho-acoustic hearing model to regularize the network. With a one-hour open-source singing voice database, we explore the impact of the PE loss on various main-stream sequence-to-sequence models, including the RNN-based, transformer-based, and conformer-based models. Our experiments show that the PE loss can mitigate the over-fitting problem and significantly improve the synthesized singing quality reflected in objective and subjective evaluations. Jiatong Shi, Nan Huo, Yuekai Zhang, Qin Jin |
ICASSP | 4 |
| 2021 | Tiny Transducer: A Highly-Efficient Speech Recognition Model on Edge DevicesabstractThis paper proposes an extremely lightweight phone-based transducer model with a tiny decoding graph on edge devices. First, a phone synchronous decoding (PSD) algorithm based on blank label skipping is first used to speed up the transducer decoding process. Then, to decrease the deletion errors introduced by the high blank score, a blank label deweighting approach is proposed. To reduce parameters and computation, deep feedforward sequential memory network (DFSMN) layers are used in the transducer encoder, and a CNN-based stateless predictor is adopted. SVD technology compresses the model further. WFST-based decoding graph takes the context-independent (CI) phone posteriors as input and allows us to flexibly bias user-specific information. Finally, with only 0.9M parameters after SVD, our system could give a relative 9.1% - 20.5% improvement compared with a bigger conventional hybrid system on edge devices. Yuekai Zhang, Sining Sun |
ICASSP | 1 |
| 2021 | SPGISpeech: 5, 000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech RecognitionabstractIn the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization of non-standard words) is imputed by separate post-processing models.This adds complexity and limits performance, as many formatting tasks benefit from semantic information present in the acoustic signal but absent in transcription.Here we propose a new STT task: endto-end neural transcription with fully formatted text for target labels.We present baseline Conformer-based models trained on a corpus of 5,000 hours of professionally transcribed earnings calls, achieving a CER of 1.7.As a contribution to the STT research community, we release the corpus free for noncommercial use. 1 Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D. Shulman, Boris Ginsburg, Shinji Watanabe 0001, Georg Kucsko |
Interspeech | 5 |
| 2020 | x-Vectors Meet Adversarial Attacks: Benchmarking Adversarial Robustness in Speaker Verification
Jesús Villalba 0001, Yuekai Zhang, Najim Dehak |
INTERSPEECH | 2 |
| 2020 | Black-Box Attacks on Spoofing Countermeasures Using Transferability of Adversarial Examples
Yuekai Zhang, Ziyan Jiang, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 1 |
| 2019 | Information-Guided Pilot Insertion for OFDM-Based Vehicular Communications SystemsabstractOrthogonal frequency division multiplexing (OFDM) is widely considered as a promising technique for vehicular communications. In current OFDM-based vehicular communications systems, a portion of tones are dedicated for pilot symbols and transmitted along with the data-bearing tones to estimate residual frequency error and/or channel state information. However, these pilot symbols themselves do not carry any information, thus inevitably degrading the system spectral efficiency (SE). Motivated by the concept of index modulation, in this paper, we propose a novel pilot insertion technique to enhance the SE, in which the pilot positions are selected according to extra information bits. Specifically, three different types of pilot position selection schemes, which result in equal-spaced, unequal-spaced, and hybrid pilot placements, are studied, and the corresponding pilot position detection methods are designed, where the pilots can be used for either carrier phase tracking or channel estimation purpose. Computer simulation results corroborate the advantages of the proposed technique in terms of bit error rate. Qiang Li 0020, Miaowen Wen, Yuekai Zhang, Jun Li 0036, Fangjiong Chen, Fei Ji 0001 |
IEEE Internet Things J. | 3 |