EDBT 2026 Demo / reviewers in the wild / expert
Junjie Lai
dblp:127/9614
· DBLP profile ↗
15ranked-venue papers
1as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Deployment of Large Speech Recognition Models on GPUabstractLarge automatic speech recognition (ASR) models have achieved remarkable progress in recent years, but their deployment in production faces significant challenges due to large model sizes and autoregressive decoding methods. This paper presents comprehensive solutions for efficiently deploying large ASR models on GPUs using NVIDIA Triton Inference Server and TensorRT-LLM. Our deployment framework supports both encoder-decoder architectures and speech LLMs. With a modular design based on NVIDIA Triton, the framework can be easily extended to other model architectures. We implement optimized TensorRT-LLM engines for Whisper models and speech LLMs. Compared to existing implementations, our Whisper TensorRTLLM solution achieves more than $\mathbf{5 0 \%}$ throughput improvement. The complete deployment solutions are open-sourced and provide one-click deployment through docker-compose, facilitating rapid adoption in production environments.121https://github.com/k2-fsa/sherpa/tree/master/triton/whisper2https://github.com/k2-fsa/sherpa/tree/master/triton/speech_llm Yuekai Zhang, Junjie Lai |
ASRU | 3 |
| 2025 | InstantSpeech: Instant Synchronous Text-to-Speech Synthesis for LLM-driven Voice ChatbotsabstractChatbots powered by large language models (LLMs) offer natural, human-like interactions. However, traditional text-to-speech (TTS) models paired with LLMs typically wait for the entire sentence to be generated before starting synthesis, leading to increased response latency. Although word-by-word speech synthesis models have been proposed to address this issue, they still face challenges, such as relying on autoregressive architectures to maintain smooth transitions between words or conditioning on auxiliary features from LLM for naturalness. To overcome these limitations, we introduce InstantSpeech, a novel low-latency synchronous speech synthesis model. InstantSpeech employs a fully parallel architecture, combining a causal transformer-based acoustic model with a causal convolution-based vocoder, enabling it to start streaming speech synthesis immediately after the LLM generates the initial words. Furthermore, we utilize knowledge distillation to enhance speech quality under limited lookahead. Experimental results show that InstantSpeech can deliver high-quality speech with consistently low speech response latency when integrated with LLMs of various sizes. Muyang Du, Junjie Lai |
ICASSP | 3 |
| 2025 | Practical Guidance and Tutorial on Incentivizing Reasoning in LLMs using Distillation and Reinforcement LearningabstractWith reasoning models like DeepSeek-R1 and OpenAI's o1 demonstrating breakthrough capabilities in complex problem-solving, there is growing interest in the AI community about how to unlock similar capabilities in other large language models (LLMs). This hands-on tutorial dives into practical methods for building reasoning capabilities in LLMs through two primary approaches: knowledge distillation from advanced reasoning models and post-training with reinforcement learning techniques. Participants will learn how to transfer reasoning capabilities from cutting-edge models like DeepSeek-R1 into smaller LLMs such as Qwen and Llama, and then explore how reinforcement learning can take these capabilities even further. Through interactive Jupyter notebook, the participants will exercise through the entire process. By the end of this session, participants will be equipped with practical knowledge in how to incentivize reasoning capabilities into LLMs, understand how to use various frameworks for this task, and leave with hands-on experience that can be applied to their own projects. The related materials are available at https://zpqiu.github.io/reasoning-model-tutorial-kdd2025. Zhaopeng Qiu, Junjie Lai |
KDD (2) | 5 |
| 2025 | SeerAttention: Self-distilled Attention Gating for Efficient Long-context PrefillingabstractAttention is the cornerstone of modern Large Language Models (LLMs). Yet its quadratic complexity hinders efficiency and scalability, especially for long-context processing. A promising approach is to leverage sparsity in attention. However, existing sparsity-based solutions predominantly rely on predefined patterns or heuristics at the attention head level, struggling to adapt dynamically to different contexts efficiently. We propose SeerAttention, a simple yet effective attention mechanism that directly learns the block-level attention sparsity from the LLM itself. Inspired by the gating mechanism in Mixture of Experts (MoE), SeerAttention augments the conventional attention with a **learnable gate** that **selectively activates important blocks** within the attention map. Specifically, the gate first pools the query (Q) and key (K) tensors along the sequence dimension and processes them through learnable linear layers. The resulting matrices are then multiplied together to produce the gating scores, which are used to predict block-level attention sparsity. Combined with our block-sparse FlashAttention kernel, SeerAttention can achieve significant speedup on GPUs. When applied to pre-trained LLMs, SeerAttention only requires training the gate parameters in a lightweight self-distillation manner, allowing rapid convergence. Our evaluation results demonstrate that SeerAttention achieves better model accuracy and lower latency for long-context pre-filling compared to prior methods. Code is available at: https://github.com/microsoft/SeerAttention. Yizhao Gao 0002, Zhichen Zeng 0002, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004 |
NeurIPS | 7 |
| 2024 | Romanization Encoding For Multilingual ASRabstractWe introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated tokenizer within a FastConformer-RNNT framework equipped with a Roman2Char module, we significantly reduce vocabulary and output dimensions, enabling larger training batches and reduced memory consumption. Our method decouples acoustic modeling and language modeling, enhancing the flexibility and adaptability of the system. In our study, applying this method to Mandarin-English ASR resulted in a remarkable 63.51% vocabulary reduction and notable performance gains of 13.72% and 15.03% on SEAME code-switching benchmarks. Ablation studies on MandarinKorean and Mandarin-Japanese highlight our method’s strong capability to address the complexities of other script-heavy languages, paving the way for more versatile and effective multilingual ASR systems. Wen Ding 0005, Fei Jia, Hainan Xu, Yu Xi, Junjie Lai, Boris Ginsburg |
SLT | 5 |
| 2024 | Semi-Supervised Learning For Code-Switching ASR With Large Language Model FilterabstractCode-switching (CS) phenomenon occurs when words or phrases from different languages are alternated in a single sentence. Due to data scarcity, building an effective CS Automatic Speech Recognition (ASR) system remains challenging. In this paper, we propose to enhance CS-ASR systems by utilizing rich unsupervised monolingual speech data within a semi-supervised learning framework, particularly when access to CS data is limited. To achieve this, we establish a general paradigm for applying noisy student training (NST) to the CS-ASR task. Specifically, we introduce the LLM-Filter, which leverages well-designed prompt templates to activate the correction capability of large language models (LLMs) for monolingual data selection and pseudo-labels refinement during NST. Our experiments on the supervised ASRU-CS and unsupervised AISHELL-2 and LibriSpeech datasets show that our method not only achieves significant improvements over supervised and semi-supervised learning baselines for the CS task, but also attains better performance compared with the fully-supervised oracle upper-bound on the CS English part. Additionally, we further investigate the influence of accent on AESRC dataset and demonstrate that our method can get achieve additional benefits when the monolingual data contains relevant linguistic characteristic. Yu Xi, Wen Ding 0005, Kai Yu 0004, Junjie Lai |
SLT | 4 |
| 2023 | Improving Noisy Student Training on Non-Target Domain Data for Automatic Speech RecognitionabstractNoisy Student Training (NST) has recently demonstrated extremely strong performance in Automatic Speech Recognition (ASR). In this paper, we propose a data selection strategy named LM Filter to improve the performance of NST on non-target domain data in ASR tasks. Hypotheses with and without a Language Model are generated and the CER differences between them are utilized as a filter threshold. Results reveal that significant improvements of 10.4% compared with no data filtering baselines. We can achieve 3.31% CER in AISHELL-1 test set, which is best result from our knowledge without any other supervised data. We also perform evaluations on the supervised 1000 hour AISHELL-2 dataset and competitive results of 4.73% CER can be achieved. Wen Ding 0005, Junjie Lai |
ICASSP | 3 |
| 2023 | Improving WaveRNN with Heuristic Dynamic Blending for Fast and High-Quality GPU Vocoding
Muyang Du, Jiaxing Qi, Junjie Lai |
INTERSPEECH | 4 |
| 2023 | Knowledge Enhanced Graph Neural Networks for Explainable RecommendationabstractRecently, explainable recommendation has attracted increasing attentions, which can make the recommender system more transparent and improve user satisfactions by recommending products with useful explanations. However, existing methods trend to trade-off between the recommendation accuracy and the interpretability of recommendation results. In this manuscript, we propose Knowledge Enhanced Graph Neural Networks (KEGNN) for explainable recommendation. Semantic knowledge from the external knowledge base is leveraged into representation learning of three sides, respectively user, items and user-item interactions, and the knowledge enhanced semantic embedding are exploited to initialize the user/item entities and user-item relations of one constructed user behavior graph. We design a graph neural networks based user behavior learning and reasoning model to perform both semantic and relational knowledge propagation and reasoning over the user behavior graph for comprehensive understanding of user behaviors. On the top of comprehensive representations of users/items and user-item interactions, hierarchical neural collaborative filtering layers are developed for precise rating prediction, and one generation-mode and copy-mode combined generator is devised for human-like semantic explanation generation by integrating the copy mechanism into gated recurrent neural networks. Quantitative and qualitative results demonstrate the superiority of KEGNN over the state-of-art methods, and the explainability and interpretability of our method. Ziyu Lyu, Yue Wu 0013, Junjie Lai, Min Yang 0007, Chengming Li 0004, Wei Zhou 0028 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Pattern Matters: Hierarchical Correlated Strip Convolutional Network for Scene Text RecognitionabstractMany state-of-the-art scene text recognition methods leverage a pre-designed backbone for general object recognition, where general objects are usually lacking of distinctive characteristics. However, unlike general objects, texts in images are usually composed of narrow strokes and their characters often share similar attributes, e.g. texture and intensity. Such useful patterns can be deteriorated if directly employing standard convolutional networks to the scene text recognition. To solve the problem, we introduce a novel operation Strip Convolution, a specially designed convolution for extracting features of narrow strokes. And we further apply a Hierarchical Correlation strategy, adopting multi-level attention mechanisms to capture common text attributes. Based on this, we design a novel framework, named as Hierarchical Correlated Strip Convolutonal network for scene text recognition. Extensive experiments demonstrate the superiority of the proposed HCSC network, improving the accuracy of text recognition effectively. Junjie Lai |
ICME | 2 |
| 2022 | Connect to the Past: Graph-Based Latent Estimation for Blind Super-ResolutionabstractIn this paper, we focus on blind super-resolution, which is a challenging problem due to the unknown degradation process. Previous methods assume a simplified degradation model and fail to generalize to multiple real-world scenarios. We reformulate the image degradation with a more general assumption and embed it as a degradation latent vector. To predict the latent, we first construct a correlation graph between input LR images and pre-sampled LR-HR image pairs. And then through a graph propagation, the degradation latent is obtained, aggregating valuable information from the graph. We name the graph module as GALE (Graph-bAsed Latent Estimation), which is referred as a connection to the past. Furthermore, we design a novel blind SR framework GALE-SR. Extensive experiments on synthetic and real-world images show that the proposed GALE-SR can provide visually favorable and state-of-the-art performance in blind SR problem. Junjie Lai |
ICME | 3 |
| 2022 | WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureabstractGraph neural networks (GNNs) are prevalent to deal with graph-structured datasets, encoding graph data into low dimensional vectors. In this paper, we present a fast training graph neural network framework, i.e., WholeGraph, based on a multi-GPU distributed shared memory architecture. Whole-Graph consists of partitioning the graph and corresponding node or edge features to multi-GPUs, eliminating the bottleneck of communication between CPU and GPUs during the training process. And the communication between different GPUs is implemented by GPUDirect Peer-to-Peer (P2P) memory access technology. Furthermore, WholeGraph provides several optimized computing operators. Our evaluations show that on large-scale graphs WholeGraph outperforms state-of-the-art GNN frameworks, such as Deep Graph Library (DGL) and Pytorch Geometric (PyG). The speedups of WholeGraph are up to$57.32\mathrm{x}$and$242.98\mathrm{x}$compared with DGL and PyG on a single machine multi-GPU node, respectively. The ratio of GPU utilization can sustain above 95% during GNN training process. Jiaxing Qi, Junjie Lai |
SC | 4 |
| 2021 | Optimizing Winograd-Based Convolution with Tensor CoresabstractConvolution computing is one of the primary time consuming part of convolutional neural networks (CNNs). State of the art convolutional neural networks use samll, 3 × 3 filters. Recent work on Winograd convolution can reduce the computational complexity a lot, making the convolution computing fast. But existing implementations of Winograd convolution is limited to small tiles, i.e. F(4 × 4, 3 × 3) and F(2 × 2, 3 × 3) where 4 × 4 and 2 × 2 are tile sizes of output channels and 3 × 3 is the filter size, and single precision data. In this paper, we propose an optimized mixed precision F(6 × 6, 3 × 3) Winograd convolution implementation on NVIDIA Ampere GPUs using Tensor Cores. Our experiments show that the accuracy of mixed precision F(6 × 6, 3 × 3) Winograd convolution is sufficient to infer the convolutional neural networks. Besides, our method achieves up to 15.71x and 2.41x speedup on NVIDIA Ampere A100, compared with the state of the art Winograd based convolution and GEMM based convolution in cuDNN 8.1.0, respectively. Moreover, we integrate our F(6 × 6, 3 × 3) Winograd convolution implementation into NVIDIA TensorRT, which is a C++ inference library on GPUs provided by NVIDIA, as custom layer plugins. And we build the whole VGG network model using our custom Winograd convolution layers and other layers supported by TensorRT. The experiments show that the accuracy of the whole VGG network using our F(6 × 6, 3 × 3) Winograd convolution is 71.24%, while the accuracy of using FP32 computing for the VGG network is 71.22%. Junjie Lai |
ICPP | 3 |
| 2021 | EDGES: An Efficient Distributed Graph Embedding System on GPU ClustersabstractGraph embedding training models access parameters sparsely in a “one-hot” manner. Currently, the distributed graph embedding neural network is learned by data parallel with the parameter server, which suffers significant performance and scalability problems. In this article, we analyze the problems and characteristics of training this kind of models on distributed GPU clusters for the first time, and find that fixed model parameters scattered among different machine nodes are a major limiting factor for efficiency. Based on our observation, we develop an efficient distributed graph embedding system called EDGES, which can utilize GPU clusters to train large graph models with billions of nodes and trillions of edges using data and model parallelism. Within the system, we propose a novel dynamic partition architecture for training these models, achieving at least one half of communication reduction compared to existing training systems. According to our evaluations on real-world networks, our system delivers a competitive accuracy for the trained embeddings, and significantly accelerates the training process of the graph node embedding neural network, achieving a speedup of 7.23x and 18.6x over the existing fastest training system on single node and multi-node, respectively. As for the scalability, our experiments show that EDGES obtains a nearly linear speedup. Junjie Lai |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Performance upper bound analysis and optimization of SGEMM on Fermi and Kepler GPUsabstractIn this paper, we present an approach to estimate GPU applications' performance upper bound based on algorithm analysis and assembly code level benchmarking. As an example, we analyze the potential peak performance of SGEMM (Single-precision General Matrix Multiply) on Fermi (GF110) and Kepler (GK104) GPUs. We try to answer the question of how much optimization space is left for SGEMM and why. According to our analysis, the nature of Fermi (Kepler) instruction set and the limited issue throughput of the schedulers are the main limitation factors for SGEMM to approach the theoretical peak performance. The estimated upper-bound peak performance of SGEMM is around 82.5% of the theoretical peak performance on GTX580 Fermi GPU and 57.6% on GTX680 Kepler GPU. Guided by this analysis and using the native assembly language, on average, our SGEMM implementations achieve about 5% better performance than CUBLAS in CUDA 4.1 SDK for large matrices on GTX580. The achieved performance is around 90% of the estimated upper-bound performance of SGEMM on GTX580. On GTX680, the best performance we achieve is around 77.3% of the estimated performance upper bound. We also describe how to use native assembly language directly in the CUDA runtime source code. Junjie Lai, André Seznec |
CGO | 1 |