Kai Zhen

dblp:87/10130 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0003-4133-8158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 MaZO: Masked Zeroth-Order Optimization for Multi-Task Fine-Tuning of Large Language Models
abstract
Large language models have demonstrated exceptional capabilities across diverse tasks, but their fine-tuning demands significant memory, posing challenges for resource-constrained environments.Zeroth-order (ZO) optimization provides a memory-efficient alternative by eliminating the need for backpropagation.However, ZO optimization suffers from high gradient variance, and prior research has largely focused on single-task learning, leaving its application to multi-task learning unexplored.Multi-task learning is crucial for leveraging shared knowledge across tasks to improve generalization, yet it introduces unique challenges under ZO settings, such as amplified gradient variance and collinearity.In this paper, we present MaZO, the first framework specifically designed for multi-task LLM fine-tuning under ZO optimization.MaZO tackles these challenges at the parameter level through two key innovations: a weight importance metric to identify critical parameters and a multi-task weight update mask to selectively update these parameters, reducing the dimensionality of the parameter space and mitigating task conflicts.Experiments demonstrate that MaZO achieves state-of-the-art performance, surpassing even multi-task learning methods designed for firstorder optimization.
Kai Zhen, Nathan Susanj, Athanasios Mouchtaris, Siegfried Kunzmann
EMNLP3
2025 QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models
abstract
Jiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu, Yequan Zhao, Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong, Zheng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jiajun Zhou 0004, Kai Zhen, Ziyue Liu 0003, Yequan Zhao, Seyed Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong 0001, Zheng Zhang 0005
EMNLP3
2024 AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-Tuning
abstract
Fine-tuning large language models (LLMs) has achieved remarkable performance across various natural language processing tasks, yet it demands more and more memory as model sizes keep growing.To address this issue, the recently proposed Memory-efficient Zerothorder (MeZO) methods attempt to fine-tune LLMs using only forward passes, thereby avoiding the need for a backpropagation graph.However, significant performance drops and a high risk of divergence have limited their widespread adoption.In this paper, we propose the Adaptive Zeroth-order Tensor-Train Adaption (AdaZeta) framework, specifically designed to improve the performance and convergence of the ZO methods.To enhance dimension-dependent ZO estimation accuracy, we introduce a fast-forward, low-parameter tensorized adapter.To tackle the frequently observed divergence issue in large-scale ZO finetuning tasks, we propose an adaptive query number schedule that guarantees convergence.Detailed theoretical analysis and extensive experimental results on Roberta-Large and Llama-2-7B models substantiate the efficacy of our AdaZeta framework in terms of accuracy, memory efficiency, and convergence speed. 1
Kai Zhen, Seyed Ershad Banijamali, Athanasios Mouchtaris, Zheng Zhang 0005
EMNLP2
2024 Max-Margin Transducer Loss: Improving Sequence-Discriminative Training Using a Large-Margin Learning Strategy
abstract
In this work, we propose a novel sequence-discriminative training criterion for automatic speech recognition (ASR) based on the Conformer Transducer. Inspired by the large-margin classifier framework, we separate the "good" and the "bad" hypotheses in an N-best list produced from a pre-trained transducer model by a margin (τ), hence the term, Max-Margin Transducer (MMT) loss. It is observed that fine-tuning with the proposed loss achieves significant improvement over baseline transducer loss but does not outperform the state-of-the-art minimum word error rate (MWER) training. However, combining the proposed MMT loss with MWER surpasses the performance of either losses suggesting the complimentary nature of MWER and MMT losses. With the combined losses, we obtained 7.44% and 7.68% relative WER improvements on Librispeech test-clean and test-other sets, respectively, and up to 8.9% relative improvement on Multi-lingual Librispeech test sets.
Rupak Vignesh Swaminathan, Grant P. Strimel, Ariya Rastrow, Sri Harish Reddy Mallidi, Kai Zhen, Hieu Duy Nguyen, Nathan Susanj, Athanasios Mouchtaris
ICASSP5
2024 Quantitative Analysis of Eye-Tracking Data Based on Information-Theoretic Tools for Measuring Driver Drowsiness
abstract
Assessing the drowsy state of driver is important for driving safety. In this paper, three kinds of information-theoretic tools are utilized to objectively and quantitatively measure the driver's detailed degrees of drowsiness. That is, fixation cross entropy (FCE) defined as the cross entropy between fixation count and duration, stationary gaze entropies based on fixation count (SGEc) and duration (SGEd) respectively, gaze transition entropy (GTE) based on fixation count are exploited to successfully obtain the drowsiness measurements. These measurements are based on eye-tracking data focusing on a novel division approach to area of interest (AOI), particularly emphasizing the significance of gaze deviation from the center point of screen. Psychophysical results showed that all the proposed measures are strongly correlated with Karolinska Sleepiness Scale (KSS), a widely used subjective measure of drowsiness. Particularly, compared with classical indicators for measuring drowsiness, FCE shows a stronger robustness under different drowsy conditions.
Yueming Zhu, Qing Xu 0002, Kai Zhen, Runlin Zhang, Shunbo Wang
ICME3
2023 Conmer: Streaming Conformer Without Self-attention for Interactive Voice Assistants
Martin Radfar, Paulina Lyskawa, Brandon Trujillo, Kai Zhen, Jahn Heymann, Denis Filimonov, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris
INTERSPEECH5
2022 Sub-8-Bit Quantization Aware Training for 8-Bit Neural Network Accelerator with On-Device Speech Recognition
Kai Zhen, Hieu Duy Nguyen, Raviteja Chinta, Nathan Susanj, Athanasios Mouchtaris, Tariq Afzal, Ariya Rastrow
INTERSPEECH1
2022 Sub-8-Bit Quantization for On-Device Speech Recognition: A Regularization-Free Approach
abstract
For on-device automatic speech recognition (ASR), quantization aware training (QAT) is ubiquitous to achieve the trade-off between model predictive performance and efficiency. Among existing QAT methods, one major drawback is that the quantization centroids have to be predetermined and fixed. To overcome this limitation, we introduce a regularization-free, “soft-to-hard” compression mechanism with self-adjustable centroids in a$\mu$-Law constrained space, resulting in a simpler yet more versatile quantization scheme, called General Quantizer (GQ). We apply GQ to ASR tasks using Recurrent Neural Network Transducer (RNN-T) and Conformer architectures on both LibriSpeech and de-identified far-field datasets. Without accuracy degradation, GQ can compress both RNN-T and Conformer into sub-8-bit, and for some RNN-T layers, to 1-bit for fast and accurate inference. We observe a 30.73% memory footprint saving and 31.75% user-perceived latency reduction compared to 8-bit QAT via physical device benchmarking.
Kai Zhen, Martin Radfar, Hieu Duy Nguyen, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris
SLT1
2022 Scalable and Efficient Neural Speech Coding: A Hybrid Design
abstract
We present a scalable and efficient neural waveform coding system for speech compression. We formulate the speech coding problem as an autoencoding task, where a convolutional neural network (CNN) performs encoding and decoding as a neural waveform codec (NWC) during its feedforward routine. The proposed NWC also defines quantization and entropy coding as a trainable module, so the coding artifacts and bitrate control are handled during the optimization process. We achieve efficiency by introducing compact model components to NWC, such as gated residual networks and depthwise separable convolution. Furthermore, the proposed models are with a scalable architecture, cross-module residual learning (CMRL), to cover a wide range of bitrates. To this end, we employ the residual coding concept to concatenate multiple NWC autoencoding modules, where each NWC module performs residual coding to restore any reconstruction loss that its preceding modules have created. CMRL can scale down to cover lower bitrates as well, for which it employs linear predictive coding (LPC) module as its first autoencoder. The hybrid design integrates LPC and NWC by redefining LPC’s quantization as a differentiable process, making the system training an end-to-end manner. The decoder of proposed system is with either one NWC (0.12 million parameters) in low to medium bitrate ranges (12 to 20 kbps) or two NWCs in the high bitrate (32 kbps). Although the decoding complexity is not yet as low as that of conventional speech codecs, it is significantly reduced from that of other neural speech coders, such as a WaveNet-based vocoder. For wide-band speech coding quality, our system yields comparable or superior performance to AMR-WB and Opus on TIMIT test utterances at low and medium bitrates. The proposed system can scale up to higher bitrates to achieve near transparent performance.
Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, Minje Kim 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Source-Aware Neural Speech Coding for Noisy Speech Compression
abstract
This paper introduces a novel neural network-based speech coding system that can process noisy speech effectively. The proposed source-aware neural audio coding (SANAC) system harmonizes a deep autoencoder-based source separation model and a neural coding system, so that it can explicitly perform source separation and coding in the latent space. An added benefit of this system is that the codec can allocate a different amount of bits to the underlying sources, so that the more important source sounds better in the decoded signal. We target a new use case where the user on the receiver side cares about the quality of the non-speech components in the speech communication, while the speech source still carries the most important information. Both objective and subjective evaluation tests show that SANAC can recover the original noisy speech better than the baseline neural audio coding system, which is with no source-aware coding mechanism, and two conventional codecs.
Haici Yang, Kai Zhen, Seungkwon Beack
ICASSP2
2021 Sparsification via Compressed Sensing for Automatic Speech Recognition
abstract
In order to achieve high accuracy for machine learning (ML) applications, it is essential to employ models with a large number of parameters. Certain applications, such as Automatic Speech Recognition (ASR), however, require real-time interactions with users, hence compelling the model to have as low latency as possible. Deploying large scale ML applications thus necessitates model quantization and compression, especially when running ML models on resource constrained devices. For example, by forcing some of the model weight values into zero, it is possible to apply zero-weight compression, which reduces both the model size and model reading time from the memory. In the literature, such methods are referred to as sparse pruning. The fundamental questions are when and which weights should be forced to zero, i.e. be pruned. In this work, we propose a compressed sensing based pruning (CSP) approach to effectively address those questions. By reformulating sparse pruning as a sparsity inducing and compression-error reduction dual problem, we introduce the classic compressed sensing process into the ML model training process. Using ASR task as an example, we show that CSP consistently outperforms existing approaches in the literature.
Kai Zhen, Hieu Duy Nguyen, Feng-Ju Chang, Athanasios Mouchtaris, Ariya Rastrow
ICASSP1
2020 A Dual-Staged Context Aggregation Method towards Efficient End-to-End Speech Enhancement
abstract
In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in the time domain without time-frequency transformation or mask estimation. However, aggregating contextual information from a high-resolution time domain signal with an affordable model complexity still remains challenging. In this paper, we propose a densely connected convolutional and recurrent network (DCCRN), a hybrid architecture, to enable dual-staged temporal context aggregation. With the dense connectivity and cross-component identical shortcut, DCCRN consistently outperforms competing convolutional baselines with an average STOI improvement of 0.23 and PESQ of 1.38 at three SNR levels. The proposed method is computationally efficient with only 1.38 million parameters. The generalizability performance on the unseen noise types is still decent considering its low complexity, although it is relatively weaker comparing to Wave-U-Net with 7.25 times more parameters.
Kai Zhen, Mi Suk Lee
ICASSP1
2020 Efficient and Scalable Neural Residual Waveform Coding with Collaborative Quantization
abstract
Scalability and efficiency are desired in neural speech codecs, which supports a wide range of bitrates for applications on various devices. We propose a collaborative quantization (CQ) scheme to jointly learn the codebook of LPC coefficients and the corresponding residuals. CQ does not simply shoehorn LPC to a neural network, but bridges the computational capacity of advanced neural network models and traditional, yet efficient and domain-specific digital signal processing methods in an integrated manner. We demonstrate that CQ achieves much higher quality than its predecessor at 9 kbps with even lower model complexity. We also show that CQ can scale up to 24 kbps where it outperforms AMR-WB and Opus. As a neural waveform codec, CQ models are with less than 1 million parameters, significantly less than many other generative models.
Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack
ICASSP1
2020 Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
abstract
Conventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded audio signals. For neural audio codecs, however, the objective nature of the loss function usually leads to suboptimal sound quality as well as high run-time complexity due to the large model size. In this work, we present a psychoacoustic calibration scheme to re-define the loss functions of neural audio coding systems so that it can decode signals more perceptually similar to the reference, yet with a much lower model complexity. The proposed loss function incorporates the global masking threshold, allowing the reconstruction error that corresponds to inaudible artifacts. Experimental results show that the proposed model outperforms the baseline neural codec twice as large and consuming 23.4% more bits per second. With the proposed method, a lightweight neural codec, with only 0.9 million parameters, performs near-transparent audio coding comparable with the commercial MPEG-1 Audio Layer III codec at 112 kbps.
Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack, Minje Kim 0001
IEEE Signal Process. Lett.1
2019 Cascaded Cross-Module Residual Learning Towards Lightweight End-to-End Speech Coding
abstract
Speech codecs learn compact representations of speech signals to facilitate data transmission.Many recent deep neural network (DNN) based end-to-end speech codecs achieve low bitrates and high perceptual quality at the cost of model complexity.We propose a cross-module residual learning (CMRL) pipeline as a module carrier with each module reconstructing the residual from its preceding modules.CMRL differs from other DNN-based speech codecs, in that rather than modeling speech compression problem in a single large neural network, it optimizes a series of less-complicated modules in a two-phase training scheme.The proposed method shows better objective performance than AMR-WB and the state-of-the-art DNNbased speech codec with a similar network architecture.As an end-to-end model, it takes raw PCM signals as an input, but is also compatible with linear predictive coding (LPC), showing better subjective quality at high bitrates than AMR-WB and OPUS.The gain is achieved by using only 0.9 million trainable parameters, a significantly less complex architecture than the other DNN-based codecs in the literature.
Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack
INTERSPEECH1