Yong Zhang 0058

dblp:66/4615-58 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2025 GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression
abstract
Recent studies have demonstrated that many layers are functionally redundant in large language models (LLMs), enabling model compression by removing these layers to reduce inference cost.While such approaches can improve efficiency, indiscriminate layer pruning often results in significant performance degradation.In this paper, we propose GRASP (Gradient-based Retention of Adaptive Singular Parameters), a novel compression framework that mitigates this issue by preserving sensitivity-aware singular values.Unlike direct layer pruning, GRASP leverages gradient-based attribution on a small calibration dataset to adaptively identify and retain critical singular components.By replacing redundant layers with only a minimal set of parameters, GRASP achieves efficient compression while maintaining strong performance with minimal overhead.Experiments across multiple LLMs show that GRASP consistently outperforms existing compression methods, achieving 90% of the original model's performance under 20% compression ratio.The source code is available at https://github.com/LyoAI/GRASP.
Kainan Liu, Yong Zhang 0058, Ning Cheng 0001, Zhitao Li 0002, Jing Xiao 0006
EMNLP2
2025 Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning Distillation
abstract
The rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation can help small models acquire reasoning capabilities, but most existing methods focus primarily on improving teacher-generated reasoning paths. Our observations reveal that small models can generate high-quality reasoning paths during sampling, even without chain-of-thought prompting, though these paths are often latent due to their low probability under standard decoding strategies. To address this, we propose Self-Enhanced Reasoning Training (SERT), which activates and leverages latent reasoning capabilities in small models through self-training on filtered, self-generated reasoning paths under zero-shot conditions. Experiments using OpenAI’s GPT-3.5 as the teacher model and GPT-2 models as the student models demonstrate that SERT enhances the reasoning abilities of small models, improving their performance in reasoning distillation.
Yong Zhang 0058, Zhitao Li 0002, Ming Li 0010, Ning Cheng 0001, Minchuan Chen, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006
ICASSP1
2025 Logic Consistency Makes Large Language Models Personalized Reasoning Teachers
abstract
Large Language Models (LLMs) have advanced natural language processing, particularly through Chain-of-Thought (CoT) reasoning, but their high computational costs limit deployment. We propose Personalized Chain-of-Thought Distillation (PeCoTD), a method that transfers CoT reasoning from LLMs to smaller models by addressing the distribution gap—the difference in how large and small models process information. To bridge this gap, PeCoTD introduces the Self Logic Consistency (SLC) metric, which helps small models evaluate and select LLM-generated rationales that align better with their reasoning abilities. PeCoTD iteratively refines these rationales, adjusting them to better fit the learning patterns of small models while preserving their original meaning. Experiments show PeCoTD significantly enhances the reasoning abilities of small models across datasets, making CoT distillation more practical and effective.
Xulong Zhang 0001, Yong Zhang 0058, Jun Yu 0001, Jianzong Wang
IJCNN3
2024 Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
abstract
Ming Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, Tianyi Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ming Li 0010, Yong Zhang 0058, Shwai He, Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Tianyi Zhou 0001
ACL (1)2
2024 Leveraging Biases in Large Language Models: "bias-kNN" for Effective Few-Shot Learning
abstract
Large Language Models (LLMs) have shown significant promise in various applications, including zero-shot and few-shot learning. However, their performance can be hampered by inherent biases. Instead of traditionally sought methods that aim to minimize or correct these biases, this study introduces a novel methodology named "bias-kNN". This approach capitalizes on the biased outputs, harnessing them as primary features for kNN and supplementing with gold labels. Our comprehensive evaluations, spanning diverse domain text classification datasets and different GPT-2 model sizes, indicate the adaptability and efficacy of the "bias-kNN" method. Remarkably, this approach not only outperforms conventional in-context learning in few-shot scenarios but also demonstrates robustness across a spectrum of samples, templates and verbalizers. This study, therefore, presents a unique perspective on harnessing biases, transforming them into assets for enhanced model performance.
Yong Zhang 0058, Hanzhang Li, Zhitao Li 0002, Ning Cheng 0001, Ming Li 0010, Jing Xiao 0006, Jianzong Wang
ICASSP1
2024 MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
abstract
One-shot voice conversion aims to change the timbre of any source speech to match that of the unseen target speaker with only one speech sample. Existing methods face difficulties in satisfactory speech representation disentanglement and suffer from sizable networks as some of them leverage numerous complex modules for disentanglement. In this paper, we propose a model named MAIN-VC to effectively disentangle via a concise neural network. The proposed model utilizes Siamese encoders to learn clean representations, further enhanced by the designed mutual information estimator. The Siamese structure and the newly designed convolution module contribute to the lightweight of our model while ensuring performance in diverse voice conversion tasks. The experimental results show that the proposed model achieves comparable subjective scores and exhibits improvements in objective metrics compared to existing methods in a one-shot voice conversion scenario.
Pengcheng Li 0013, Jianzong Wang, Xulong Zhang 0001, Yong Zhang 0058, Jing Xiao 0006, Ning Cheng 0001
IJCNN4
2024 EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning
abstract
Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into account disentangling speech components through human-crafted bottleneck features which can not achieve sufficient information disentangling, while pitch and rhythm may still be mixed together. There is a risk of information overlap in the disentangling process which results in less speech naturalness. To overcome such limits, we propose a two-stage model to disentangle speech representations in a self-supervised manner without a human-crafted bottleneck design, which uses the Mutual Information (MI) with the designed upper bound estimator (IFUB) to separate overlapping information between speech components. Moreover, we design a Joint Text-Guided Consistent (TGC) module to guide the extraction of speech content and eliminate timbre leakage issues. Experiments show that our model can achieve a better performance than the baseline, regarding disentanglement effectiveness, speech naturalness, and similarity. Audio samples can be found at https://largeaudiomodel.com/eadvc.
Ziqi Liang, Jianzong Wang, Xulong Zhang 0001, Yong Zhang 0058, Ning Cheng 0001, Jing Xiao 0006
IJCNN4
2024 QLSC: A Query Latent Semantic Calibrator for Robust Extractive Question Answering
abstract
Extractive Question Answering (EQA) in Machine Reading Comprehension (MRC) often faces the challenge of dealing with semantically identical but format-variant inputs. Our work introduces a novel approach, called the "Query Latent Semantic Calibrator (QLSC)", designed as an auxiliary module for existing MRC models. We propose a unique scaling strategy to capture latent semantic center features of queries. These features are then seamlessly integrated into traditional query and passage embeddings using an attention mechanism. By deepening the comprehension of the semantic queries-passage relationship, our approach diminishes sensitivity to variations in text format and boosts the model’s capability in pinpointing accurate answers. Experimental results on robust Question-Answer datasets confirm that our approach effectively handles format-variant but semantically identical queries, highlighting the effectiveness and adaptability of our proposed method.
Sheng Ouyang, Jianzong Wang, Yong Zhang 0058, Zhitao Li 0002, Ziqi Liang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006
IJCNN3
2024 From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
abstract
Ming Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Ming Li 0010, Yong Zhang 0058, Zhitao Li 0002, Jiuhai Chen, Lichang Chen, Ning Cheng 0001, Jianzong Wang, Tianyi Zhou 0001, Jing Xiao 0006
NAACL-HLT2
2023 PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual Adapter
abstract
The Retrieval Question Answering (ReQA) task employs the retrieval-augmented framework, composed of a retriever and generator.The generator formulates the answer based on the documents retrieved by the retriever.Incorporating Large Language Models (LLMs) as generators is beneficial due to their advanced QA capabilities, but they are typically too large to be fine-tuned with budget constraints while some of them are only accessible via APIs.To tackle this issue and further improve ReQA performance, we propose a trainable Pluggable Reward-Driven Contextual Adapter (PRCA), keeping the generator as a black box.Positioned between the retriever and generator in a Pluggable manner, PRCA refines the retrieved information by operating in a tokenautoregressive strategy via maximizing rewards of the reinforcement learning phase.Our experiments validate PRCA's effectiveness in enhancing ReQA performance on three datasets by up to 20% improvement to fit black-box LLMs into existing frameworks, demonstrating its considerable potential in the LLMs era.
Zhitao Li 0002, Yong Zhang 0058, Jianzong Wang, Ning Cheng 0001, Ming Li 0010, Jing Xiao 0006
EMNLP3
2023 Boosting Chinese ASR Error Correction with Dynamic Error Scaling Mechanism
Yong Zhang 0058, Hanzhang Li, Jianzong Wang, Zhitao Li 0002, Sheng Ouyang, Ning Cheng 0001, Jing Xiao 0006
INTERSPEECH2
2023 Prompt Guided Copy Mechanism for Conversational Question Answering
abstract
Conversational Question Answering (CQA) is a challenging task that aims to generate natural answers for conversational flow questions.In this paper, we propose a pluggable approach for extractive methods that introduces a novel prompt-guided copy mechanism to improve the fluency and appropriateness of the extracted answers.Our approach uses prompts to link questions to answers and employs attention to guide the copy mechanism to verify the naturalness of extracted answers, making necessary edits to ensure that the answers are fluent and appropriate.The three prompts, including a question-rationale relationship prompt, a question description prompt, and a conversation history prompt, enhance the copy mechanism's performance.Our experiments demonstrate that this approach effectively promotes the generation of natural answers and achieves good results in the CoQA challenge.
Yong Zhang 0058, Zhitao Li 0002, Jianzong Wang, Yiming Gao 0010, Ning Cheng 0001, Fengying Yu, Jing Xiao 0006
INTERSPEECH1
2022 Self-Attention for Incomplete Utterance Rewriting
abstract
Incomplete utterance rewriting (IUR) has recently become an essential task in NLP, aiming to complement the incomplete utterance with sufficient context information for comprehension. In this paper, we propose a novel method by directly extracting the coreference and omission relationship from the self-attention weight matrix of the transformer in-stead of word embeddings and edit the original text accordingly to generate the complete utterance. Benefiting from the rich information in the self-attention weight matrix, our method achieved competitive results on public IUR datasets.
Yong Zhang 0058, Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006
ICASSP1