Huanran Zheng

dblp:325/3065 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2026
0009-0005-6261-6702ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Guiding LLMs to decode text via aligning semantics in EEG signals and language
Huanran Zheng, Yuanbin Wu, Tianwen Qian, Wenjing Yue, Xiaoling Wang 0004
Expert Syst. Appl.1
2025 Faster Speculative Decoding via Effective Draft Decoder with Pruned Candidate Tree
abstract
Speculative Decoding (SD) is a promising method for reducing the inference latency of large language models (LLMs). A well-designed draft model and an effective draft candidate tree construction method are key to enhancing the acceleration effect of SD. In this paper, we first propose the Effective Draft Decoder (EDD), which treats the LLM as a powerful encoder and generates more accurate draft tokens by leveraging the encoding results as soft prompts. Furthermore, we use KL divergence instead of the standard cross-entropy loss to better align the draft model’s output with the LLM. Next, we introduce the Pruned Candidate Tree (PCT) algorithm to construct a more efficient candidate tree. Specifically, we found that the confidence scores predicted by the draft model are well-calibrated with the acceptance probability of draft tokens. Therefore, PCT estimates the expected time gain for each node in the candidate tree based on confidence scores and retains only the nodes that contribute to acceleration, pruning away redundant nodes. We conducted extensive experiments with various LLMs across four datasets. The experimental results verify the effectiveness of our proposed method, which significantly improves the performance of SD and reduces the inference latency of LLMs.
Huanran Zheng
ACL (1)1
2025 TS-FourierLLM: Frozen Frequency-Domain Large Language Blocks for Enhancing Time-Series Modeling
Pengfei Wang 0009, Huanran Zheng, Wenjing Yue, Xiaoling Wang 0004
DASFAA (2)2
2024 Chimera Model of Candidate Soups for Non-Autoregressive Translation
Huanran Zheng, Wei Zhu 0016, Xiaoling Wang 0004
DASFAA (2)1
2024 NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and Better
abstract
With the increasing number of web documents, the demand for translation has increased dramatically. Non-autoregressive translation (NAT) models can significantly reduce decoding latency to meet the growing translation needs, but they sacrifice translation quality. And there is still an irreparable performance gap between NAT models and strong autoregressive translation (AT) models at the corpus level. However, more fine-grained comparative experiments on AT and NAT are currently lacking. Therefore, in this paper, we first conducted analysis experiments at the sentence level and found complementarity and high similarity between the translations generated by AT and NAT. Then, based on this observation, we propose a general and effective method called NAT4AT, which can not only use NAT to speed up the inference speed of AT significantly but also improve its final translation quality. Specifically, NAT4AT first uses a NAT model to generate an original translation in parallel and then uses an AT model as a correction model to revise errors in the original translation. In this way, the AT model no longer needs to predict the entire translation but only needs to predict a small number of error parts in the NAT result. Extensive experimental results on major WMT benchmarks verify the generality and effectiveness of our method, whose translation quality is superior to the strong AT model and achieves a 5.0x speedup.
Huanran Zheng, Wei Zhu 0016, Xiaoling Wang 0004
WWW1
2022 Candidate Soups: Fusing Candidate Results Improves Translation Quality for Non-Autoregressive Translation
abstract
Non-autoregressive translation (NAT) model achieves a much faster inference speed than the autoregressive translation (AT) model because it can simultaneously predict all tokens during inference.However, its translation quality suffers from degradation compared to AT.And existing NAT methods only focus on improving the NAT model's performance but do not fully utilize it.In this paper, we propose a simple but effective method called "Candidate Soups," which can obtain high-quality translations while maintaining the inference speed of NAT models.Unlike previous approaches that pick the individual result and discard the remainders, Candidate Soups (CDS) can fully use the valuable information in the different candidate translations through model uncertainty.Extensive experiments on two benchmarks (WMT'14 EN-DE and WMT'16 EN-RO) demonstrate the effectiveness and generality of our proposed method, which can significantly improve the translation quality of various base models.More notably, our best variant outperforms the AT model on three translation tasks with 7.6× speedup. 1
Huanran Zheng, Wei Zhu 0016, Pengfei Wang 0009, Xiaoling Wang 0004
EMNLP1