Huaixiu Steven Zheng

dblp:307/3201 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 60% Learning paradigms · 15% Efficient and distributed learning · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
prompting
1.522024
SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models · ICLR 2024
Machine learning › Learning paradigms
multi-task learning
1.222023
UL2: Unifying Language Learning Paradigms · ICLR 2023
HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022
Natural language and speech › Language models and text generation
pre-trained language model
1.222023
UL2: Unifying Language Learning Paradigms · ICLR 2023
ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning · ICLR 2022
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting · ICLR 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
0.912025
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting · ICLR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model reasoning
self-correction
0.812024
Large Language Models Cannot Self-Correct Reasoning Yet · ICLR 2024
Machine learning › Deep learning architectures and training
scaling laws
0.712023
Transcending Scaling Laws with 0.1% Extra Compute · EMNLP 2023
Machine learning › Learning paradigms › multi-task learning
multi-task transfer learning
0.612022
ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning · ICLR 2022
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.612022
HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022
Natural language and speech › Language models and text generation
prompt tuning
0.612022
HyperPrompt: Prompt-based Task-Conditioning of Transformers · ICML 2022
Machine learning › Trustworthy machine learning
interpretability
0.212024
Large Language Models Cannot Self-Correct Reasoning Yet · ICLR 2024
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.212024
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models · ICLR 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
reasoning structure
0.212024
SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures · NeurIPS 2024
Machine learning › Efficient and distributed learning › model reuse
model upcycling
0.212023
Transcending Scaling Laws with 0.1% Extra Compute · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

verification · 1.7drafting · 1.7distillation · 1.7self-discovery · 0.8self-consistency · 0.8large language model prompting · 0.8intrinsic self-correction evaluation · 0.8in-context learning · 0.8chain-of-thought · 0.8knowledge distillation · 0.7
YearPublicationVenuePosition
2025 Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
abstract
Retrieval augmented generation (RAG) combines the generative abilities of large language models (LLMs) with external knowledge sources to provide more accurate and up-to-date responses. Recent RAG advancements focus on improving retrieval outcomes through iterative LLM refinement or self-critique capabilities acquired through additional instruction tuning of LLMs. In this work, we introduce Speculative RAG - a framework that leverages a larger generalist LM to efficiently verify multiple RAG drafts produced in parallel by a smaller, distilled specialist LM. Each draft is generated from a distinct subset of retrieved documents, offering diverse perspectives on the evidence while reducing input token counts per draft. This approach enhances comprehension of each subset and mitigates potential position bias over long context. Our method accelerates RAG by delegating drafting to the smaller specialist LM, with the larger generalist LM performing a single verification pass over the drafts. Extensive experiments demonstrate that Speculative RAG achieves state-of-the-art performance with reduced latency on TriviaQA, MuSiQue, PopQA, PubHealth, and ARC-Challenge benchmarks. It notably enhances accuracy by up to 12.97% while reducing latency by 50.83% compared to conventional RAG systems on PubHealth.
Zilong Wang 0002, Zifeng Wang 0002, Long T. Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang 0001, Anush Mattapalli, Ankur Taly, Jingbo Shang, Chen-Yu Lee, Tomas Pfister
ICLR4
2024 Large Language Models Cannot Self-Correct Reasoning Yet
abstract
Large Language Models (LLMs) have emerged as a groundbreaking technology with their unparalleled text generation capabilities across various applications. Nevertheless, concerns persist regarding the accuracy and appropriateness of their generated content. A contemporary methodology, self-correction, has been proposed as a remedy to these issues. Building upon this premise, this paper critically examines the role and efficacy of self-correction within LLMs, shedding light on its true potential and limitations. Central to our investigation is the notion of intrinsic self-correction, whereby an LLM attempts to correct its initial responses based solely on its inherent capabilities, without the crutch of external feedback. In the context of reasoning, our research indicates that LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction. Drawing from these insights, we offer suggestions for future research and practical applications in this field.
Jie Huang 0009, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou
ICLR4
2024 Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
abstract
We present STEP-BACK PROMPTING, a simple prompting technique that enables LLMs to do abstractions to derive high-level concepts and first principles from instances containing specific details. Using the concepts and principles to guide reasoning, LLMs significantly improve their abilities in following a correct reasoning path towards the solution. We conduct experiments of STEP-BACK PROMPTING with PaLM-2L, GPT-4 and Llama2-70B models, and observe substantial performance gains on various challenging reasoning-intensive tasks including STEM, Knowledge QA, and Multi-Hop Reasoning. For instance, STEP-BACK PROMPTING improves PaLM-2L performance on MMLU (Physics and Chemistry) by 7% and 11% respectively, TimeQA by 27%, and MuSiQue by 7%.
Huaixiu Steven Zheng, Swaroop Mishra, Heng-Tze Cheng, Ed H. Chi, Quoc V. Le, Denny Zhou
ICLR1
2024 SELF-DISCOVER: Large Language Models Self-Compose Reasoning Structures
abstract
We introduce SELF-DISCOVER, a general framework for LLMs to self-discover the task-intrinsic reasoning structures to tackle complex reasoning problems that are challenging for typical prompting methods. Core to the framework is a self-discovery process where LLMs select multiple atomic reasoning modules such as critical thinking and step-by-step thinking, and compose them into an explicit reasoning structure for LLMs to follow during decoding. SELF-DISCOVER substantially improves GPT-4 and PaLM 2’s performance on challenging reasoning benchmarks such as BigBench-Hard, grounded agent reasoning, and MATH, by as much as 32% compared to Chain of Thought (CoT). Furthermore, SELF-DISCOVER outperforms inference-intensive methods such as CoT-Self-Consistency by more than 20%, while requiring 10-40x fewer inference compute. Finally, we show that the self-discovered reasoning structures are universally applicable across model families: from PaLM 2-L to GPT-4, and from GPT-4 to Llama2, and share commonalities with human reasoning patterns.
Jay Pujara, Xiang Ren 0001, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, Denny Zhou, Swaroop Mishra, Huaixiu Steven Zheng
NeurIPS10
2023 Transcending Scaling Laws with 0.1% Extra Compute
abstract
Yi Tay, Jason Wei, Hyung Chung, Vinh Tran, David So, Siamak Shakeri, Xavier Garcia, Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, Denny Zhou, Donald Metzler, Slav Petrov, Neil Houlsby, Quoc Le, Mostafa Dehghani. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q. Tran 0002, David R. So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, Denny Zhou, Donald Metzler, Slav Petrov, Neil Houlsby, Quoc V. Le, Mostafa Dehghani 0001
EMNLP8
2023 UL2: Unifying Language Learning Paradigms
Yi Tay, Mostafa Dehghani 0001, Vinh Q. Tran 0002, Xavier Garcia, Jason Wei, Xuezhi Wang 0002, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, Donald Metzler
ICLR10
2022 ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran 0002, Dara Bahri, Jianmo Ni, Jai Gupta 0001, Kai Hui 0001, Sebastian Ruder, Donald Metzler
ICLR5
2022 HyperPrompt: Prompt-based Task-Conditioning of Transformers
abstract
Prompt-Tuning is a new paradigm for finetuning pre-trained language models in a parameter efficient way. Here, we explore the use of HyperNetworks to generate hyper-prompts: we propose HyperPrompt, a novel architecture for prompt-based task-conditioning of self-attention in Transformers. The hyper-prompts are end-to-end learnable via generation by a HyperNetwork. HyperPrompt allows the network to learn task-specific feature maps where the hyper-prompts serve as task global memories for the queries to attend to, at the same time enabling flexible information sharing among tasks. We show that HyperPrompt is competitive against strong multi-task learning baselines with as few as 0.14% of additional task-conditioning parameters, achieving great parameter and computational efficiency. Through extensive empirical experiments, we demonstrate that HyperPrompt can achieve superior performances over strong T5 multi-task learning baselines and parameter-efficient adapter variants including Prompt-Tuning and HyperFormer++ on Natural Language Understanding benchmarks of GLUE and SuperGLUE across many model sizes.
Huaixiu Steven Zheng, Yi Tay, Jai Gupta 0001, Vamsi Aribandi, Zhe Zhao 0001, YaGuang Li, Donald Metzler, Heng-Tze Cheng, Ed H. Chi
ICML2