Zane Cao

dblp:424/0891 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 67% Language models and text generation · 33%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
2.022026
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding · ACL (1) 2026
ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification · ACL (1) 2026
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
2.022026
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding · ACL (1) 2026
ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification · ACL (1) 2026
Natural language and speech › Language models and text generation › decoding
autoregressive decoding
1.012026
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding · ACL (1) 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

entropy-based verification complexity estimation · 1.0data-driven stratification · 1.0confidence-gated verification · 1.0cascaded verification · 1.0
YearPublicationVenuePosition
2026 ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification
abstract
Chain-of-Thought reasoning significantly improves the performance of large language models on complex tasks, but incurs high inference latency due to long generation traces.Steplevel speculative reasoning aims to mitigate this cost, yet existing approaches face a longstanding trade-off among accuracy, inference speed, and resource efficiency.We propose ConfSpec, a confidence-gated cascaded verification framework that resolves this trade-off.Our key insight is an asymmetry between generation and verification: while generating a correct reasoning step requires substantial model capacity, step-level verification is a constrained discriminative task for which small draft models are well-calibrated within their competence range, enabling high-confidence draft decisions to be accepted directly while selectively escalating uncertain cases to the large target model.Evaluation across diverse workloads shows that ConfSpec achieves up to 2.24× end-to-end speedups while matching target-model accuracy.Our method requires no external judge models and is orthogonal to token-level speculative decoding, enabling further multiplicative acceleration.
Siran Liu, Zane Cao, Yongchao He
ACL (1)2
2026 HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
abstract
Autoregressive decoding inherently limits the inference throughput of Large Language Model (LLM) due to its sequential dependency.Speculative decoding mitigates this by verifying multiple predicted tokens in parallel, but its efficiency remains constrained by what we identify as verification heterogeneity-the uneven difficulty of verifying different speculative candidates.In practice, a small subset of high-confidence predictions accounts for most successful verifications, yet existing methods treat all candidates uniformly, leading to redundant computation.We present HeteroSpec, a heterogeneity-adaptive speculative decoding framework that allocates verification effort in proportion to candidate uncertainty.Het-eroSpec estimates verification complexity using a lightweight entropy-based quantifier, partitions candidates via a data-driven stratification policy, and dynamically tunes speculative depth and pruning thresholds through coordinated optimization.Across five benchmarks and four LLMs, HeteroSpec delivers an average 4.24× decoding speedup over state-of-the-art methods such as EAGLE-3, while preserving exact output distributions.Crucially, HeteroSpec requires no model retraining and remains compatible with other inference optimizations, making it a practical direction for improving speculative decoding efficiency.
Siran Liu, Qianchao Zhu, Zane Cao, Yongchao He
ACL (1)4