Yongchao He

dblp:246/9290 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated Verification
abstract
Chain-of-Thought reasoning significantly improves the performance of large language models on complex tasks, but incurs high inference latency due to long generation traces.Steplevel speculative reasoning aims to mitigate this cost, yet existing approaches face a longstanding trade-off among accuracy, inference speed, and resource efficiency.We propose ConfSpec, a confidence-gated cascaded verification framework that resolves this trade-off.Our key insight is an asymmetry between generation and verification: while generating a correct reasoning step requires substantial model capacity, step-level verification is a constrained discriminative task for which small draft models are well-calibrated within their competence range, enabling high-confidence draft decisions to be accepted directly while selectively escalating uncertain cases to the large target model.Evaluation across diverse workloads shows that ConfSpec achieves up to 2.24× end-to-end speedups while matching target-model accuracy.Our method requires no external judge models and is orthogonal to token-level speculative decoding, enabling further multiplicative acceleration.
Siran Liu, Zane Cao, Yongchao He
ACL (1)3
2026 HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
abstract
Autoregressive decoding inherently limits the inference throughput of Large Language Model (LLM) due to its sequential dependency.Speculative decoding mitigates this by verifying multiple predicted tokens in parallel, but its efficiency remains constrained by what we identify as verification heterogeneity-the uneven difficulty of verifying different speculative candidates.In practice, a small subset of high-confidence predictions accounts for most successful verifications, yet existing methods treat all candidates uniformly, leading to redundant computation.We present HeteroSpec, a heterogeneity-adaptive speculative decoding framework that allocates verification effort in proportion to candidate uncertainty.Het-eroSpec estimates verification complexity using a lightweight entropy-based quantifier, partitions candidates via a data-driven stratification policy, and dynamically tunes speculative depth and pruning thresholds through coordinated optimization.Across five benchmarks and four LLMs, HeteroSpec delivers an average 4.24× decoding speedup over state-of-the-art methods such as EAGLE-3, while preserving exact output distributions.Crucially, HeteroSpec requires no model retraining and remains compatible with other inference optimizations, making it a practical direction for improving speculative decoding efficiency.
Siran Liu, Qianchao Zhu, Zane Cao, Yongchao He
ACL (1)5
2023 A Generic Service to Provide In-Network Aggregation for Key-Value Streams
abstract
Key-value stream aggregation is a common operation in distributed systems, which requires intensive computation and network resources. We propose a generic in-network aggregation service for key-value streams, ASK, to accelerate the aggregation operations in diverse distributed applications. ASK is a switch-host co-designed system, where the programmable switch provides a best-effort aggregation service, and the host runs a daemon to interact with applications. ASK makes in-depth optimization tailored to traffic characteristics, hardware restrictions, and network unreliable natures: it vectorizes multiple key-value tuples’ aggregation of one packet in one switch pipeline pass, which improves the per-host’s goodput; it develops a lightweight reliability mechanism for key-value stream’s asynchronous aggregation, which guarantees computation correctness; it designs a hot-key agnostic prioritization for key-skewed workloads, which improves the switch memory utilization. We prototype ASK and use it to support Spark and BytePS. The evaluation shows that ASK could accelerate pure key-value aggregation tasks by up to 155 times and big data jobs by 3-5 times, and be backward compatible with existing INA-empowered distributed training solutions with the same speedup.
Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu 0027, ChonLam Lao
ASPLOS (2)1
2023 RateSheriff: Multipath Flow-aware and Resource Efficient Rate Limiter Placement for Data Center Networks
abstract
Emerging cloud services and applications request different Quality of Service (QoS) in Data Center Networks (DCNs). To meet these various requirements, programmable switch-based rate limiters are introduced to provide performance isolation and benefit from easy control and fast deployment. However, existing programmable switch-based rate limiters have two limitations: (1) multipath flows (i.e., MultiPath TCP) cannot be precisely limited, and (2) rate limiter placement solutions in DCNs are missing. These limitations could lead to poor rate limiting performance and low bandwidth utilization. In this paper, we propose RateSheriff to improve rate limiting performance by providing multipath flow-aware and resource efficient rate limiter placement for programmable switch-enabled DCNs. We identify and associate subflows to a multipath flow by extracting and comparing specific packets and header fields. By solving the formulated resource efficient rate limiter placement problem, we can improve rate limiting performance and balance memory utilization among programmable switches in DCNs. Simulation results show that RateSheriff can correctly limit the rate of multipath flows, improve rate limiting performance by up to 46%, and improve memory balancing performance by up to 79% with low computation time, compared with baselines.
Songshi Dou, Yongchao He, Sen Liu 0002, Wenfei Wu, Zehua Guo 0001
IWQoS2
2022 SFP: Service Function Chain Provision on Programmable Switches for Cloud Tenants
abstract
Recent progress in programmable switches provides opportunities for service function chains (SFCs) provision to cloud tenants, which has the advantage of flexible deployment and high performance. We devise SFP for such SFC provision in the cloud. SFP's data plane installs physical NFs and is virtualized to host logical SFCs from multiple tenants. SFP's control plane uses a relaxed integer programming model to jointly optimize the placement of physical and logical NFs, which can achieve resource efficiency and high tenant traffic processing throughput within efficient execution time. Our prototype and evaluation shows that SFP can significantly offload NFV computation from server to the switch and maximize the switch resource utilization.
Hongyi Huang, Wenfei Wu, Yongchao He, Zehua Guo 0001
IPDPS3
2022 Consistent and Fine-Grained Rule Update with In-Network Control for Distributed Rate Limiting
abstract
Coexisting applications contend for limited WAN bandwidth when communicating over distributed data centers in private clouds. Online service providers deploy distributed rate limiting systems to dynamically estimate each application instance’s bandwidth demand, and update rate limiting rules to provide performance isolation and bandwidth guarantees for applications with different priorities. However, the isolation violation caused by the inconsistent update of rate limiting rules among servers would eventually violate the Service Level Agreement requirements. Motivated by the observation that InNetwork Control can update rate limiting rules consistently with ultra-low latency for hundreds of thousands of end-hosts through in-band control messages, this paper presents DistRL, a dynamic distributed rate limiting system that can achieve fine-grained consistent updates. The main idea of DistRL is to replace the traditional controller with programmable switches to improve communication efficiency, thereby achieving finer-grained consistent updates. The evaluation shows that DistRL can support sub-second distributed updates of rate limiting rules without isolation violation for O(105) servers.
Yongchao He, Wenfei Wu
IWQoS1
2021 Scalable On-Switch Rate Limiters for the Cloud
abstract
While most clouds use on-server rate limiters for bandwidth allocation, we propose to implement them on switches. On-switch rate limiters can simplify network management and promote the performance of control-plane rate limiting applications. We leverage the recent progress of programmable switches to implement on-switch rate limiters, named SwRL. In the design of SwRL, we make design choices according to the programmable hardware characteristics, we deeply optimize the memory usage of the algorithm so as to fit a cloud-scale (one million) rate limiters in a single switch, and we complement the missing computation primitives of the hardware using a pre-computed approximate table. We further developed three control-plane applications and integrate them with SwRL, showing the control-plane interoperability of SwRL. We prototype and evaluate SwRL in both testbed and production environments, demonstrating its good properties of precision rate control, scalability, interoperability, and manageability (execution environmental isolation).
Yongchao He, Wenfei Wu, Xuemin Wen, Yongqiang Yang
INFOCOM1
2021 NFD: Using Behavior Models to Develop Cross-Platform Network Functions
abstract
NFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs.
Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005
INFOCOM3
2019 SpeedyBox: Low-Latency NFV Service Chains with Cross-NF Runtime Consolidation
abstract
Software-based service chains in Network Function Virtualization (NFV) typically suffers high processing latency. This latency grows as chain lengths increase and possibly violates application requirements. Previous efforts focus on reducing latency while maintaining the perspective of each NF being an independent, isolated module. This results in processing redundancy that could eventually become the performance bottleneck. In this paper, we propose a low-latency NFV framework called SpeedyBox, that innovatively enables cross-NF runtime optimizations in a service chain to eliminate processing redundancy. SpeedyBox builds a fast data path for flows at runtime by consolidating the aggregate actions across diverse network functions (NFs) in a service chain. In SpeedyBox, each NF is instrumented with a stateful Local Match-Action Table (MAT), and leverages our easy-to-use APIs to record its per-flow behavior in the Local MAT. Next, SpeedyBox uses a Global MAT to build the fast data path by consolidating actions from each Local MAT, while providing the ability to express the stateful NF behaviors with an Event Table. We have implemented a prototype of SpeedyBox on the BESS and OpenNetVM NFV platforms. Our trace-driven evaluation on common NFs shows that SpeedyBox achieves significant latency reduction under real world scenarios.
Yong Cui 0001, Wenfei Wu, Jiahan Gu, K. K. Ramakrishnan, Yongchao He, Xuehai Qian
ICDCS7