Jintao He

dblp:187/9864 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Rearchitecting Programmable Networks For In-Network Computing: From Hardware To Language
abstract
In-network computing (INC) offers orders-of-magnitude performance gains for various applications. However, the existing pipeline-based switch architecture, though well-suited for stateless packet forwarding, imposes three constraints on stateful INC applications: (i) control-plane table management, (ii) scarce memory resources under per-stage layout, and (iii) a strict memory access model. To address these constraints, this paper provides a full-stack solution, covering a programmable chipset Solar-NP, a programming language NPC, and a complete toolchain XuanWu. First, Solar-NP adopts a run-to-completion-based (RTC-based) architecture with three new hardware features: data-plane table management, a hierarchical memory pool, and built-in data structures with atomic access guarantees. Second, to fully exploit the hardware features, our language NPC introduces a new abstraction, namely operation-action table (OAT), that allows table manipulations in the data plane. Finally, our XuanWu toolchain provides various utility software covering the complete workflow of developing, testing, debugging, and deployment. To show the power of our solution, we implement three types of INC cases, including in-network control, in-network telemetry, and in-network storage. Experimental results demonstrate that our solution allows more operations to be offloaded and achieves higher performance than today's pipeline-based solutions.
Haifeng Sun 0002, Taixu Tian, Jinbo Sun, Jintao He, Qun Huang 0001, Luyou He, Xiangcan Xu, Junyi Guo, Yongqiang Yang
EuroSys5
2026 SketchPlan: Full-Visibility Sketch-Based Telemetry with Limited Programmable Switch Coverage
Jinbo Sun, Haifeng Sun 0004, Jintao He, Qun Huang 0001, Sa Wang, Yungang Bao
IWQoS3
2026 Enabling General and Efficient Window Mechanism for In-Network Telemetry
abstract
Recent network telemetry solutions typically target programmable switches to achieve high performance and in-network visibility. They partition the packet stream into windows and then apply various stream processing techniques to summarize flow-level statistics. However, existing studies focus on the measurement within each window. Window management is still a missing piece due to the resource limitation of programmable switches. In this paper, we propose OmniWindow, a general and efficient window mechanism framework. OmniWindow splits the original window into fine-grained sub-windows such that the sub-windows can be merged into various window types. To deal with the resource restriction, OmniWindow carefully designs its data plane memory layout and proposes a window synchronization method. It also employs a collaborative architecture that can collect and reset stateful data in sub-windows within a limited time. We prototype OmniWindow on Tofino. We incorporate OmniWindow into a SOTA query-driven telemetry system and eight sketch-based telemetry algorithms. Our experiments demonstrate that OmniWindow enables these telemetry solutions to achieve higher accuracy than conventional window mechanism.
Haifeng Sun 0004, Jintao He, Jie Gui, Qun Huang 0001
IEEE Trans. Netw.3
2025 FD-Filter: A Compact Data Structure for Fine-Grained Intra-Flow Packet Delay Monitoring
Jintao He, Jie Gui, Tian Lv, Qun Huang 0001
INFOCOM1
2025 Delay-Aware Joint Microservice Deployment and Request Routing in Multi-Edge Environments Based on Reinforcement Learning
abstract
The service modules of the traditional Mobile Edge Computing (MEC) are difficult to deploy, extend, and maintain in real networks because of the highly sophisticated systems. To promote the generalization, openness, and flexibility of the network edge environment, an increasing number of studies are exploring the integration of microservices with MEC. However, the existing work usually treats microservice deployment and request routing as two separate issues, ignoring the interaction between them. Therefore, this paper focuses on the joint optimization of microservice deployment and request routing in the multi-edge cloud scenarios. We establish a problem model for minimizing the average response latency, considering the transmission of requests across edge clouds. Then, in view of the complexity of the scene, this paper proposes a joint training strategy of microservice deployment and request routing based on deep reinforcement learning and Best Fit Decreasing algorithm. The algorithm takes the change of microservice deployment scheme as the action of the agent, introduces the Best Fit Decreasing algorithm to construct request routing based on the deployment scheme, and calculates rewards using the complete joint microservice deployment and request routing scheme for subsequent network training. Finally, experimental results show that the proposed algorithm can effectively reduce the response time delay and system running power compared with other algorithms.
Kai Peng 0001, Jialu Guo, Hao Wang 0152, Jintao He, Zhiqing Zou, Tianping Deng, Menglan Hu
IEEE Trans. Netw. Serv. Manag.4
2024 A frequency and two-hop configuration checking-driven local search algorithm for the minimum weakly connected dominating set problem
Jintao He, Cuisong Lin, Shuli Hu, Minghao Yin
Neural Comput. Appl.2
2024 Joint Optimization of Service Deployment and Request Routing for Microservices in Mobile Edge Computing
abstract
Microservices as an emerging architecture are creating new opportunities to enable superior network services in Mobile Edge Computing (MEC). In the presence of huge amounts of user requests, the massive communications among microservices have become notoriously complicated. Due to the intricate data dependencies of the microservices, the overall performance of large-scale MEC applications simultaneously depends on both service deployment and request routing. However, most existing work ignores the interdependencies of microservices and studies the deployment and routing as two isolated problems. In this case, this paper investigates the joint optimization of service deployment and request routing in edge computing. We first formulate a delay minimization problem via mixed integer linear programming and queuing analysis, and then provide a hardness proof on the problem. In addition, this paper presents a 2-approximation algorithm, followed with rigorous mathematical proofs to demonstrate the approximation ratio. The proposed two-phase algorithm consists of rounding based service deployment and adaptive-scaling-based request routing policies, which employ fine grained joint optimization to minimize service response delay. Finally, we illustrate the near-optimal performance of the proposed algorithm via comprehensive experiments.
Kai Peng 0001, Liangyuan Wang, Jintao He, Chao Cai 0001, Menglan Hu
IEEE Trans. Serv. Comput.3
2023 HistSketch: A Compact Data Structure for Accurate Per-Key Distribution Monitoring
abstract
Stream processing is critical to data analytics. However, one important class of characteristics namely per-key distribution (i.e., the item distribution of every key) remains unsolved. Traditional stream processing methods such as sampling and histogram do not focus on per-key distribution. Though sketch is widely applied to deal with huge and high-speed streaming data, it mainly computes singular-value characteristics. However, per-key distribution needs to deal with multiple values for each key, which amplifies the needed resources.To this end, we present a novel sketch-based algorithm HistSketch for per-key distribution. Its key idea is to differentiate hot keys from infrequent keys and use different components to deal with them. For hot keys, HistSketch allocates dedicated counters. For infrequent keys, HistSketch allows counter sharing to alleviate memory usage. In addition, we propose two optimization mechanisms for HistSketch: the histogram shedding mechanism further reduces the storage overheads, while the equation-based decoding compensates for the error caused by counter sharing. Our evaluation compares HistSketch with nine state-of-the-art sketch-based solutions using five datasets. Our results show that HistSketch achieves both high accuracy and low resource usage.
Jintao He, Qun Huang 0001
ICDE1
2023 OmniWindow: A General and Efficient Window Mechanism Framework for Network Telemetry
abstract
Recent network telemetry solutions typically target programmable switches to achieve high performance and in-network visibility. They partition the packet stream into windows and then apply various stream processing techniques to summarize flow-level statistics. However, existing studies focus on the measurement within each window. Window management is still a missing piece due to the resource limitation of programmable switches. In this paper, we propose OmniWindow, a general and efficient window mechanism framework. OmniWindow splits the original window into fine-grained sub-windows such that the sub-windows can be merged into various window types. To deal with the resource restriction, OmniWindow carefully designs its data plane memory layout and proposes a window synchronization method. It also employs a collaborative architecture that can collect and reset stateful data in sub-windows within a limited time. We prototype OmniWindow on Tofino. We incorporate OmniWindow into a SOTA query-driven telemetry system and eight sketch-based telemetry algorithms. Our experiments demonstrate that OmniWindow enables these telemetry solutions to achieve higher accuracy than conventional window mechanism.
Haifeng Sun 0004, Jintao He, Jie Gui, Qun Huang 0001
SIGCOMM3
2023 Noah: Reinforcement-Learning-Based Rate Limiter for Microservices in Large-Scale E-Commerce Services
abstract
Modern large-scale online service providers typically deploy microservices into containers to achieve flexible service management. One critical problem in such container-based microservice architectures is to control the arrival rate of requests in the containers to avoid containers from being overloaded. In this article, we present our experience of rate limit for the containers in Alibaba, one of the largest e-commerce services in the world. Given the highly diverse characteristics of containers in Alibaba, we point out that the existing rate limit mechanisms cannot meet our demand. Thus, we design Noah, a dynamic rate limiter that can automatically adapt to the specific characteristic of each container without human efforts. The key idea of Noah is to use deep reinforcement learning (DRL) that automatically infers the most suitable configuration for each container. To fully embrace the advantages of DRL in our context, Noah addresses two technical challenges. First, Noah uses a lightweight system monitoring mechanism to collect container status. In this way, it minimizes the monitoring overhead while ensuring a timely reaction to system load changes. Second, Noah injects synthetic extreme data when training its models. Thus, its model gains knowledge on unseen special events and hence remains highly available in extreme scenarios. To guarantee model convergence with the injected training data, Noah adopts task-specific curriculum learning to train the model from normal data to extreme data gradually. Noah has been deployed in the production of Alibaba for two years, serving more than 50000 containers and around 300 types of microservice applications. Experimental results show that Noah can well adapt to three common scenarios in the production environment. It effectively achieves better system availability and shorter request response time compared with four state-of-the-art rate limiters.
Zhao Li 0007, Haifeng Sun 0004, Zheng Xiong, Qun Huang 0001, Zehong Hu, Shasha Ruan, Hai Hong, Jie Gui, Jintao He, Zebin Xu
IEEE Trans. Neural Networks Learn. Syst.10
2022 Elastic Bloom Filter: Deletable and Expandable Filter Using Elastic Fingerprints
abstract
The Bloom filter, answering whether an item is in a set, has achieved great success in various fields, including networking, databases, and bioinformatics. However, the Bloom filter has two main shortcomings: no support of item deletion and no support of expansion. Existing solutions either support deletion at the cost of using additional memory, or support expansion at the cost of increasing the false positive rate and decreasing the query speed. Unlike existing solutions, we propose the Elastic Bloom filter (EBF) to address the two shortcomings simultaneously. Importantly, when EBF expands, the false positives decrease. Our key technique isElastic Fingerprints, which dynamically absorb and release bits during compression and expansion. To support deletion, EBF can first delete the corresponding fingerprint and then update the corresponding bit in the Bloom filter. To support expansion, Elastic Fingerprints release bits and insert them to the Bloom filter. Our experimental results show that the Elastic Bloom filter significantly outperforms existing works.
Yuhan Wu 0001, Jintao He, Shen Yan 0004, Tong Yang 0003, Olivier Ruas, Gong Zhang 0001, Bin Cui 0001
IEEE Trans. Computers2
2021 SketchINT: Empowering INT with TowerSketch for Per-flow Per-switch Measurement
abstract
1Network measurement is indispensable to network operations. Two most promising measurement solutions are In-band Network Telemetry (INT) solutions and sketching solutions. INT solutions provide fine-grained per-switch per-packet information at the cost of high network overhead. Sketching solutions have low network overhead but fail to achieve both simplicity and accuracy for per-flow measurement. To keep their advantages, and at the same time, overcome their shortcomings, we first design SketchINT to combine INT and sketches, aiming to obtain all per-flow per-switch information with low network overhead. Second, for deployment flexibility and measurement accuracy, we design a new sketch for SketchINT, namely TowerSketch, which achieves both simplicity and accuracy. The key idea of TowerSketch is to use different-sized counters for different arrays under the property that the number of bits used for different arrays stays the same. TowerSketch can automatically record larger flows in larger counters and smaller flows in smaller counters. We have fully implemented our SketchINT prototype on a testbed consisting of 10 switches. We also implement our TowerSketch on P4, single-core CPU, multi-core CPU, and FPGA platforms to verify its deployment flexibility. Extensive experimental results verify that 1) TowerSketch achieves better accuracy than prior art on various tasks, outperforming the state-of-the-art ElasticSketch up to 13.9 times in terms of error; 2) Compared to INT, SketchINT reduces the number of packets in the collection process by 3 4 orders of magnitude with an error smaller than 5%.
Kaicheng Yang 0001, Yuanpeng Li 0002, Zirui Liu 0002, Tong Yang 0003, Yu Zhou 0008, Jintao He, Jing'an Xue, Zhengyi Jia, Yongqiang Yang
ICNP6