VLDB 2026 Research / reviewers in the wild / expert
Hongtao Guan
dblp:73/10888
· DBLP profile ↗
19ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0007-3290-4767ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 6 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Moirai: Dependency-Impact-Based Communication Scheduling for Multi-job Distributed Deep Learning Cluster
Jingbin Yang, Jinglei Pei, Qinghua Wu 0004, Hongtao Guan |
Euro-Par (2) | 5 |
| 2026 | RAN Parameters Under the Microscope: A Deep Dive into 5G Handover Performance in Dense Urban Network
Chengxiang Lin, Hongtao Guan |
IWQoS | 4 |
| 2026 | Graph-based fast-flux domain detection using graph neural networks
Haiyang Jiang 0001, Hongtao Guan |
Comput. Networks | 4 |
| 2025 | SNARY: A High-Performance and Generic SmartNIC-accelerated Retrieval System
Qiaoyin Gan, Hongtao Guan, Zhaohua Wang, Zhenyu Li 0001, Gaogang Xie |
USENIX ATC | 5 |
| 2023 | AD2S: Adaptive anomaly detection on sporadic data streams
Yang Wang 0147, Zhenyu Li 0001, Hongtao Guan, Gaogang Xie |
Comput. Commun. | 4 |
| 2022 | MicroCBR: Case-Based Reasoning on Spatio-temporal Fault Knowledge Graph for Microservices Troubleshooting
Yang Wang 0147, Zhenyu Li 0001, Hongtao Guan, Gaogang Xie |
ICCBR | 5 |
| 2022 | Muses: Enabling Lightweight Learning-Based Congestion Control for Mobile DevicesabstractVarious congestion control (CC) algorithms have been designed to target specific scenarios. To automate this process, researchers have begun to use machine learning to automatically control the congestion window. These, however, often rely on heavyweight learning models (e.g., neural networks). This can make them unsuitable for resource-constrained mobile devices. On the other hand, lightweight models (e.g., decision trees) are often incapable of reflecting the complexity of diverse mobile wireless environments. To address this, we present Muses, a learning-based approach for generating lightweight congestion control algorithms. Muses relies on imitation learning to train a universal (heavy) LSTM model, which is then used to extract (lightweight) decision tree models that are each targeted at an individual environment. Muses then dynamically selects the most appropriate decision tree on a per-flow basis. We show that Muses can generate high throughput policies across a diverse set of environments, and it is sufficiently light to operate on mobile devices. Zhiren Zhong, Wei Wang 0334, Yiyang Shao, Zhenyu Li 0001, Hongtao Guan, Gareth Tyson, Gaogang Xie, Kai Zheng 0003 |
INFOCOM | 6 |
| 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation OffloadingabstractProgrammable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance. Penglai Cui, Zhenyu Li 0001, Penghao Zhang, Tianhao Miao, Jianer Zhou, Hongtao Guan, Gaogang Xie |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2021 | NetFC: Enabling Accurate Floating-point Arithmetic on Programmable SwitchesabstractProgrammable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded to the network on programmable switches. These tasks may require the support of on-the-fly floatingpoint operations. Unfortunately, the computational capacity of programmable switches is limited to simple integer arithmetic operations. To address this issue, prior approaches either adopt a float-to-integer method or rely on local CPUs of switches, incurring accuracy loss and delayed processing.To this end, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. NetFC adopts a divide-and-conquer mechanism that converts the original huge table into several much smaller tables that are operated by the built-in integer operations. NetFC further leverages a scaling-factor mechanism for improving computational accuracy, and a prefix-based lossless table compression method to reduce memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.94% with only 448KB memory consumption. Furthermore, we integrate NetFC into Sonata [12] for detecting Slowloris attack, yielding significant decrease of detection delay. Penglai Cui, Zhenyu Li 0001, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Hongtao Guan, Gaogang Xie |
ICNP | 7 |
| 2021 | DSQNet: Domain SeQuence based Deep Neural Network for AGDs DetectionabstractModern botnets widely rely on Algorithmically Generated Domains (AGDs) to contact with Command-and-Control (C&C) servers. Existing AGD detection solutions check the domains one by one based on the structural differences between AGD and benign ones, e.g., some AGD families show much more random character composition than legitimate ones. These methods can hardly deal with the newly emerged camouflage technology based AGD types, as each individual AGD seems benign in domain structure features of itself. In this work, the structural correlations among AGDs are analyzed and we find the inter-AGD correlation can be adopted for the AGD detection. We then propose DSQNet, a Domain SeQuence based Deep Neural Network AGD detection model, that simultaneously checks the domains in batch to take the inter-AGD correlation into consideration during the detection. Experiments on the public and real-world dataset show the superiority of the proposed approach. Haiyang Jiang 0001, Hongtao Guan |
ISCC | 3 |
| 2020 | DOS-GAN: A Distributed Over-Sampling Method Based on Generative Adversarial Networks for Distributed Class-Imbalance Learning
Hongtao Guan, Xingkong Ma |
ICA3PP (3) | 1 |
| 2020 | Logchain: Cloud workflow reconstruction & troubleshooting with unstructured logs
Pengpeng Zhou, Yang Wang 0147, Zhenyu Li 0001, Gareth Tyson, Hongtao Guan, Gaogang Xie |
Comput. Networks | 5 |
| 2019 | A Massively Multi-Tenant Virtualized Network Intrusion Prevention Service on NFV PlatformabstractMulti-Tenancy (MT) is critical for Network Function Virtualization (NFV) platform as it reduces the cost of having network services by sharing expensive server resource among customers. This is especially critical for memory and CPU intensive services like Network Intrusion Prevention System (NIPS). In this work, we explore the issue of deploying a large-scale virtualized NIPS service on a commercial NFV platform. We observe that the scalability of NIPS service is not good when based on independent Virtual Machines (VMs). We propose a Multi-Tenant Aho-Corasick state machine data structure (MT-AC) and adapt it into NIPS to solve the issue. One MT-AC based NIPS service simultaneously checks traffic belonging to different tenants against a merged ruleset. The MT-AC data structure is very efficient as it eliminates the redundancies among tenants' signatures during the rulesets merging. Experimental results with real-world ruleset show that, in comparison with an independent VM-based solution, the MT-AC based NIPS service can support 2 to 4 times more tenants. Moreover, the throughput and latency performance of MT-AC based NIPS engine only degrades by 1%, when the tenant count increases from 8 to 128. The results validate that, the proposed MT-AC based NIPS service on NFV platform can support a large amount of tenants with a very low cost. Haiyang Jiang 0001, Hongtao Guan, Gaogang Xie, Kavé Salamatian |
ICCCN | 3 |
| 2018 | An Overview on the Convergence of High Performance Computing and Big Data ProcessingabstractAt present, with the rapid development of big data processing technology, streaming data processing and real-time data analysis have gradually become new research hotspots. Both the industry and the academia have invested a lot of research into the efficient processing methods of massive data generated in the environment such as the Internet and e-commerce. Meanwhile, high-performance computing technology and supercomputers are also looking for new business growth points. The convergence of big data processing and high performance computing technology is the general trend of big data analysis in the future. This paper will give a brief overview of typical technologies in the fusion process of big data processing and high-performance computing. Songzhu Mei, Hongtao Guan |
ICPADS | 2 |
| 2018 | Efficient Action Computation for Compositional SDN PoliciesabstractSoftware-defined networking envisions the support of multiple applications collaboratively operating on the same traffic. Policies of applications therefore require composition into a rule list that represents the union of application intents. In this context, ensuring the correctness and efficiency of composition for match fields as well as the associated actions is the fundamental requirement. Prior work however focuses only on the composition of match fields and assumes simple concatenation for action composition. We show in this paper that simple concatenation can result in incorrect behavior and inefficiency of packet processing. To address this issue, we formalize the action composition problem and propose two graph-based computation models to facilitate efficient composition of action lists. Our proposed approach has been integrated into the CoVisor code base and the evaluation results show its fitness for purpose. Zhenyu Li 0001, Gaogang Xie, Peng He 0003, Hongtao Guan, Laurent Mathy |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2018 | Partial Order Theory for Fast TCAM UpdatesabstractTernary content addressable memories (TCAMs) are frequently used for fast matching of packets against a given ruleset. While TCAMs can achieve fast matching, they are plagued by high update costs that can make them unusable in a high churn rate environment. We present, in this paper, a systematic and in-depth analysis of the TCAM update problem. We apply partial order theory to derive fundamental constraints on any rule ordering on TCAMs, which ensures correct checking against a given ruleset. This theoretical insight enables us to fully explore the TCAM update algorithms design space, to derive the optimal TCAM update algorithm (though it might not be suitable to be used in practice), and to obtain upper and lower bounds on the performance of practical update algorithms. Having lower bounds, we checked if the smallest update costs are compatible with the churn rate observed in practice, and we observed that this is not always the case. We therefore developed a heuristic based on ruleset splitting, with more than a single TCAM chip, that achieves significant update cost reductions (1.05~11.3×) compared with state-of-the-art techniques. Peng He 0003, Hongtao Guan, Kavé Salamatian, Gaogang Xie |
IEEE/ACM Trans. Netw. | 3 |
| 2013 | Toward predictable performance in decision tree based packet classification algorithmsabstractPacket classification has been studied extensively in the past decade. While many efficient algorithms have been proposed, the lack of deterministic performance has hindered the adoption and deployment of these algorithms: the expensive and power-hungry TCAM is still the de facto standard solution for packet classification. In this work, in contrast to proposing yet another new packet classification algorithm, we present the first steps to understand this unpredictability in performance for the existing algorithms. We focus on decision-tree based algorithms in this paper. In order to achieve the predictability, we firstly revisit the classical and many state-of-art packet classification algorithms. Through a detailed analysis, we conclude that two features of ruleset usually dominate the performance results: 1) the uniformity of the range distribution in different dimensions of the rules; 2) the existence and the number of “orthogonal structure” and wildcard rules in the ruleset. We conduct experiments to show the correctness of these observations, and discribe some potential applications for those results. Our work provides some insight to make the packet classification algorithms a credible alternative to the TCAM-only solutions. Peng He 0003, Hongtao Guan, Laurent Mathy, Kavé Salamatian, Gaogang Xie |
LANMAN | 2 |
| 2012 | Evaluating and Optimizing IP Lookup on Many Core ProcessorsabstractIn recent years, there has been a growing interest in multi/many core processors as a target architecture for high performance software router. This is a clear difference from the previous trend to use dedicated network processors and hardware components. Because of its key position in routers, hardware IP lookup implementation has been intensively studied with TCAM and FPGA based architecture. However, increasing interest in software implementation has also been observed. In this paper, we evaluate the performance of software only IP lookup on a many core chip, the TILEPro64 processor. For this purpose we have implemented two widely used IP lookup algorithms, DIR-24-8-BASIC and Tree Bitmap. We evaluate the performance of these two algorithms over the TILEPro64 processor with both synthetic and real-world traces. After a detailed analysis, we propose a hybrid scheme which provides high lookup speed and low worst case update overhead. Our work shows how to exploit the architectural features of TILEPro64 to improve the performance, including many optimization in both single-core and parallelism aspects. Experiment results show by using only 18 cores, we can achieve a lookup throughput of 60Mpps (almost 40Gbps) with low power consumption, which demonstrates great performance potentials in many core processor. Peng He 0003, Hongtao Guan, Gaogang Xie, Kavé Salamatian |
ICCCN | 2 |
| 2012 | Building a Flexible and Scalable Virtual Hardware Data Plane
Yingke Xie, Gaogang Xie, Layong Luo, Fuxing Zhang, Qingsong Ning, Hongtao Guan |
Networking (1) | 8 |