Dan Zhao 0003

dblp:10/3489-3 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Security and privacy · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Imagine with Layout and Sketch: Enhancing Vision-Language Retrieval with Dual-Stream Multi-Modal Query Refinement
abstract
Vision-Language Retrieval (VLR) aims to retrieve relevant visual or textual information from multimodal data using language or image queries. However, traditional VLR methods often rely on data-driven shallow semantic alignment and fail to understand the deeper structural and fine-grained entity features of queries, resulting in poor performance on multi-entity layouts and challenging entities. In this paper, we propose the Layout-Aware and Sketch-Enhanced (LASE) VLR framework, which refines query representations by incorporating multimodal layout and sketch knowledge. Specifically, layout knowledge encodes the spatial arrangement of entities, while sketch knowledge refines entity perception by capturing essential structural details. To extract these knowledge representations, we leverage Large Language Models' (LLMs) powerful semantic understanding for layout generation, and Diffusion Models' (DMs) fine-grained cross-modal generative capabilities for sketch generation. However, integrating knowledge into queries may introduce biases and query-specific preferences due to varying visual content and knowledge demands. To address this, we propose the Gated Dual-Stream Knowledge Module (GDKM), which consists of a multi-instance fusion network with a sample-aware gating network. The fusion network aggregates diverse knowledge using multi-head attention to reduce bias, while the gating network adjusts knowledge weights based on query characteristics. Extensive experiments demonstrate that the LASE significantly enhances VLR performance across multiple benchmarks, with superior generalization and transferability.
Guanghao Meng, Jinpeng Wang 0002, Qian-Wei Wang, Xudong Ren, Dan Zhao 0003
AAAI5
2026 Suit the Remedy to the Retriever: Interpretable Query Optimization with Retriever Preference Alignment for Vision-Language Retrieval
abstract
Vision-language retrieval (VLR), which uses text or image queries to retrieve corresponding cross-modal content, plays a crucial role in multimedia and computer vision tasks. However, challenging concepts in queries often confuse retrievers, limiting their ability to align concepts with visual content. Existing query optimization methods neglect retrievers’ preferences (i.e., text descriptions that better match their corresponding visual content), resulting in unadapted to the retriever and leading to suboptimal performance. To address this, we propose the Retriever-Adaptive Query Optimization (RAQO), an interpretable framework that rewrites queries based on retriever-specific preferences. Specifically, we first leverages multimodal large language Models (MLLMs) and retrieval's feedback to construct the MLLMs-Driven Preference-Aware Dataset Engine (MPADE), which automatically refine queries offline, capturing the retriever’s implicit preferences. Then, we introduce a ``detect-then-rewrite" chain-of-thought rewriting (ReCoT) strategy equipped with a progressive preference alignment pipeline, including three stages: ambiguity detection fine-tuning, query rewriting fine-tuning, and preference rank optimization. This design enables the rewriter to focus on confusing concepts and produce retriever-adapted, high-quality queries. Extensive VLR benchmark experiments have demonstrated the superiority of RAQO in cross-modal retrieval, as well as its interpretability, generalizability and transferability.
Guanghao Meng, Jinpeng Wang 0002, Jieming Zhu, Yong Jiang 0001, Dan Zhao 0003, Qing Li 0006
AAAI6
2026 SmartGen: Synthesizing Context-Aware User Behavior Data for Adaptive Smart Home Intelligence
abstract
As smart homes become increasingly prevalent, intelligent models are widely used for tasks such as anomaly detection and behavior prediction. These models are typically trained on static datasets, making them brittle to behavioral drift caused by seasonal changes, lifestyle shifts, or evolving routines. However, collecting new behavior data for retraining is often impractical due to its slow pace, high cost, and privacy concerns. In this paper, we propose SmartGen, an LLM-based framework that synthesizes context-aware user behavior data to support continual adaptation of downstream smart home models. SmartGen consists of four key components. First, we design a Time and Semantic-aware Split module to divide long behavior sequences into manageable, semantically coherent subsequences under dual time-span constraints. Second, we propose Semantic-aware Sequence Compression to reduce input length while preserving representative semantics by clustering behavior mapping in latent space. Third, we introduce Graph-guided Sequence Synthesis, which constructs a behavior relationship graph and encodes frequent transitions into prompts, guiding the LLM to generate data aligned with contextual changes while retaining core behavior patterns. Finally, we design a Two-stage Outlier Filter to identify and remove implausible or semantically inconsistent outputs, aiming to improve the factual coherence and behavioral validity of the generated sequences. Experiments on three real-world datasets demonstrate that SmartGen significantly enhances model performance on anomaly detection and behavior prediction tasks under behavioral drift, with anomaly detection improving by 85.43% and behavior prediction by 70.51% on average. The code is available at https://github.com/xzyvoid/SmartGen.
Zhiyao Xu, Dan Zhao 0003, Qingsong Zou, Qing Li 0006, Yong Jiang 0001, Yuhang Wang 0036, Jingyu Xiao
KDD (1)2
2026 A Multimodal Multi-Drone Cooperation System for Real-Time Human Searching
abstract
Aerial images from drones have been used to search individuals in the crowd. However, using a single drone for human searching faces challenges including low accuracy and long latency, due to poor visibility and limited on-board computing resources. In this paper, we propose SkyNet, a multi-drone cooperation system for real-time human searching, including locating and identifying. To locate a person, SkyNet uses Multi-View Cross Search with only 2D images. To achieve accurate identification, SkyNet processes faces in images from multi-view drones in three steps. First, a Multi-Modal Face Correction is designed to transform less useful face into desired target face, guided by text instructions. Second, an Angle Masking Network is developed to minimize invalid data of a single profile face. Third, the multiple face from drones are fused by a Fusion Weight Network. Moreover, by predicting the estimated finishing time of tasks, SkyNet schedules and balances workloads among edge devices and the cloud server to minimize processing latency. We implement SkyNet in real life, and evaluate the performance with 20 human participants. The results show that SkyNet can locate people within 0.18m error. The identification accuracy reaches 95.87%, and the system process is completed within 0.84s.
Junkun Peng, Qing Li 0006, Yuanzheng Tan, Dan Zhao 0003, Yong Jiang 0001
IEEE Trans. Mob. Comput.4
2025 Learning-Enhanced High-Throughput Pattern Matching Based on Programmable Data Plane
Guanglin Duan, Qing Li 0006, Dan Zhao 0003, Zili Meng, Dirk Kutscher, Ruoyu Li 0003, Yong Jiang 0001, Mingwei Xu 0001
USENIX ATC5
2025 Minos : A Lightweight and Dynamic Defense against Traffic Analysis in Programmable Data Planes
Qing Li 0006, Guorui Xie, Dan Zhao 0003, Zhuochen Fan, Lianbo Ma 0004, Yong Jiang 0001
USENIX ATC4
2025 Helios: Learning and Adaptation of Matching Rules for Continual In-Network Malicious Traffic Detection
abstract
Network Intrusion Detection Systems (NIDS) are critical for web security by identifying and blocking malicious traffic. In-network NIDS leverage programmable switches for high-speed traffic processing. However, they are unable to reconcile the fine-grained classification of known classes and the identification of unseen attacks. Moreover, they lack support for incremental updates. In this paper, we propose Helios, an in-network malicious traffic detection system, for continual adaptation in attack-incremental scenarios. First, we design a novel Supervised Mixture Prototypical Learning (SMPL) method combined with clustering initialization to learn prototypes that encapsulate the knowledge, based on the weighted infinity norm distance. SMPL enables known class classification and unseen attack identification through similarity comparison between prototypes and samples. Then, we design boundary calibration and overlap refinement to transform learned prototypes into priority-guided matching rules, ensuring precise and efficient in-network deployment. Additionally, Helios supports incremental prototype learning and rule updates, achieving low-cost hardware reconfiguration. We implement Helios on a Tofino switch and evaluation on three datasets shows that Helios achieves superior performance in classifying known classes (92%+ in ACC and F1) as well as identifying unseen attacks (62% - 98% in TPR). Helios has also reduced resource consumption and reconfiguration time, demonstrating its scalability and efficiency for real-world deployment.
Zhenning Shi, Dan Zhao 0003, Yijia Zhu, Guorui Xie, Qing Li 0006, Yong Jiang 0001
WWW2
2025 DNSGuard: In-Network Defense Against DNS Attacks
abstract
The Domain Name System (DNS) is a growing center of cyber attacks, including both volumetric and non-volumetric attacks. Programmable switches provide a new opportunity for more efficient defense against DNS attacks since they can offer better cost, performance, and flexibility trade-offs compared to traditional defense systems. However, programmable switches have strict limitations on the operations and storage space supported to ensure line-speed packet processing. In this paper, we propose DNSGuard, an intelligent in-network defense framework that can handle volumetric and non-volumetric DNS attacks on programmable switches. We propose a recursive incremental parsing algorithm that can effectively extract variable-length domain names. To achieve real-time and accurate detection against two types of DNS attacks, we design a switch-optimized and resource-efficient algorithm to extract both independent features of each packet and domain-based cumulative features. Then, we propose a multi-phase hybrid model architecture to perform dynamic packet analysis at different time phases of a domain. Further, we design efficient model representation mechanisms to deploy tree-based ensemble models in the data plane. Experimental results show that DNSGuard can defend against diverse DNS attacks at the line rate. In addition, DNSGuard introduces a minimal nanosecond latency to normal traffic in heavily loaded networks.
Guanglin Duan, Qing Li 0006, Dan Zhao 0003, Guorui Xie, Yuan Yang 0001, Zhenhui Yuan, Yong Jiang 0001, Mingwei Xu 0001
IEEE Trans. Dependable Secur. Comput.4
2025 Stateless and Proactive Routing for Dynamic Multicast With Deep Reinforcement Learning
abstract
Stateful multicast protocols manage multicast group memberships by maintaining state information about active groups and their members. They have seen limited adoption in the modern internet due to lack of scalability, simplicity, and flexibility. Although stateless multicast protocols, like BIER, eliminate extensive state management, they still face complex tree computation and limited scalability for concurrent requests. In this paper, we propose Hawkeye, a stateless multicast mechanism with deep reinforcement learning (DRL) for real-time responses to dynamic multicast requests with near-optimal multicast TE performance. This mechanism is suited for Software-Defined Networking (SDN) environment where the controller has a global view of the network and supports flexible configuration of network resources for traffic engineering. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership, and thus are able to build multicast trees proactively for upcoming requests. We develop a novel source aggregation mechanism to facilitate the convergence of the DRL agent under high volume of multicast requests. Moreover, to improve the practicality and robustness of Hawkeye, we design incremental deployment and single failure handling mechanisms, which take advantages of source aggregation and fit well with multicast routing. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye responds effectively to dynamic multicast requests. Itoffers rapid routing decisions, e.g., making routing decisions in under 5ms on a tested topology, and reduces path latency variation by up to 89.5%, with less than a 10% increase in bandwidth consumption compared to the offline theoretical minimum.
Qing Li 0006, Lie Lu, Dan Zhao 0003, Zeyu Luan, Yuan Yang 0001, Yong Jiang 0001, Jingpu Duan, Ruobin Zheng, Shaoteng Liu, Dingding Chen
IEEE Trans. Netw.3
2024 Quantized Side Tuning: Fast and Memory-Efficient Tuning of Quantized Large Language Models
abstract
Zhengxin Zhang, Dan Zhao, Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Qing Li, Yong Jiang, Zhihao Jia. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dan Zhao 0003, Xupeng Miao, Gabriele Oliaro, Zhihao Zhang 0001, Qing Li 0006, Yong Jiang 0001
ACL (1)2
2024 Proteus: A Difficulty-Aware Deep Learning Framework for Real-Time Malicious Traffic Detection
abstract
Deep learning (DL) has been recently used for malicious traffic detection. However, DL models are often faced with a dilemma between model size and performance: larger models have better accuracy, but suffer from high detection latency, which severely impacts realtime traffic performance, while lightweight models have low detection latencies, but sacrifice accuracy. In this paper, we introduce Proteus, a swift and precise attack detection framework that adaptively adjusts DL models in real-time based on sample detection difficulty. To address diverse detection difficulties in traffic data, we devise a Double Dynamic Convolutional Neural Network (DDCN) with two pivotal modules: the Dynamic Feature Campaign (DFC) and the Tailor Module (TM). DFC enables the model to discern and accentuate the most influential features, while TM autonomously gauges sample difficulty, cropping the overall model. We further design an auxiliary detection module to streamline the detection, especially for network devices like routers lacking GPUs but equipped with multiple CPU cores. Experiments on different network devices show that Proteus completes the detection of each flow within 0.6 ms, and achieves$\mathbf{9 9. 3 4 \%}$detection accuracy, outperforming other solutions.
Chupeng Cui, Qing Li 0006, Guorui Xie, Ruoyu Li 0003, Dan Zhao 0003, Zhenhui Yuan, Yong Jiang 0001
ICNP5
2024 Genos: General In-Network Unsupervised Intrusion Detection by Rule Extraction
abstract
Anomaly-based network intrusion detection systems (A-NIDS) use unsupervised models to detect unforeseen attacks. However, existing A-NIDS solutions suffer from low throughput, lack of interpretability, and high maintenance costs. Recent in-network intelligence (INI) exploits programmable switches to offer line-rate deployment of NIDS. Nevertheless, current in-network NIDS are either model-specific or only apply to supervised models. In this paper, we propose Genos, a general in-network framework for unsupervised A-NIDS by rule extraction, which consists of a Model Compiler, a Model Interpreter, and a Model Debugger. Specifically, observing benign data are multi-modal and usually located in multiple subspaces in the feature space, we utilize a divide-and-conquer approach for model-agnostic rule extraction. In the Model Compiler, we first propose a tree-based clustering algorithm to partition the feature space into subspaces, then design a decision boundary estimation mechanism to approximate the source model in each subspace. The Model Interpreter interprets predictions by important attributes to aid network operators in understanding the predictions. The Model Debugger conducts incremental updating to rectify errors by only fine-tuning rules on affected subspaces, thus reducing maintenance costs. We implement a prototype using physical hardware, and experiments demonstrate its superior performance of 100 Gbps throughput, great interpretability, and trivial updating overhead.
Ruoyu Li 0003, Qing Li 0006, Dan Zhao 0003, Xi Xiao 0001, Yong Jiang 0001
INFOCOM4
2024 Make Your Home Safe: Time-aware Unsupervised User Behavior Anomaly Detection in Smart Homes via Loss-guided Mask
abstract
Smart homes, powered by the Internet of Things, offer great convenience but also pose security concerns due to abnormal behaviors, such as improper operations of users and potential attacks from malicious attackers. Several behavior modeling methods have been proposed to identify abnormal behaviors and mitigate potential risks. However, their performance often falls short because they do not effectively learn less frequent behaviors, consider temporal context, or account for the impact of noise in human behaviors. In this paper, we propose SmartGuard, an autoencoder-based unsupervised user behavior anomaly detection framework. First, we design a Loss-guided Dynamic Mask Strategy (LDMS) to encourage the model to learn less frequent behaviors, which are often overlooked during learning. Second, we propose a Three-level Time-aware Position Embedding (TTPE) to incorporate temporal information into positional embedding to detect temporal context anomaly. Third, we propose a Noise-aware Weighted Reconstruction Loss (NWRL) that assigns different weights for routine behaviors and noise behaviors to mitigate the interference of noise behaviors during inference. Comprehensive experiments on three datasets with ten types of anomaly behaviors demonstrates that SmartGuard consistently outperforms state-of-the-art baselines and also offers highly interpretable results.
Jingyu Xiao, Zhiyao Xu, Qingsong Zou, Qing Li 0006, Dan Zhao 0003, Ruoyu Li 0003, Wenxin Tang, Xudong Zuo, Penghui Hu, Yong Jiang 0001, Zixuan Weng, Michael R. Lyu
KDD5
2024 Themis: A passive-active hybrid framework with in-network intelligence for lightweight failure localization
Jingyu Xiao, Qing Li 0006, Dan Zhao 0003, Xudong Zuo, Wenxin Tang, Yong Jiang 0001
Comput. Networks3
2024 SeIoT: Detecting Anomalous Semantics in Smart Homes via Knowledge Graph
abstract
Existing IoT Network Anomaly Detection Systems (NADSes) typically treat IoT devices as independent entities and model them by Euclidean space features. These approaches suffer from low accuracies on new attacks (e.g., platform-based attacks and evasion attacks), since they do not fully consider the semantic information including traffic periodicity and device/environment interactions. In this paper, we propose SeIoT, a knowledge graph-based bimodal anomaly detection framework for smart homes. We propose a knowledge graph structure to represent the semantic information of a smart home. First, we propose the Action Fingerprint module, an efficient and effective traffic classification approach to extract the device actions and features required by the knowledge graph. Then, we propose a bimodal anomaly detection framework including interaction-related and time-related detectors to detect the knowledge graph. We propose a feature separation-based heterogeneous graph attention network that can accurately model the interactions among devices and environments, and a method to represent traffic periodicity for the time-related detector. For evaluation, we set up a real-world testbed and evaluate the detection performance of both device-targeted attacks and platform-based attacks. Experiment results show that SeIoT can achieve better detection capability than prior work on both of the attacks.
Ruoyu Li 0003, Qing Li 0006, Qingsong Zou, Dan Zhao 0003, Yong Jiang 0001, Fa Zhu, Athanasios V. Vasilakos
IEEE Trans. Inf. Forensics Secur.5
2024 IoTGemini: Modeling IoT Network Behaviors for Synthetic Traffic Generation
abstract
Synthetic traffic generation can produce sufficient data for model training of various traffic analysis tasks for IoT networks with few costs and ethical concerns. However, with the increasing functionalities of the latest smart devices, existing approaches can neither customize the traffic generation of various device functions nor generate traffic that preserves the sequentiality among packets as the real traffic. To address these limitations, this paper proposes IoTGemini, a novel framework for high-quality IoT traffic generation, which consists of a Device Modeling Module and a Traffic Generation Module. In the Device Modeling Module, we propose a method to obtain the profiles of the device functions and network behaviors, enabling IoTGemini to customize the traffic generation like using a real IoT device. In the Traffic Generation Module, we design a Packet Sequence Generative Adversarial Network (PS-GAN), which can generate synthetic traffic with high fidelity of both per-packet fields and sequential relationships. We set up a real-world IoT testbed to evaluate IoTGemini. The experiment result shows that IoTGemini can achieve great effectiveness in device modeling, high fidelity of synthetic traffic generation, and remarkable usability to downstream tasks on different traffic datasets and downstream traffic analysis tasks.
Ruoyu Li 0003, Qing Li 0006, Qingsong Zou, Dan Zhao 0003, Xiangyi Zeng, Yong Jiang 0001, Feng Lyu 0001, Gaston Ormazabal, Henning Schulzrinne
IEEE Trans. Mob. Comput.4
2024 A Cooperative Caching System in Heterogeneous Edge Networks
abstract
Recently, the rapid growth of video content and the increasing demand for high Quality of Experience (QoE) have significantly strained the backbone network. Edge caching is a promising approach to alleviate the strain by caching content closer to users. However, it confronts challenges stemming from the low capability of individual edge nodes and the high density of their distribution, resulting in low hit ratios and unbalanced workloads. In this paper, we conduct in-depth analyses of these challenges and formulate a typical cooperative edge caching problem. Based on the insights, we introduce MagNet, a cooperative edge caching system featuring two key mechanisms: Automatic Content Congregating (ACC) and Mutual Assistance Group (MAG). ACC improves hit ratios by intelligently guiding requests to their optimal edges, thereby facilitating content aggregation. Complementing this, Quick Cache is implemented to accelerate this congregation process by prefetching content and optimizing cache space, effectively boosting hit ratios. MAG, on the other hand, achieves workload balance by dynamically forming groups to augment edge capabilities and redistribute requests on overloaded edges. To elucidate the design principles of MagNet, we conduct detailed component-level comparisons and quantitative analyses. To validate the overall performance, we compare it with various caching solutions using real-world datasets, demonstrating significant performance improvements.
Junkun Peng, Qing Li 0006, Dan Zhao 0003, Chuang Hu, Yong Jiang 0001
IEEE Trans. Mob. Comput.4
2024 DeviceRadar: Online IoT Device Fingerprinting in ISPs Using Programmable Switches
abstract
Device fingerprinting can be used by Internet Service Providers (ISPs) to identify vulnerable IoT devices for early prevention of threats. However, due to the wide deployment of middleboxes in ISP networks, some important data, e.g., 5-tuples and flow statistics, are often obscured, rendering many existing approaches invalid. It is further challenged by the high-speed traffic of hundreds of terabytes per day in ISP networks. This paper proposes DeviceRadar, an online IoT device fingerprinting framework that achieves accurate, real-time processing in ISPs using programmable switches. We innovatively exploit “key packets” as a basis of fingerprints only using packet sizes and directions, which appear periodically while exhibiting differences across different IoT devices. To utilize them, we propose a packet size embedding model to discover the spatial relationships between packets. Meanwhile, we design an algorithm to extract the “key packets” of each device, and propose an approach that jointly considers the spatial relationships and the key packets to produce a neighboring key packet distribution, which can serve as a feature vector for machine learning models for inference. Last, we design a model transformation method and a feature extraction process to deploy the model on a programmable data plane within its constrained arithmetic operations and memory to achieve line-speed processing. Our experiments show that DeviceRadar can achieve state-of-the-art accuracy across 77 IoT devices with 40 Gbps throughput, and requires only 1.3% of the processing time compared to GPU-accelerated approaches.
Ruoyu Li 0003, Qing Li 0006, Qingsong Zou, Dan Zhao 0003, Gareth Tyson, Guorui Xie, Yong Jiang 0001
IEEE/ACM Trans. Netw.5
2024 FlexNF: Flexible Network Function Orchestration for Scalable On-Path Service Chain Serving
abstract
Programmable Data Plane (PDP) has been leveraged to offload Network Functions (NFs). Due to its high processing capability, the PDP improves the performance of NFs by more than one order of magnitude. However, the coarse-grained NF orchestration on the PDP makes it hard to fulfill the dynamic service chain demands and unreasonable network function deployment causes long end-to-end delays. In this paper, we propose the Flexible Network Function (FlexNF) deployment on the PDP. First, we design an NF Selection Framework, leveraging the service selection label and re-entering operations for flexible NF orchestration. Second, to support runtime NF reconfiguration to meet the dynamic flow demands, we propose the Per-Flow On-Demand servicing mechanism, where one Match-Action Table with multiple mixed NFs works as different NFs for different flows. Third, to ensure the QoS of flows, on the one hand, we design an SP-aware NF Placement Algorithm to find a near-optimal placement solution that accommodates peak traffic volume while minimizing the overall routing path lengths of all the requests, on the other hand, we design a Two-Stage Service Path Construction Algorithm to provide on-path service while considering load balancing. We implement 15 types of network functions on the P4 switch, based on which we construct the comprehensive experiments. FlexNF reduces the traffic delay by 42.6% while increasing the service chain acceptance rate by five times compared with current solutions. Besides, when switching functions, the FlexNF improves the throughput by 2.04Gbps and reduces the packet loss by 8.269% compared with current solutions.
Jingyu Xiao, Xudong Zuo, Qing Li 0006, Dan Zhao 0003, Yong Jiang 0001, Jiyong Sun, Bin Chen 0011
IEEE/ACM Trans. Netw.4
2024 Empowering In-Network Classification in Programmable Switches by Binary Decision Tree and Knowledge Distillation
abstract
Given the high packet processing efficiency of programmable switches (e.g., P4 switches of Tbps), several works are proposed to offload the decision tree (DT) to P4 switches for in-network classification. Although the DT is suitable for the match-action paradigm in P4 switches, the range match rules used in the DT may not be supported across devices of different P4 standards. Additionally, emerging models including neural networks (NNs) and ensemble models, have shown their superior performance in networking tasks. But their sophisticated operations pose new challenges to the deployment of these models in switches. In this paper, we propose Mousikav2 to address these drawbacks successfully. First, we design a new tree model, i.e., the binary decision tree (BDT). Unlike the DT, our BDT consists of classification rules in the form of bits, which is a good fit for the standard ternary match supported by different hardware/software switches. Second, we introduce a teacher-student knowledge distillation architecture in Mousikav2, which enables the general transfer from other sophisticated models to the BDT. Through this transfer, sophisticated models are indirectly deployed in switches to avoid switch constraints. Finally, a lightweight P4 program is developed to perform classification tasks in switches with the BDT after knowledge distillation. Experiments on three networking tasks and three commodity switches show that Mousikav2 not only improves the classification accuracy by 3.27%, but also reduces the switch stage and memory usage by$2.00\times $and 28.67%, respectively. Code is available athttps://github.com/xgr19/Mousika.
Guorui Xie, Qing Li 0006, Guanglin Duan, Jiaye Lin, Yutao Dong, Yong Jiang 0001, Dan Zhao 0003, Yuan Yang 0001
IEEE/ACM Trans. Netw.7
2023 Gleaning the Consensus for Linearizable and Conflict-Free Per-Replica Local Reads
abstract
The optimal read strategy for strong consistent key-value applications is to enable the per-replica local reads that each replica has the ability to serve reads locally. Unfortunately, current schemes for the per-replica local reads are perplexed by two issues. First, some schemes have to violate the per-replica local reads when the workload is skewed, degrading the throughput. Second, most of current schemes rely on leases or a specialized hardware to guarantee the linearizability, bringing difficulties to the deployment.
Jian Yi, Qing Li 0006, Bin Zhang 0048, Yong Jiang 0001, Dan Zhao 0003, Yuan Yang 0001, Zhenhui Yuan
APNet5
2023 Efficient Attack Detection with Multi-Latency Neural Models on Heterogeneous Network Devices
abstract
To achieve fast and accurate attack detection, some works manually tailor neural networks (NNs) for deployment on CPUs of gateways, routers, or even programmable switches. However, with such solutions, NNs must be custom-tailored across different devices to meet the heterogeneous settings (e.g., OS and CPU types). Even worse, a model may require frequent adjustments to adapt to the same device's varying traffic rates. In this paper, we present Soteria, an automated multi-latency NN generation and scheduling system for fast and accurate detection against fluctuating traffic rates across heterogeneous hardware. Soteria first uses an evolutionary training algorithm to evolve the Pareto front, i.e., the set of NNs with a good spread on accuracy and model size. Then, for each device, Soteria filters the optimal multi-latency NNs by non-dominating sorting on the NNs' test latency on the device. Finally, to cope with the dynamic traffic rate, we design a heuristic scheduling scheme that adaptively selects NN s to maintain a balance between the detection accuracy and latency.
Guorui Xie, Qing Li 0006, Haolin Yan, Dan Zhao 0003, Gianni Antichi, Yong Jiang 0001
ICNP4
2023 Dryad: Deploying Adaptive Trees on Programmable Switches for Networking Classification
abstract
Decision trees (DT) have been used for high-speed networking classification on programmable switches. Most DT solutions, however, are static and cannot be deployed once the switch resource changes. In this paper, we propose Dryad to fast reprogram tree models when resource budgets change. In Dryad, we first develop a large and accurate “one-training-for-all“ DT (ODT) that can be quickly resized without computational retraining. ODTs are deployed in switches using a progressive search algorithm that searches the adaptations according to their resources. To achieve high accuracy and low packet latency, the adaptation leverages 1) innovative hard and soft pruning methods to compress the ODT rapidly with minimal performance loss; and 2) P4 scaling operations of match-action table arrangement and joint range-ternary match, which allow the switch to accommodate a larger (i.e., more accurate) ODT. Finally, an ODTCompiler is proposed to automatically convert the adapted ODT into a P4 program and then install it. Experimental results on three commodity switches under different resource scenarios show that Dryad achieves a higher classification F1-score (3.78 % higher), and completes the adaptation 161 × faster than other solutions.
Guorui Xie, Qing Li 0006, Jiaye Lin, Gianni Antichi, Dan Zhao 0003, Zhenhui Yuan, Ruoyu Li 0003, Yong Jiang 0001
ICNP5
2023 Hawkeye: A Dynamic and Stateless Multicast Mechanism with Deep Reinforcement Learning
abstract
Multicast traffic is growing rapidly due to the development of multimedia streaming. Lately, stateless multicast protocols, such as BIER, have been proposed to solve the excessive routing states problem of traditional multicast protocols. However, the high complexity of multicast tree computation and the limited scalability for concurrent requests still pose daunting challenges, especially under dynamic group membership. In this paper, we propose Hawkeye, a dynamic and stateless multicast mechanism with deep reinforcement learning (DRL) approach. For real-time responses to multicast requests, we leverage DRL enhanced by a temporal convolutional network (TCN) to model the sequential feature of dynamic group membership and thus is able to build multicast trees proactively for upcoming requests. Moreover, an innovative source aggregation mechanism is designed to help the DRL agent converge when faced with a large amount of multicast requests, and relieve ingress routers from excessive routing states. Evaluation with real-world topologies and multicast requests demonstrates that Hawkeye adapts well to dynamic multicast: it reduces the variation of path latency by up to 89.5% with less than 12% additional bandwidth consumption compared with the theoretical optimum.
Lie Lu, Qing Li 0006, Dan Zhao 0003, Yuan Yang 0001, Zeyu Luan, Jianer Zhou, Yong Jiang 0001, Mingwei Xu 0001
INFOCOM3
2023 SkyNet: Multi-Drone Cooperation for Real-Time Person Identification and Localization
Junkun Peng, Qing Li 0006, Yuanzheng Tan, Dan Zhao 0003, Zhenhui Yuan, Hanling Wang, Yong Jiang 0001
INFOCOM4
2023 Counterfactual Video Recommendation for Duration Debiasing
abstract
Duration bias widely exists in video recommendations, where models tend to recommend short videos for the higher ratio of finish playing and thus possibly fail to capture users' true interests. In this paper, we eliminate the duration bias from both data and model. First, based on the extensive data analysis, we observe that play completion rate of videos with the same duration presents a bimodal distribution. Hence, we propose to perform threshold division to construct binary labels as training labels for alleviating the drawback of finish playing labels overly biased towards short videos. Algorithmically, we resort to causal inference, which enables us to inspect causal relationships of video recommendations with a causal graph. We identify that duration has two kinds of effect on prediction: direct and indirect. Duration bias lies in the direct effect, while the indirect effect benefits prediction. To this end, we design a model-agnostic Counterfactual Video Recommendation for Duration Debiasing (CVRDD) framework, which incorporates multi-task learning to estimate different causal effect during training. In the inference phase, we perform counterfactual inference to remove the direct effect of duration for unbiased prediction. We conduct experiments on two industrial datasets, and in addition to achieving highly promising results on traditional top-k recommendation metrics, CVRDD also improves the user watch time.
Shisong Tang, Qing Li 0006, Dingmin Wang, Ci Gao, Wentao Xiao, Dan Zhao 0003, Yong Jiang 0001, Aoyang Zhang
KDD6
2023 Interpreting Unsupervised Anomaly Detection in Security via Rule Extraction
abstract
Many security applications require unsupervised anomaly detection, as malicious data are extremely rare and often only unlabeled normal data are available for training (i.e., zero-positive). However, security operators are concerned about the high stakes of trusting black-box models due to their lack of interpretability. In this paper, we propose a post-hoc method to globally explain a black-box unsupervised anomaly detection model via rule extraction. First, we propose the concept of distribution decomposition rules that decompose the complex distribution of normal data into multiple compositional distributions. To find such rules, we design an unsupervised Interior Clustering Tree that incorporates the model prediction into the splitting criteria. Then, we propose the Compositional Boundary Exploration (CBE) algorithm to obtain the boundary inference rules that estimate the decision boundary of the original model on each compositional distribution. By merging these two types of rules into a rule set, we can present the inferential process of the unsupervised black-box model in a human-understandable way, and build a surrogate rule-based model for online deployment at the same time. We conduct comprehensive experiments on the explanation of four distinct unsupervised anomaly detection models on various real-world datasets. The evaluation shows that our method outperforms existing methods in terms of diverse metrics including fidelity, correctness and robustness.
Ruoyu Li 0003, Qing Li 0006, Dan Zhao 0003, Yong Jiang 0001, Yong Yang 0001
NeurIPS4
2023 Metis: Understanding and Enhancing In-Network Regular Expressions
abstract
Regular expressions (REs) offer one-shot solutions for many networking tasks, e.g., network intrusion detection. However, REs purely rely on expert knowledge and cannot utilize labeled data for better accuracy. Today, neural networks (NNs) have shown superior accuracy and flexibility, thanks to their ability to learn from rich labeled data. Nevertheless, NNs are often incompetent in cold-start scenarios and too complex for deployment on network devices. In this paper, we propose Metis, a general framework that converts REs to network device affordable models for superior accuracy and throughput by taking advantage of REs' expert knowledge and NNs' learning ability. In Metis, we convert REs to byte-level recurrent neural networks (BRNNs) without training. The BRNNs preserve expert knowledge from REs and offer adequate accuracy in cold-start scenarios. When rich labeled data is available, the performance of BRNNs can be improved by training. Furthermore, we design a semi-supervised knowledge distillation to transform the BRNNs into pooling soft random forests (PSRFs) that can be deployed on network devices. To the best of our knowledge, this is the first method to employ model inference as an alternative to RE matching in network scenarios. We collect network traffic data on our campus for three weeks and evaluate Metis on them. Experimental results show that Metis is more accurate than original REs and other baselines, achieving superior throughput when deployed on network devices.
Guanglin Duan, Qing Li 0006, Dan Zhao 0003, Yong Jiang 0001, Lianbo Ma 0004, Xi Xiao 0001, Hengyang Xu
NeurIPS5
2023 HorusEye: A Realtime IoT Malicious Traffic Detection Framework using Programmable Switches
Yutao Dong, Qing Li 0006, Kaidong Wu, Ruoyu Li 0003, Dan Zhao 0003, Gareth Tyson, Junkun Peng, Yong Jiang 0001, Shutao Xia, Mingwei Xu 0001
USENIX Security Symposium5
2023 Deployment of UAV-BSs for on-demand full communication coverage
Xingwei Wang 0001, Min Huang 0001, Jie Jia 0001, Novella Bartolini, Qing Li 0006, Dan Zhao 0003
Ad Hoc Networks7
2023 Pontus: Finding Waves in Data Streams
abstract
The bumps and dips in data streams are valuable patterns for data mining and networking scenarios such as online advertising and botnet detection. In this paper, we define the wave, a data stream pattern with a serious deviation from the stable arrival rate for a period of time. We then propose Pontus, an efficient framework for wave detection and estimation. In Pontus, a lightweight data structure is utilized for the preliminary processing of incoming packets in the data plane to take advantage of its high processing speed; then, the powerful control plane carries out computationally intensive wave detection and estimation. In particular, we propose the Multi-Stage Progressive Tracking strategy which detects waves in stages and removes any disqualified items promptly to save memory. Hash collisions are addressed by a Stage Variance Maximization technique to reduce estimation error. Moreover, we prove the theoretical error bound and establish upper bounds of false positive and false negative. Experiment results show that the software version of Pontus can achieve around 97% F1-Score even under scarce memory when baselines fail. Furthermore, the implemented prototype of Pontus based on P4 achieves 842x higher throughput than the baseline strawman solution.
Qing Li 0006, Guanglin Duan, Dan Zhao 0003, Jingyu Xiao, Guorui Xie, Yong Jiang 0001
Proc. ACM Manag. Data4
2022 Drift-bottle: a lightweight and distributed approach to failure localization in general networks
abstract
Network failure severely impairs network performance, affecting latency and throughput of data transmission. Existing failure localization solutions for general networks face problems such as difficulty in acquiring data from end hosts, need for extra infrastructure, and excessive resource consumption. Meanwhile, solutions designed for data center networks are hard to apply in general networks, as they usually rely on the topology regularity of DCNs. In this paper, we propose Drift-Bottle, a lightweight and distributed approach to failure localization in general networks. In Drift-Bottle, each switch judges the status of flows and makes a local inference for suspicious links. We design a distributed localization scheme where each normal packet is used as a "drift-bottle" that carries a "letter", i.e., a lightweight inference header, while traversing the network. Each switch along the path updates the inference header by aggregating it with its local inference. Whenever the inference is evident enough to identify the culprit links of failures, a warning is sent to the operator immediately. Drift-Bottle implements its function mainly on the data plane of programmable switches and thus reduce the overhead brought to switches significantly. Evaluation based on simulation on different topologies demonstrates that Drift-Bottle provides fast, precise and lightweight failure localization to operators of general networks.
Xudong Zuo, Qing Li 0006, Jingyu Xiao, Dan Zhao 0003, Jiang Yong
CoNEXT4
2022 Soter: Deep Learning Enhanced In-Network Attack Detection Based on Programmable Switches
abstract
Though several deep learning (DL) detectors have been proposed for the network attack detection and achieved high accuracy, they are computationally expensive and struggle to satisfy the real-time detection for high-speed networks. Recently, programmable switches exhibit a remarkable throughput efficiency on production networks, indicating a possible deployment of the timely detector. Therefore, we present Soter, a DL enhanced in-network framework for the accurate real-time detection. Soter consists of two phases. One is filtering packets by a rule-based decision tree running on the Tofino ASIC. The other is executing a well-designed lightweight neural network for the thorough inspection of the suspicious packets on the CPU. Experiments on the commodity switch demonstrate that Soter behaves stably in ten network scenarios of different traffic rates and fulfills per-flow detection in 0.03s. Moreover, Soter naturally adapts to the distributed deployment among multiple switches, guaranteeing a higher total throughput for large data centers and cloud networks.
Guorui Xie, Qing Li 0006, Chupeng Cui, Peican Zhu, Dan Zhao 0003, Wanxin Shi, Zhuyun Qi, Yong Jiang 0001, Xi Xiao 0001
SRDS5
2015 Resource Allocation for Multiple Access Channel With Conferencing Links and Shared Renewable Energy Sources
abstract
This paper investigates the resource allocation problem for the Gaussian multiple access channel (MAC) with conferencing links, where the two transmitters can talk to each other via wired rate-limited channels. Moreover, the two transmitters are powered by a shared energy harvester which captures energy from the environment. We consider both the non-causal (the energy arrival levels at future time slots are known before transmissions) and the causal (only the energy arrival levels of past and present slots are known) energy-harvesting (EH) models. For the non-causal case, we formulate a resource allocation problem over a finite horizon ofNtime slots to characterize the boundary of the maximum departure region. We then develop the optimal offline power and rate allocation scheme by exploiting the hidden convexity of this problem. Interestingly, it is shown that there exists a maximum transmission rate (named the capping rate) for one of the transmitters. For the causal case, we examine the performance of the greedy scheme, in which the energy is depleted within each slot. In particular, we measure the utility of this scheme against the optimal offline one by competitive analysis, where the competitive ratio of the online greedy scheme, i.e., the maximum ratio between the profits obtained by the offline and online schemes over arbitrary energy arrival profiles, is derived.
Dan Zhao 0003, Chuan Huang 0001, Yue Chen 0002, Fuad E. Alsaadi, Shuguang Cui
IEEE J. Sel. Areas Commun.1
2013 Optimal resource allocation for multiple access channel with conferencing links and a shared renewable energy source
abstract
This paper investigates the optimal resource allocation for the Gaussian multiple access channel (MAC) with conferencing links, where the two transmitters could talk to each other via some wired rate-limited channels. Moreover, the two transmitters are assumed to be powered by a shared energy harvester, and a deterministic energy-harvesting (EH) model is adopted by assuming that the energy arrival times and the corresponding harvested amounts are non-causally known prior to transmissions. We formulate a continuous-time power allocation problem to characterize the maximum departure region over a finite time horizon. By exploiting its convexity, this problem is simplified as a discrete-time problem and the optimal solution is obtained. In particular, it is shown that there exists a certain maximum possible transmission rate (the capping rate) at one of the transmitters. Finally, we compare the performance of the optimal offline algorithm against that of the online one.
Dan Zhao 0003, Chuan Huang 0001, Yue Chen 0002, Shuguang Cui
ICASSP1
2013 Optimal resource allocation for multiple access channels with a shared renewable energy source
abstract
This paper investigates the optimal resource allocation for a Gaussian multiple access channel (MAC) with two transmitters powered by a shared energy harvester. A deterministic energy-harvesting (EH) model is adopted, which assumes that the energy arrival amounts and timing are non-causally known before transmissions. Besides, packets for both transmitters are assumed always ready before transmissions. We first formulate the resource allocation problem to characterize the maximum departure region over a finite time horizon as a convex optimization problem. The structural properties of the sum power profile is then studied by exploiting the convexity of the sum power function, which simplifies the optimization problem. Finally, the optimal resource allocation between the two transmitters, in which there exists a cut-off rate at the stronger transmitter, is obtained. We also demonstrate that under the same energy arrival profile, MAC with shared energy harvester achieves the same maximum departure region as its dual broadcast channel (BC).
Dan Zhao 0003, Chuan Huang 0001, Yue Chen 0002, Shuguang Cui
PIMRC1