VLDB 2026 Research / reviewers in the wild / expert
Zirui Zhuang
dblp:235/7014
· DBLP profile ↗
82ranked-venue papers
4as first author
76since 2021 · last 2026
0000-0003-3345-1732ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 30 · 4 first-author · 25 since 2021Artificial intelligence and machine learning · 22 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Systems, architecture and hardware · 10 · 9 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Example Quality Matters: Multi-Aspects Example Augmentation for Private Library ProgrammingabstractYuhao Li, Haifeng Sun, Xuesong Zhang, Shu Yao, Haoyu Zheng, Yvchuan Wang, Huazheng Wang, Zirui Zhuang, Qi Qi, Jianxin Liao, Jingyu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haifeng Sun 0001, Shu Yao, Yvchuan Wang, Huazheng Wang, Zirui Zhuang, Qi Qi 0001, Jianxin Liao, Jingyu Wang 0001 |
ACL (1) | 8 |
| 2026 | From Generation to Guarantee: Intent-Based Configuration Update with Verification Feedback
Lingqi Guo, Qi Qi 0001, Haifeng Sun 0001, Yuxing Peng 0007, Zirui Zhuang, Bo He 0003, Shaoling Sun, Jianxin Liao, Jingyu Wang 0001 |
INFOCOM | 6 |
| 2026 | RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching
Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao, Jing Wang 0039 |
INFOCOM | 7 |
| 2026 | OIPR: Evaluation for Time-Series Anomaly Detection Inspired by Operator InterestabstractWith the growing adoption of time-series anomaly detection (TAD) technology, numerous studies have employed deep learning-based detectors to analyze time-series data in the fields of Internet services, industrial systems, and sensors. The selection and optimization of anomaly detectors strongly rely on the availability of an effective evaluation for TAD performance. Since anomalies in time-series data often manifest as a sequence of points, conventional metrics that solely consider the detection of individual points are inadequate. Existing TAD evaluators typically employ point-based or event-based metrics to capture the temporal context. However, point-based evaluators tend to overestimate detectors that excel only in detecting long anomalies, while event-based evaluators are susceptible to being misled by fragmented detection results. To address these limitations, we propose OIPR1, a novel TAD evaluator with area-based metrics. It models the process of operators receiving detector alarms and handling anomalies, utilizing area under the operator interest curve to evaluate TAD performance. Furthermore, we build a special scenario dataset to compare the characteristics of different evaluators. Through experiments conducted on the special scenario dataset and five real-world datasets, we demon-strate the remarkable performance of OIPR in extreme and complex scenarios. It achieves a balance between point and event perspectives, overcoming their primary limitations and offering applicability to broader situations. Yuhan Jing, Jingyu Wang 0001, Lei Zhang 0094, Haifeng Sun 0001, Bo He 0003, Zirui Zhuang, Chengsen Wang, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | HyperWay: Proactively Mitigating Transient Congestion With Edge Capsule Tunnel in Massive IoT
Bo He 0003, Jinsheng Zhang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Yuhan Jing, Jing Shang 0001, Jianxin Liao |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | LLM-Powered Intent-Driven Configuration Generation for Multi-Vendor Networks
Jingyu Wang 0001, Bo He 0003, Jinyu Zhao, Yixin Xuan, Haifeng Sun 0001, Qi Qi 0001, Junzhe Liang, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Netw. Serv. Manag. | 8 |
| 2026 | Region Partitioning-Based Scalable Real-Time Network Verification via Native Distributed Architecture
Bo He 0003, Lingqi Guo, Chenyang Zhao 0005, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Gong Zhang 0001, Jianxin Liao, Cheng Huang 0001, Jingyu Wang 0001 |
IEEE Trans. Netw. | 7 |
| 2026 | Transient Resource Provisioning for Connected Autonomous Vehicles-Oriented Edge Slicing: A Learning-Based Two-Timescale ApproachabstractEdge slicing is envisioned to support connected autonomous vehicle (CAV) applications with diverse key performance indicator (KPI) requirements by splitting the shared physical infrastructure into several virtual networks. Unfortunately, existing provisioning approaches struggle to accommodate the spatiotemporal dynamics of CAV traffic, leading to significant violated KPIs or soared resource usage. In this paper, we introduce the transient sharing mechanism among edge slices to obtain reused gains without generating harmful performance interference, in which a slice is allowed to access to the under-utilized reserved resources of other slices but may experience interruptions at any time. Considering the heterogeneity and uncertainty of transient resources, we further develop a two-timescale provisioning scheme. Specifically, slices proactively make reservation decisions based on multi-armed bandit architectures at the beginning of large timescales, while hinging on cost-incentive auction mechanisms selectively preempt transient resources in terms of real-time application demands at each small timescale. With extensive experiments based on real traffic traces, we demonstrate that the proposed scheme can improve 10.43% resource utilization and make slices reduce 42.92% cost than state-of-the-art works, which verifies its high assurance and adaptability. Yu Liu 0016, Jingyu Wang 0001, Qi Qi 0001, Dezhi Chen, Zirui Zhuang, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Netw. | 5 |
| 2026 | Hammurabi: Establish Cooperative Order From Pre-Trained Policies in Multi-UAV Networks
Dezhi Chen, Hongchuan He, Qi Qi 0001, Jingyu Wang 0001, Rongxin Han, Bo He 0003, Zirui Zhuang, Qianlong Fu, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual DataabstractHuman experts typically integrate numerical and textual multimodal information to analyze time series. However, most traditional deep learning predictors rely solely on unimodal numerical data, using a fixed-length window for training and prediction on a single dataset, and cannot adapt to different scenarios. The powered pre-trained large language model has introduced new opportunities for time series analysis. Yet, existing methods are either inefficient in training, incapable of handling textual information, or lack zero-shot forecasting capability. In this paper, we innovatively model time series as a foreign language and construct ChatTime, a unified framework for time series and text processing. As an out-of-the-box multimodal time series foundation model, ChatTime provides zero-shot forecasting capability and supports bimodal input/output for both time series and text. We design a series of experiments to verify the superior performance of ChatTime across multiple tasks and scenarios, and create four multimodal datasets to address data gaps. The experimental results demonstrate the potential and utility of ChatTime. Chengsen Wang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao |
AAAI | 5 |
| 2025 | ClusterAttn: KV Cache Compression under Intrinsic Attention ClusteringabstractSparse attention can effectively alleviate the significant demands on memory when large language models (LLMs) process long contexts. Existing methods typically apply the same sparse pattern across different attention heads and inputs. However, this uniform approach fails to capture the inherent diversity of attention patterns within LLMs — the intrinsic attention clustering. To address this, we propose ClusterAttn, a training-free sparse attention method that provides an efficient prompt cache compression scheme under intrinsic attention clustering for efficient LLM inference.Our findings show that attention heads consistently focus on specific clusters of the prompt during decoding, a pattern detectable from an observation window at the prompt’s end. ClusterAttn adaptively fits these clusters utilizing a density-based attention clustering algorithm, thus compressing the KV cache of the prompt. Evaluations on different models across various benchmarks demonstrate ClusterAttn’s superior compression rates and efficiency. By utilizing only 1024 tokens, it can reduce memory usage by 10%–65%, resulting in a latency reduction of 12%–23% and a throughput increase of 2.6–4.8 times, all with nearly no accuracy loss. Additionally, ClusterAttn can handle up to 128k context on a single A100-80GB GPU, outperforming existing methods. Minwei Zhang, Haifeng Sun 0001, Jingyu Wang 0001, Shaolong Li, Wanyi Ning, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
ACL (1) | 7 |
| 2025 | Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand Reconstructionabstract3D hand reconstruction is essential in non-contact human-computer interaction applications, but existing methods struggle with low-resolution images, which occur in slightly distant interactive scenes. Leveraging temporal information can mitigate the limitations of individual low-resolution images that lack detailed appearance information, enhancing the robustness and accuracy of hand reconstruction. Existing temporal methods typically use joint features to represent temporal information, avoiding interference from redundant background information. However, joint features excessively disregard the spatial context of visual features, limiting hand reconstruction accuracy. We propose to integrate temporal joint features with visual features to construct a robust low-resolution visual representation. We adopt Triplane Features, a dense representation with 3D spatial awareness, to bridge the gap between the joint features and visual features that are misaligned in terms of representation form and semantics. Triplane Features are obtained by orthogonally projecting joint features, embedding hand structure information into the 3D spatial context. Furthermore, we compress the spatial information of the three planes into a 2D dense feature thourgh Spatial-Aware Fusion to enhance the visual features. By using enhanced visual features enriched with temporal information for hand reconstruction, our method achieves competitive performance at much lower resolutions compared to state-of-the-art methods operating at high resolution on DexYCB, HanCo and H2O. Code is available at https://github.com/NewbieFan/Temp-LowRes-hand. Kaixin Fan, Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
CVPR | 6 |
| 2025 | From Static to Dynamic: GNNs-Driven Clinical Decision-Making Assistance
Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao, Jiachang Hao, Haifeng Sun 0001 |
DASFAA (2) | 2 |
| 2025 | Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioabstractYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi, Zirui Zhuang, Huazheng Wang, Pengfei Ren, Jing Wang, Jianxin Liao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Huazheng Wang, Pengfei Ren 0001, Jing Wang 0039, Jianxin Liao |
EMNLP | 5 |
| 2025 | Atlas: Towards Real-Time Verification in Large-Scale Networks via a Native Distributed ArchitectureabstractData plane verification (DPV) can be critical in ensuring the network operates correctly. To be useful in practice, they need to be: (1) fast so as to prevent significant packet loss or security violations; (2) scalable so as to accommodate today's large-scale network architecture. Current DPV tools struggle to meet these requirements due to their centralized architecture. To be concrete, there is a bottleneck for a single-point server to perform real-time DPV tasks. Furthermore, a single-point server makes it hard to collect real-time data plane updates from every device in large-scale networks. Jingyu Wang 0001, Bo He 0003, Chenyang Zhao 0005, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Lingqi Guo, Yuebin Guo, Gong Zhang 0001, Jianxin Liao |
EuroSys | 7 |
| 2025 | Spy Inside: Scalable Verification of Dependable Transformers for Event Time Series SystemsabstractEvent time series appear in many software scenarios and are a necessary data type in data analytics systems. Transformers are the preferred type of sequential neural network for advanced analytics on event time series, particularly due to their significant contributions to the recent surge of large language models (LLMs). Event series analytics heavily depends on the quality of input data, which may contain natural measurement errors or adversarial noises. Since the input data deviates from the true state, the opaque nature of neural networks presents a challenge in ensuring the reliability of output, which might be deemed untrustworthy. In this paper, we introduce an innovative formal verification framework for Transformer-based event series systems, leveraging sampling, linear programming, and the extreme value theorem. This framework can support the verification of the dependability of Transformers in managing inputs characterized by unpredictability and uncertainty. To exemplify its utility, we apply our verification approach to verify natural requirements from a real-world event series environments: network traffic classification. It outperforms the current state-of-the-art verifier in terms of effectiveness, providing more stringent verified bounds. Our experimental findings provide valuable benchmarks for guaranteeing reliable deployment of systems in scenarios where the credibility of event data is compromised, and for exposing specific cases in which the expected requirements are not satisfied. Haodong Deng, Qi Qi 0001, Lu Lu 0015, Zirui Zhuang, Xingyu Zeng, Jinguang Wang, Bo He 0003, Wei Li 0119, Jingyu Wang 0001 |
ICASSP | 4 |
| 2025 | QoKV: Comprehending and Surpassing the Hurdles of KV Cache QuantizationabstractLarge language models (LLMs) have demonstrated outstanding performance in various tasks. However, the memory footprint of the key-value (KV) cache generated during model inference poses significant challenges for efficient model deployment. This paper presents a detailed analysis of the KV cache and identifies potential sources of quantization errors. We find that the attention scores exhibit a strong power-law distribution while showing an additional attention to the initial and recent tokens, and that only a small subset of outlier channels in KV cache span a wide numerical range. Based on these findings, we propose QoKV, an efficient post-training quantization (PTQ) method designed for the KV cache, which facilitates accurate low-bit quantization through a novel group evaluation strategy and numerical scaling. Extensive experimental results across various tasks demonstrate that our proposed method outperforms existing methods, achieving near-floating-point performance while saving 73.75% of the cache footprint. Jinguang Wang, Yuexi Yin, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jingyu Wang 0001 |
ICASSP | 6 |
| 2025 | Mitigating Object Hallucination in Large Vision-Language Models via Visual Attention Direct Preference OptimizationabstractLarge Vision-Language Models (LVLMs) suffer from severe object hallucinations, leading them to frequently generate outputs that do not correspond to the image content, significantly reducing the credibility and reliability of their responses. Recent research has attempted to enhance LVLMs by employing Direct Preference Optimization (DPO) to reduce hallucinations and improve response quality. However, these approaches typically utilize text-only preference response pairs, neglecting the influence of visual input in optimizing LVLMs. In this paper, we propose VA-DPO, a multimodal optimization objective. VA-DPO leverages the LVLMs' attention to select and corrupt critical parts of images, thereby constructing visual preference image pairs. This approach integrates both text and visual preference optimization objectives to achieve effective alignment optimization of LVLMs. We conduct extensive experiments on LVLMs of different sizes, and the results demonstrate that VA-DPO effectively reduces hallucinations across various tasks. Compared to other hallucination mitigation approaches, VA-DPO achieves more competitive results. Yixiao He, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Pengfei Ren 0001, Huazheng Wang, Yafeng Nan, Jingyu Wang 0001 |
ICME | 4 |
| 2025 | Beyond Statistical Analysis: Multimodal Framework for Time Series Forecasting with LLM-Driven Temporal PatternabstractAccurate forecasting of time series is crucial for many applications in the real world. Conventional methods primarily rely on statistical analysis of historical data, often leading to overfitting and failing to account for background information and constraints imposed by external events. Therefore, introducing large language models (LLMs) with robust textual capabilities holds significant potential. However, due to the inherent limitations of LLMs in handling numerical data, they do not exhibit advantages in precise numerical prediction tasks. Therefore, we propose a framework to integrate LLMs with conventional methods synergistically. Rather than directly outputting numerical predictions, we leverage the capabilities of the LLMs to generate textual temporal patterns, thereby fully utilizing their inherent knowledge and reasoning abilities. Additionally, we introduce a memory network designed to decode these textual representations into a format that numerical models can effectively interpret. This approach not only capitalizes on the strengths of the LLM in text processing but also bridges the gap between textual and numerical data, enhancing the overall predictive performance of the model. Our experimental results demonstrate the framework's effectiveness, achieving state-of-the-art performance on various benchmark datasets. Jiahong Xiong, Chengsen Wang, Haifeng Sun 0001, Yuhan Jing, Qi Qi 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
IJCAI | 6 |
| 2025 | Network CoPilot: Intent-Driven Network Configuration Updating for Service Guarantee
Rongxin Han, Jingyu Wang 0001, Haifeng Sun 0001, Zengteng Jiang, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
INFOCOM | 6 |
| 2025 | Foresail: LLM Sensor Knowledge Empowered Status-guided Network for Multivariate Time-series ClassificationabstractMultivariate time-series (MTS) classification tasks play a key role in data-driven applications spanning healthcare, finance, and mobile communication. As MTS data are typically collected from multiple interdependent sensors, the resulting temporal patterns inherently reflect the characteristics of the underlying sensing systems. Despite this connection, conventional MTS classification models predominantly focus on raw time-series data while disregarding valuable sensor-specific prior knowledge, which fundamentally constrains their classification accuracy. The emergence of large language models (LLMs) has encoded extensive sensor-related knowledge within their parameter spaces. However, effectively harnessing such knowledge to enhance MTS classification networks remains an open challenge. To address this, we propose Foresail, a status-guided neural framework that bridges this gap through systematic integration of LLM-derived sensor knowledge via the status relationship matrix and fine-grained status labels. Foresail can be seamlessly integrated with existing MTS networks to optimize performance and generate interpretable intermediate results. Experiments on irregularly and regularly sampled MTS data demonstrate that Foresail outperforms state-of-the-art approaches, achieving a notable improvement in F1-score of up to 10.9% compared to the basic MTS network. Yuhan Jing, Bo He 0003, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
ACM Multimedia | 5 |
| 2025 | Evaluating and Mitigating Object Hallucination in Large Vision-Language Models: Can They Still See Removed Objects?abstractYixiao He, Haifeng Sun, Pengfei Ren, Jingyu Wang, Huazheng Wang, Qi Qi, Zirui Zhuang, Jing Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixiao He, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Huazheng Wang, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039 |
NAACL (Long Papers) | 7 |
| 2025 | Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationabstractMulti-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings. Calibration-free approaches address these limitations by learning adaptive view interactions, typically leveraging expressive and flexible continuous representations. However, as the multiview interaction relationship is learned entirely from data without constraint, they are vulnerable to noisy input, which can propagate, amplify and accumulate errors across all views, severely corrupting the final estimated pose.
To mitigate this, we propose a novel framework that integrates a noise-resilient discrete prior into the continuous representation-based model. Specifically, we introduce the \textit{UniCodebook}, a unified, compact, robust, and discrete representation complementary to continuous features, allowing the model to benefit from robustness to noise while preserving regression capability.
Furthermore, we further propose an attribute-preserving and complementarity-enhancing Discrete-Continuous Spatial Attention (DCSA) mechanism to facilitate interaction between discrete priors and continuous pose features.
Extensive experiments on three representative datasets demonstrate that our approach outperforms both calibration-required and calibration-free methods, achieving state-of-the-art performance. Geng Chen 0006, Pengfei Ren 0001, Xufeng Jian, Haifeng Sun 0001, Menghao Zhang 0004, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 7 |
| 2025 | Generalizable Hand-Object Modeling from Monocular RGB Images via 3D GaussiansabstractRecent advances in hand-object interaction modeling have employed implicit representations, such as Signed Distance Functions (SDF) and Neural Radiance Fields (NeRF) to reconstruct hands and objects with arbitrary topology and photo-realistic detail. However, these methods often rely on dense 3D surface annotations, or are tailored to short clips constrained in motion trajectories and scene contexts, limiting their generalization to diverse environments and movement patterns. In this work, we present HOGS, an adaptively perceptive 3D Gaussian Splatting (3DGS) framework for generalizable hand-object modeling from unconstrained monocular RGB images. By integrating photometric cues from the visual modality with the physically grounded structure of 3D Gaussians, HOGS disentangles inherent geometry from transient lighting and motion-induced appearance changes. This endows hand-object assets with the ability to generalize to unseen environments and dynamic motion patterns. Experiments on two challenging datasets demonstrate that HOGS outperforms state-of-the-art methods in monocular hand-object reconstruction and photo-realistic rendering. Pengfei Ren 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 5 |
| 2025 | Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence PatternsabstractLarge Vision-Language Models (LVLMs) pretrained on large-scale multimodal data have shown promising capabilities in Video Anomaly Detection (VAD). However, their ability to reason about abnormal events based on scene semantics remains underexplored. In this paper, we investigate LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, focusing on whether their decisions are driven by statistical shortcuts between visual instances and textual phrases. By analyzing visual-textual co-occurrence in pretraining data and conducting experiments under different data settings, we reveal a hallucination phenomenon: LVLMs tend to rely on co-occurrence patterns between visual instances and textual phrases associated with either normality or abnormality, leading to incorrect predictions when these high-frequency objects appear in semantically mismatched contexts. To address this issue, we propose VAD-DPO, a direct preference optimization method supervised with counter-example pairs. By constructing visually similar but semantically contrasting video clips, VAD-DPO encourages the model to align its predictions with the semantics of scene rather than relying on co-occurrence patterns. Extensive experiments on six benchmark datasets demonstrate the effectiveness of VAD-DPO in enhancing both anomaly detection and reasoning performance, particularly in scene-dependent scenarios. Menghao Zhang 0004, Huazheng Wang, Pengfei Ren 0001, Kangheng Lin, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 7 |
| 2025 | FedMI: Reliable and Privacy-Aware Vertical Federated Learning for Anomaly Detection in Distributed Edge SystemsabstractAnomaly detection in distributed systems faces critical challenges from feature heterogeneity—where incomplete or divergent feature sets across nodes degrade detection relia-bility—and privacy risks under regulations like GDPR. While federated learning (FL) enables collaborative training without raw data sharing, existing solutions fail to address both challenges simultaneously: traditional FL suffers from performance drops under feature-missing scenarios, and differential privacy techniques introduce utility penalties. This paper proposes FedMI, a vertical federated learning framework that achieves provable privacy preservation and robust anomaly detection in feature-heterogeneous environments. FedMI's key innovations include a novel framework for vertical federated learning in anomaly detection for distributed systems that maintains high detection accuracy, mimicking real-world distributed system conditions, and a mutual information-guided training mechanism that quantifies and minimizes privacy leakage during federated updates. Evaluations on healthcare, financial, and industrial sensor datasets demonstrate FedMI's robustness: it achieves performance comparable to centralized methods in F1-score under data-island scenarios while ensuring compliance with privacy constraints. By unifying privacy quantification and robustness to feature heterogeneity, FedMI advances the development of dependable AI-driven monitoring for distributed systems. Zirui Zhuang, Qi Qi 0001, Haifeng Sun 0001, Shaoxiong Zhu, Xiaoyuan Fu, Jing Wang 0039 |
SRDS | 2 |
| 2025 | NetKeeper: Enhancing Network Resilience with Autonomous Network Configuration Update on Traffic Patterns and Anomalies
Zhaoyang Wan, Rongxin Han, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Bo He 0003, Jianxin Liao, Jingyu Wang 0001 |
USENIX ATC | 5 |
| 2025 | Robustness Verification of Deep Graph Neural Networks Tightened by Linear ApproximationabstractRecent research indicates that adding residual connections in Graph Neural Networks (GNNs) would amplify susceptibility to anomalous nodes, consequently undermining the robustness of deep GNNs in practical settings. However, existing verification methods encounter challenges with the increasing number of parameters and computational overhead in deep GNNs. In this paper, we derive the general form of the residual connections and apply the dual backpropagation network to deep GNNs. Considering the heightened computational errors arising from the increased number of layers in deep GNNs, we propose a new method for calculating intermediate activation bounds of GNNs based on linear approximation. Experimental results show that new method can effectively enhance the verification accuracy. Notably, the maximum perturbation value of nodes correctly classified shows an average improvement of 119.5%. To showcase the the efficacy and scalability of our method, we verify robustness of deep GNNs on six different graph datasets, and our method can effectively verify the robustness of deep GNNs even with 32 layers of residual connections, i.e. verify over 87.29% of nodes in the Citeseer dataset. Furthermore, we analyse the influence of the graph structural properties on the robustness of the model. Xingyu Zeng, Qi Qi 0001, Jingyu Wang 0001, Haodong Deng, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
WSDM | 7 |
| 2025 | Towards Bare-Hand Interaction for Whiteboard Collaboration in Virtual RealityabstractWhiteboard collaboration in virtual reality (VR) is an important task in collaborative virtual environments. The current research mainly relies on the use of controllers or dedicated pens but additional devices will cause inconvenience to users. Bare-hand writing offers rich collaborative semantics through natural gestures but remains underexplored. This paper addresses challenges and solutions for bare-hand whiteboard collaboration. We analyze the input process and identify key challenges in determining pen-drop, writing, and pen-lift intentions while maintaining user control over their avatar. Our approach addresses two VR scenarios: one without and one with physical planes. The method for the first case is called Air-writing, which dynamically adjusts the distance between the avatar's torso and the virtual whiteboard during the processes of pen-drop and pen-lift to ensure a consistent writing experience in VR. The method for the second case is called Physical-writing, which allows users to write smoothly with passive haptic feedback and physical constraints provided by the real surface by remapping the whiteboard in VR with a plane in reality. A comprehensive user study is conducted to evaluate communication efficiency, input accuracy, collaboration efficiency, and user experience of the two methods. The experimental results indicate that bare-hand interaction improves communication efficiency by 8% over controllers and performs similarly to real-world whiteboard collaboration. The Physical-writing method also demonstrates higher accuracy and user satisfaction compared to the Air-writing method. Guangtian Liu, Haonan Su, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Jianxin Liao |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2025 | DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge ClusterabstractParallelizing CNN inference on heterogeneous edge clusters with data parallelism has gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms struggle to find optimal parallel granularity for complex CNNS, the structure of which is a directed acyclic graph (DAG) rather than a chain, and the parallel dimension is inflexible. To distribute the workload of modern CNNs on heterogeneous devices is also proven as NP-hard problem. In this article, we introduce DeepZoning , a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution. Our experimental evaluations show that DeepZoning is effective, achieving up to a 3.02× speed improvement on our experimental prototype compared to state-of-the-art algorithms. Jingyu Wang 0001, Ruilong Ma, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Song Guo 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | Trust Model-Based Consensus Optimization for Vehicle Platooning Networks: A Novel Deep Reinforcement Learning Approach With GenAIabstractVehicle platooning has emerged as a promising solution for efficient traffic management. Multiple platoons traveling in a cooperative way can alleviate congestion and enhance driving safety by information sharing and consensus. To address the data security and privacy concerns, blockchain could be applied to enable secure data sharing and consensus across multiple platoons. However, existing performance of blockchain is insufficient to ensure reliable and efficient data consensus among multiple platoons. First, the hierarchical structure of platoons with different roles of vehicles complicates the trust establishment between platoons, making it challenging to evaluate their trustworthiness and ensure consensus reliability. Additionally, data sharing in vehicle platooning networks demands timely information and efficient consensus-building. To tackle above challenges, we design a role-adaptive trust model for trust evaluation of platoons in consideration of different roles of vehicles within a platoon. Based on the proposed model, we formulate a blockchain consensus optimization problem to facilitate both reliability and efficiency of data consensus among multiple platoons. Leveraging Generative Artificial Intelligence (GenAI) techniques, we then propose the Diffusion Enhanced Soft Actor-Critic (DESAC) by integrating the diffusion model and SAC, to further improve the performance of blockchain consensus. Experiment results demonstrate the effectiveness and efficiency of the proposed consensus optimization approach. Xiaoyuan Fu, Quan Yuan 0004, Zirui Zhuang, Jiawen Kang 0001, Zhiquan Liu 0001, Jingyu Wang 0001, Dusit Niyato |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Understanding and Guiding Weakly Supervised Entity Alignment with Potential Isomorphism PropagationabstractWeakly Supervised Entity Alignment (EA) is the task of identifying equivalent entities across diverse knowledge graphs (KGs) using only a limited number of seed alignments. Despite substantial advances in aggregation-based weakly supervised EA, the underlying mechanisms in this setting remain unexplored. In this article, we present a propagation perspective to analyze weakly supervised EA and explain the existing aggregation-based EA models. Our theoretical analysis reveals that these models essentially seek propagation operators for pairwise entity similarities. We further prove that, despite the structural heterogeneity across different KGs, the potentially aligned entities within aggregation-based EA models exhibit isomorphic subgraphs, a fundamental yet underexplored premise of EA. Leveraging this insight, we introduce a potential isomorphism propagation operator to enhance the propagation of neighborhood information across KGs. We develop a general EA framework, PipEA, incorporating this operator to improve the accuracy of every type of aggregation-based model without altering the learning process. Extensive experiments substantiate our theoretical findings and demonstrate PipEA’s significant performance gains over state-of-the-art weakly supervised EA methods. Our work advances the field and enhances our comprehension of aggregation-based weakly supervised EA. Haifeng Sun 0001, Yuanyi Wang, Wei Tang 0013, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2025 | MCAKE: Memory-Augmented Autoencoder with Contrastive Learning for Unsupervised Anomaly DetectionabstractRecently, reconstruction-based deep models have gained widespread usage in unsupervised anomaly detection. However, they may overlook some anomalies owing to the over-generalization of neural networks. Several studies have incorporated memory networks to mitigate this problem. Nonetheless, some of them lack an explicit memory updating process, while others rely on data-driven updating methods that are sensitive to initial values and unsuitable for end-to-end training. Additionally, the traditional criterion for detection computed in the high-dimensional input space may collapse as the spike in the deviation score is averaged across numerous dimensions. To address these challenges, we propose MCAKE, a M emory-augmented C ontrastive A utoencoder with K NN-Based E xtraction. It is designed to highlight the deviation score for anomalies by reconstructing input using fixed normal prototypes recorded in the memory. We explicitly encourage the memory to be autonomously learned and effectively allocated through contrastive learning with multiple positive and multiple negative samples. Furthermore, we introduce a bivariate detection criterion that calculates anomaly scores considering both input and latent space to tackle the collapse. Extensive experiments on 50 datasets across various categories demonstrate the superiority of our approach, with a 2% relative improvement over the previous state-of-the-art models. Chengsen Wang, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Yuhan Jing, Lianyuan Li, Jingyu Wang 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2025 | Hierarchical Index Retrieval-Driven Wireless Network Intent Translation With LLMabstractIntent-Based Networking (IBN) represents an emerging network management concept that is designed to fulfill user service requirements through automation. At its core, IBN is capable of translating user intent into network policies, thereby enabling automated configuration and management. However, the application of IBN has been limited by challenges associated with automation and intelligence. The recent widespread adoption of Large Language Model (LLM) has partially mitigated these issues. Nonetheless, hardware heterogeneity and high dynamic networks remain significant challenges for IBN: (i) Devices from different vendors are challenging to manage uniformly; (ii) Aligning service demands with rapidly changing network status is difficult. To address these challenges, we propose LIT, a framework of LLM-empowered Intent Translation with manual guidance. LIT incorporates Retrieval-Augmented Generation (RAG) to reference hardware manuals and enhance the generation results of LLMs. To reduce noise from retrieval results, we optimized the general RAG process. Additionally, LIT introduces MoE (Mixture of Experts) to adjust parameter values according to network status by synthesizing results from multiple expert models. Experiments demonstrate that LIT alleviates the challenges faced by IBN, achieving a 57.5% improvement in F1 score compared to the baseline. Jingyu Wang 0001, Lingqi Guo, Caijun Yan, Haifeng Sun 0001, Lei Zhang 0094, Zirui Zhuang, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | Federated Fine-Tuning on Heterogeneous LoRAs With Error-Compensated AggregationabstractFederated learning (FL) has recently been applied to the parameter-efficient fine-tuning (PEFT) of large language models (LLMs). While promising, client resource heterogeneity has imposed the challenge of the "bucket effect" to FL, where model configuration must cater to the client with the fewest resources. To tackle this issue, heterogeneous low-rank adaptation (LoRA) has recently emerged in FL, which enables clients to do local fine-tuning with different LoRA ranks. However, existing works in this area typically adopt zero-padding, stacking, or singular value decomposition (SVD) for LoRA aggregation, which often incur precision loss or significant overhead, limiting their practicality. In this article, we propose ECLoRA, a novel method for federated fine-tuning with heterogeneous LoRA settings across clients. ECLoRA employs randomized SVD (RSVD) to dramatically reduce aggregation overhead while introducing an error compensation (EC) mechanism that incorporates the decomposition error from previous rounds to improve aggregation precision. Extensive experiments on four widely used foundation models across six public tasks demonstrate the effectiveness of ECLoRA. Specifically, ECLoRA is: (1) accurate, significantly improving the final model performance; (2) fast, accelerating convergence with an average speedup of $1.54\times $ to $3.01\times $ ; and (3) practical, reducing aggregation time by approximately $40\times $ compared to classical SVD. Wanyi Ning, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Daixuan Cheng, Cong Liu 0046, Lei Zhang 0094, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Fast and Scalable Data Plane Verification for Burst Updates With Edge-PredicateabstractThere is an increasing interest in data plane verification, which is designed to automatically verify network correctness through directly analyzing the data plane. Recent data plane verifiers have been able to do real-time sub-millisecond per rule verification. However, we observe that in real-world networks, individual data plane updates rarely occur. On the contrary, there are always a certain number of updates generated in a short period of time, called asburst updates, due to high-level user intend or uncertain network events. When it comes to this real-life scenario, yet, the current equivalence class (EC) based methods are unable to solve themodel-wide changesproblem caused by the EC itself, which significantly slows down the verification speed of burst updates. To overcome this limitation, we present EPVerifier, a fast, scalable data plane verifier accelerating burst updates verification with edge-predicate (EP). Instead of classifying packets into ECs according to global forwarding behavior, the EPVerifier uses one EP per edge to represent all packets that can pass through. Furthermore, with EPs that clearly have localized properties, we introduce a rule type extension that does not require a change in the granularity of the network model to support ACLs and NATs that are prevalent in real devices, and obtain better-performing parallelism by dividing the verification task based on switches. Experiments on both dataset simulations and real-life deployments show that EPVerifier achieves 2-$10\times $faster data plane verification than the state-of-the-art and such advantage expand with the data plane’s complexity and update scale growth. Jingyu Wang 0001, Chenyang Zhao 0005, Zirui Zhuang, Qi Qi 0001, Yuebin Guo, Haifeng Sun 0001, Lingqi Guo, Jianxin Liao |
IEEE Trans. Netw. | 3 |
| 2025 | LogNotion: Highlighting Massive Logs to Assist Human Reading and Decision MakingabstractMassive logs contain crucial information about the working status of software systems, which contributes to anomaly detection and troubleshooting. For engineers, it is a laborious task to manually inspect raw logs to know the system running status, and therefore an automated log summarization tool can be helpful. However, due to the specificity of logs in terms of grammar, vocabulary and semantics, existing natural language-based methods cannot perform well in log analysis. To address these issues, we propose LogNotion, a general log summarization framework that highlights the log messages to assist human reading and decision making. We first explore the role played by triplets in log analysis, and propose a triplet extraction method based on sequence tagging and component alignment, in which the specificity of logs is fully taken into account. Then, we propose an unsupervised log summarization method to extract both regular and noteworthy information based on triplets. Comprehensive experiments are conducted on seven real-world log datasets and the results show that LogNotion improves the average ROUGE-1 by 0.26, recall by 0.12, and compression ratio by 2.13%, compared to state-of-the-art log summarization tools. The helpfulness, readability and generalizability are also verified through human evaluation and cross-dataset tests. Guojun Chu, Jingyu Wang 0001, Tao Sun 0010, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Serv. Comput. | 6 |
| 2025 | Anomaly Detection on Interleaved Log Data With Semantic Association Mining on Log-Entity GraphabstractLogs record crucial information about runtime status of software system, which can be utilized for anomaly detection and fault diagnosis. However, techniques struggle to perform effectively when dealing with interleaved logs and entities that influence each other. Although manually specifying a grouping field for each dataset can handle the single grouping scenario, the problems of multiple and heterogeneous grouping still remain unsolved. To break through these limitations, we first design a log semantic association mining approach to convert log sequences into Log-Entity Graph, and then propose a novel log anomaly detection model named Lograph. The semantic association can be utilized to implicitly group the logs and sort out complex dependencies between entities, which have been overlooked in existing literature. Also, a Heterogeneous Graph Attention Network is utilized to effectively capture anomalous patterns of both logs and entities, where Log-Entity Graph serves as a data management and feature engineering module. We evaluate our model on real-world log datasets, comparing with nine baseline models. The experimental results demonstrate that Lograph can improve the accuracy of anomaly detection, especially on the datasets where entity relationships are intricate and grouping strategies are not applicable. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Bo He 0003, Yuhan Jing, Lei Zhang 0094, Jianxin Liao |
IEEE Trans. Software Eng. | 5 |
| 2024 | Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationabstractPrevious 3D hand pose estimation methods primarily rely on a single modality, either RGB or depth, and the comprehensive utilization of the dual modalities has not been extensively explored. RGB and depth data provide complementary information and thus can be fused to enhance the robustness of 3D hand pose estimation. However, there exist two problems for applying existing fusion methods in 3D hand pose estimation: redundancy of dense feature fusion and ambiguity of visual features. First, pixel-wise feature interactions introduce high computational costs and ineffective calculations of invalid pixels. Second, visual features suffer from ambiguity due to color and texture similarities, as well as depth holes and noise caused by frequent hand movements, which interferes with modeling cross-modal correlations. In this paper, we propose Keypoint-Fusion for RGB-D based 3D hand pose estimation, which leverages the unique advantages of dual modalities to mutually eliminate the feature ambiguity, and performs cross-modal feature fusion in a more efficient way. Specifically, we focus cross-modal fusion on sparse yet informative spatial regions (i.e. keypoints). Meanwhile, by explicitly extracting relatively more reliable information as disambiguation evidence, depth modality provides 3D geometric information for RGB feature pixels, and RGB modality complements the precise edge information lost due to the depth noise. Keypoint-Fusion achieves state-of-the-art performance on two challenging hand datasets, significantly decreasing the error compared with previous single-modal methods. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
AAAI | 7 |
| 2024 | NetRen: Service Migration-Driven Network Renascence with Synthesizing Updated ConfigurationabstractChanges in enterprise networks require updated configurations. However, manual configurations with slow update efficiency, poor performance, and handling limitations, lead to the unavailability of updated networks. Therefore, we propose an efficient network renascence framework, NetRen, which synthesizes OSPF/BGP configurations driven by service and traffic migration. We follow the workflow of sketch extraction, configuration synthesis, and repair. Initially, comprehensive graphs are constructed to represent configuration sketches. We propose a GraphTrans synthesizer with Transformer's benefits of long-range focus and parallel reasoning. Training samples with the optimization relationship enable the synthesizer to achieve a mapping that optimizes performance based on configurations. To overcome the satisfiability barrier, configurations from the synthesizer are input to the stepwise configuration repairer as well-initialized solutions, achieving rapid configuration repair. Experiments demonstrate that the consistency of network configurations output by the GraphTrans synthesizer averages 98%. NetRen achieves a 312.4× increase in synthesis efficiency and a 5.83% improvement in network performance. Rongxin Han, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Chaowei Xu, Zhaoyang Wan, Zirui Zhuang, Yichuan Yu, Jianxin Liao |
ASPLOS (3) | 7 |
| 2024 | Distantly Supervised Contrastive Learning for Low-Resource Scripting Language SummarizationabstractCode summarization provides a natural language description for a given piece of code. In this work, we focus on scripting code—programming languages that interact with specific devices through commands. The low-resource nature of scripting languages makes traditional code summarization methods challenging to apply. To address this, we introduce a novel framework: distantly supervised contrastive learning for low-resource scripting language summarization. This framework leverages limited atomic commands and category constraints to enhance code representations. Extensive experiments demonstrate our method’s superiority over competitive baselines. Junzhe Liang, Haifeng Sun 0001, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
LREC/COLING | 3 |
| 2024 | Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation LearningabstractRecent progress in video anomaly detection suggests that the features of appearance and motion play crucial roles in distinguishing abnormal patterns from normal ones. However, we note that the effect of spatial scales of anomalies is ignored. The fact that many abnormal events occur in limited localized regions and severe background noise in-terferes with the learning of anomalous changes. Mean-while, most existing methods are limited by coarse-grained modeling approaches, which are inadequate for learning highly discriminative features to discriminate subtle differences between small-scale anomalies and normal patterns. To this end, this paper address multi-scale video anomaly detection by multi-grained spatiotemporal representation learning. We utilize video continuity to design three proxy tasks to perform feature learning at both coarse-grained and fine-grained levels, i.e., continuity judgment, discontinuity localization, and missing frame estimation. In particular, we formulate missing frame estimation as a contrastive learning task in feature space instead of a reconstruction task in RGB space to learn highly discriminative features. Experiments show that our proposed method outperforms state-of-the-art methods on four datasets, especially in scenes with small-scale anomalies. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Ruilong Ma, Jianxin Liao |
CVPR | 5 |
| 2024 | Coarse-to-Fine Implicit Representation Learning for 3D Hand-Object Reconstruction from a Single RGB-D Image
Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
ECCV (51) | 6 |
| 2024 | FLUK: Protecting Federated Learning Against Malicious Clients for Internet of Vehicles
Mengde Zhu, Wanyi Ning, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao |
Euro-Par (2) | 5 |
| 2024 | Interdependency Matters: Graph Alignment for Multivariate Time Series Anomaly DetectionabstractAnomaly detection in multivariate time series (MTS) is crucial for various applications in data mining and industry. Current industrial methods typically approach anomaly detection as an unsupervised learning task, aiming to identify deviations by estimating the normal distribution in noisy, label-free datasets. These methods increasingly incorporate interdependencies between channels through graph structures to enhance accuracy. However, the role of interdependencies is more critical than previously understood, as shifts in interdependencies between MTS channels from normal to anomalous data are significant. This observation suggests that anomalies could be detected by changes in these interdependency graph series. To capitalize on this insight, we introduce MADGA (MTS Anomaly Detection via Graph Alignment), which redefines anomaly detection as a graph alignment (GA) problem that explicitly utilizes interdependencies for anomaly detection. MADGA dynamically transforms subsequences into graphs to capture the evolving interdependencies, and Graph alignment is performed between these graphs, optimizing an alignment plan that minimizes cost, effectively minimizing the distance for normal data and maximizing it for anomalous data. Uniquely, our GA approach involves explicit alignment of both nodes and edges, employing Wasserstein distance for nodes and Gromov-Wasserstein distance for edges. To our knowledge, this is the first application of GA to MTS anomaly detection that explicitly leverages interdependency for this purpose. Extensive experiments on diverse real-world datasets validate the effectiveness of MADGA, demonstrating its capability to detect anomalies and differentiate interdependencies, consistently achieving state-of-the-art across various scenarios. Yuanyi Wang, Haifeng Sun 0001, Chengsen Wang, Mengde Zhu, Jingyu Wang 0001, Wei Tang 0013, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
ICDM | 8 |
| 2024 | Following the Compass: LLM-Empowered Intent Translation with Manual GuidanceabstractIntent-Based Networking (IBN) represents a novel paradigm of network automation and intelligence that has gradually been applied to network management. While the emergence of Large Language Models (LLMs) has improved the current state of IBN, hardware heterogeneity and high network dynamics remain significant challenges. Hardware heterogeneity requires that IBN effectively manage a diverse range of devices. The high network dynamics demands that IBN align service needs with rapidly changing network resources. We propose LIT, a framework of LLM-empowered Intent Translation with manual guidance. Given the outstanding language understanding and generation capabilities of LLM, LIT utilizes it in intent translation task. To further address two prevalent problems encountered in IBN, we introduce manual guidance and Mixture of Experts (MoE). Under the guidance of the manual, LLM improves its ability to generate high-quality policies that comply with syntax. After introducing MoE, it makes fine-grained adjustments to the parameters of policies based on network status and service requirements. The experimental outcomes demonstrate that LIT considerably alleviates numerous current challenges confronted by IBN and excels in intent translation, attaining an F1 score that is$\mathbf{5 6. 7 \%}$higher than the baseline model. Lingqi Guo, Jingyu Wang 0001, Caijun Yan, Haifeng Sun 0001, Zirui Zhuang, Qi Qi 0001, Haibao Ren, Jianxin Liao |
ICNP | 6 |
| 2024 | Safeguarding Sustainable Cities: Unsupervised Video Anomaly Detection through Diffusion-based Latent Pattern Learning
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao |
IJCAI | 6 |
| 2024 | Fast Policy Convergence for Traffic Engineering with Proactive Distributed Message-PassingabstractNowadays, the rise of various network applications makes network traffic become increasingly complex, which brings more stringent requirements to traffic engineering (TE). Although the state-of-the-art TE approaches based on deep reinforcement learning (DRL) or traditional methods can generate optimal solutions for fixed traffic matrices, they cannot converge fast enough to provide real-time optimization in real networks either because of excessive computation times or high communication overheads. Moreover, due to the dynamically changing traffic load on the network, it is also challenging to achieve optimization of maximum link utilization (MLU) and end-to-end delay at the same time since these two optimization objectives may be conflicting, especially when the network is under a low traffic load, which makes the modeling very difficult. To meet these challenges, we present RT-TE, a TE system based on DRL and distributed message-passing between intelligent agents that can achieve real-time optimization for both MLU and end-to-end delay. To reduce the communication time due to link propagation delay during the optimization process, we design a proactive message-passing mechanism that allows agents to use partial messages to compute the routing policy while maintaining the optimization performance. Additionally, to achieve the tradeoff between the two optimization objectives, we model the propagation delay into the DRL model and design a multi-objective training framework with parameter transfer for training. Based on theoretical modeling, we can find the best tradeoff between the two objectives. Moreover, to improve the model's generalization for various traffic flows, we use a GNN model to generate the rewards of the DRL model, which greatly speeds up the training phase and allows us to feed massive amounts of data into the model. Through evaluations of real-world network topologies, our approach shows a 10%-20% improvement in optimizing MLU under short traffic-changing intervals and yields a 9%-13% improvement in optimizing end-to-end delay compared to state-of-the-art approaches. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
IPDPS | 2 |
| 2024 | Beyond Throughput-Optimal: Second-Order Smooth Backpressure Algorithm for Reducing Jitter and DelayabstractIn the imminent era of 6G, Quality of Service (QoS) emerges as a pivotal concern in wireless communications. The prescribed transmission rates and vast access demands mandated by 6G standards impose heightened requirements on network throughput and delay. However, the highly dynamic and often bursty nature of application demands presents challenges for routing and congestion control. The backpressure-based joint rate and routing control algorithm adaptively adjusts network traffic to achieve optimal throughput. However, varying traffic conditions hinder the algorithm’s convergence to ideal states. Additionally, relying solely on first-order backlog differences for forwarding can lead to poor convergence and high delays. In this study, we propose a Second-Order Smooth Backpressure (SoSBP) algorithm, leveraging second-order backlog metrics and dual-level queue mapping, to address throughput, delay, and jitter issues in dynamic network environments. We validate the efficacy of this novel backlog metric using Lyapunov optimization techniques. Simulation results demonstrate that our approach significantly reduces end-to-end delay and data jitter while preserving throughput and eliminating routing loops. Yuexi Yin, Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Xiaoyuan Fu, Jianxin Liao |
IWQoS | 2 |
| 2024 | STAR-VP: Improving Long-term Viewport Prediction in 360° Videos via Space-aligned and Time-varying FusionabstractAccurate long-term viewport prediction in tile-based 360° video adaptive streaming helps pre-download tiles for a further future, thus establishing a longer buffer to cope with network fluctuations. Long-term viewport motion is mainly influenced by Historical viewpoint Trajectory (HT) and Video Content information (VC). However, HT and VC are difficult to align in space due to their different modalities, and their relative importance in viewport prediction varies across prediction time steps. In this paper, we propose STAR-VP, a model that fuses HT and VC in a Space-aligned and Time-vARying manner for Viewport Prediction. Specifically, we first propose a novel saliency representation salxyz and a Spatial Attention Module to solve the spatial alignment of HT and VC. Then, we propose a two-stage fusion approach based on Transformer and gating mechanisms to capture their time-varying importance. Visualization of attention scores intuitively demonstrates STAR-VP's capability in space-aligned and time-varying fusion. Evaluation on three public datasets shows that STAR-VP achieves state-of-the-art accuracy for long-term (2-5s) viewport prediction without sacrificing short-term (<1s) prediction performance. Baoqi Gao, Daoxu Sheng, Lei Zhang 0094, Qi Qi 0001, Bo He 0003, Zirui Zhuang, Jingyu Wang 0001 |
ACM Multimedia | 6 |
| 2024 | Video Anomaly Detection via Progressive Learning of Multiple Proxy TasksabstractLearning multiple proxy tasks is a popular training strategy in semi-supervised video anomaly detection. However, the traditional method of learning multiple proxy tasks simultaneously is prone to suboptimal solutions, and simply executing multiple proxy tasks sequentially cannot ensure continuous performance improvement. In this paper, we thoroughly investigate the impact of task composition and training order on performance enhancement. We find that ensuring continuous performance improvement in multi-task learning requires different but continuous optimization objectives in different training phases. To this end, a training strategy based on progressive learning is proposed to enhance the multi-task learning in VAD. The learning objectives of the model in previous phases contribute to the training in subsequent phases. Specifically, we decompose video anomaly detection into three phases: perception, comprehension, and inference, continuously refining the learning objectives to enhance model performance. In the three phases, we perform the visual task, the semantic task and the open-set task in turn to train the model. The model learns different levels of features and focuses on different types of anomalies in different phases. Extensive experiments demonstrate the effectiveness of our method, highlighting that the benefits derived from the progressive learning transcend specific proxy tasks. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Huazheng Wang, Lei Zhang 0094, Jianxin Liao |
ACM Multimedia | 6 |
| 2024 | QUIC-Enabled Framework for Alleviating Transient Congestion in Time-Critical IoTabstractThe real-time control capability of IoT devices is contingent upon the transmission of packets. However, due to the influence of multiple devices accessing the network, the bandwidth available to IoT devices from access points may decline significantly, which causes a surge in the queuing latency and interrupts the transmission. This phenomenon is referred to as transient congestion. To achieve stable high-quality network service, this paper designs RushWay, a QUIC-enabled framework for alleviating transient congestion. RushWay employs stream multiplexing to compress the original packet into the stream frame, thereby reducing the bandwidth required for transmission. Furthermore, RushWay employs an adaptive decision-making algorithm to assess uplink queue conditions and packet latency requirements, thereby alleviating transient congestion. The simulation results demonstrate that RushWay can improve key performance by 13% to 94%. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
MobiCom | 6 |
| 2024 | Rethinking the Power of Timestamps for Robust Time Series Forecasting: A Global-Local Fusion PerspectiveabstractTime series forecasting has played a pivotal role across various industries, including finance, transportation, energy, healthcare, and climate. Due to the abundant seasonal information they contain, timestamps possess the potential to offer robust global guidance for forecasting techniques. However, existing works primarily focus on local observations, with timestamps being treated merely as an optional supplement that remains underutilized. When data gathered from the real world is polluted, the absence of global information will damage the robust prediction capability of these algorithms. To address these problems, we propose a novel framework named GLAFF. Within this framework, the timestamps are modeled individually to capture the global dependencies. Working as a plugin, GLAFF adaptively adjusts the combined weights for global and local information, enabling seamless collaboration with any time series forecasting backbone. Extensive experiments conducted on nine real-world datasets demonstrate that GLAFF significantly enhances the average performance of widely used mainstream forecasting models by 12.5\%, surpassing the previous state-of-the-art method by 5.5\%. Chengsen Wang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
NeurIPS | 5 |
| 2024 | EPVerifier: Accelerating Update Storms Verification with Edge-Predicate
Chenyang Zhao 0005, Yuebin Guo, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Lingqi Guo, Yuming Xie, Jianxin Liao |
NSDI | 5 |
| 2024 | UAV Trajectory Tracking via RNN-Enhanced IMM-KF with ADS-B DataabstractWith the increasing use of autonomous unmanned aerial vehicles (UAVs), it is critical to ensure that they are continuously tracked and controlled, especially when UAVs op-erate beyond the communication range of ground stations (GSs). Conventional surveillance methods for UAVs, such as satellite communications, ground mobile networks and radars are subject to high costs and latency. The automatic dependent surveillance-broadcast (ADS-B) emerges as a promising method to monitor UAVs, due to the advantages of real-time capabilities, easy deployment and affordable cost. Therefore, we employ the ADS-B for UAV trajectory tracking in this work. However, the inherent noise in the transmitted data poses an obstacle for precisely tracking UAVs. Hence, we propose the algorithm of recurrent neural network-enhanced interacting multiple model-Kalman filter (RNN-enhanced IMM-KF) for UAV trajectory filtering. Specifically, the algorithm utilizes the RNN to capture the maneuvering behavior of UAVs and the noise level in the ADS-B data. Moreover, accurate UAV tracking is achieved by adaptively adjusting the process noise matrix and observation noise matrix of IMM-KF with the assistance of the RNN. The proposed algorithm can facilitate GSs to make timely decisions during trajectory deviations of UAVs and improve the airspace safety. Finally, via comprehensive simulations, the total root mean square error of the proposed algorithm decreases by 28.56%, compared to the traditional IMM-KF. Ziye Jia, Qihui Wu 0001, Chao Dong 0001, Zirui Zhuang, Huiling Hu |
WCNC | 5 |
| 2024 | TacNet: A Tactic-Interactive Resource Allocation Method for Vehicular NetworksabstractTo support safety driving and various on-board services, efficient resource allocation is crucial for the promising implement of vehicle platooning in intelligent transportation systems (ITSs). The resource allocation of vehicle-to-everything (V2X) communications for vehicular platoons is studied in this article. First, a multiobjective function is formulated to jointly optimize sub-band and power allocation to satisfy Quality-of- Service (QoS) in vehicular networks. With the advantage of dealing with complex decision-making problems in multiagent systems, distributed multiagent deep reinforcement learning (MADRL) stands out for resource allocation of vehicular networks. However, it faces the challenge of cooperation aging when every agent is only learning from information of others to form a cooperation model in the training process. Considering the random and dynamic combination of vehicles in vehicle platooning, a tactic-interactive MADRL method named as TacNet is then proposed to improve the cooperation efficiency of multiple agents. In TacNet, the tactics of other agents will be encoded and transmitted through interactive communications among agents. In addition, with the development of vehicular edge computing (VEC), digital twin (DT) networks are constructed to assist offloading computation-intensive resource allocation tasks in vehicles to the edge. The superiority of the proposed method is verified through extensive simulation results, which refers to convergence and performance of satisfying diversified QoS requirements compared with state-of-the-art MADRL methods. Xiaoyuan Fu, Quan Yuan 0004, Zirui Zhuang, Jianxin Liao, Dongmei Zhao |
IEEE Internet Things J. | 3 |
| 2024 | Slice Sandwich: Jagged Slicing Multi-Tier Dynamic Resources for Diversified V2X ServicesabstractWith the advancement of intelligent transportation systems, a series of diversified V2X applications come into being, which have different key performance indicators (KPIs) and transmission features. Moreover, multi-tier computing as a new system-level architecture distributes computing and communication capabilities anywhere between the cloud and the end-user. Unfortunately, the existing network paradigm for V2X services adopts a one-shot allocation of resources ignoring the inherent differences of V2X service. To cope with these problems, three types of refined network slices for V2X services are first proposed to simultaneously support heterogeneous service characteristics without excessively splitting resources. Considering the spatiotemporal correlation between service traffic and physical resources, a jagged slicing in multi-tier dynamic resources, which forms a “slice sandwich” brightly, is realized by a dual timescale intelligent resource management scheme. The inter-slice resource configuration is based on neural bandits with upper confidence bounds at each large-time period, while the exclusive resources are managed elastically by deep Q-learning in terms of the real-time changing network state in the small slot. We developed a simulation environment by Simulation of Urban Mobility (SUMO) including real-world road conditions and traffic models. The experiment results demonstrate that the proposed scheme can effectively guarantee KPIs of V2X services and improve the system revenue compared with benchmark algorithms. Yu Liu 0016, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Dezhi Chen, Lu Lu 0015, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Dynamic Network Slice for Bursty Edge TrafficabstractEdge network slicing promises better utilization of network resources by dynamically allocating resources on demand. However, addressing the imbalance between slice resources and user demands becomes challenging when complex user behaviors lead to bursty traffic within the edge network. Hence, we propose a comprehensive dynamic slice strategy with two coupled sub-strategies (i) bursty-sensitive slice resource coordination and (ii) proactive demand resource matching to find an optimal balance. For obtaining stable strategies, the edge network with bursty traffic is formulated as a bi-level Lyapunov optimization problem. Then we propose a resource allocation and request redirection (RA-RR) algorithm with polynomial complexity by introducing deep reinforcement learning to guarantee real-time. Specifically, two agents are trained to solve two sub-strategies, and the Lyapunov drift-plus-penalty function is used as the reward to keep queues stable. RA-RR is responsive to fluctuations in demand and realizes an efficient interaction of coupled decision-making. Moreover, a training method based on alternating optimization is designed to ensure convergence of the RA-RR algorithm. Experiments demonstrate that the proposal can maximize network revenue while ensuring the stability of slice services when edge traffic bursts, and has an average improvement of 20.4% compared with comparisons. Rongxin Han, Jingyu Wang 0001, Qi Qi 0001, Dezhi Chen, Zirui Zhuang, Haifeng Sun 0001, Xiaoyuan Fu, Jianxin Liao, Song Guo 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Fast and Scalable ACL Policy Solving Under Complex Constraints With Graph Neural NetworksabstractNetwork operators often need to modify Access Control List (ACL) policies to align with to network upgrades. An essential part of the ACL update task is reachability satisfaction. Previous studies formalize reachability requirements as a set of constraints and then use Boolean Satisfiability (SAT) or Satisfiability Modulo Theories (SMT) solvers to search for solutions. However, as today’s networks grow in size and complexity, the constraints derived from the requirements become increasingly complex, leading to an unacceptable time cost to obtain a correct policy. The sluggish updating of ACL policies can affect the properties of a network, such as connectivity and security. This paper presents a novel approach for fast and scalable ACL policy synthesis under complex constraints. We utilize Graph Neural Networks (GNNs) to learn the relations between nodes and reason the solution that satisfies the update requirements. We further integrate global position encoding into the GNN architecture, which allows for better differentiation of nodes in ACL update tasks. Additionally, an enhanced stochastic local search solver is introduced to address incorrect predictions made by the GNN. Experiments on real-world topologies show that GNN saves up$278\times $time costs compared to advanced SAT/SMT solvers on a 125-node network, and this advantage expands with the network size. Furthermore, our model extrapolates well when faced with different requirements and topologies, demonstrating its ability to handle frequent network upgrades. Haifeng Sun 0001, Xingjian Liao, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao, Dapeng Oliver Wu |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Diner: Interpretable Anomaly Detection for Seasonal Time Series in Web ServicesabstractMonitoring and anomaly detection of key performance indicators (KPIs) are crucial for large Internet companies to maintain the reliability of their Web services. Influenced by human behavior and schedules, the KPIs of Web services typically exhibit seasonal characteristics. These characteristics may be complex as different KPIs exhibit differences in trend, multiple periods, and noise behaviors. However, existing anomaly detection methods typically only model one fixed pattern of seasonal KPIs, which may lead to performance degradation when dealing with diverse seasonal KPIs. In this work, we propose a novel anomaly detection model for seasonal KPIs,Diner, which incorporates multiple interpretable components. It is able to capture the additive and multiplicative trends, multiple periods, and seasonal noise in intricate seasonal KPIs, making it easily adaptable to different types of seasonal KPIs. Additionally, we present a set of evaluation criteria for generic time series anomaly detection tasks, which prove more effective in handling ambiguous manual labels and various anomaly events. Experiments are conducted on three real-world datasets, and the performanceDinersurpassed both the statistical baseline and the state-of-the-art deep learning baselines. Yuhan Jing, Jingyu Wang 0001, Ji Qi 0005, Qi Qi 0001, Bo He 0003, Zirui Zhuang, Naixing Wu, Jianxin Liao |
IEEE Trans. Serv. Comput. | 6 |
| 2024 | Cognition Guided Video Anomaly Detection Framework for Surveillance ServicesabstractThe aim of surveillance services is to detect anomalous events that occur in given surveillance videos. Most existing video anomaly detection methods rely on minimizing reconstruction or prediction errors due to the lack of abnormal data, which results in poor generalization and overfitting. In fact, cognitions for anomalies in surveillance videos mainly relies on crucial relationships, including ones between objects and ones between objects and scenes. Focusing on this property of anomaly detection, aCognitionGuidedVideoAnomalyDetection framework based on prior knowledge is proposed, calledCG-VAD. CG-VAD introduces both explicit and implicit prior knowledge into the frame prediction network to let the model exploit crucial relationships. Explicit knowledge containing crucial relationships related to anomaly is introduced into the anomaly detection model through a proposed embedding network based on multi-layer Graph Convolutional Networks. Implicit knowledge in the form of learnable parameters enhances the ability of the model to learn crucial relationships through prompt tuning. By integrating prior knowledge to focus the model on the relationships associated with the anomaly, we find that CG-VAD is not only quick to adapt to new real-world scenarios, but it is also able to recognize the type of anomaly. We have conducted extensive experiments on four benchmark datasets and the results indicate that the proposed method outperforms previous methods. Specifically, CG-VAD achieves an AUROC score of 87.2$\%$on the ShanghaiTech dataset. Code is available athttps://github.com/zmh0124/CG-VAD. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | Robust Video Anomaly Detection Framework via Prior Knowledge and Multi-Path Frame PredictionabstractVideo anomaly detection aims to automatically detect abnormal objects or behaviors. Most existing methods tackle the problem by minimizing the reconstruction errors stemming from the lack of anomalous data, which leads to poor interpretability and robustness. Focus on the context-dependent nature of anomaly detection, a robust unsupervised Video Anomaly Detection framework based on Knowledge and Frame Prediction is proposed, called VAD-KFP. Prior knowledge which contains the context of anomaly is introduced into the multi-path frame prediction network through multi-layer Graph Convolutional Networks. By integrating the prior knowledge to accurately define anomalies, VAD-KFP is robust to different scenarios and is able to recognize the type of anomaly. An extensive range of experiments have been conducted on three benchmarks, the results of which indicate that our method outperforms strong baselines. Specifically, VAD-KFP obtains an AUROC score of 91.6% for the Avenue dataset. Menghao Zhang 0004, Jingyu Wang 0001, Jing Wang 0039, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001 |
ICASSP | 5 |
| 2023 | Deep Reinforcement Learning Based Fast Anomaly Detection and Localization for Programmable NetworksabstractThe fast anomaly detection and localization is essential for network management, however, it is also very challenging for the current networks due to the lack of flexible control and telemetry capabilities. Fortunately, the maturity of Deep Reinforcement Learning (DRL) and programmable networking technologies could shed a light on realizing fast and intelligent anomaly detection and localization. In the paper, we design a fast anomaly detection and localization system for programmable networks by leveraging the in-band network telemetry and flexible control capabilities of programmable networks. Based on the system, we propose a DRL-based abnormal link detection and localization algorithm. It can iteratively infer abnormal links based on the ingress-to-egress performance metrics of flows and the one-hop performance metrics of the flows on the already identified abnormal links. The simulation results show that our proposals can detect and localize link anomalies in a matter of seconds to tens of seconds with low network telemetry overhead. Peng Zhan, Guangyi Qin, Xingxin Qian, Xiong Wang 0001, Jing Ren 0002, Zirui Zhuang, Shizhong Xu |
ICC | 6 |
| 2023 | CONFPILOT: A Pilot for Faster Configuration by Learning from Device ManualsabstractThe command line interface (CLI) is widely used to configure and manage network devices. However, as heterogeneous devices are introduced into the network, the CLI-based method is becoming time-consuming and inefficient because much effort is required to learn proprietary configuration languages of different vendors or consult online documents. In this work, we present CONFPILOT, an assistant system that can accelerate configuration by automatically converting natural language intents into commands. Our solution is based on a retrieval-augmented generation framework that features a unified parser that parses device manuals into a searchable configuration library, a vendor-agnostic retriever that finds the most relevant$k$syntaxes through a two-stage coarse-to-fine process, and a reliable generator that predicts syntactically correct commands via a pointer-generator network and syntax-guided decoding. In a nutshell, CONFPILOT frees engineers from most time-consuming efforts by learning directly from device manuals to generate configuration commands. Our evaluation and user study show, CONFPILOT can speed up the configuration process by 60x compared to manual lookup while maintaining acceptable exact match accuracy. Furthermore, CONFPILOT can quickly adapt to new vendors and devices with little human effort and time cost. Jinyu Zhao, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Shimin Tao, Jianxin Liao |
ICDCS | 5 |
| 2023 | Solving Distributed ACL Policies Under Complex Constraints with Graph Neural NetworksabstractAccess Control List (ACL) policies often need to be updated due to upgrades in network architecture and services. A critical part of ACL update tasks is reachability satisfaction, which is typically handled using Boolean Satisfiability (SAT) or Satisfiability Modulo Theories (SMT) solvers. However, as modern networks grow in size and complexity, the constraints derived from reachability requirements become increasingly complex, resulting in a considerable time cost to obtain a satisfying policy. The slow update of ACL policies can endanger network connectivity and security. This paper presents a new approach for fast and scalable ACL policy synthesis under complex constraints. We leverage Graph Neural Networks (GNNs) to learn the relations between nodes and reason the solution that satisfies the update requirements. In addition, an enhanced stochastic local search solver is introduced to deal with erroneous predictions of the GNN. Evaluations show that the proposed method guarantees 100% accuracy on real-world topologies. GNN outperforms modern SAT/SMT solvers in speed, saving up to 278x time costs on a 125-node topology. Furthermore, our method extrapolates well when faced with different requirements and topologies. Xingjian Liao, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
ICNP | 5 |
| 2023 | Not Only Pairwise Relationships: Fine-Grained Relational Modeling for Multivariate Time Series ForecastingabstractRecent graph-based methods achieve significant success in multivariate time series modeling and forecasting due to their ability to handle relationships among time series variables. However, only pairwise relationships are considered in most existing works. They ignore beyond-pairwise relationships and their potential categories in practical scenarios, which leads to incomprehensive relationship learning for multivariate time series forecasting. In this paper, we present ReMo, a Relational Modeling-based method, to promote fine-grained relational learning among multivariate time series data. Firstly, by treating time series variables and complex relationships as nodes and hyperedges, we extract multi-view hypergraphs from data to capture beyond-pairwise relationships. Secondly, a novel hypergraph message passing strategy is designed to characterize both nodes and hyperedges by inferring the potential categories of relationships and further distinguishing their impacts on time series variables. By integrating these two modules into the time series forecasting framework, ReMo effectively improves the performance of multivariate time series forecasting. The experimental results on seven commonly used datasets from different domains demonstrate the superiority of our model. Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Zhikang Wu, Zirui Zhuang, Jianxin Liao |
IJCAI | 6 |
| 2023 | Fine-Grained Flow Control Agent on Path MTU for IoT SoftwareabstractInternet of Things (IoT) software is used to control the distributed hardware of the underlying network and provide a reliable operating platform for various services. In production system, diversity IoT software provides multiple services, flows of different software run in parallel on the same IoT platform. Thus, system-level network parameter configurations may not be suitable for all service needs. In this paper, we focus on the challenge of the personally parameterizing transmission unit size and congestion windows (CWND) in flow control. We propose deeper flow control software model (DeepFC) for finer-grained flow control than traditional algorithms. DeepFC consist of two parts: (i) Since system-level transmission unit size may degrade network performance due to frequent fragmentation, we combine path MTU (PMTU) and deep reinforcement learning (DRL) to predict fine-grained flow-level transmission unit size. (ii) Transmission unit size is related to CWND in flow control. The fine-grained transmission unit size needs fine-grained congestion control solution. In DeepFC, we consider the mutual coupling between transmission unit size and CWND parameter configuration to further improve network performance. Experimental results show that DeepFC can reduce fragmentation by 67.8% compared to the protocols with system-level transmission unit size, flow completion time can be reduced by 20.93%, and throughput can be increased by 19.96% compared to the average of benchmark algorithms. Hongchuan He, Dezhi Chen, Zirui Zhuang, Qi Qi 0001, Lejian Zhang, Tong Xu 0002, Jingyu Wang 0001 |
Internetware | 3 |
| 2023 | Drift doesn't Matter: Dynamic Decomposition with Diffusion Reconstruction for Unstable Multivariate Time Series Anomaly DetectionabstractMany unsupervised methods have recently been proposed for multivariate time series anomaly detection. However, existing works mainly focus on stable data yet often omit the drift generated from non-stationary environments, which may lead to numerous false alarms. We propose **D**ynamic **D**ecomposition with **D**iffusion **R**econstruction (D$^3$R), a novel anomaly detection network for real-world unstable data to fill the gap. D$^3$R tackles the drift via decomposition and reconstruction. In the decomposition procedure, we utilize data-time mix-attention to dynamically decompose long-period multivariate time series, overcoming the limitation of the local sliding window. The information bottleneck is critical yet difficult to determine in the reconstruction procedure. To avoid retraining once the bottleneck changes, we control it externally by noise diffusion and directly reconstruct the polluted data. The whole model can be trained end-to-end. Extensive experiments on various real-world datasets demonstrate that D$^3$R significantly outperforms existing methods, with a 11% average relative improvement over the previous SOTA models. Chengsen Wang, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
NeurIPS | 2 |
| 2023 | Poster: PipeLLM: Pipeline LLM Inference on Heterogeneous Devices with Sequence SlicingabstractLarge Language Models (LLMs) has fostered the creation of innovative requirements. Locally deployed LLMs for micro-enterprise mitigates potential issues such as privacy infringements and sluggish response. However, they are hampered by the limitations in computing capability and memory space of possessed devices. We introduce PipeLLM, which allocates the model across devices commensurate with their computing capabilities. It enables the parallel execution of layers with slicing input sequence along the token dimension. PipeLLM demonstrates the potential to accelerate LLM inference with heterogeneity devices, offering a solution for LLM deployment in micro-enterprise hardware environment. Ruilong Ma, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
SIGCOMM | 6 |
| 2023 | Brief Announcement: Accelerate CNN Inference with Zoning Graph at Dynamic GranularityabstractPartitioning a CNN and parallel executing inference with multiple IoT devices have gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms have struggled to find the optimal model partitioning granularity for complex CNNs. Additionally, executing inference with heterogeneous IoT devices is NP-hard when the structure of the CNN is a directed acyclic graph (DAG) rather than a chain. In this paper, we introduce a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution. Ruilong Ma, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Jing Wang 0039 |
SPAA | 5 |
| 2023 | Unsupervised Portrait Drawing Generation for Free StylesabstractArtistic portrait drawing (APDrawing) generation has seen progress in recent years. However, due to the naturally high scarcity and artistry, it is difficult to collect large‐scale labeled and paired data and generally divide drawing styles into several specific recognized categories. Existing works suffer from the limited labeled data and naive manual division of drawing styles according to the corresponding artists. They cannot adapt to the actual situations, for example, a single artist might have multiple drawing styles and APDrawings from different artists might share similar styles. In this paper, we propose to use unlabeled and unpaired data and perform the task in an unsupervised manner. Without manual division of drawing styles, we take each portrait drawing as a unique style and introduce self‐supervised feature learning to learn free styles for unlabeled portrait drawings. Besides, we devise a style bank and a decoupled cycle structure to take over two main considerations in the task: generation quality and style control. Extensive experiments show that our model is more adaptable to different style inputs than state‐of‐the‐art methods. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Cong Liu 0046 |
Int. J. Intell. Syst. | 6 |
| 2023 | Standing on the Shoulders of Giants: Cross-Slice Federated Meta Learning for Resource Orchestration to Cold-Start SliceabstractNetwork slicing is a key technology in 6G communication systems to support numerous vertical applications for all scenes while providing resources on demand. Due to more time-varying and dynamic traffic flows, it is difficult for traditional methods to manage complex and highly dynamic 6G networks. Therefore, intelligent method such as Deep Reinforcement Learning (DRL) is employed into network management since DRL is a model-free and experience-driven approach. However, it is difficult to leverage one DRL model to provide customized intra-slice orchestration for various applications because of their diverse flow characteristics and service requirements. Moreover, training a DRL model is notoriously time-consuming so that it is not realistic for network operator to individually orchestrate customized network slice for each application with the DRL algorithm. Additionally, the data privacy of each application should be considered into the training process. In this paper, we propose a Federated Meta Reinforcement Learning (FedMRL) approach to tackle the cold-start problem in network slice orchestration, while reserving the data privacy. The Meta Reinforcement Learning (MRL) is leveraged to train a meta policy for rapidly learn a local policy for a specific slice orchestration task by finding a common initialization that allows for a quick adaptation towards each optimal solution. With the help of federated learning setting, the training process of meta policy is not required to collect raw data of applications to the centralized server. Experimental results show that FedMRL outperforms three baselines in terms of overall costs, end-to-end latency and convergence speed. Tianjian Dong, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao, Zhu Han 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | FedNKD: A Dependable Federated Learning Using Fine-tuned Random Noise and Knowledge DistillationabstractMultimedia retrieval models need the ability to extract useful information from large-scale data for clients. As an important part of multimedia retrieval, image classification model directly affects the efficiency and effect of multimedia retrieval. We need a lot of data to train a image classification model applied to multimedia retrieval task. However, with the protection of data privacy, the data used to train the model often needs to be kept on the client side. Federated learning is proposed to use data from all clients to train one model while protecting privacy. When federated learning is applied, the distribution of data across different clients varies greatly. Disregarding this problem yields a final model with unstable performance. To enable federated learning to work dependably in the real world with complex data environments, we propose FedNKD, which utilizes knowledge distillation and random noise. The superior knowledge of each client is distilled into a central server to mitigate the instablity caused by Non-IID data. Importantly, a synthetic dataset is created by some random noise through back propagation of neural networks. The synthetic dataset will contain the abstract features of the real data. Then we will use this synthetic dataset to realize the knowledge distillation while protecting users' privacy. In our experimental scenarios, FedNKD outperforms existing representative algorithms by about 1.5% in accuracy. Shaoxiong Zhu, Qi Qi 0001, Zirui Zhuang, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
ICMR | 3 |
| 2021 | Mean Field Deep Reinforcement Learning for Fair and Efficient UAV ControlabstractUnmanned aerial vehicles (UAVs) can provide flexible network coverage services. UAVs can be applied in a large number of scenarios, such as emergency communication and network access in areas without terrestrial network coverage. However, UAVs are limited to relatively short communication range and restricted energy resources. In extreme conditions such as disasters, there may also be a problem that the communication bandwidth is limited and the UAV cannot communicate with the server with a large amount of information, so a decentralized solution is expected. In addition, the interaction between multiple objectives and multiple UAVs leads to a huge state space, which makes large-scale practical applications difficult. To simplify complex interactions, we modeled the UAV control problem with mean-field game (MFG). We propose a new UAV control method, the mean-field trust region policy optimization (MFTRPO), which uses the MFG method to construct the Hamilton-Jacobi-Bellman/Fokker-Planck-Kolmogorov equation that obtains the optimal solution and solves the difficulties in the practical application through the trust region policy optimization and neural network feature embedding methods. The proposed method: 1) maximizes communication efficiency while ensuring fair communication range and network connectivity; 2) fuses the mean-field theory with deep reinforcement learning techniques; and 3) is scalable and adaptive. We conduct extensive simulations for performance evaluation. The simulation results have shown that MFTRPO significantly and consistently outperforms two commonly used baseline methods in terms of coverage, fairness, and energy consumption. Dezhi Chen, Qi Qi 0001, Zirui Zhuang, Jingyu Wang 0001, Jianxin Liao, Zhu Han 0001 |
IEEE Internet Things J. | 3 |
| 2021 | Adaptive and Robust Routing With Lyapunov-Based Deep RL in MEC Networks Enabled by BlockchainsabstractThe most recent development of the Internet of Things brings massive timely sensitive and bursty data flows. Also, joint optimization on storage, computation, and communication is in need for multiaccess edge computing frameworks. The adaptive network control has been explored using deep reinforcement learning (RL), but it is not sufficient for bursty network traffic flows, especially when the network traffic pattern may change over time. We formulate the routing control in an environment with time-variant link delays as a Lyapunov optimization problem. We identify that there is a tradeoff between optimization performance and modeling accuracy when the propagation delays are included. We propose a novel deep RL (DRL)-based adaptive network routing method to tackle the issues mentioned above. A Lyapunov optimization technique is used to reduce the upper bound of the Lyapunov drift, improving queuing stability in networked systems. By modeling the network traffic pattern using the Markovian arrival process, we show that network routing problems can be modeled as Markov decision processes and value-iteration-based RL methods can be used to solve them. We design a blockchain-based protocol using proof of elapsed time consensus mechanism to ensure a trustworthy network statistics information exchange for the routing framework. Experiment results show that the proposed method can learn a routing policy and adapt to the changing environment. The proposed method outperforms the baseline backpressure method in multiple settings and converges faster than existing methods. Moreover, the DRL module can effectively learn a better estimation of the long-term Lyapunov drift and penalty functions, providing superior results in terms of the backlog size, end-to-end latency, age of information, and throughput. Furthermore, the blockchain-based network statistics exchange can provide the routing framework against malicious nodes. In addition, the proposed model performs well under various topologies, and thus can be used in general cases. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Zhu Han 0001 |
IEEE Internet Things J. | 1 |
| 2021 | Generative Adversarial Network-Based Transfer Reinforcement Learning for Routing With Prior KnowledgeabstractWith the incremental deployment of software defined networking, the routing algorithms have gained more power on observability and controllability. Deep reinforcement learning, as an experience-driven approach, shows considerable potential in routing problem with the help of the centralized controller. It is an adaptive, lightweight, and model-free approach to coping with dynamic runtime status, large-scale traffic, and heterogeneous objective of SDN routing. However, it is still not suitable for the variable and complex emerging networks, because the huge training cost prevents fast convergence in a varying or discrepant environment. In this paper, we propose a transfer reinforcement learning algorithm to improve the training efficiency, and handle the variation in network status and topology. Specifically, we leverage the generative adversarial network to learn domain-invariant features that is suitable for deep reinforcement learning-based routing in different network environments. This mechanism utilizes the previous model and accelerates the training process. We implement our routing algorithm in the production level software switches and controller, while evaluating it comprehensively with many topologies and network status distributions. The experimental results show that our work not only outperforms the state-of-the-art deep reinforcement learning-based routing frameworks, but also has more training efficiency than the naive transfer learning algorithm both on different topologies and network status distributions. Tianjian Dong, Qi Qi 0001, Jingyu Wang 0001, Alex X. Liu, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2020 | DeepHop on Edge: Hop-by-hop Routing byDistributed Learning with Semantic AttentionabstractMulti-access Edge Computing (MEC) and ubiquitous smart devices help serve end-users efficiently and optimally through providing emerging edge-deployed services. Meanwhile, heavy and time-varying traffic loads are produced in the edge network, so that an efficient traffic forwarding mechanism is required. In this paper, we propose a parallel and distributed learning approach, DeepHop, to adapt to the volatile environments and realize hop-by-hop routing. The Multi-Agent Deep Reinforcement Learning (MADRL) is used to alleviate the edge network congestion and maximize the utilization of network resources. DeepHop determines the routing among edge network nodes for heterogeneous types of traffic according to the current workload and capability. By joining with an attention mechanism, DeepHop obtains the semantics from the elements of the network state to help the agents learn the importance of each element on routing. Experiment results show that DeepHop achieves the increase of successfully transmitted packets by 15% compared with the state-of-the-art algorithms. Besides, DeepHop with an attention mechanism reduces convergence time by nearly half compared with the common-used structures of neural networks. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Cong Liu 0046, Jianxin Liao |
ICPP | 5 |
| 2020 | Adaptive and Robust Network Routing Based on Deep Reinforcement Learning with Lyapunov OptimizationabstractThe most recent development of the Internet of Things brings massive timely-sensitive and yet bursty data flows. The adaptive network control has been explored using deep reinforcement learning, but it is not sufficient for extremely bursty network traffic flows, especially when the network traffic pattern may change over time. We model the routing control in an environment with time-variant link delays as a Lyapunov optimization problem. We identify that there is a tradeoff between optimization performance and modeling accuracy when the propagation delays are included. We propose a novel deep reinforcement learning-based adaptive network routing method to tackle the issues mentioned above. A Lyapunov optimization technique is used to reduce the upper bound of the Lyapunov drift, which leads to improved queuing stability in networked systems. Experiment results show that the proposed method can learn a routing control policy and adapt to the changing environment. The proposed method outperforms the baseline backpressure method in multiple settings, and converges faster than existing methods. Moreover, the deep reinforcement learning module can effectively learn a better estimation of the longterm Lyapunov drift and penalty functions, and thus it provides superior results in terms of the backlog size, end-to-end latency, age of information, and throughput. Extensive experiments also show that the proposed model performs well under various topologies, and thus the proposed model can be used in general cases. Also the user can adjust the preference parameter at ant time without the need to retrain the neural networks. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Zhu Han 0001 |
IWQoS | 1 |
| 2020 | Multi-Agent Deep Reinforcement Learning for Secure UAV CommunicationsabstractIn this paper, we investigate a multi-unmanned aerial vehicle (UAV) cooperation mechanism for secure communications, where the UAV transmitter moves around to serve the multiple ground users (GUs) while the UAV jammers send the 3D jamming signals to the ground eavesdroppers (GEs) to protect the UAV transmitter from being wiretapped. The 3D jamming guarantees the GEs not being interfered by the jamming signals. It is challenging to make a joint trajectory design and power control for a UAV team without central control. To this end, we propose a multi-agent deep reinforcement learning approach to achieve the maximum sum secure rate by designing the dynamic trajectory of each UAV. The proposed multi-agent deep deterministic policy gradient (MADDPG) technique is centralized training at high altitude platforms (HAPs) and distributed execution at each UAV, which enables the fully distributed cooperation among UAVs. Finally, the simulation results show the proposed method can efficiently solve the multi-UAV cooperation trajectory design problem in secure communication scenarios. Yu Zhang 0047, Zirui Zhuang, Feifei Gao 0001, Jingyu Wang 0001, Zhu Han 0001 |
WCNC | 2 |
| 2018 | A Case-Based Decision System for Routing in Packet-Switched NetworksabstractRoute planning with global optimization objectives in graphs is a challenging task with enormous computational complexity and finding the best solution is NP-complete. In addition, the network's operational performance varies whenever the environment changes. Traditional routing schemes fail to deal with these situations. We propose a case-based decision system for routing in packet switched networks to track the networking status. We also design a graph-aware neural network to suggest and revise the solutions from the past cases. The low-level structure of the neural network is learned by fitting with the features not only from each standalone vertex but also from the neighbors of each vertex. Experiments show that the proposed system outperforms state-of-art traffic-split and traffic-engineered routing schemes. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
IPCCC | 1 |
| 2018 | Common Knowledge Based Transfer Learning for Traffic ClassificationabstractDeep neural networks have been used for traffic classification, and promising results are obtained. However, most previous work confined to one specific classification task and restricted the classifiers potential performance and applications. As the traffic flows can be labeled from different perspectives, the performance of the classifier might be improved by exploring more meaningful latent features. For this purpose, we adopted a multi-output DNN model that simultaneously learns different traffic classification tasks. The common knowledge of traffic is exploited by the synergy among the tasks and boosts the individual performances of the tasks. Experiments show that this structure has the potential to meet new future demands and achieve the classification with advanced speeds and fair accuracies. Yet, due to the heavy training cost, the neural networks, though achieving good performance, are hard to implement in the real environment. We further show that few-shot learning could be a viable approach. Yunming Xiao, Haifeng Sun 0001, Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001 |
LCN | 3 |
| 2018 | Graph-Aware Deep Learning Based Intelligent Routing StrategyabstractSoftware defined networking decouples the control plane and data plane, which grants more computing power for routing computations. Traditional routing methods suffer from the complex dynamics in networking, and they are facing issues such as slow convergence and performance decline. Deep learning techniques have shown preliminary results on solving the routing problem, bring more accuracy and precision compared with traditional modeling techniques. However, the deep learning architecture needs to be specially customized to learn the topological relations between switches in an efficient way. Thus, we propose a deep learning based intelligent routing strategy with revised graph-aware neural networks and we design a set of features suitable for network routing. Then we demonstrate the performance of our works by using a real-world topology and the production level software switch. The simulation result shows our work is more accurate and efficient compared to state-of-art routing strategy. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
LCN | 1 |