Shiming He

dblp:89/10696 · DBLP profile ↗
← Back
44ranked-venue papers
20as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 4 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TD-RCA: Topology-Aware and Dual-Perspective Decoupling for Root Cause Analysis in Microservices
Shiming He, Mengyao Wei, Lingyun Xiang, Kun Xie 0001
IWQoS1
2026 Multivariate Time Series Anomaly Detection Using Learnable Spatial-Temporal Graph Ordinary Differential Equations Network
abstract
Multivariate time series anomaly detection (MTSAD) plays a critical role in the Internet of Things (IoT) by identifying malfunctions and attacks. Graph Neural Networks (GNNs) have been widely employed in MTSAD to capture spatial features but require predefined and explicit graph structures. Graph Structure Learning (GSL) addresses this limitation by jointly learning the graph structure and downstream tasks. However, existing GSL-based MTSAD methods fail to effectively leverage prior knowledge and struggle with insufficient GNN depth, limiting their ability to capture long-range dependencies. To address these challenges, we propose a multivariate time series anomaly detection method based on a learnable spatio-temporal graph ordinary differential equation network (STGODE), named MAD-ODE. Our approach leverages hybird graph learning, which includes two types of graph structures: a static similarity graph and a learnable graph. The static similarity graph is constructed using prior knowledge and provides a stable, interpretable representation of sensor dependencies. In contrast, the learnable graph captures complex relationships between sensors by optimizing its structure through backpropagation. This hybird graph learning effectively incorporates both prior knowledge and learned dependencies, ensuring robust and flexible modeling of sensor relationships. Furthermore, we design a STGODE predictor, which operates on both graph structures and employs continuous graph convolutional networks, enabling it to capturing long-range spatio-temporal dependencies for forecasting the next timestamp. Extensive experiments conducted on five datasets demonstrate that MAD-ODE achieves the best average performance and maintains stable results compared to existing methods.
Shiming He, Keyao Feng, Diqing Liang, Kun Xie 0001, Pradip Kumar Sharma
IEEE Trans. Dependable Secur. Comput.1
2026 Multimodal Anomaly Detection for Microservice Systems via Grassmann Manifolds-Based Graph Fusion
abstract
Microservice architecture decomposes complex applications into multiple small, independent services, enabling independent deployment and scaling. However, microservice systems introduce complex and dynamic interactions between service instances. When a service instance fails, it can significantly degrade the performance of the entire system, harm the user experience, and cause substantial economic losses. Therefore, effective anomaly detection for service instances is crucial to ensure system reliability. Metrics, logs, and traces provide complementary insights into microservice operations, and recent approaches have attempted to jointly exploit these modalities for anomaly detection. However, the varying data volumes per interval and heterogeneous data types across modalities, combined with the complex and elusive relationships between them, make effective representation and learning challenging. To address these limitations, we propose MGFusion, an unsupervised, end-to-end multimodal anomaly detection method for microservice systems using Grassmann manifolds-based graph fusion. MGFusion employs a robust method to unify metrics, logs, and traces into time series, enabling consistent processing across modalities. It leverages multiple graph structure learning (GSL) techniques to explore multi-view relationships between the modalities. To mitigate the impact of data noise, we further introduce prior knowledge based on the Adamic-Adar Index (AAI) and employ a Grassmann manifolds-based graph fusion method to combine multiple basic graph structures. Finally, anomaly detection is achieved using Diffusion Convolutional Recurrent Neural Network (DCRNN) predictors and anomaly score calculation. Extensive experiments on a public dataset demonstrate that MGFusion effectively fuses multimodal data and captures complex relationships, significantly improving the accuracy and efficiency of anomaly detection. Compared to the state-of-the-art unsupervised multimodal anomaly detection method, AnoFusion, MGFusion achieves an improvement in the F1 score ranging from 11.1% to 15.3%.
Shiming He, Keyao Feng, Kaixuan Meng, Kun Xie 0001, Xibin Zhao
IEEE Trans. Serv. Comput.1
2025 To Split or to Merge? Exploring Multi-modal Data Flexibly for Failure Classification in Microservices
abstract
Classifying failures for automated diagnosis is critical for maintaining the reliability of ever-increasing microservice systems.Existing failure classification methods focus solely on analyzing one specific data modality (e.g., metrics) or feeding multi-modal data (e.g., logs and traces) in a one-fit-all model.This would lead to biased inference as different modalities have differentiated proportions and manifest distinct sensitivities and patterns in failures of varied runtime contexts.This work proposes a novel failure classification framework for microservices, named AgileMFC, with the basic idea of flexibly learning and assembling the expertise of multi-modal monitoring data.The flexibility is attained with the split-and-merge design of data representation, where it assigns each modality a different gate that routes it to dedicated expert networks and uses a sharing gate to facilitate controllable knowledge interaction among modalities.Meanwhile, AgileMFC trains the representation networks and a classifier in consecutive stages with concatenated and weighted feature fusion, thus decoupling task-agnostic feature extraction with task-specific classification.Experiments on three datasets show that AgileMFC outperforms four baselines w.r.t., precision, recall, and F1-score on failure classification.
Xiuhong Tan, Yuan Yuan 0034, Tongqing Zhou, Shiming He
Internetware4
2025 TCMS: A Multi-Sequence Log Parsing Method Based on Token Conversion
abstract
Detailed system operations are recorded in logs. To ensure system reliability, developers can detect system anomalies through log anomaly detection. Log parsing, which converts semi-structured log messages into structured data, is a crucial step in log anomaly detection and advanced program analysis and verification. Despite the availability of various log parsing tools, they generally suffer from low parsing accuracy and slow efficiency due to the ignorance of variable characteristics and the use of costly pairwise comparison methods. In this paper, we propose a TCMS framework to parse logs, consisting of two main technologies. First, by studying 16 public log datasets, we find that most log variable tokens are structured variable tokens. Based on this discovery, we propose a token conversion algorithm to improve parsing accuracy. This algorithm converts the changed parts in structured variable tokens into wildcards (‘$< *> $’), preventing these tokens from being directly identified as constant tokens. Second, to improve efficiency, we propose the LogMLCS algorithm, which intelligently constructs a graph to facilitate the extraction of common parts from multiple log messages at once, instead of using pairwise comparisons. Comprehensive experiments conducted on 16 log datasets reveal that our TCMS outperforms seven other parsing methods, achieving the highest parsing accuracy at the fastest speed. Furthermore, experimental results from running a log anomaly detection algorithm in conjunction with different log parsing methods demonstrate that TCMS significantly boosts detection accuracy. For instance, on the OpenStack dataset, our TCMS-facilitated log anomaly detection algorithm achieves a perfect F1-score, precision, and recall of 100% each, surpassing the best peer method by 32.2, 0.8, and 19.5 percentage points, respectively.
Mingkuan Wei, Jigang Wen, Shiming He, Kun Xie 0001, Wei Liang 0005, Gaogang Xie, Kenli Li 0001
IEEE Trans. Dependable Secur. Comput.3
2025 An Unsupervised Malicious Web Request Detection Based on Transformer and Contrastive Learning
abstract
The World Wide Web (Web) is a crucial part of the Internet. Web attacks are becoming more and more serious and complex. Malicious Web request detection aims to rapidly and accurately identify abnormal attacks on the network. Deep learning is being applied to malicious Web request detection, resulting in high detection performance. However, most deep learning-based methods are supervised and ignore special characters, which are hard to detect unknown malicious Web requests. The labels of Web request are fewer and Web request data is insufficient. Therefore, we propose an unsupervised malicious Web request detection based on transformer and contrastive learning (UTCDetector). UTCDetector exploits preprocessing and 2-gram word segmentationto preserve special characters, extracts semantic feature by Transformer, and leverages hypersphere loss function and contrastive learning to handle insufficient Web data without abnormal label. Since the public Web request datasets (CSIC 2010, CSIC TORPEDA 2012, and ECML/PKDD 2007) were created before 2012, we collected Web requests from a university Web application server in 2023 to build a private dataset named School 2023. This dataset contains more modern and complex attacks. The experimental results on the four datasets demonstrate that our method achieves a higher F1-score than other existing methods and ablation variants.
Shiming He, Diqing Liang, Pradip Kumar Sharma
IEEE Trans. Netw. Serv. Manag.1
2025 Delay and Load Fairness Optimization With Queuing Model in Multi-AAV Assisted MEC: A Deep Reinforcement Learning Approach
abstract
Autonomous aerial vehicles (AAV) can alleviate the computational burden on edge devices through assisted computing. However, with the increase in the number of Internet of Things Devices (IoTDs), it is essential to establish a task queue on the AAV to schedule computing tasks from IoTDs. In addition, the load fairness of AAVs should be optimized to fully utilize the computing resources. Therefore, a multi-AAV-assisted mobile edge computing (MEC) network framework based on the queuing model is proposed, which aims at optimizing the average delay of all user devices and the load fairness of AAVs. Firstly, we prove that the arrangement of tasks with different computing delays on the AAV queue can affect the user’s average delay, so a short-job-first (SJF) queuing model is proposed to minimize the average delay of users. On this basis, a joint optimization problem related to the AAV’s three-dimensional trajectory and user connection scheduling is formulated. A SJF based low-complexity connection scheduling algorithm is proposed and combined in a deep reinforcement learning (DRL) to solve this NP-hard problem. To evaluate the performance of the proposed algorithm, we compare it with deep deterministic policy gradient (DDPG), particle swarm optimization (PSO), random moving (RM), and local computing (LC). Simulation results show that our algorithm effectively reduces user average delay and enhances AAV load fairness. Finally, SJF is compared with the traditional first-come-first-served (FCFS) queuing model on different algorithms. The results indicate that the average delay of SJF is significantly lower than that of FCFS.
Qiang Tang 0006, Bao Li 0008, Halvin Yang, Shiming He, Kun Yang 0001
IEEE Trans. Netw. Serv. Manag.5
2025 Simulation-Aided Policy Tuning for Black-Box Robot Learning
abstract
How can robots learn and adapt to new tasks and situations with little data? Systematic exploration and simulation are crucial tools for efficient robot learning. We present a novel black-box policy search algorithm focused on data-efficient policy improvements. The algorithm learns directly on the robot and treats simulation as an additional information source to speed up the learning process. At the core of the algorithm, a probabilistic model learns the dependence between the policy parameters and the robot learning objective not only by performing experiments on the robot, but also by leveraging data from a simulator. This substantially reduces interaction time with the robot. Using the model, we can guarantee improvements with high probability for each policy update, thereby facilitating fast, goal-oriented learning. We evaluate our algorithm on simulated fine-tuning tasks and demonstrate the data-efficiency of the proposed dual-information source optimization algorithm. In a real robot learning experiment, we show fast and successful task learning on a robot manipulator with the aid of an imperfect simulator.
Shiming He, Alexander von Rohr, Dominik Baumann, Ji Xiang, Sebastian Trimpe
IEEE Trans. Robotics1
2025 Unsupervised Multi-Target Cross-Service Log Anomaly Detection
abstract
Log analysis, especially log anomaly detection, can help debug systems and analyze root causes to provide reliable services. Deep learning is a promising technology for log anomaly detection. However, deep learning methods need a large amount of training data, which is hard for a newly deployed system to collect sufficient logs. Transfer learning becomes a possible method to solve the problem that can apply the knowledge from a long-term deployed system (source) to a newly deployed system (target). Existing transfer learning methods focus on transferring the knowledge from a source system to a single target system within the same service, in which the source and the target belong to the same service (e.g. operating system, supercomputer, or distributed system). They achieve low performance when applied to multiple target and different services systems because of the obvious differences in log format, syntax, semantics, and component call between different services and the individual training of multiple models for each target system. To tackle the problems, we propose an unsupervised multi-target cross-service log anomaly detection method based on transfer learning and contrastive learning (LogMTC). LogMTC exploits contrastive learning to learn a single model on combined data from the source and multiple target systems, which can fit multiple target systems simultaneously and improve efficiency. LogMTC exploits a hypersphere loss and two contrastive losses to minimize the feature differences crossing different services. Our experiments on two services (supercomputer and distributed system) and three log datasets show that our method is superior to the existing transfer learning methods in the same service, cross-service, and multi-target log anomaly detection. Compared with the best peer accurate transfer learning algorithm LogTAD, LogMTC improves 1.14%-8.28$\%$F1 score in multi-target transfer and is 1.12-1.22 times faster.
Shiming He, Kun Xie 0001, Jigang Wen
IEEE Trans. Sustain. Comput.1
2024 Demonstration Retrieval-Augmented Generative Event Argument Extraction
abstract
We tackle Event Argument Extraction (EAE) in the manner of template-based generation. Based on our exploration of generative EAE, it suffers from several issues, such as multiple arguments of one role, generating words out of context and inconsistency with prescribed format. We attribute it to the weakness of following complex input prompts. To address these problems, we propose the demonstration retrieval-augmented generative EAE (DRAGEAE), containing two components: event knowledge-injected generator (EKG) and demonstration retriever (DR). EKG employs event knowledge prompts to capture role dependencies and semantics. DR aims to search informative demonstrations from training data, facilitating the conditional generation of EKG. To train DR, we use the probability-based rankings from large language models (LLMs) as supervised signals. Experimental results on ACE-2005, RAMS and WIKIEVENTS demonstrate that our method outperforms all strong baselines and it can be generalized to various datasets. Further analysis is conducted to discuss the impact of diverse LLMs and prove that our model alleviates the above issues.
Shiming He, Yu Hong 0001, Jianmin Yao 0001, Guodong Zhou 0001
LREC/COLING1
2024 Word-level Commonsense Knowledge Selection for Event Detection
abstract
Event Detection (ED) is a task of automatically extracting multi-class trigger words. The understanding of word sense is crucial for ED. In this paper, we utilize context-specific commonsense knowledge to strengthen word sense modeling. Specifically, we leverage a Context-specific Knowledge Selector (CKS) to select the exact commonsense knowledge of words from a large knowledge base, i.e., ConceptNet. Context-specific selection is made in terms of the relevance of knowledge to the living contexts. On this basis, we incorporate the commonsense knowledge into the word-level representations before decoding. ChatGPT is an ideal generative CKS when the prompts are deliberately designed, though it is cost-prohibitive. To avoid the heavy reliance on ChatGPT, we train an offline CKS using the predictions of ChatGPT over a small number of examples (about 9% of all). We experiment on the benchmark ACE-2005 dataset. The test results show that our approach yields substantial improvements compared to the BERT baseline, achieving the F1-score of about 78.3%. All models, source codes and data will be made publicly available.
Yu Hong 0001, Shiming He, Qingting Xu
LREC/COLING3
2024 Memory-efficient anomaly detection for online data streams
abstract
Network Intrusion Detection refers to the use of anomaly detection to protect network security, that is, detecting and identifying anomalies and malicious behavior in the network by monitoring network traffic and other metrics. Identifying anomalies and intrusions plays a crucial role in network security, as it helps organizations to timely detect and respond to various network attacks and threats. Unlike static data, intrusion detection data is usually in the form of data streams. However, existing methods have not fully considered the inherent characteristics of data streams, such as infinity, real-time nature, and concept drift, which leads to lower detection accuracy and significant memory waste. Considering these issues, we propose an online anomaly detection method called MEO-AD, based on Locality Sensitive Hashing (LSH), Isolation Forest, and an adaptive updating. It can handle the concept drift problem of data streams. We evaluate MEO-AD on two public network intrusion detection datasets. Experimental results demonstrate the accuracy of the proposed method and its lower memory consumption. We consider three types of LSH. The LSH with Euclidean distance presents the best performance.
Shiming He
CSCWD1
2024 Multi-Graph Structure Learning-based Multivariate Time Series Anomaly Detection with Extended Prior Knowledge
abstract
In the Internet of Things (IoT), substantial time series data is recorded by sensors and other devices. Multivariate time series anomaly detection (MTSAD) identifies anomalies derived from device malfunctioning or system attacks to reduce economic losses. Graph structure learning (GSL)-based anomaly detection method learns an optimal graph structure joint with the downstream anomaly detection task, which achieves superior performance. However, the existing GSL-based methods only learn a single graph structure and can not represent multiple and complex relationships. Therefore, we propose a multi-graph structure learning-based multivariate time series anomaly detection with extended prior knowledge (MEGLAD). MEGLAD selects three kinds of typical graph structure learners to learn as many relationship types among sensors as possible. Extensive experiments show that our approach has better detection performance than state-of-the-art single graph structure learning techniques on four public and real-world datasets.
Shiming He, GenXin Li, Qinqing Guo, Kun Xie 0001
CSCWD1
2024 Coverage Probability of Distributed CoMP UAV-Assisted Cellular Networks
abstract
It is well-established that terrestrial communication systems may fail during emergencies such as earthquakes, tsunamis, and floods. Fortunately, with the rapid advancement of unmanned aerial vehicle (UAV) network technology, deploying UAV nodes as aerial base stations (BSs) is assuming an increasingly crucial role in facilitating downlink transmissions and restoring ground communication capabilities. However, a single UAV node is not sufficient to meet the requirements. Inspired by distributed communication, we introduce a performance analysis framework based on stochastic geometry to analyze the distributed coordinated multi-point (CoMP) UAV-assisted communication network. Specifically, we assume that all UAV nodes follow a homogeneous Poisson point process (PPP) and maintain a constant altitude. The entire space is tessellated by multiple hexagons, with multiple UAV nodes within each hexagonal region working together to serve terrestrial user equipments (UEs). For this region-centric cooperative model, we derive an exact expression for the coverage probability to quantify the performance improvement enabled by UAVs, analyze the upper bound of the coverage probability, and provide a simplified approximation. We then compare this model to a user-centric model. Our numerical findings demonstrate that the cooperation of UAV nodes can significantly enhance the coverage probability and save spectrum resources.
Qingmin Long, Qiang Tang 0006, Shiming He, Bing Xiong 0001
ISPA4
2024 Collaborative Filtering-based Fast Delay-aware algorithm for joint VNF deployment and migration in edge networks
Zhuofan Liao, Wenqiang Deng, Shiming He, Qiang Tang 0006
Comput. Networks3
2024 ActiveGuardian: An accurate and efficient algorithm for identifying active elephant flows in network traffic
Bing Xiong 0001, Jinyuan Zhao, Shiming He, Baokang Zhao, Kun Yang 0001, Keqin Li 0001
J. Netw. Comput. Appl.5
2024 A multi-UAV assisted non-orthogonal multiple access based relay system for minimal average receiving rate maximization
Qiang Tang 0006, Xinyu Qu, Jin Wang 0001, Shiming He
Soft Comput.4
2024 Fusion Graph Structure Learning-Based Multivariate Time Series Anomaly Detection With Structured Prior Knowledge
abstract
Multivariate time series anomaly detection (MTSAD) plays a crucial role in the Internet of Things (IoT) to identify device malfunction or system attacks. Graph neural networks (GNN) are widely applied in MTSAD to capture the spatial features among sensors. However, GNNs depend on a graph structure and explicit graph structures are not always available. To solve the problem of missing explicit graph structure, graph structure learning is introduced to learn an accurate graph structure joint with a GNNs-based anomaly detection task. However, the existing GSL-based methods provide only a partial view of the graph structure and cannot represent multiple and complex relationships. The noise of data also brings noisy edges. Therefore, we propose a fusion graph structure learning-based multivariate time-series anomaly detection with structured prior knowledge (FuGLAD). To the best of our knowledge, it appears to be the first application of fusion graphs in time series anomaly detection. FuGLAD selects three kinds of typical graph structure learners to learn as many relationship types among sensors as possible and exploits the prior similarity to evaluate the importance of all learned graphs and adaptively learn the fusion weight instead of the direct average weight. To handle noise in raw data, FuGLAD compares the neighbors of nodes by Jaccard similarity to identify and remove the noisy edges in the prior graph. Extensive experiments demonstrate that our approach outperforms state-of-the-art single-graph structure learning techniques in detection performance across four public and real-world datasets.
Shiming He, GenXin Li, Kun Xie 0001, Pradip Kumar Sharma
IEEE Trans. Inf. Forensics Secur.1
2024 Trajectory Tracking Control for Differential-Driven Unmanned Surface Vessels Considering Propeller Servo Loop
abstract
This article addresses the robust trajectory tracking control strategy of differential-driven unmanned surface vessels, with the propeller servo loop taken into consideration. The proposed strategy takes duty cycles of propeller motors as control input and thereby can be directly applied in practice without any modification. Compared to the existing methods, the proposed method provides the framework of position control considering propeller servo loop and improves robustness to external disturbances. The controller design procedure is divided into three stages through backstepping technique. Disturbance observers are constructed to provide estimations of composite disturbances, which guarantee robustness in the presence of external disturbances and model uncertainties. An auxiliary system is introduced to handle the input saturation of duty cycles. With a linear growth condition of the output of propeller motors, a rigorous proof is presented to show tracking errors are uniformly ultimately bounded. Simulation and experiments illustrate the effectiveness of the proposed control strategy.
Zishi Xu, Shiming He, Ji Xiang
IEEE Trans. Ind. Informatics4
2024 GraphIoT: Lightweight IoT Device Detection Based on Graph Classifiers and Incremental Learning
abstract
The rapid expansion of the Internet of Things (IoT) has led to growing concerns about the security of IoT devices. A crucial aspect of ensuring their security is IoT device identification, which involves pinpointing the specific type of device. Existing solutions, however, either necessitate complex feature engineering or struggle to handle the ever-increasing number of new devices in open IoT environments. To tackle these challenges, this paper introduces GraphIoT, a lightweight IoT device detection method based on graph classifiers. GraphIoT leverages lightweight flow information, such as packet length, direction, and timestamp, to create an IoT Device Traffic Graph Representation (IoT-DTGR). This representation offers a comprehensive view of IoT device flows while preserving features in bidirectional IoT Device-Gateway interactions. By transforming the IoT device detection problem into a graph classification problem, GraphIoT employs a powerful Graph Neural Network that takes into account both node and edge features, as well as subgraph structures in IoT-DTGRs, to classify graphs and consequently identify device types. Additionally, the paper proposes an incremental learning framework called CL-GraphIoT that continuously learns features of new IoT device flows without forgetting previously learned device features. This is achieved through two strategies: parameter sharing and sample replaying. The paper gathers a real-world dataset from 18 IoT devices and conducts experiments on two datasets: the gathered real-world dataset and an open-source dataset covering 21 IoT device types. The experimental results demonstrate that both GraphIoT and CL-GraphIoT outperform state-of-the-art methods, achieving high accuracy in device detection with fast processing speed.
Yansong Yin, Kun Xie 0001, Shiming He, Yanbiao Li 0001, Jigang Wen, Zulong Diao, Da-Fang Zhang 0001, Gaogang Xie
IEEE Trans. Serv. Comput.3
2023 Joint Optimization of Multi-Type Caching Placement and Multi-User Computation Offloading for Vehicular Edge Computing
abstract
With the rapid development of Artificial Intelligence (AI) and Internet of Vehicles (IoV), the types of vehicular applications are becoming more diverse. And Vehicular Edge Computing (VEC) can provide the computing resource and caching resource for the diverse applications with the lower latency compared with the cloud. However, due to the limited resource of VEC and the long haul transmission from the cloud, the multi-type caching of the diverse applications from multi-users bring the huge challenges. In this paper, we propose a joint optimization problem of multi-type caching placement and multi-user computation offloading in the three-layer end-edge-cloud architecture to minimize the overall system latency. As the resolution of the NP-Hard problem, a Caching and Offloading Framework for Multi-user Multi-type Requests (COF-MMR) based on Deep Deterministic Policy Gradient (DDPG) algorithm is explored. Simulation results show that our proposed COFMMR framework has achieved an up to 20% improvement in reducing the overall system latency compared to the baseline scheme.
Dun Cao, Shiming He
GLOBECOM4
2023 Adaptive Routing for Datacenter Networks Using Ant Colony Optimization
Jinbin Hu 0001, Man He, Shuying Rao, Jing Wang 0209, Shiming He
ICA3PP (3)6
2023 HAECN: Hierarchical Automatic ECN Tuning with Ultra-Low Overhead in Datacenter Networks
Jinbin Hu 0001, Youyang Wang, Zikai Zhou, Shuying Rao, Rundong Xin, Jing Wang 0209, Shiming He
ICA3PP (3)7
2023 A Mixture of Experts with Adaptive Semantic Encoding for Event Detection
abstract
Event Detection (ED) is a challenging but valuable task. It aims to identify the words that trigger the events in text and classify them into pre-defined types. The previous works utilize entity information as supplementary clues for ED. However, word semantics in different contexts always have subtle variations. Capturing the connection between entity information and changing semantics is challenging. To tackle the issue, we propose the Mixture of Experts (MoE) technique with context clues to adaptively model semantics. We term the framework as CMoE. Concretely, the CMoE simultaneously performs entity and event detection with the MoE technique which are flexible components embedded transformer block. Furthermore, we apply multi-task learning to mine the shared knowledge between entity and event. We conduct experiments on the public ACE 2005 and KBP 2017 datasets. The results show that our model achieves competitive performance on ACE 2005 without using external knowledge, yielding an improvement of about 2.3% F1-score for ED. More importantly, our model outperforms all State-of-The-Art models on KBP 2017.
Zhongqiu Li, Yu Hong 0001, Shiming He, Guodong Zhou 0001
IJCNN3
2023 Parameter-Efficient Log Anomaly Detection based on Pre-training model and LORA
abstract
Logs record both the normal and abnormal system operating status at any time, which are crucial data during system operation. Log anomaly detection can help with system debugging and analyzing root causes, such as system fault, shutdown fault, null-pointer exception, illegal-argument exception, and class cast exception. Deep learning is widely applied to log anomaly detection to enhance detection accuracy. However, the deep learning model requires a lot of label logs, which consume large amounts of labor and time. To tackle this label requirement problem, the pre-training model is introduced, for instance, the Bidirectional Encoder Representations from Transformers (BERT). However, the pre-training model brings new problems. The parameters of BERT needed to be fine-tuned are huge, resulting in a high training overhead. Besides, the direct word sequence input representation of BERT ignores the semantic information among logs. Therefore, we propose a parameter-efficient log anomaly detection scheme (LogBP-LORA) based on BERT and Low-Rank Adaptation (LORA). LORA is an effective parameter-tuning strategy. LogBP-LORA increases bypass weight matrices and only updates the bypass parameters instead of all the original parameters to reduce the training overhead. Additionally, LogBP-LORA exploits log event sequence representation to obtain more semantic information with a shorter sequence length. Extensive experiments carry on three public log datasets, BGL, Thunderbird, and HDFS, demonstrate LogBP-LORA can obtain favorable performance with lower resource consumption. When fewer label data is available, LogBP-LORA achieves about 10%-99% higher F1-score compared with Neurallog, Deeplog, MADDC, and Loganomaly. The training parameters of LogBP-LoRA are only 0.06% of the original parameters of BERT.
Shiming He, Kun Xie 0001, Pradip Kumar Sharma
ISSRE1
2023 A joint matrix factorization and clustering scheme for irregular time series data
Shiming He, Zhuozhou Li, Kun Xie 0001, Naixue Xiong
Inf. Sci.1
2023 PMP: A partition-match parallel mechanism for DNN inference acceleration in cloud-edge collaborative environments
Zhuofan Liao, Shiming He, Qiang Tang 0006
J. Netw. Comput. Appl.3
2023 A cooperative MEC framework based on multi-UAV and AP to minimize weighted energy consumption
Qiang Tang 0006, Linjiang Li, Shiming He, Jin Wang 0001
Pervasive Mob. Comput.4
2022 Unregulated Chinese-to-English Data Expansion Does NOT Work for Neural Event Detection
abstract
We leverage cross-language data expansion and retraining to enhance neural Event Detection (abbr., ED) on English ACE corpus. Machine translation is utilized for expanding English training set of ED from that of Chinese. However, experimental results illustrate that such strategy actually results in performance degradation. The survey of translations suggests that the mistakenly-aligned triggers in the expanded data negatively influences the retraining process. We refer this phenomenon to “trigger falsification”. To overcome the issue, we apply heuristic rules for regulating the expanded data, fixing the distracting samples that contain the falsified triggers. The supplementary experiments show that the rule-based regulation is beneficial, yielding the improvement of about 1.6% F1-score for ED. We additionally prove that, instead of transfer learning from the translated ED data, the straight data combination by random pouring surprisingly performs better.
Zhongqiu Li, Yu Hong 0001, Shiming He, Jianmin Yao 0001, Guodong Zhou 0001
COLING4
2022 Event Detection with Cross-Sentence Graph Convolutional Networks
abstract
The goal of Event Detection (ED) task is to identify the words that mark the occurrence of events in text, and classify them into a set of event types. To model informative word semantics, some researchers apply Graph Convolutional Network (GCN) to exploit the syntactic graph transformed from the dependency tree, within one sentence. We are motivated to simultaneously leverage syntactic clues and context information across sentences. To this end, we propose a novel ED model with Cross-Sentence Graph Convolutional Networks (CSGCN). The CSGCN contains two main components, including a tree extension module and the syntax-aware graph convolution. Each sentence is parsed to a dependency tree by an automatic toolkit. The first module merges entity-specific subtrees from neighbor sentences into the dependency tree of current sentence, which constructs a cross-sentence dependency tree. On this basis, we transform the tree into an undirected graph. After that, a syntax-aware attention mechanism is employed in the computation of graph convolution. This mechanism dynamically captures syntax-relevant information from neighbor nodes via the graph structure. Finally, we devise an entity aggregation module to aggregate key entity information for trigger candidates. We conduct experiments on the ACE 2005 and KBP 2017 datasets. The results show that our model achieves satisfactory and competitive performance on ACE 2005, and outperforms all State-of-The-Art models on KBP 2017.
Shiming He, Yu Hong 0001, Zhongqiu Li, Jianmin Yao 0001, Guodong Zhou 0001
ICTAI1
2022 Multiple Strategies Differential Privacy on Sparse Tensor Factorization for Network Traffic Analysis in 5G
abstract
Due to high capacity and fast transmission speed, 5G plays a key role in modern electronic infrastructure. Meanwhile, sparse tensor factorization (STF) is a useful tool for dimension reduction to analyze high-order, high-dimension, and sparse tensor (HOHDST) data, which is transmitted on 5G Internet-of-things (IoT). Hence, HOHDST data relies on STF to obtain complete data and discover rules for real time and accurate analysis. From another view of computation and data security, the current STF solution seeks to improve the computational efficiency but neglects privacy security of the IoT data, e.g., data analysis for network traffic monitor system. To overcome these problems, this article proposes a multiple-strategies differential privacy framework on STF (MDPSTF) for HOHDST network traffic data analysis.MDPSTFcomprises three differential privacy (DP) mechanisms, i.e.,$\varepsilon -$DP, concentrated DP, and local DP. Furthermore, the theoretical proof of privacy bound is presented. Hence,MDPSTFcan provide general data protection for HOHDST network traffic data with high-security promise. We conduct experiments on two real network traffic datasets ($Abilene$and$G\grave{E}ANT$). The experimental results show thatMDPSTFhas high universality on the various degrees of privacy protection demands and high recovery accuracy for the HOHDST network traffic data.
Jin Wang 0001, Hao Li 0025, Shiming He, Pradip Kumar Sharma, Lydia Y. Chen
IEEE Trans. Ind. Informatics4
2021 Intelligent Detection for Key Performance Indicators in Industrial-Based Cyber-Physical Systems
abstract
Intelligent anomaly detection for key performance indicators (KPIs) is important for keeping services reliable in industrial-based cyber-physical systems (CPS). However, it is common in practice for various KPI sampling strategies to be utilized. We experimentally verify that anomaly detection is highly sensitive to irregular sampling, and accordingly go on to investigate low-cost anomaly detection for large-scale irregular KPIs. Irregular KPIs can be classified into four types: equal interval and unequal quantity (EIUQ) KPIs, unequal interval (UI) KPIs, unequal interval with equal duration (UIED) KPIs, and segmented irregular KPIs. In this article, we propose an anomaly detection framework based on these irregular types. Moreover, to handle the various lengths and phase shifts among EIUQ KPIs, we propose a normalized version of unequal cross-correlation, which slides the KPIs to enable finding the most similar position. To avoid high computational costs, we analyze the low-rank feature of KPIs data and propose a matrix factorization-based alignment algorithm for UIED KPIs; this algorithm treats UIED KPIs as an incomplete matrix and recovers the KPIs to align them before performing anomaly detection. Extensive simulations using three public datasets and two real-world datasets demonstrate that our algorithm can achieve a larger F1-score than Minkowski distance and less time than dynamic time warping distance.
Shiming He, Zhuozhou Li, Jin Wang 0001, Naixue Xiong
IEEE Trans. Ind. Informatics1
2019 Multi-Source Multicast Routing with QoS Constraints in Network Function Virtualization
abstract
The emergence of Network Function Virtualization (NFV) greatly improves the convenience of network layout, and at the same time reduces the service cost for operators. Currently, many studies on NFV layout focus on one-to-one unicast communication and cannot be extended to multicast. Therefore, in this paper, we are committed to exploring the joint problem of Virtual Network Function (VNF) layout and path selection in multi-source multicast. In addition, we take into account bandwidth and latency constraints, for they are the vital indices of Quality of Service (QoS). To solve this problem, we design a heuristic algorithm, named Multi-Source Multicast Tree Construction (MMTC). The algorithm aims to find a common link to place the Service Function Chain (SFC), a chain composed by multiple VNFs in rotation manner, so that the deployed SFC can be shared by all users, thereby improving the resource utilization. We then evaluate the performance of the proposed algorithm with different methods in four real topologies. Simulation results indicate that, compared to other heuristic algorithms, our design effectively reduce the total cost of services.
Kun Xie 0001, Thabo Semong, Shiming He
ICC4
2017 Simultaneous Wireless Information and Power Transfer for Multi-hop Energy-Constrained Wireless Network
Shiming He, Kun Xie 0001, Weiwei Chen 0004, Da-Fang Zhang 0001, Jigang Wen
WASA1
2017 A novel fault diagnosis method based on optimal relevance vector machine
Shiming He, Long Xiao, Yalin Wang 0003, Xinggao Liu, Chunhua Yang 0001, Jiangang Lu, Weihua Gui 0001, Youxian Sun
Neurocomputing1
2017 An efficient privacy-preserving compressive data gathering scheme in WSNs
Kun Xie 0001, Xueping Ning, Xin Wang 0001, Shiming He, Zuoting Ning, Jigang Wen, Zheng Qin 0001
Inf. Sci.4
2017 Opportunistic Routing and Scheduling for Wireless Networks
abstract
In spatial time division multiple access wireless mesh networks, not all links can be activated simultaneously, as links scheduled for transmission must satisfy the specified SINR requirements. Previously, slot assignment has been done on a link basis, where a set of links is selected for transmission in a given slot. However, if selected links are in deep fade or have no traffic to transmit, the slot is wasted. Thus, a node-based scheme was proposed, where a set of nodes is selected for transmission. Which link to be used by a node depends on the links' instantaneous traffic load. Although this allows us to exploit multi-user diversity, it creates a planning discrepancy: slot assignment is designed based on long-term channel statistics, but scheduling on short-term channel fading conditions. Consequently, the performance gain of the node-based scheme is not consistent: it is marginal under certain scenarios. To avoid the design discrepancy, we develop a new slot-assignment and routing framework in this paper. The new approach incorporates short-term channel fading statistics to optimize the long term slot assignment, routing and scheduling simultaneously. Hence, multi-user diversity can be exploited more efficiently. Not only is the performance gain of the resulting system significant (can be as much as 64% higher throughput than the scheme introduced by Chen and Lea), it is also less topology dependent compared with the one by Chen and Lea.
Weiwei Chen 0004, Chin-Tau A. Lea, Shiming He, Zhe Xuanyuan
IEEE Trans. Wirel. Commun.3
2015 Completion Time-Aware Flow Scheduling in Heterogenous Networks
Shiming He, Kun Xie 0001, Da-Fang Zhang 0001
ICA3PP (1)1
2015 Privacy Preserving for Network Coding in Smart Grid
Shiming He, Weini Zeng, Kun Xie 0001
ICA3PP (3)1
2015 A Distributed Joint Cooperative Routing and Channel Assignment in Multi-radio Wireless Mesh Network
Hong Qiao, Da-Fang Zhang 0001, Kun Xie 0001, Shiming He
ICA3PP (1)5
2015 An Efficient Privacy-Preserving Compressive Data Gathering Scheme in WSNs
Kun Xie 0001, Xueping Ning, Xin Wang 0001, Jigang Wen, Shiming He, Daqiang Zhang 0001
ICA3PP (1)6
2014 Routing and channel assignment in wireless cooperative networks
abstract
In recent years, cooperative communication has attracted researchers' attention as it showed a good capability to increase network performance. On the other hand, a cooperative transmission may cause more interference. This can cause difficulty in achieving cooperative diversity gain in multi-flow and multi-hop networks. In this paper, we propose a novel interference aware-cooperative-routing metric which will lead to creating a new scheme for cooperative routing and channel assignment in multi-flow and multi-hop networks. Then, we show through preliminary simulation results, by investigating the impact of the node density and impact of the flow number, that the proposed scheme minimizes interference while achieving maximum cooperative diversity gain.
Kun Xie 0001, Xin Wang 0001, Shiming He, Jigang Wen, Mohsen Guizani
IWCMC4
2014 Channel Aware Opportunistic Routing in Multi-Radio Multi-Channel Wireless Mesh Networks
Shiming He, Da-Fang Zhang 0001, Kun Xie 0001, Hong Qiao
J. Comput. Sci. Technol.1
2011 A Simple Channel Assignment for Opportunistic Routing in Multi-radio Multi-channel Wireless Mesh Networks
abstract
Opportunistic routing (OR) involves multiple forwarding candidates to relay packets by taking advantage of the broadcast nature and multi-user diversity of the wireless medium. Compared with Traditional Routing (TR), OR is more suitable for the unreliable wireless link, and can evidently improve the end to end throughput of Wireless Mesh Networks (WMNs). At present, there are many achievements concerning OR in the single radio wireless network. However, the study of OR in multi radio wireless network stays the beginning stage. In this paper, we focus on OR in multi-radio multi-channel WMNs. We validate the advantage of OR in multi-radio multi-channel WMNs, and propose a Simple Channel Assignment for Opportunistic Routing (SCAOR), which assigns channel to flows. According to interference state of every node, SCAOR assigns a channel with minimum interference to each flow to balance channel load. The simulation result shows OR of dual-radio dual-channel WMNs can promote throughput evidently, specifically, 16.8% higher than throughput of TR in the dual-radio dual-channel WMNs, 87.11% and 111.8% higher than throughput of OR and TR in single-radio single-channel, respectively.
Shiming He, Da-Fang Zhang 0001, Kun Xie 0001, Hong Qiao
MSN1