EDBT 2026 Demo / reviewers in the wild / expert
Shuhui Chen
dblp:31/9980
· DBLP profile ↗
63ranked-venue papers
6as first author
47since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 29 · 25 since 2021Security and privacy · 14 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorSystems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving Unified Memory for FPGA-Based String MatchingabstractString matching serves as a critical module for network security systems. To meet escalating network bandwidth demands, recent studies have transitioned to hardware platforms like FPGA, leveraging the parallel processing ability to accelerate string matching. However, existing hardware solutions face a critical challenge in parallel matching of variable-length string patterns. They require length-specific memory blocks, such as separate hash tables to store patterns of different lengths. This distributed memory architecture causes a memory fragmentation issue when patterns are unevenly distributed, which impacts the scalability of prior works. To address the memory fragmentation issue of the distributed memory architecture, this paper proposes (1) a unified memory architecture that enables unified storage of variable-length patterns, and (2) a collision-free hash scheme that supports parallel matching of variable-length strings in the architecture. We implemented the proposed memory-efficient scheme on an FPGA-based prototype. Extensive evaluations demonstrate that the unified memory architecture achieves 4-5× lower memory usage compared to state-of-the-art alternatives while achieving comparable throughput. Meanwhile, the architecture can store patterns across arbitrary length distributions within a specified length range until its maximum capacity. Zhuoxuan Sun, Jincheng Zhong, Jiayao Wang 0002, Shuhui Chen |
APNet | 4 |
| 2026 | AdaptTree: A Practical and Adaptive Packet Classification Scheme on FPGA
Jincheng Zhong, Gaofeng Lv, Shuhui Chen |
SECON | 7 |
| 2026 | Doubling the speed of large-scale packet classification through compressing decision tree nodes
Jincheng Zhong, Gaofeng Lv, Shuhui Chen |
Comput. Networks | 4 |
| 2026 | A survey on network traffic analysis with incomplete data
Zhengpeng Li, Shuhui Chen, Biying Wang, Minxin Wang |
Comput. Commun. | 3 |
| 2026 | Exploring the android TLS certificate ecosystem in ChinaabstractAbstract The HTTPS certificate ecosystem has long been a key topic in cybersecurity, yet the certificate landscape of Android applications remains insufficiently studied. In particular, while China has actively promoted the adoption of China’s national cryptographic algorithms in recent years, their actual deployment within the Chinese Android certificate ecosystem remains unclear. In this study, we analyzed TLS traffic from 19,980 applications in the Huawei App Market and extracted 131,933 certificate chains. While most certificates are properly configured, we identified 530 certificates with security risks, affecting 2043 applications. Notably, three SDK-related risk certificates were propagated across 1462 applications, substantially widening their security impact. Only 94 certificates using China’s national cryptographic algorithms were found, all within 89 financial applications, indicating deployment driven mainly by regulatory compliance. Furthermore, nearly 99% of leaf certificates chain back to foreign root Certificate Authorities, underscoring a strong dependency that may pose digital sovereignty risks under geopolitical uncertainty. This study highlights the existing challenges in the Chinese Android certificate ecosystem, particularly in terms of security and digital sovereignty, and offers relevant recommendations for improvement. Shuhui Chen, Ziling Wei, Fei Wang 0076, Zhenhao Luo |
Cybersecur. | 2 |
| 2026 | Blocking Is Not Stagnation: A Synchronous FPGA-CPU Architecture for Regular Expression Matching in Real-Time DPIabstractRegular expression matching is a crucial step in traffic analysis. Many hardware-based architectures are proposed to improve the matching throughput, such as FPGA. To date, however, the existing FPGA-CPU architectures are difficult to implement in DPI systems due to the following two reasons. First, existing architectures use asynchronous workflows to interact data between FPGA and CPU, making them difficult to be compatible with synchronous DPI systems. Second, asynchronous architectures require batch input, which does not meet the requirements of real-time environments. In this paper, we concentrate on the real-time deployment of a regular expression matching architecture. To improve the deployment throughput, we propose an FPGA-CPU architecture with a parallel layer between the driver and DPI systems. Then, coroutines are introduced and proved to have significant advantages. Meanwhile, some optimization methods are proposed to address idle time, memory allocation, and MMIO control. Our experiments demonstrate that directly deploying an asynchronous architecture on a synchronous DPI would result in a throughput degradation of 3 orders of magnitude. Our approach enhances throughput by 2-3 orders of magnitude. This indicates that we reach a throughput in synchronous mode that is comparable to that in asynchronous mode, and it is over 10 times faster than the software solution, making the direct deployment of asynchronous architectures on mainstream DPI systems feasible. To the best of our knowledge, this is the first attempt to improve hardware-based regular expression matching under synchronous logic, achieving both high throughput and usability. Shuhui Chen, Ziling Wei, Jincheng Zhong, Puguang Liu |
IEEE Trans. Netw. | 2 |
| 2025 | Cellular-Snooper: A General and Real-Time Mobile Application Fingerprinting Attack in LTE Networks
Wenao Zhang, Shuhui Chen, Ziling Wei, Qianqian Xing, Jinshu Su |
ICIC (4) | 2 |
| 2025 | ByteGT: A Hybrid Sequential-Attention Network for Enhancing File Fragment Classification on Raw DataabstractIn digital forensics practice, the precise determination of file fragment types serves as an essential prerequisite for successful file carving. Recent advancements in neural network-based methods have shown promise in this area, though challenges remain regarding temporal pattern capture in byte data and feature representation scalability within individual architectures. We propose ByteGT, the first hybrid neural network that integrates sequential modeling and attention mechanisms to further enhance the classification performance for file fragments. The model operates end-to-end on raw byte data without manual preprocessing through two novel components. The first component is a deep sequence perception module combining byte embeddings with bidirectional GRU to capture comprehensive temporal dependencies, and the second component is a fine-grained feature enhancement module using convolution-based attention layers to amplify discriminative features. Extensive evaluations on standard datasets demonstrate ByteGT’s superiority. Specifically, in most complex classification scenarios, we achieve 6.9% and 7% accuracy gains over state-of-the-art methods for 512-byte and 4096-byte sector sizes, respectively. When tested in other scenarios, ByteGT exhibits strong generalizability and robustness. Shuhui Chen, Ziling Wei |
IJCNN | 3 |
| 2025 | TFMana: A Traffic Feature Calibration Method to Empower Reliable Network Traffic AnalysisabstractIn recent years, network traffic analysis solutions that are driven by artificial intelligence models have achieved impressive performance. The “magic spells” of these solutions come from the knowledge that they learn from large amounts of network traffic data. However, these solutions neglect the impact of the real-world network's complexity on data quality, which makes the knowledge they learn from regular network traffic data difficult to be effective on low-quality data. Considering the packet loss in real-world network environments, this paper presents TFMana to calibrate the inaccurate packet length features extracted from incomplete network traffic data. TFMana utilizes an encoder-based masked language model to predict features of lost packets, incorporating network traffic feature embeddings to enhance prediction accuracy. This approach enables the calibrated features to approximate those extracted from loss-free network traffic asymptotically. Comprehensive experiments are conducted to verify the effectiveness of the proposed method. The evaluation demonstrates that TFMana's calibration achieves recovery accuracy between 83.94 % and 85.66 %, with minimal sensitivity to packet loss rates. Integrated with four benchmark application identification models, TFMana significantly improves classification accuracy under packet loss conditions. Notably, the analysis models maintains reliable performance even at high packet loss rates of 30 %. Zhengpeng Li, Shuhui Chen, Biying Wang, Minxin Wang |
IPCCC | 3 |
| 2025 | TrafficBM: A Dual-Modality Pre-Training Framework for Network Traffic ClassificationabstractNetwork traffic classification is critical for ensuring network quality, security, and stability. However, the increasing complexity of network environments and the growth of encrypted traffic bring significant challenges. Traditional rule-based, machine learning-based, and deep learning-based approaches are limited by the scarcity of plaintext, reliance on handcrafted features, and the need for large labeled datasets. Pre-training methods have alleviated these issues, but existing models mainly focus on payload semantics and lack dedicated learning of traffic behavior patterns essential for encrypted traffic characterization. Motivated by this, we propose TrafficBM, a dual-modality pre-training framework that jointly models semantic features and traffic behavior patterns. Our approach extracts dualmodality features from network traffic and applies modalityspecific data augmentation to mitigate data imbalance and scarcity. During pre-training, BERT leverages masked bigram modeling (MBM) to capture semantic information, while Mamba uses a masked autoencoder (MAE) architecture to learn traffic behavior patterns. An adaptive gating network, together with a parameter-preserving warm-up strategy, fuses features from both pre-trained models during fine-tuning to improve downstream classification performance. TrafficBM achieves state-of-the-art results on six tasks across eight datasets, including over 0.99 accuracy on five datasets and a 10 % improvement over the best baseline on Datacon2021 Part 2, demonstrating strong generalization and robustness in network traffic classification. Minxin Wang, Junhong Liao, Jinshu Su, Ziling Wei, Shuhui Chen, Zhengpeng Li, Biying Wang |
IPCCC | 5 |
| 2025 | I Know Who You are: An Identity Mapping Attack Based on Time Series Similarity in Mobile NetworksabstractIdentity privacy leakage through the wireless interface in mobile networks represents a persistent security challenge and a long-standing concern that network designers have aimed to address. Despite the remediation introduced in 5G standards, identity privacy attacks targeting the wireless interface continue to pose a potential threat. In this paper, we present an identity mapping attack based on time series similarity in 4G and 5G networks. This attack enables an adversary with no privileges to map a victim's social media account to their RNTI by sending a single image message to the victim and measuring the similarity of time series extracted from the generated downlink traffic. To improve the attack success rate, we specifically design an elastic similarity measure for time series, tailored to the properties of the data collected during the attack. We investigate the feasibility of the attack under various scenarios, achieving a success rate of 83% for a single attempt and nearly 100% when conducting two or three attempts. Our work provides new insights into the vulnerability of 4G/5G standards to identity privacy attacks. Wenao Zhang, Shuhui Chen, Junhong Liao, Ziling Wei, Mengyi Gong |
IPCCC | 2 |
| 2025 | PSSketch: Finding Persistent and Sparse Flow with High Accuracy and EfficiencyabstractFinding persistent sparse (PS) flow is critical to early warning of various threats. Previous works have predominantly focused on either heavy or persistent flows, with limited attention given to PS flows. Although some recent studies pay attention to PS flows, they struggle to establish an objective criterion due to insufficient data-driven observations, resulting in reduced accuracy. In this paper, we define a new criterion ''anomaly boundary'' to distinguish PS flows from regular flows. Specifically, a flow whose persistence exceeds a threshold will be protected, while a protected flow with a density lower than a threshold is reported as a PS flow. We then introduce PSSketch, a high-precision layered sketch, to find PS flows. PSSketch employs variable-length bitwise counters, where the first layer tracks the frequency and persistence of all flows, and the second layer protects potential PS flows and records overflow counts from the first layer. Some optimizations have also been implemented to reduce memory consumption further and improve accuracy. The experiments show that PSSketch reduces memory consumption by 1-2 orders of magnitude compared to the strawman solution combined with existing work. Compared with SOTA solutions for finding PS flows, it outperforms up to 2.94x higher in F1 score and reduces ARE by 1-2 orders of magnitude. Meanwhile, PSSketch achieves a higher throughput than these solutions. Qilong Shi, Xiyan Liang, Han Wang 0022, Wenjun Li 0004, Ziling Wei, Weizhe Zhang, Shuhui Chen |
KDD (2) | 8 |
| 2025 | SeRed: A Selective Reduction Method for Efficient Network Flow Storage
Shuhui Chen |
WASA (2) | 3 |
| 2025 | A survey of existing attacks on 5G SA
Mengyi Gong, Ziling Wei, Shuhui Chen, Wanrong Yu, Fei Wang 0076 |
Comput. Networks | 3 |
| 2025 | MFSI: Multi-flow based service identification for encrypted network traffic
Biying Wang, Ziling Wei, Shuhui Chen, Zhengpeng Li, Minxin Wang |
Comput. Networks | 5 |
| 2025 | EATIS: an environmentally adaptive traffic identification system for open world networks
Yulong Liang, Shuhui Chen, Yunjiao Bo |
Int. J. Inf. Comput. Secur. | 3 |
| 2025 | Toward an Effective Few-Shot Website Fingerprinting Attack With Quadruplet Networks and Deep Local Fingerprinting FeaturesabstractWebsite fingerprinting (WF) attacks can reveal the users' online privacy by the traffic analysis technique, even with the protection of the Tor anonymity network. Recent WF attacks tend to leverage the deep learning (DL) models, which require a large number of traffic samples for training. In this case, it is impractical for low-resource adversaries in reality. Thus, we propose a lightweight WF attack to tackle this challenge, i.e., Deep Quadruplet Fingerprinting (DQF), which only needs one training sample to obtain an accuracy of 87.1%. Regarding the overall design, DQF first combines the metric learning and meta-learning schemes. To improve the generalization ability of the trained model, DQF leverages the quadruplet networks as the architecture and modifies the quadruplet loss function. Besides, by taking the deep local fingerprinting features (DLFFs), DQF avoids losing a lot of discriminative information, which is a problem with previous attacks. To evaluate DQF, we use multiple typical datasets and conduct 11 different experiments. In closed-world settings, the accuracy of DQF can exceed the best baseline attack by 10%. In open-world settings, DQF steadily performs the best even in the most challenging scenario, namely, 1-shot learning, where previous attacks significantly degrade the performance or even fail. Hongcheng Zou, Jinshu Su, Ziling Wei, Shuhui Chen, Chunfang Yang, Mantun Chen |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | BK-Index: A Multi-attribute Index algorithm for Network TrafficabstractArchiving raw network traffic is fundamental for network forensics and troubleshooting. In order to efficiently retrieve specific entries from large-scale archives, multi-attribute queries are essential. However, most previous work mainly focus on single-attribute indexing, leading to limited performance for multi-attribute queries and index updates. To this end, we propose BK-Index, a novel multi-attribute index algorithm for network traffic. BK-Index combines multi-dimensional k-ary search trees and bitmap structures to efficiently handle both high-cardinality and low-cardinality attributes. This hybrid approach significantly reduces multi-attribute retrieval response time. Additionally, BK-Index employs a periodic batch construction scheme for real-time index updates in high-speed network links, eliminating the need for costly dynamic index maintenance. A comprehensive suite of experiments conducted on actual campus network traffic data has elucidated the superior performance of the BK-Index in terms of multi-attribute indexing and retrieval compared with other state-of-the-art index methods. Linghao Zhao, Shuhui Chen |
HPCC | 4 |
| 2024 | KP-WF: Cross-Domain Few-Shot Website Fingerprinting
Lin Liu 0018, Ziling Wei, Shuhui Chen, Jinshu Su |
ICDF2C (2) | 3 |
| 2024 | FingerMamba: Mamba-based Efficient Multi-tab Website FingerprintingabstractNowadays, protecting user privacy on the Internet is paramount, especially with the increasing use of the Tor network to anonymize online activities. However, Tor is vulnerable to website fingerprinting (WF), where patterns in encrypted traffic are analyzed to infer visited websites. It can be utilized to monitor and investigate illegal activities on the dark web. Existing website fingerprinting techniques typically assume single-tab browsing, which is unrealistic as users often open multiple tabs consecutively or within a short period due to Tor’s slow loading speeds and typical user habits. Moreover, current multi-tab approaches face challenges in classification speed, which is crucial for high-throughput networks. FingerMamba, our proposed model, addresses these gaps by efficiently extracting local information and establishing long-range dependencies using a Mamba-based structured state-space model. It significantly enhances the accuracy and speed of multi-tab website fingerprinting. Extensive experiments on the largest real-world multi-tab dataset demonstrate that FingerMamba effectively improves classification accuracy in both closed-world and open-world settings. Furthermore, with maintaining similar accuracy performance, FingerMamba can increase inference speed by up to four times compared to the existing methods. To our knowledge, FingerMamba is the first model to tailor the Mamba architecture for website fingerprinting. Lin Liu 0018, Ziling Wei, Shuhui Chen, Zixuan Dong, Jinshu Su |
IPCCC | 3 |
| 2024 | AST-Trans: Detecting Web Tracking using Transformer-based Deep Learning with Abstract Syntax TreeabstractWeb tracking has become a key tool for service providers to collect online data and analyze user behaviors, raising concerns about the privacy of Internet users. In this paper, we propose a new web tracking detection method, namely AST-Trans, which detects and removes web tracking behavior using Transformer-based deep learning with abstract syntax trees. In the method, the abstract syntax tree is built for the detected website codes. Then, a sequence generation algorithm is proposed to convert tree-like code structures into one-dimensional sequences for deep models. To enhance training efficiency, we devise a reduction strategy to simplify the code tree structure by introducing equivalent nodes. After that, a Transformer-based deep learning algorithm is introduced to realize web tracking detection. By the proposed method, the exact tracking code blocks can be identified, and thus, we can implement the tracking code removal with minimum website breakage. To verify the effectiveness of the proposed method, an HTTPS proxy with AST-Trans on it is implemented to detect and remove tracking codes. We evaluate AST-Trans with the TrackSign-labeled dataset. The results show that the proposed method can detect the tracking behavior with high precision. In addition, we validate the feasibility of the method by measuring the website page breakage. Ziling Wei, Lin Liu 0018, Shuhui Chen, Jinshu Su |
IPCCC | 4 |
| 2024 | From n to n+x: A Novel Open World Traffic Classification Framework Based on Multiple Classifiers EnsembleabstractApplication-level traffic classification is a critical component in network management and security. In recent years, intelligent classification methods have been proven effective, especially for encrypted network traffic. However, these methods face several limitations in real networks, such as a lack of interpretability and inefficiency of complex models. More importantly, they seldom consider unknown applications and only identify them as unknown. This paper proposes a novel traffic classification framework named APPrint, which identifies the traffic of trained applications but can also differentiate flows generated from diverse unknown applications. The core of APPrint is to trade off recognition accuracy for enhanced generalization capability. By integrating the results of multiple high-accuracy classifiers, APPrint effectively addresses the n+x classification problem using n known applications. Through comparative experiments with two State-of-the-Art(SOTA) solutions, APPprint demonstrates a significant accuracy advantage when there are a large number of unknown applications. Additionally, unlike existing solutions that only perform n+1 classification using n known applications, APPrint is capable of classifying unknown applications, achieving n+x classification. Kaixing Liu, Shuhui Chen |
ISPA | 3 |
| 2024 | TriNT: A Framework for ROV Identification Based on Triplet
Jiangbin Chen, Shuhui Chen |
ISPEC | 3 |
| 2024 | A Large-Scale Mobile Traffic Dataset For Mobile Application IdentificationabstractAbstract With Internet access shifting from desktop-driven to mobile-driven, application-level mobile traffic identification has become a research hotspot. Although considerable progress has been made in this research field, two obstacles are hindering its further development. Firstly, there is a lack of sharable labeled mobile traffic datasets. Although it is easy to capture mobile traffic, labeling traffic at the application level is non-trivial. Besides, researchers usually hold a conservative attitude toward publishing their datasets for privacy concerns. Secondly, most of the datasets used by existing studies are inadequate to evaluate the proposed methods, since they usually have the problems of inaccurate labels, small scale and simple collection configurations. To tackle these two obstacles, a mobile traffic collection is carried out in this paper. The collected traffic has the advantages of large-scale data size, accurate application-level labels and diverse collection configurations. Then, the collected traffic is anonymized carefully to make it public. Several mobile traffic identification methods are compared based on our anonymized dataset, which proves the applicability of our dataset. Shuhui Chen, Fei Wang 0076, Ziling Wei, Jincheng Zhong, Jianbing Liang |
Comput. J. | 2 |
| 2024 | Protecting unauthenticated messages in LTE/5G mobile networks: A two-level Hierarchical Identity-Based Signature (HIBS) solution
Chuan Yu 0003, Shuhui Chen, Qianqian Xing, Ziling Wei |
Comput. Networks | 2 |
| 2024 | STI: A self-evolutive traffic identification system for unknown applications based on improved random forest
Yulong Liang, Shuhui Chen, Beier Chen, Yunjiao Bo |
Comput. Commun. | 3 |
| 2024 | Relation-CNN: Enhancing website fingerprinting attack with relation features and NFS-CNN
Hongcheng Zou, Ziling Wei, Jinshu Su, Shuhui Chen |
Expert Syst. Appl. | 4 |
| 2024 | Statistic Ratio Attention-Guided Siamese U-Net for SAR Image Semantic Change DetectionabstractSemantic change detection, which aims to locate land cover changes and identify their categories using pixel-level boundaries, has promising applications in Earth vision, including precise urban planning and natural resource management. This paper proposes a novel Siamese U-Net architecture for semantic change detection in synthetic aperture radar (SAR) images, incorporating a residual network with weight-sharing as the backbone network. The network is capable of simultaneously yielding binary change detection and semantic change detection results. Additionally, we have designed a statistic ratio attention module that utilizes statistical features from the original image as spatial ratio attention, coupled with channel attention, to extract change information from the bi-temporal SAR images. Furthermore, as there is currently no existing dataset for semantic change detection in SAR images, we have constructed a dedicated dataset to facilitate model training and evaluation. Our experiment results demonstrate the superiority of our proposed model over other comparison algorithms. Shuhui Chen, Xin Su 0003, Li Zheng 0004, Qiangqiang Yuan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Toward a Truly Secure Telecom Network: Analyzing and Exploiting Vulnerable Security Configurations/ Implementations in Commercial LTE/IMS NetworksabstractAuthentication and data protection (both integrity and confidentiality) between the network and cellular devices are two fundamental security features in LTE and IMS networks. The first is implemented via authentication and key agreement mechanisms and can be compromised by relaying authentication parameters. The second security feature builds on the first one and is activated through corresponding security setup procedures. This work intends to investigate whether these basic security procedures are securely implemented and deployed in commercial networks. We analyzed the de facto situation of these security features in three major operators in China and found several new and previously disclosed configuration and implementation flaws that do not conform to specifications. These vulnerabilities allow attackers to disable LTE and IMS data protection mechanisms. We further propose novel proof-of-concept attacks to exploit the identified vulnerabilities includingIMEIandPhone Number Catching,SMSandCall ImpersonationandInterceptionattacks. To show the urgency of addressing these security issues and thus secure the real-world telecom networks, we successfully demonstrated these attacks in practice using open-source SDR tools as they have serious implications. For instance, the interception attacks undermine the widely-used SMS verification code security mechanism. We also discuss countermeasures to resist the proposed attacks. Chuan Yu 0003, Shuhui Chen, Ziling Wei, Fei Wang 0076 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | FPGA-CPU Architecture Accelerated Regular Expression Matching With Fast PreprocessingabstractAbstract Regular Expression Matching (REM) is the core of Deep Packet Inspection (DPI), which is important for various network security applications. The burgeoning Software Defined Network and Network Function Virtualization technologies make the network evolve more dynamic, which brings serious challenges for DPI engines to achieve high matching performance with fast rule-set update capability. To meet these challenges, this paper proposes a heterogeneous Field Programmable Gate Array (FPGA)-Central Processing Unit (CPU) architecture to accelerate Deterministic Finite Automaton (DFA)-based REM with high preprocessing performance. Firstly, a novel regex decomposition technique is proposed to solve the DFA state explosion problem, which splits each regex into one prefix and several postfixes. Secondly, heterogeneous architecture is presented to collaboratively handle regex matching, in which prefixes are matched in parallel in an FPGA and postfixes are matched in a CPU. To further improve the matching performance, several well-designed DFA compression techniques and regex decomposition optimizations are proposed. Our design has been implemented in a DPI prototype employing a medium-end FPGA. Extensive experiments are conducted to evaluate the performance. Results reveal that our proposed architecture achieves 6.33 Gbps matching throughput on the Snort rule-set (v3.0), which is close to state-of-the-art FPGA NFA-based schemes. However, the rule-set preprocessing time is significantly reduced to <7 minutes, compared with up to several hours of FPGA NFA-based countermeasures. Jincheng Zhong, Shuhui Chen, Biao Han 0003 |
Comput. J. | 2 |
| 2023 | FECC: DNS tunnel detection model based on CNN and clustering
Jianbing Liang, Suxia Wang, Shuhui Chen |
Comput. Secur. | 4 |
| 2023 | SecChecker: Inspecting the security implementation of 5G Commercial Off-The-Shelf (COTS) mobile devices
Chuan Yu 0003, Shuhui Chen, Ziling Wei, Fei Wang 0076 |
Comput. Secur. | 2 |
| 2023 | TupleTree: A High-Performance Packet Classification Algorithm Supporting Fast Rule-Set UpdatesabstractPacket classification plays a crucial role in various network functions such as access control and routing. In recent years, the rapid development of SDN and NFV poses new challenges for packet classification to support fast rule-set updates as introducing strong dynamics for the structure of networks. To this end, this paper proposes a novel scheme, TupleTree, to perform high-speed packet classification while providing fast rule-set update ability. TupleTree is a hybrid scheme combining decision tree and tuple space. In TupleTree, it organizes rules in a decision tree-like structure, but distributes rules in each node into child nodes through hashing rather than cutting or splitting. With the decision tree structure, for each classification, one leaf node containing a few rules can be rapidly indexed. Hence, a high classification performance can be achieved. Meanwhile, with hashing instead of cutting or splitting, it is easy to support fast rule-set updates due to having avoided the rule replication problem. Compared to state-of-the-art schemes that support fast rule-set updates, experimental results show that our proposed scheme achieves a classification performance improvement of 85% to 237% while retaining close update performance for large rule-sets. Jincheng Zhong, Ziling Wei, Shuhui Chen |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | Multi-Level Text Importance Classification Architecture Based on Deep LearningabstractIn the era of information explosion, the Internet is full of spam and false information, making it more difficult for people to obtain effective information. Since text data is the main carrier for disseminating information and knowledge, we propose a multi-level text importance classification architecture based on deep learning to enable Internet users to quickly and accurately access text content of interest. Experiments demonstrate that the proposed architecture can achieve a good performance. Meizhen Huang, Jinshu Su, Zhong Liao, Shuhui Chen, Ziling Wei |
APNet | 4 |
| 2022 | RLCS: A Classification System Based on Random Forest and Logistic Regression for Hybrid Zero-day TrafficabstractTraffic classification has attracted public attention for a long time because of its essential role in network management. However, the presence of zero-day traffic, network traffic generated by previously unknown applications, leads to a significant reduction in the practicability and effectiveness of conventional traffic classification schemes. This poster innovatively proposes a traffic classification scheme named RLCS to accomplish the high accurate traffic classification task in hybrid zero-day traffic. The evaluations with real-world traffic verify the effectiveness and broad applicability of RLCS. Yulong Liang, Fei Wang 0076, Shuhui Chen |
APNet | 3 |
| 2022 | Reducing Network Traffic Storage Overhead: A Hardware-Accelerated Lossless Data Compression SystemabstractSkyrocketing network traffic volume brings significant overhead of traffic storage. We propose a hardware-accelerated lossless data compression system to reduce traffic storage overhead. The prototype evaluation validates the excellent resource efficiency and performance scalability of our traffic compression system. Puguang Liu, Shuhui Chen |
APNet | 2 |
| 2022 | FATSS: Filter-Assisted Tuple Space Search for Packet ClassificationabstractPacket Classification is a key part of supporting lots of network functions. Various algorithms have been proposed over the years to meet the increasing performance requirements of packet classification. Tuple space search (TSS) is one of the most popular algorithms and well-suited to scenarios requiring efficient online updates. However, the huge number of tuples in the algorithm leads to numerous memory accesses during packet classification, which limits the classification performance. This paper proposes a novel model named FATSS, which uses Filters to Assist the Tuple Space Search algorithm and reduces the number of tuple accesses. We first create the ImCuckoo Filter by improving the Cuckoo Filter from its structure, capacity and hash calculation. Then, we embed ImCuckoo Filter into TSS in two ways (online and offline) to adapt to diverse scenarios and requirements. By the experiments, it can be found that the ImCuckoo Filter can reduce more than 80% of tuple accesses. Furthermore, the access time of the filter is no more than 60% compared with that of the hash table. The experimental results show that the classification time of FATSS is 17%–19% faster than that of existing widely used algorithms. Jiayao Wang 0002, Ziling Wei, Jincheng Zhong, Shuhui Chen |
IPCCC | 5 |
| 2022 | Research on the derivation of AS hidden links and the Discovery of Critical ASabstractAS relationships are the basis for studying Internet security, route hijacking, route leakage, etc. To obtain more complete AS relationships, using machine learning (ML) models to learn the similarity between adjacent link groups and predict hidden links is a method that can obtain more complete AS relationships. The features selected by the ML model have a large impact on the accuracy of the prediction results, and we extract 10 ML features by combining the actual geographic location information of AS. After our optimization, the accuracy of the prediction model reaches 91.57%. In the classification of hidden link types, we oversample the small sample type data and optimize the classifier, and the classification accuracy of hidden link categories reaches 97.42%. The recall rate of p2c and c2p links improved by 24.29% and 7.17%, respectively. We found that the hidden links caused the change of network traffic transmission routes by the change of "Critical AS" in the AS network. AS 3549 has the highest number of effective paths, and the network traffic prefers to choose the AS with a lower hierarchy for forwarding. Jiangbin Chen, Shuhui Chen, Xiangquan Shi |
LCN | 3 |
| 2022 | RTSS: Robust Tuple Space Search for Packet ClassificationabstractPacket classification shows an essential role in net-work functions. Traditional classification algorithms assume that all field values are available and valid. However, such a premise is being challenged as networks become more complex now. Scenarios with field-missing poses great challenges to packet classifiers. Existing approaches can only list all possible situations in such cases, increasing the workload exponentially. RFC algorithm is proved to be helpful for this issue in our previous work, but its spacial performance is much poor. In this paper, we propose a novel classification scheme using Tuple Space Search (TSS) to deal with missing fields. We redesign the hash calculation method and raise a new data structure to recover field-missing packets. The experiment shows that RTSS reduce the memory consumption and construction time by several orders of magnitude. At the same time, RTSS has better classification performance than previous work, while supporting fast updates. Jiayao Wang 0002, Ziling Wei, Shuhui Chen, Jincheng Zhong |
MSN | 4 |
| 2022 | Comprehensive Mobile Traffic Characterization Based on a Large-Scale Mobile Traffic Dataset
Jincheng Zhong, Shuhui Chen, Jianbing Liang |
NSS | 3 |
| 2022 | An efficient cross-domain few-shot website fingerprinting attack with Brownian distance covariance
Hongcheng Zou, Jinshu Su, Ziling Wei, Shuhui Chen, Baokang Zhao |
Comput. Networks | 4 |
| 2022 | HAGDetector: Heterogeneous DGA domain name detection modelabstractThe botnet relies on the Command and Control (C&C) channels to conduct its malicious activities remotely. The Domain Generation Algorithm (DGA) is often used by botnets to hide their Command and Control (C&C) server and evade take-down attempts, which allows the bot to generate a large number of domain names until it finds its C&C server. The lengths of domain names generated by DGAs are different. Our research finds that the length of the domain name has an impact on the performance of the DGA domain name detection model. In other words, the model is sensitive to the length of the domain name. In this case, attackers can evade detection simply by designing domain names of specific lengths. Moreover, the detection accuracy of DGA domain names still needs to be further improved. To solve these problems, three feature extraction methods adapted to the length of the domain name are proposed in this paper. For extra-short domain names, we use the attention-based method to extract features, which can make use of the character-level semantic feature. For moderate-length domain names, a two-dimensional structure, namely Right Shifted Tensor (RST), is constructed to make the domain name present apparent features similar to images. For the extra-long domain name, the effective classification of domain names can be achieved by manually crafted easy-to-calculate features. Then, different detection structures are designed based on these tree feature extraction methods to form a heterogeneous DGA detection model, namely HAGDetector. In addition, the public suffix is an important part of the domain name. We further analyze the public suffix to evaluate its impact on the detection of DGA domain names. Finally, the experiments are conducted to assess the validity of HAGDetector, as well as compare our approach with the current state-of-the-art and highlight the impact of the domain name length. The experimental results show that our method greatly improves the detection performance. Jianbing Liang, Shuhui Chen, Ziling Wei |
Comput. Secur. | 2 |
| 2022 | Entity and relation collaborative extraction approach based on multi-head attention and gated mechanismabstractEntity and relation extraction has been widely studied in natural language processing, and some joint methods have been proposed in recent years. However, existing studies still suffer from two problems. Firstly, the token space information has been fully utilized in those studies, while the label space information is underutilized. However, a few preliminary works have proven that the label space information could contribute to this task. Secondly, the performance of relevant entities detection is still unsatisfactory in entity and relation extraction tasks. In this paper, a new model GANCE (Gated and Attentive Network Collaborative Extracting) is proposed to address these problems. Firstly, GANCE exploits the label space information by applying a gating mechanism, which could improve the performance of the relation extraction. Then, two multi-head attention modules are designed to update the token and token-label fusion representation. In this way, the relevant entities detection could be solved. Experimental results demonstrate that GANCE has better accuracy than several competitive approaches in terms of entity recognition and relation extraction on the CoNLL04 dataset at 90.32% and 73.59%, respectively. Moreover, the F1 score of relation extraction increased by 1.24% over existing approaches in the ADE dataset. Shuhui Chen, Tien-Hsiung Weng, Wenjie Kang |
Connect. Sci. | 3 |
| 2021 | CMT: An Efficient Algorithm for Scalable Packet ClassificationabstractAbstract Packet classification plays an essential role in diverse network functions such as quality of service, firewall filtering and load balancer. However, implementing an efficient packet classifier is a challenging problem. The problem even gets worse in the era of software-defined network, in which frequent rule updates are performed, and complex flow tables are used. This paper proposes CMT, a new software algorithm named by its novel data structure—common mask tree—to implement an efficient multi-field packet classifier. The core idea of CMT is to combine the strengths of both decision-tree and tuple-space schemes by employing tree-like structures and hash tables simultaneously. The objective of CMT is to achieve both high classification performance and fast rule updates. In the evaluation section, CMT is compared with decision-tree and tuple-space schemes. Compared to the state-of-the-art decision-tree methods, CMT performs rule updates at two orders of magnitude faster. CMT has a stable performance on different rulesets and achieves a 40% improvement in memory access compared to the state-of-the-art tuple-space method. Shuhui Chen, Jincheng Zhong, Ziling Wei |
Comput. J. | 1 |
| 2021 | Improving 4G/5G air interface security: A survey of existing attacks on different LTE layers
Chuan Yu 0003, Shuhui Chen, Fei Wang 0076, Ziling Wei |
Comput. Networks | 2 |
| 2021 | Efficient multi-category packet classification using TCAM
Jincheng Zhong, Shuhui Chen |
Comput. Commun. | 2 |
| 2021 | RF-RISA: A novel flexible random forest accelerator based on FPGA
Shuhui Chen, Fei Wang 0076, Ziling Wei |
J. Parallel Distributed Comput. | 2 |
| 2020 | Real Network Traffic Collection and Deep Learning for Mobile App IdentificationabstractThe proliferation of mobile devices over recent years has led to a dramatic increase in mobile traffic. Demand for enabling accurate mobile app identification is coming as it is an essential step to improve a multitude of network services: accounting, security monitoring, traffic forecasting, and quality-of-service. However, traditional traffic classification techniques do not work well for mobile traffic. Besides, multiple machine learning solutions developed in this field are severely restricted by their handcrafted features as well as unreliable datasets. In this paper, we propose a framework for real network traffic collection and labeling in a scalable way. A dedicated Android traffic capture tool is developed to build datasets with perfect ground truth. Using our established dataset, we make an empirical exploration on deep learning methods for the task of mobile app identification, which can automate the feature engineering process in an end-to-end fashion. We introduce three of the most representative deep learning models and design and evaluate our dedicated classifiers, namely, a SDAE, a 1D CNN, and a bidirectional LSTM network, respectively. In comparison with two other baseline solutions, our CNN and RNN models with raw traffic inputs are capable of achieving state-of-the-art results regardless of TLS encryption. Specifically, the 1D CNN classifier obtains the best performance with an accuracy of 91.8% and macroaverage F-measure of 90.1%. To further understand the trained model, sample-specific interpretations are performed, showing how it can automatically learn important and advanced features from the uppermost bytes of an app’s raw flows. Xin Wang 0076, Shuhui Chen, Jinshu Su |
Wirel. Commun. Mob. Comput. | 2 |
| 2019 | TCoD: A Traveling Companion Discovery Method Based on Clustering and Association AnalysisabstractWidely used mobile locating equipment, like phones, generates extensive spatio-temporal data every day. Since the data implicitly reflects behavior characteristics of moving objects, many applications focus on trajectory data mining while traveling companion discovery is one of the most fundamental techniques in these areas. Previous work based on time snapshot slicing have yielded some success on finding companions in regularly and frequently sampled trajectory data. However, it does not work well when data is sparse. This situation commonly appears in real world because of equipment failure or manual intervention. In this paper, we propose a novel Traveling Companion Discovery (TCoD) method that can discovery travelling companions even when time gaps between sample data are more than hours. TCoD combines density clustering and association analysis, while density clustering mine potential sets from perspective of location and association analysis identify related objects in potential sets. The evaluation on two real trajectory data sets shows that TCoD effectively overwhelms previous work with satisfying discovery of long-term companion patterns in sparse trajectory data. Ruihong Yao, Shuhui Chen |
IJCNN | 3 |
| 2019 | On Effects of Mobility Management Signalling Based DoS Attacks Against LTE TerminalsabstractLong Term Evolution (LTE) has become the most mature and stable mobile communication network technology worldwide so far. Billions of mobile subscribers are using LTE networks to surf the Internet, make phone calls, and send text messages every day. Meanwhile, Denial-of- Service (DoS) attacks seriously threaten the availability of LTE networks. In this paper, we focus on the research into the effects of reject signalling based DoS attacks against LTE terminals. We revealed a new DoS attack vulnerability in the authentication procedure after a detailed exploration in 3GPP standard specifications and verified it using software radio tools. Moreover, we specifically tested the impacts of such device-targeted DoS attacks on users under different test conditions (e.g. different operators, chip vendors, etc and we classified the actual test results into 6 different impact levels to better evaluate the effects on subscribers. Several possible countermeasures are also discussed to defend the attacks. Chuan Yu 0003, Shuhui Chen |
IPCCC | 2 |
| 2019 | Practical privacy-preserving deep packet inspection outsourcingabstractSummary Hardware‐based middleboxes are ubiquitous in computer networks, which usually incur high deployment and management expenses. A recently arising trend aims to address those problems by outsourcing the functions of traditional hardware‐based middleboxes to high volume servers in a cloud. This technology is promising but still faces a few challenges from different aspects, including privacy concerns, middlebox functionality, and performance. In this paper, we propose two practical approaches to implementing a cloud‐based DPI middlebox. The outsourced DPI middlebox performs payload inspection over encrypted traffic while preserving the privacy of both communication data and inspection rules. Our first approach employs a modified reversible sketch structure, which is used for efficient error‐free membership testing, and our second approach extends the famous AC pattern matching algorithm to the cipher text domain. We utilize unkeyed one‐way hash functions instead of complex cryptographic protocols to achieve the privacy preservation requirements. Our system supports a wide range of real‐world inspection rules. We conduct evaluations on the ClamAV rule set, and the experiment results demonstrate the effectiveness of our proposals. Jie Li 0041, Jinshu Su, Rongmao Chen, Xiaofeng Wang 0002, Shuhui Chen |
Concurr. Comput. Pract. Exp. | 5 |
| 2019 | LTE Phone Number Catcher: A Practical Attack against Mobile PrivacyabstractPhone number is a unique identity code of a mobile subscriber, which plays a more important role in the mobile social network life than another identification number IMSI. Unlike the IMSI, a mobile device never transmits its own phone number to the network side in the radio. However, the mobile network may send a user’s phone number to another mobile terminal when this user initiating a call or SMS service. Based on the above facts, with the help of an IMSI catcher and 2G man-in-the-middle attack, this paper implemented a practicable and effective phone number catcher prototype targeting at LTE mobile phones. We caught the LTE user’s phone number within a few seconds after the device camped on our rogue station. This paper intends to verify that mobile privacy is also quite vulnerable even in LTE networks as long as the legacy GSM still exists. Moreover, we demonstrated that anyone with basic programming skills and the knowledge of GSM/LTE specifications can easily build a phone number catcher using SDR tools and commercial off-the-shelf devices. Hence, we hope the operators worldwide can completely disable the GSM mobile networks in the areas covered by 3G and 4G networks as soon as possible to reduce the possibility of attacks on higher-generation cellular networks. Several potential countermeasures are also discussed to temporarily or permanently defend the attack. Chuan Yu 0003, Shuhui Chen, Zhiping Cai |
Secur. Commun. Networks | 2 |
| 2019 | Identifying Known and Unknown Mobile Application Traffic Using a Multilevel ClassifierabstractDue to the proliferation of mobile applications, mobile traffic identification plays a crucial role in understanding the network traffic. However, the pervasive unconcerned apps and the emerging apps pose great challenges to the mobile traffic identification method based on supervised machine learning, since such method merely identifies and discriminates several apps of interest. In this paper we propose a three-layer classifier using machine learning to identify mobile traffic in open-world settings. The proposed method has the capability of identifying traffic generated by unconcerned apps and zero-day apps; thus it can be applied in the real world. A self-collected dataset that contains 160 apps is used to validate the proposed method. The experimental results show that our classifier achieves over 98% precision and produces a much smaller number of false positives than that of the state of the art. Shuhui Chen, Yipin Sun, Zhiping Cai, Jinshu Su |
Secur. Commun. Networks | 2 |
| 2019 | SeCo-LDA: Mining Service Co-Occurrence Topics for Composition RecommendationabstractService composition remains an important topic where recommendation is widely recognized as a core mechanism. Existing works on service recommendation typically examine either association rules from mashup-service usage records, or latent topics from service descriptions. This paper moves one step further, by studying latent topic models over service collaboration history. A concept of service co-occurrence topic is coined, equipped with a mechanism developed to construct service co-occurrence documents. The key idea is to treat each service as a document and its co-occurring services as the bag of words in that document. Four gauges are constructed to measure self-co-occurrence of a specific service. A theoretical approach, Service Co-occurrence LDA (SeCo-LDA), is developed to extract latent service co-occurrence topics, including representative services and words, temporal strength, and services' impact on topics. Such derived knowledge of topics will help to reveal the trend of service composition, understand collaboration behaviors among services and lead to better service recommendation. To verify the effectiveness and efficiency of our approach, experiments on a real-world data set were conducted. Compared with methods of Apriori, content matching based on service description, and LDA using mashup-service usage records, our experiments show that SeCo-LDA can recommend service composition more effectively, i.e., 5% better in terms of Mean Average Precision than baselines. Zhenfeng Gao, Yushun Fan, Cheng Wu 0002, Wei Tan 0001, Jia Zhang 0001, Yayu Ni, Shuhui Chen |
IEEE Trans. Serv. Comput. | 8 |
| 2018 | Privacy-Preserving Mining of Association Rule on Outsourced Cloud Data from Multiple Parties
Lin Liu 0018, Jinshu Su, Rongmao Chen, Ximeng Liu, Xiaofeng Wang 0002, Shuhui Chen, Ho-fung Leung |
ACISP | 6 |
| 2017 | Service Recommendation Based on Separated Time-aware Collaborative Poisson Factorization
Shuhui Chen, Yushun Fan, Wei Tan 0001, Jia Zhang 0001, Zhenfeng Gao |
J. Web Eng. | 1 |
| 2016 | Time-Aware Collaborative Poisson Factorization for Service RecommendationabstractWith the booming number of web services, it is a challenge for inexperienced developers to select suitable services and make service compositions. Therefore, recommending services based on user queries becomes a necessity. For modeling the queries and services' descriptions, many recent studies are based on LDA (Latent Dirichlet Allocation). However, some previous empirical works indicate that LDA model doesn't gain high accuracy in generating latent presentation which is subject to the restrictive assumption of the Dirichlet-Multinomial distribution. In this paper, we propose a Time-aware Collaborative Poisson Factorization (TCPF) to tackle the problem. TCPF takes Poisson Factorization as the foundation to model mashup queries and service descriptions separately, and incorporate them with the historical usage data together using collective matrix factorization. Experiments on the real-world ProgrammableWeb dataset show that our model outperforms the state-of-the-art methods (e.g., Time-aware collaborative domain regression) by 7.7% in terms of mean average precision, and costs much less time on the sparse, massive and long-tailed data set. Shuhui Chen, Yushun Fan, Wei Tan 0001, Jia Zhang 0001, Zhenfeng Gao |
ICWS | 1 |
| 2016 | SeCo-LDA: Mining Service Co-occurrence Topics for RecommendationabstractService ecosystem consists of all kinds of services, and some of them may be composed by developers to create new mashups. Existing work on service recommendation and composition mine either frequent patterns from mashup-service usage records, or latent topics from service metadata. In this paper, we propose Service Co-occurrence LDA (SeCo-LDA), a novel approach that mines latent topic models over service co-occurrence patterns. The key idea is to treat each service as a document, and its bag of co-occurring services as the bag of words in that document. Using this model, we can analyze such service co-occurrence documents with a probabilistic topic model. We show how to derive service co-occurrence topics, and then validate our model on the real-world ProgrammableWeb.com dataset. We illustrate that SeCo-LDA can discover meaningful latent service composition patterns including their temporal strength and services' impacts, which conventional Apriori can not reveal. Comparing with Apriori, content matching based on service description and LDA directly using mashup-service usage records, we have demonstrated that SeCo-LDA can recommend service composition more effectively, 5% better in terms of MAP than the baseline approach. Zhenfeng Gao, Yushun Fan, Cheng Wu 0002, Wei Tan 0001, Jia Zhang 0001, Yayu Ni, Shuhui Chen |
ICWS | 8 |
| 2016 | A 60Gbps DPI Prototype based on Memory-Centric FPGAabstractDeep packet inspection (DPI) is widely used in content-aware network applications to detect string features. It is of vital importance to improve the DPI performance due to the ever-increasing link speed. In this demo, we propose a novel DPI architecture with a hierarchy memory structure and parallel matching engines based on memory-centric FPGA. The implemented DPI prototype is able to provide up to 60Gbps full-text string matching throughput and fast rules update speed. Jinshu Su, Shuhui Chen, Biao Han 0003, Xin Wang 0076 |
SIGCOMM | 2 |
| 2015 | Service Recommendation for Mashup Creation Based on Time-Aware Collaborative Domain RegressionabstractMash up has emerged as a promising way to compose web APIs and create value-added compositions. The increasing of APIs demands more accurate recommendation algorithms. However, service domain evolution, mash up-side cold-start and information evaporation are somehow overlooked by existing work. In this paper, by extending the collaborative topic regression (CTR) model, the procedure of service selection is modeled with a generative process, and the mash up-side cold-start problem that cannot be dealt with by naïve CTR is resolved. By learning the maximum a posteriori estimates of the whole generative process, both content information and historical usage are taken into consideration to extract service domains, thus the service domains can evolve with the evaluation of historical usage pattern. Meanwhile, information evaporation is also considered by giving time-related confidence levels to historical usage to track the evolution of service ecosystem. Experiments on the real-world Programmable Web data set show that compared with the state-of-the-art methods, our approach gains a 6.8% improvement in terms of recommendation accuracy. Yushun Fan, Keman Huang, Wei Tan 0001, Bofei Xia, Shuhui Chen |
ICWS | 6 |
| 2015 | A Novel Location Privacy Mining Threat in Vehicular Internet Access Service
Yipin Sun, Shuhui Chen, Biao Han 0003, Bofeng Zhang, Jinshu Su |
WASA | 2 |
| 2014 | SRC: a multicore NPU-based TCP stream reassembly card for deep packet inspectionabstractABSTRACT Stream reassembly is the premise of deep packet inspection, regarded as the core function of network intrusion detection system and network forensic system. As moving packet payload from one block of memory to another is essential for the reason of packet disorder, throughput performance is very vital in stream reassembly design. In this paper, a stream reassembly card (SRC) is designed to improve the stream reassembly throughput performance. The designed SRC adjusts the sequence of packets on the basis of the multicore network processing unit by managing and reassembling streams through an additional level of buffer. Specifically, three optimistic techniques, namely stream table dispatching, no‐locking timeout, and multichannel virtual queue, are introduced to further improve the throughput. To address the critical role of memory size in SRC, the relationship between the system throughput and memory size is analyzed. Extensive experiments demonstrate that the proposed SRC achieves more than 3 Gbps in terms of reassembly and submission throughput and triply outperforms the traditional server‐based architecture with a lower cost. Copyright © 2013 John Wiley & Sons, Ltd. Shuhui Chen, Rongxing Lu, Xuemin Shen |
Secur. Commun. Networks | 1 |
| 2013 | Study of the regularities in the treatment of psoriasis vulgaris by TCM: Applying association rule mining to TCM literatureabstractThe purpose of this study was to find out the association among herbs in specific prescriptions for psoriasis vulgaris and to evaluate the application of Association Rule Mining(ARM) in this field. Datas including herbs, prescriptions, and syndromes were extracted from 3,136 Traditional Chinese medicine (TCM) publications. An improved A priori algorithm of the association rule mining (ARM) method using WEKA 3.6.6 software was performed to mine herb pairs and groups, and rules that emerged between syndromes and herb pairs. Results revealed common herbal treatments and herb associations for psoriasis vulgaris in TCM Single herbs as well as herb combinations identified in the study should be considered when tailoring prescriptions for psoriasis based on symptom differentiation. Our work also demonstrated the practicality of improved AR in studies using a large TCM literature database. Shuhui Chen, Xiuli Xie, Zhao Zeng, Chuanjian Lu |
BIBM | 1 |