Chao Liu 0020

dblp:15/5923-20 · DBLP profile ↗
← Back
44ranked-venue papers
9as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 7 since 2021Computer networks · 11 · 3 first-author · 4 since 2021Security and privacy · 10 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 DASA-Trans-STM: Adaptive Efficient Transformer for Short Text Matching using Data Augmentation and Semantic Awareness
abstract
Rencent advancements in large language models (LLM) have shown impressive versatility across various tasks.Short text matching is one of the fundamental technologies in natural language processing.In previous studies, the common approach to applying them to Chinese is segmenting each sentence into words, and then taking these words as input.However, existing approaches have three limitations: 1) Some Chinese words are polysemous, and semantic information is not fully utilized.2) Some models suffer potential issues caused by word segmentation and incorrect recognition of negative words affects the semantic understanding of the whole sentence.3) Fuzzy negation words in ancient Chinese are difficult to recognize and match.In this work, we propose a novel adaptive Transformer for Chinese short text matching using Data Augmentation and Semantic Awareness (DASA), which can fully mine the information expressed in Chinese text to deal with word ambiguity.DASA is based on a Graph Attention Transformer Encoder that takes two word lattice graphs as input and integrates sense information from N-HowNet to moderate word ambiguity.Specially, we use an LLM to generate similar sentences for the optimal text representation.Experimental results show that the augmentation done using DASA can considerably boost the performance of our system and achieve significantly better results than previous state-of-theart methods on four available datasets, namely MNS, LCQMC, AFQMC, and BQ.
Jiguo Liu, Chao Liu 0020, Meimei Li, Shihao Gao, Dali Zhu
EMNLP2
2025 DiB-ETC: A Distillation Framework with BERT Teacher Model for Imbalanced Encrypted Traffic Classification
abstract
Encrypted Traffic Classification (ETC) plays a crucial role in managing the mobile Internet and ensuring the quality of service(QoS), especially with the explosion of mobile applications using encrypted communications. Although some existing ETC methods have shown good analysis results, they still grapple with the serious classification bias problem inherent in encrypted traffic classification methods. This issue suggests that models tend to classify the majority of classes while ignoring the minority class. To tackle these challenges, most existing models attempt to enhance their feature extraction capabilities by increasing the number of model parameters. However, in order to deploy models in real-world environments with limited resources, how to enable smaller models to achieve such capabilities remains an unresolved challenge. In this paper, we propose a novel distillation-based ETC framework called DiB-ETC. Our primary insight involves designing three training tasks for the BERT-based teacher model to aid in the deployment of strong encrypted traffic classification feature extraction capabilities. We employ distillation technology to transfer its classification capabilities to our student model, enabling it to achieve performance on par with the teacher model, even with smaller-scale parameters. DiB-ETC demonstrates strong performance on four classification tasks on three public datasets with significant class imbalances, and significantly improves the accuracy and F1 score in the two tasks of ISCX-VPN-app and ISCX-Tor. Furthermore, we validated the effectiveness of auxiliary training in the teacher model and augmentation techniques in the student model within our framework. Ablation experiments proved that these components contribute to improving the overall classification performance of the framework.
Xingchen Zhan, Meimei Li, Chao Liu 0020
HPCC4
2025 SWV: A Large-Scale Sensitive Word Variants Dataset for Semantic Text Matching
Jiguo Liu, Chao Liu 0020, Meimei Li, Shihao Gao, Dali Zhu
KSEM (1)2
2023 LADA-Trans-NER: Adaptive Efficient Transformer for Chinese Named Entity Recognition Using Lexicon-Attention and Data-Augmentation
abstract
Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the semantic relationship before and after the sentence after integrating lexical information. Therefore, the regularity of word length information has not been fully explored in various word-character fusion methods. In this work, we propose a Lexicon-Attention and Data-Augmentation (LADA) method for Chinese NER. We discuss the challenges of using existing methods in incorporating word information for NER and show how our proposed methods could be leveraged to overcome those challenges. LADA is based on a Transformer Encoder that utilizes lexicon to construct a directed graph and fuses word information through updating the optimal edge of the graph. Specially, we introduce the advanced data augmentation method to obtain the optimal representation for the NER task. Experimental results show that the augmentation done using LADA can considerably boost the performance of our NER system and achieve significantly better results than previous state-of-the-art methods and variant models in the literature on four publicly available NER datasets, namely Resume, MSRA, Weibo, and OntoNotes v4. We also observe better generalization and application to a real-world setting from LADA on multi-source complex entities.
Jiguo Liu, Chao Liu 0020, Shihao Gao, Mingqi Liu, Dali Zhu
AAAI2
2023 Intrusion Detection Method for SCADA System Based on Spatio-Temporal Characteristics
abstract
Supervisory Control and Data Acquisition (SCADA) systems are one of the most common industrial control systems (ICS). As the security threat of SCADA systems has been rising in recent years, intrusion detection has become indispensable. Among SCADA systems, there is a lack of research on intrusion detection of temporal and spatial characteristics, and the effectiveness of the existing intrusion detection approaches could be improved. This paper proposes an intrusion detection model based on spatio-temporal characteristics of SCADA systems, combining the attention mechanism, called STAM, allows a full understanding of the correlation between sensor and controller parameters. Experiments on three typical SCADA system datasets show that STAM proposed in this paper achieves state-of-the-art results and can be better applied to intrusion detection in SCADA systems. The effectiveness of STAM is evaluated by accuracy, precision, recall, and F1-score. The accuracy rates on new gas pipeline, water storage tank, and Secure Water Treatment (SWaT) datasets can reach 95.34%, 98.92% and 99.95% respectively.
Meimei Li, Zhongfeng Jin, Jiguo Liu, Chao Liu 0020
CSCWD6
2023 GHGA-Net: Global Heterogeneous Graph Attention Network for Chinese Short Text Classification
Meimei Li, Yuzhi Bao, Jiguo Liu, Chao Liu 0020, Shihao Gao
PRICAI (2)4
2023 Auto-Tuning with Reinforcement Learning for Permissioned Blockchain Systems
abstract
In a permissioned blockchain, performance dictates its development, which is substantially influenced by its parameters. However, research on auto-tuning for better performance has somewhat stagnated because of the difficulty posed by distributed parameters; thus, it is possible only with difficulty to propose an effective auto-tuning optimization scheme. To alleviate this issue, we lay a solid basis for our research by first exploring the relationship between parameters and performance in Hyperledger Fabric, a permissioned blockchain, and we propose Athena, a Fabric-based auto-tuning system that can automatically provide parameter configurations for optimal performance. The key of Athena is designing a new Permissioned Blockchain Multi-Agent Deep Deterministic Policy Gradient (PB-MADDPG) to realize heterogeneous parameter-tuning optimization of different types of nodes in Fabric. Moreover, we select parameters with the most significant impact on accelerating recommendation. In its application to Fabric, a typical permissioned blockchain system, with 12 peers and 7 orderers, Athena achieves a throughput improvement of 470.45% and a latency reduction of 75.66% over the default configuration. Compared with the most advanced tuning schemes (CDBTune, Qtune, and ResTune), our method is competitive in terms of throughput and latency.
Yazhe Wang, Shuai Ma 0001, Chao Liu 0020, Dongdong Huo, Yu Wang 0243, Zhen Xu 0009
Proc. VLDB Endow.4
2022 A highly reliable cross-domain identity authentication protocol based on blockchain in edge computing environment
abstract
Aiming at the urgent need for a cross-domain authentication scheme with low computational cost, low latency, and high security in the edge computing environment, this paper proposes a blockchain-based cross-domain authentication protocol. In this protocol, the IoT devices in the edge computing environment themselves control the generation, encrypted storage, registration, use or cancellation of their identities, and the authenticity and credibility of these identities can be guaranteed through a trusted identity checking mechanism. The public information of these identities is shared among edge nodes of different domains using the consensus mechanism of blockchain. IoT devices can perform cross-domain authentication at the edge. The authentication process not only makes use of the advantages of low latency of edge computing, but also minimizes the disclosure of detailed identity information. Security analysis shows that the protocol can resist common attacks such as DDoS Attacks and Impersonation Attacks. Compared with several previous cross-domain authentication schemes, our protocol is more secure. Besides, this protocol has the advantage of low computational cost, suitable for low-performance IoT devices.
Penghui Lv, Yu Wang 0243, Yazhe Wang, Chao Liu 0020, Qihui Zhou, Zhen Xu 0009
CSCWD4
2021 CallChain: Identity Authentication Based on Blockchain for Telephony Networks
abstract
Telephony networks serve as a reliable communication channel, and phone numbers are usually used to authenticate users' identities in telephony networks. However, modern telephony networks does not provide end-to-end authentication for phone numbers, making it difficult for a receiver to confirm a caller's real identity. Moreover, due to the lack of effective authentication solutions for users, phone numbers are usually exploited to commit identity theft, credit card fraud, or telecom harassment by phone number spoofing attacks. In this paper, we propose a blockchain-based identity authentication system for telephony networks, CallChain. CallChain issues verifiable digital identities to users which bound a user to his/her phone number. Based on the digital identity, CallChain provides end-to-end identity authentication between the caller and receiver in the data channel without modifying the core technologies in telephony networks. In the proposed system, users first register decentralized identities (DIDs) and store them in the blockchain distributed ledger. Then Call Center issues verifiable PhoneNumber Credentials to users. And finally users conduct end-to-end identity authentication based on these PhoneNumber Credentials before the call is established. We implement a prototype as an App for Android-based phones based on an open source blockchain platform Hyperledger Indy. Moreover, we detect that CallChain only adds a worst-case 2.1 seconds call establishment overhead and can resist DDoS attacks effectively.
Yazhe Wang, Yu Wang 0243, Guochao Dong, Chao Liu 0020
CSCWD6
2021 A Novel Trojan Attack against Co-learning Based ASR DNN System
abstract
ASR (Automatic Speech Recognition) technology is a key technology for human-computer interaction. Especially the DNN models of wake-up-word speech recognition, which enables the smart device to recognize wake-up words spoken by users when they are in the sleep or lock screen state, allowing the device to directly enter the wait command state, and start the first step of voice interaction. ASR technology no wonder provides great convenience for people's daily life, but it's security problem has always been a hot topic of further research. The rapid development of ASR models has made these models very vulnerable which seriously affect their performance in real scenarios. This paper proposes a new backdoor attack method TNN (Trojan Neural Network) for deep learning models in voice wake-up scenarios. The attacker leaves backdoor in the deep learning model of the ASR. Besides specific wakeup words, attackers can also use other words to force the device to wake up and get the IoT smart device's highest privileges. And when the smart device is in the awake state, any audio containing a particular vocabulary will be recognized as a specific command and executed. In this papaer, we propose a novel method to attack ASR model, and without affecting the performance of clean samples, by applying the attack method, the success rate in compulsory recognition of specific words can be up to 100%. The experimental results prove that our attack method is very effective and poses a great security problem of the voice-interactive IoT device.
Dongdong Huo, Chao Liu 0020, Yazhe Wang, Yu Wang 0243, Zhen Xu 0009
CSCWD5
2021 Achieve Better Endorsement Balance on Blockchain Systems
abstract
At the beginning of the transaction flow in blockchain system Hyperedger Fabric, client have to select endorsement peer nodes to send proposals. Fabric provides two ways for client: static configuration and random selection using service discovery. However, when a large number of transaction requests arrive, the assignment of tasks to endorsement peer nodes may become unfair. The result of the above situation may lead to the high utilization of computing resources and the working pressure of Fabric nodes, which will eventually affect the transaction performance and throughput. In this paper, we propose a calculation method to accurately measure peer node endorsement load and a dynamic stochastic load minimization algorithm, DSLM. Using DSLM to dynamically determine endorsed peer nodes based on real-time work load will make the distribute of endorsement tasks more balanced. In addition, we implement our approach on Hyperledger Fabric v1.4 LTS and carry out a series of evaluation experiments. By observing the experimental results, the proposed method can effectively balance and reduce the overall workload, and improve the performance of the blockchain system in the case of high concurrency.
Chao Liu 0020, Yazhe Wang, Yu Wang 0243, Dongdong Huo
CSCWD1
2021 Detecting Malicious PDF Documents Using Semi-Supervised Machine Learning
Nan Song, Min Yu 0001, Kam-Pui Chow, Gang Li 0009, Chao Liu 0020, Weiqing Huang
IFIP Int. Conf. Digital Forensics6
2021 FUNC-ESIM: A Dual Pairwise Attention Network for Cross-version Binary Function Matching
abstract
Binary function matching compares two pieces of binary functions to identify their similarities, which has wide applications in the field of malware origin tracing, vulnerability searching, binary level plagiarism detection, etc. Up-to-date methods commonly independently map each function to an embedding and rarely consider fine-grained pairwise semantic similarity, which influences the accuracy of matching. Moreover, few methods are available to detect similarities between versions spanning a long period for cross-version vulnerability detection or patch positioning. To solve these issues, we propose a novel binary function matching method, which takes a pair of binary functions as input, and then computes a similarity score jointly on the pair through a specifical dual pairwise cross-attention network. Specially, we apply our method to detecting similarities between cross-version binaries. The experimental analysis demonstrates that FUNC-ESIM achieves promising results on the cross-version binary matching task, where the average recall@1 reaches 85.98%.
Degang Sun, Yunting Guo, Min Yu 0001, Gang Li 0009, Chao Liu 0020, Weiqing Huang
IJCNN5
2021 A Self-Sovereign Decentralized Identity Platform Based on Blockchain
abstract
Traditional accounts and passwords mechanism usually needs to register multiple accounts for adopting different scenarios, which makes it difficult to manage passwords and handle privacy leakage issues. On the other hand, users have no real control over their identities under centralized trusted authorities or insecure third-party operators. In this paper, we propose a self-sovereign decentralized identity platform called SSIChain. SSIChain uses distributed infrastructure blockchain to replace authorities and third-party operators. In the proposed platform, the DID and a credential represent an user's digital identity. We define DID chaincode and credential protocol to cover the full lifecycle of a user's digital identity. We build a complete consortium blockchain and application environment, and implement a prototype as an App for Android-based phones. Moreover, we detect that SSIChain only takes at most 2.1 seconds for authentication with high TPS, which can meet users' demands for identity management well.
Chao Liu 0020, Yu Wang 0243, Yazhe Wang
ISCC2
2021 UTANSA: Static Approach for Multi-Language Malicious Web Scripts Detection
abstract
In order to detect malicious web scripts automatically, many detection methods using static features and machine learning are proposed. However, the existing detection methods can only detect web scripts of specific programming languages. This paper proposes the unified text features and abstract syntax tree(AST) node sequence features algorithm(UTANSA) that exploits the text feature classification method and AST node classification method, together with the corresponding unified method to enhance the generalization ability of the model. Through the algorithm, two unified approaches are proposed based on text features and AST node features respectively, so that the detection model can detect multi-language web scripts. We choose scripts written in the JavaScript(JS) and PHP languages for experimentation to evaluate our approach. The results show that the detection model trained with the proposed method has a similar detection effect as trained with only JS samples or PHP samples.
Weiqing Huang, Chenggang Jia, Min Yu 0001, Gang Li 0009, Chao Liu 0020
ISCC5
2021 Adaptive Smooth L1 Loss: A Better Way to Regress Scene Texts with Extreme Aspect Ratios
abstract
In recent years, scene text detection has experienced rapid development. Regression-based methods are currently a mainstream method for scene text detection, and the effect of bounding box regression is a major factor limiting their detection performance. The regression of bounding boxes is greatly affected by the aspect ratio of texts since the text in natural scenes varies greatly in height and width. However, the existing methods ignore the difference between the height and width of the text in the bounding box regression, which leads to an imperfect regression effect and thus suppresses the performance of the scene text detection. In this paper, we propose an Adaptive Smooth L1 Loss function (abbreviated as ASLL) for bounding box regression, which can adaptively determine the weight of each regression variable according to the current state of the model during the training process, so as to guide the bounding box to regress in a more critical direction. The experimental results demonstrate that ASLL achieves promising performance on scene text detection. Specially, an F-measure of 84.56% is achieved on CTW-1500 dataset, surpassing the state-of-the-art detectors, and the detection results on TotalText and ICDAR2015 datasets are competitive to those of state-of-the-art methods.
Chao Liu 0020, Min Yu 0001, Baole Wei, Boquan Li 0002, Gang Li 0009, Weiqing Huang
ISCC1
2021 Landscape-Enhanced Graph Attention Network for Rumor Detection
Min Yu 0001, Gang Li 0009, Mingqi Liu, Chao Liu 0020, Weiqing Huang
KSEM6
2021 Aspect and Opinion Terms Co-extraction Using Position-Aware Attention and Auxiliary Labels
Chao Liu 0020, Xintong Wei, Min Yu 0001, Gang Li 0009, Xiangmei Ma, Weiqing Huang
KSEM1
2021 NFDD: A Dynamic Malicious Document Detection Method Without Manual Feature Dictionary
Chenghao Wang 0010, Min Yu 0001, Chenggang Jia, Gang Li 0009, Chao Liu 0020, Weiqing Huang
WASA (2)6
2021 An end-to-end text spotter with text relation networks
abstract
Abstract Reading text in images automatically has become an attractive research topic in computer vision. Specifically, end-to-end spotting of scene text has attracted significant research attention, and relatively ideal accuracy has been achieved on several datasets. However, most of the existing works overlooked the semantic connection between the scene text instances, and had limitations in situations such as occlusion, blurring, and unseen characters, which result in some semantic information lost in the text regions. The relevance between texts generally lies in the scene images. From the perspective of cognitive psychology, humans often combine the nearby easy-to-recognize texts to infer the unidentifiable text. In this paper, we propose a novel graph-based method for intermediate semantic features enhancement, called Text Relation Networks. Specifically, we model the co-occurrence relationship of scene texts as a graph. The nodes in the graph represent the text instances in a scene image, and the corresponding semantic features are defined as representations of the nodes. The relative positions between text instances are measured as the weights of edges in the established graph. Then, a convolution operation is performed on the graph to aggregate semantic information and enhance the intermediate features corresponding to text instances. We evaluate the proposed method through comprehensive experiments on several mainstream benchmarks, and get highly competitive results. For example, on the , our method surpasses the previous top works by 2.1% on the word spotting task.
Baole Wei, Min Yu 0001, Gang Li 0009, Boquan Li 0002, Chao Liu 0020, Weiqing Huang
Cybersecur.6
2021 FakeFilter: A cross-distribution Deepfake detection system with domain adaptation
abstract
Abuse of face swap techniques poses serious threats to the integrity and authenticity of digital visual media. More alarmingly, fake images or videos created by deep learning technologies, also known as Deepfakes, are more realistic, high-quality, and reveal few tampering traces, which attracts great attention in digital multimedia forensics research. To address those threats imposed by Deepfakes, previous work attempted to classify real and fake faces by discriminative visual features, which is subjected to various objective conditions such as the angle or posture of a face. Differently, some research devises deep neural networks to discriminate Deepfakes at the microscopic-level semantics of images, which achieves promising results. Nevertheless, such methods show limited success as encountering unseen Deepfakes created with different methods from the training sets. Therefore, we propose a novel Deepfake detection system, named FakeFilter, in which we formulate the challenge of unseen Deepfake detection into a problem of cross-distribution data classification, and address the issue with a strategy of domain adaptation. By mapping different distributions of Deepfakes into similar features in a certain space, the detection system achieves comparable performance on both seen and unseen Deepfakes. Further evaluation and comparison results indicate that the challenge has been successfully addressed by FakeFilter.
Boquan Li 0002, Baole Wei, Gang Li 0009, Chao Liu 0020, Weiqing Huang, Meimei Li, Min Yu 0001
J. Comput. Secur.5
2020 CIDetector: Semi-Supervised Method for Multi-Topic Confidential Information Detection
abstract
Confidential information firewalling with text classifier is to identify the text containing confidential information whose publication might be harmful to national security, business trade, or personal life. Traditional methods, e.g., listing a set of suspicious keywords together with regular-expression based filter, fail to solve the multi-topic phenomenon, i.e., one text containing the confidential information with different topics. In this paper, we propose a semi-supervised method, CIDetector, for multi-topic confidential information detection. We introduce coarse confidential polarity as prior knowledge into word embeddings, which can regularize the distribution of words to have a clear task classification boundary. Then we introduce a multi-attention network classifier to extract task-related features and model dependencies between features for multi-topic classification. Experiments are conducted by real-world data from WikiLeaks and demonstrated the superiority of our proposed method.
Min Yu 0001, Yantao Jia, Jiafeng Guo, Chao Liu 0020, Weiqing Huang
ECAI6
2020 Similarity of Binaries Across Optimization Levels and Obfuscation
Gengwang Li, Min Yu 0001, Gang Li 0009, Chao Liu 0020, Zhiqiang Lv, Weiqing Huang
ESORICS (1)5
2020 A Machine Learning-Assisted Compartmentalization Scheme for Bare-Metal Systems
Dongdong Huo, Chao Liu 0020, Yu Wang 0243, Yazhe Wang, Peng Liu 0005, Zhen Xu 0009
ICICS2
2020 Enhancing the Feature Profiles of Web Shells by Analyzing the Performance of Multiple Detectors
Weiqing Huang, Chenggang Jia, Min Yu 0001, Kam-Pui Chow, Jiuming Chen, Chao Liu 0020
IFIP Int. Conf. Digital Forensics6
2020 CES2Vec: A Confidentiality-Oriented Word Embedding for Confidential Information Detection
abstract
Confidential information firewalling with text classifiers is to recognize the text containing confidential information whose publication might pose a threat to national security, business trade, or personal life. Word embedding is a component of the detector and plays an important role. Existing word embeddings, e.g., Word2Vec, fail to learn a clear task classification boundary, i.e., the confidential polarities of words are opposite but the embedding vectors of the words are close to each other. We propose a confidentiality-oriented word embedding, CES2Vec, for confidential information detection. We embed confidentiality into semantics to catch both of them together, which can learn the word embedding with a clear task classification boundary. We use real-world data from WikiLeaks and conduct the comparison experiments of our CES2Vec and popular methods. The experimental results show that our proposed method is better than the previously reported methods in detecting confidential information.
Min Yu 0001, Gang Li 0009, Chao Liu 0020, Shaohua An, Weiqing Huang
ISCC5
2020 SCX-SD: Semi-supervised Method for Contextual Sarcasm Detection
Meimei Li, Chen Lang, Min Yu 0001, Chao Liu 0020, Weiqing Huang
KSEM (2)5
2020 A Robust Representation with Pre-trained Start and End Characters Vectors for Noisy Word Recognition
Chao Liu 0020, Xiangmei Ma, Min Yu 0001, Xinghua Wu, Mingqi Liu, Weiqing Huang
KSEM (1)1
2020 Depthwise Separable Convolutional Neural Network for Confidential Information Analysis
Min Yu 0001, Chao Liu 0020, Chaochao Liu, Weiqing Huang, Zhiqiang Lv
KSEM (2)4
2019 Restoration as a Defense Against Adversarial Perturbations for Spam Image Detection
Boquan Li 0002, Min Yu 0001, Chao Liu 0020, Weiqing Huang, Lejun Fan, Jianfeng Xia
ICANN (3)4
2019 EASVD: A Modified Method to Enhance the Authentication for SPICE Virtual Desktop
abstract
Virtual Desktop (VD) provides a convenient way for users to operate their virtual machines on different types of terminals. SPICE is an excellent protocol for the VD display. However, there are still several threats when using the original SPICE in the high-level scenarios. The threats are mainly due to the fact that the original SPICE ignores the confidentiality of the VD. The threats contain the data interaction like copying, falsifying, and deleting between VDs & different physical terminals. Furthermore, the original SPICE authentication is a one-factor method based on the IP address and password of the user. We bind each VD to a specific authorized terminal to reduce the threat posed by the data interaction between VDs and physical terminals in the high-level security scenarios proposed in our paper. At the same time, we propose a two-factor authentication method which includes physical terminals authentication and user-password authentication. We also reinforce the security of the VD by putting forward a modified way to storage the password. Our approach to intensify the VDs authentication has been implemented in terminals based on SPICE and Libvirt. Our method can effectively prevent the unauthorized users from illegal logging into the VD on the illegal physical terminals, even when the correct password is intercepted by unauthorized users. We demonstrate the effectiveness of our method by experimental results.
Chao Liu 0020, Xinling Shen
ICPADS1
2019 Android Malware Family Classification Based on Sensitive Opcode Sequence
abstract
Android malware family classification is an advanced task in Android malware analysis, detection and forensics. Existing methods and models have achieved a certain success for Android malware detection, but the accuracy and the efficiency are still not up to the expectation, especially in the context of multiple class classification with imbalanced training data. To address those challenges, we propose an Android malware family classification model by analyzing the code's specific semantic information based on sensitive opcode sequence. In this work, we construct a sensitive semantic feature-sensitive opcode sequence using opcodes, sensitive APIs, STRs and actions, and propose to analyze the code's specific semantic information, generate a semantic related vector for Android malware family classification based on this feature. Besides, aiming at the families with minority, we adopt an oversampling technique based on the sensitive opcode sequence. Finally, we evaluate our method on Drebin dataset, and select the top 40 malware families for experiments. The experimental results show that the Total Accuracy and Average AUC (Area Under Curve, AUC) reach 99.50% and 98.86% with 45. 17s per Android malware, and even if the number of malware families increases, these results remain good.
Min Yu 0001, Gang Li 0009, Chao Liu 0020, Weiqing Huang
ISCC5
2019 A Two-Stage Model Based on BERT for Short Fake News Detection
Chao Liu 0020, Xinghua Wu, Min Yu 0001, Gang Li 0009, Weiqing Huang
KSEM (2)1
2019 Machine Tools Fingerprinting for Distributed Numerical Control Systems
abstract
As machine tools are connected to Industrial Ethernet and external interfaces in the wave of the fourth industrial revolution, new attacks and vulnerabilities are emerging. However, there is little security analysis on Distributed Numerical Control (DNC) system and Computerized Numerical Control (CNC) system. Researchers have demonstrated how to combine the characteristics of Industrial Control System (ICS) to augment existing Intrusion Detection System (IDS) solutions. To the best of our knowledge, there is no such work on DNC network. In response to this situation, a fingerprinting method is proposed as an enhancement technology to existing IDS for DNC systems. The first step is to extract the number of data collection points of each machine tool and the length of TCP payload of each packet. And the second step is to use data response processing times of machine tools to construct unique fingerprint for each machine. Finally, the optimum period slice k is selected and classification accuracy is evaluated using a real-world dataset from a small-scale smart factory. It is demonstrated that our fingerprinting method can be a valuable tool to enhance IDS for DNC network.
Weiqing Huang, Zhongfeng Jin, Chao Liu 0020, Meimei Li
LCN4
2019 Malicious documents detection for business process management based on multi-layer abstract model
Min Yu 0001, Gang Li 0009, Chenzhe Lou, Yunzheng Liu, Chao Liu 0020, Weiqing Huang
Future Gener. Comput. Syst.6
2018 URefFlow: A Unified Android Malware Detection Model Based on Reflective Calls
abstract
In Android malware detection, sensitive data-flows provide more accurate information on the application's behavior than regular features such as signatures and permissions. Currently, Android static taint analysis is widely adopted to identify sensitive data-flows because of its high code coverage and low false negative rate. However, existing static taint analysis tools cannot effectively analyze applications that adopt Android reflection mechanism. Reflection mechanism can block the control-flows and data-flows of the application. When constructing a call graph, the call information will point directly to the system's reflection processing method, rather than the actual method invoked by the application. This significantly affects the accurate representation of the application's behavior. To address this issue, this paper proposes a unified Android malware detection model based on reflective calls named URefFlow, in which the reflective call statement is replaced by the non-reflective call statement to make the reflective calls explicit by combining the parameters of the reflective calls into standard function calls. After extracting the complete sensitive data-flows with reflective calls from an application, we analyze the characteristics of these data-flows to determine whether the application is malicious. Evaluation results on thousands of applications show that URefFlow can achieve an impressive detection accuracy of 95.6% with a false positive rate of 0.8%. In addition, the proposed approach complements well with existing static stain analysis techniques.
Chao Liu 0020, Min Yu 0001, Gang Li 0009, Bo Luo, Weiqing Huang
IPCCC1
2018 MPP: A Join-dividing Method for Multi-table Privacy Preservation
abstract
In regard to relational databases, studies in this area typically focus on individual privacy leakage in one table. However, in reality, a database usually has many tables, some of them contain correlation information about individual, which can provide additional implication as background knowledge to attacker. In this paper, we innovatively propose a new method named MPP (Multi-table Privacy Preservation) which combines Lossy-join with Bucketization to enhance the individual privacy in database. We consider the privacy disclosure problem from the global sight of the entire dataset instead of a table. Based on this method, we not only solve the correlation information leakage by other tables, but also improve the data utility. Extensive experiments on 32.8GB real-world Express data demonstrate the effectiveness and efficiency of our approach in terms of data utility and computational cost.
Weiqing Huang, Jianfeng Xia, Min Yu 0001, Chao Liu 0020
ISCC4
2018 PSDEM: A Feasible De-Obfuscation Method for Malicious PowerShell Detection
abstract
PowerShell is so extremely powerful that we have seen that attackers are increasingly using PowerShell in their attack methods lately. In most cases, PowerShell malware arrives via spam email, using a combination of Microsoft Word documents to infect victims with its deadly payload. Nowadays, the de-obfuscation and analysis of PowerShell are still based on the manual analysis. However, as the number of malicious samples and obfuscation methods growing quickly, it is so slow that can't satisfy the demand. In this paper, we propose a de-obfuscation method of PowerShell called PSDEM which has two layers de-obfuscation to get original PowerShell scripts. One is extracting PowerShell scripts from much obfuscated document code. The other is de-obfuscating scripts including encoding, strings manipulation and code logic obfuscation. Meanwhile, we design an automatic de-obfuscation and analysis tool for malicious PowerShell scripts in Word documents based on PSDEM. We test the performance of the tool from the accuracy of de-obfuscation and the efficiency of time, and evaluation results show that it has a satisfactory performance. PSDEM improves the efficiency and accuracy rate for analyzing malicious PowerShell Scripts in Word documents, as well as provides a path in which further analysis for security experts to get more information about attacks.
Chao Liu 0020, Min Yu 0001, Yunzheng Liu
ISCC1
2018 Sentiment Embedded Semantic Space for More Accurate Sentiment Analysis
Min Yu 0001, Gang Li 0009, Chao Liu 0020, Weiqing Huang, Fangtao Zhang
KSEM (2)5
2018 MRDroid: A Multi-act Classification Model for Android Malware Risk Assessment
abstract
Risk Score (RS) on Android is aiming at offering measurement to users for evaluating the apps' trustworthiness. Much work has been done to assess Android app's risk, but few jobs use various assessment systems to analyze Android apps with various malicious acts. However, it is hard for a single system to analyze those multiple categories Android apps. To overcome such limitations, we propose a multi-act classification model MRDroid for Android malware risk assessment in this paper, which presorts an app to one category, then uses the most suitable subsystem corresponding to that category to analyze the app for giving a RS. Base on this model, we implement an Android malware risk assessment system utilizing a machine learning solution with k-means algorithm for clustering benign and malware samples to various categories and the supervised algorithms for generating specific subsystems. It can be also used for Android malware detection under the condition of human confirmation. Experiments show that MRDroid provides high detection precision and offers stable and reliable risk assessment. Though testing our system using the dataset different from the system used, the result indicates it is also effective in detecting some unknown samples.
Min Yu 0001, Chao Liu 0020, Weiqing Huang, Gang Li 0009
MASS5
2018 FGFDect: A Fine-Grained Features Classification Model for Android Malware Detection
Chao Liu 0020, Min Yu 0001, Bo Luo, Weiqing Huang
SecureComm (1)1
2017 A Visualization Scheme for Network Forensics Based on Attribute Oriented Induction Based Frequent Item Mining and Hyper Graph
Jiuming Chen, Kim-Kwang Raymond Choo, Chao Liu 0020, Kunying Liu, Min Yu 0001
ICDF2C4
2017 A Deep Learning Based Online Malicious URL and DNS Detection Scheme
Jiuming Chen, Kim-Kwang Raymond Choo, Chao Liu 0020, Kunying Liu, Min Yu 0001
SecureComm4
2017 A New Digital Watermarking Method for Data Integrity Protection in the Perception Layer of IoT
abstract
Since its introduction, IoT (Internet of Things) has enjoyed vigorous support from governments and research institutions around the world, and remarkable achievements have been obtained. The perception layer of IoT plays an important role as a link between the IoT and the real world; the security has become a bottleneck restricting the further development of IoT. The perception layer is a self-organizing network system consisting of various resource-constrained sensor nodes through wireless communication. Accordingly, the costly encryption mechanism cannot be applied to the perception layer. In this paper, a novel lightweight data integrity protection scheme based on fragile watermark is proposed to solve the contradiction between the security and restricted resource of perception layer. To improve the security, we design a position random watermark (PRW) strategy to calculate the embedding position by temporal dynamics of sensing data. The digital watermark is generated by one-way hash function SHA-1 before embedding to the dynamic computed position. In this way, the security vulnerabilities introduced by fixed embedding position can not only be solved effectively, but also achieve zero disturbance to the data. The security analysis and simulation results show that the proposed scheme can effectively ensure the integrity of the data at low cost.
Guoyin Zhang, Liang Kou, Chao Liu 0020, Qingan Da
Secur. Commun. Networks4