VLDB 2026 Research / reviewers in the wild / expert
Li Guo 0001
dblp:02/929-1
· DBLP profile ↗
152ranked-venue papers
0as first author
20since 2021 · last 2026
0000-0002-2529-7643ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 8 since 2021Databases, data management, data science and information retrieval · 44 · 6 since 2021Computer networks · 23 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 since 2021Security and privacy · 16 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 since 2021Systems, architecture and hardware · 8Human-computer interaction and ubiquitous computing · 2Theory of computation · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HyperMem: Hypergraph Memory for Long-Term ConversationsabstractJuwei Yue, Chuanrui Hu, Jiawei Sheng, Zuyi Zhou, Wenyuan Zhang, Tingwen Liu, Li Guo, Yafeng Deng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Juwei Yue, Chuanrui Hu, Jiawei Sheng, Zuyi Zhou, Wenyuan Zhang 0002, Tingwen Liu, Li Guo 0001, Yafeng Deng |
ACL (1) | 7 |
| 2026 | Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text DetectionabstractThe rapid advancement of large language models has increasingly blurred the boundary between human-written and AI-generated text, raising societal risks such as misinformation dissemination, authorship ambiguity, and threats to intellectual property rights.These concerns highlight the urgent need for effective and reliable detection methods.While existing training-free approaches often achieve strong performance by aggregating token-level signals into a global score, they typically assume uniform token contributions, making them less robust under short sequences or localized token modifications.To address these limitations, we propose Exons-Detect, a training-free method for AI-generated text detection based on an exon-aware token reweighting perspective.Exons-Detect identifies and amplifies informative exonic tokens by measuring hiddenstate discrepancy under a dual-model setting, and computes an interpretable translation score from the resulting importance-weighted token sequence.Empirical evaluations demonstrate that Exons-Detect achieves state-of-the-art detection performance and exhibits strong robustness to adversarial attacks and varying input lengths.In particular, it attains a 2.2% relative improvement in average AUROC over the strongest prior baseline on DetectRL.Code and data are available at https://github.com/ Xiaoweizhu57/Exons-Detect. Yubing Ren, Fang Fang 0009, Shi Wang 0002, Yanan Cao 0001, Li Guo 0001 |
ACL (1) | 6 |
| 2026 | Multi-modal prompt codebook learning: Achieving adaptive and generalizable prompting for CLIP-based visual recognition
Geyuan Zhang, Xiaofei Zhou 0002, Gaopeng Gou, Gang Xiong 0001, Li Guo 0001 |
Inf. Sci. | 5 |
| 2026 | BAPTISM: A Robust Framework for Encrypted Malicious Traffic Identification With Low-Quality Training DataabstractMachine learning (ML) is highly effective for accurate encrypted malicious traffic identification by using highquality training data. In fact, obtaining such data is costly and challenging. As a result, many ML-based models are inevitably trained on low-quality data and perform poorly. To enhance performance, some methods utilize various sample selection techniques to choose confident samples for model training. However, they often rely on a single metric for this selection, which restricts their adaptability across diverse datasets and noise conditions. In this paper, we propose a robust framework BAPTISM for identifying encrypted malicious traffic with low-quality training data. Particularly, BAPTISM selects a suitable base model for each task, and trains it with early stopping to generate traffic representation before overfitting occurs. Then, we devise an adaptive metric selection strategy to select confident samples. By employing two metrics (JSD and CSD) to assess the characteristic of traffic representation from distinct perspective, we find the more proper metric for each class and apply it for confident sample selection. According to the confident samples and selected metric for each class, we develop a label correction tactic which adapts to class nature to improve the quality of training data. Finally, we employ parallel training strategy to train the base model with the corrected data, further mitigating the impact of low-quality data. We conduct experiments across three real-world malicious traffic datasets with various noise settings. The results demonstrate that BAPTISM is compatible with different base models and outperforms across noise ratios ranging from 20% to 90%. Meanwhile, BAPTISM consistently selects the confident samples with the highest purity and volume under each setting. Chang Liu 0049, Gang Xiong 0001, Gaopeng Gou, Zhen Li 0011, Junzheng Shi, Li Guo 0001, Binxing Fang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2026 | D2TCDR: Disentangled Diffusion-Based Transfer for Cross-Domain RecommendationabstractCross-Domain Recommendation (CDR) aims to alleviate data sparsity in the target domain by incorporating knowledge from external domains. Existing approaches typically rely on overlapping users between the source and target domains as a bridge for knowledge transfer. However, in practice, user information across domains is often unavailable due to privacy protection, platform isolation, and data sharing restrictions, rendering most methods ineffective. In this article, we propose the D2TCDR, a two-stage generative CDR framework to address this critical limitation. By modeling the domain-level distribution that captures user preferences shared across domains, we extract transferable knowledge and guide its transfer through a generative process, reducing reliance on overlapping users and alleviating data sparsity in the target domain. D2TCDR first proposes a domain disentanglement module to extract the domain-invariant representations, capturing shared preferences across domains by eliminating domain-specific interference. Subsequently, a guided diffusion model is designed to model the domain-level distribution of these domain-invariant representations. By injecting target-domain signals into the guided diffusion model, we further steer the learned distribution toward the target domain, achieving knowledge transfer without relying on overlapping users. Extensive experiments on multiple cross-domain datasets show the superior performance of D2TCDR, validating its recommendation capabilities in complex transfer scenarios. Code is available at: https://github.com/Red-Week/D2TCDR . Xixun Lin, Yanan Cao 0001, Renqi Jia, Xiangyu Zhao 0001, Guandong Xu, Li Guo 0001 |
ACM Trans. Inf. Syst. | 8 |
| 2025 | Unraveling DoH Traces: Padding-Resilient Website Fingerprinting via HTTP/2 Key Frame Sequences
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Li Guo 0001 |
ESORICS (3) | 5 |
| 2025 | Hyperbolic-PDE GNN: Spectral Graph Neural Networks in the Perspective of A System of Hyperbolic Partial Differential EquationsabstractGraph neural networks (GNNs) leverage message passing mechanisms to learn the topological features of graph data. Traditional GNNs learns node features in a spatial domain unrelated to the topology, which can hardly ensure topological features. In this paper, we formulates message passing as a system of hyperbolic partial differential equations (hyperbolic PDEs), constituting a dynamical system that explicitly maps node representations into a particular solution space. This solution space is spanned by a set of eigenvectors describing the topological structure of graphs. Within this system, for any moment in time, a node features can be decomposed into a superposition of the basis of eigenvectors. This not only enhances the interpretability of message passing but also enables the explicit extraction of fundamental characteristics about the topological structure. Furthermore, by solving this system of hyperbolic partial differential equations, we establish a connection with spectral graph neural networks (spectral GNNs), serving as a message passing enhancement paradigm for spectral GNNs.We further introduce polynomials to approximate arbitrary filter functions. Extensive experiments demonstrate that the paradigm of hyperbolic PDEs not only exhibits strong flexibility but also significantly enhances the performance of various spectral GNNs across diverse graph tasks. Juwei Yue, Haikuo Li, Jiawei Sheng, Xiaodong Li 0012, Taoyu Su, Tingwen Liu, Li Guo 0001 |
ICML | 7 |
| 2025 | 6RIS: IPv6 Address Correlation Attacks on TLS Encrypted Traffic Using Joint Representation of Interaction and Sequential BehaviorabstractIPv6 address correlation attacks determine whether two temporary addresses belong to the same user, compromising user privacy. Particularly, existing works have shown that methods based on TLS traffic analysis can be used to perform correlation attacks. However, they suffer from inaccurate differentiation of complex user behaviors and low correlation efficiency, leading to limitations in practical applications. In this paper, we propose a 6RIS model to improve IPv6 address correlation attacks on TLS-encrypted traffic. 6RIS learns the joint representation of interaction and sequential behavior from traffic, which is used to construct a KD-Tree for efficient correlation. Statistical aggregation and semantic preference modules are designed to extract generalized features from complex interaction behavior. To model sequential behavior, we utilize a sequence learning module to capture service dependencies, enhancing behavior representation. Experiments on a real-world IPv6 dataset show that 6RIS ($\mathbf{9 1. 8 6 \%}$TPR,$\mathbf{0. 8 3 \%}$FPR) outperforms state-of-theart methods. The correlation efficiency of 6RIS improves by at least 57 % compared to existing methods. Additionally, we further confirm through 6RIS that persistent session IDs in TLS session resumption can directly expose IPv6 temporary addresses to correlation attacks. Yang Li 0002, Chang Liu 0049, Gaopeng Gou, Tianyu Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
IWQoS | 8 |
| 2025 | Graph Wave NetworksabstractDynamics modeling has been introduced as a novel paradigm in message passing (MP) of graph neural networks (GNNs). Existing methods consider MP between nodes as a heat diffusion process, and leverage heat equation to model the temporal evolution of nodes in the embedding space. However, heat equation can hardly depict the wave nature of graph signals in graph signal processing. Besides, heat equation is essentially a partial differential equation (PDE) involving a first partial derivative of time, whose numerical solution usually has low stability, and leads to inefficient model training. In this paper, we would like to depict more wave details in MP, since graph signals are essentially wave signals that can be seen as a superposition of a series of waves in the form of eigenvector. This motivates us to consider MP as a wave propagation process to capture the temporal evolution of wave signals in the space. Based on wave equation in physics, we innovatively develop a graph wave equation to leverage the wave propagation on graphs. In details, we demonstrate that the graph wave equation can be connected to traditional spectral GNNs, facilitating the design of graph wave networks (GWNs) based on various Laplacians and enhancing the performance of the spectral GNNs. Besides, the graph wave equation is particularly a PDE involving a second partial derivative of time, which has stronger stability on graphs than the heat equation that involves a first partial derivative of time. Additionally, we theoretically prove that the numerical solution derived from the graph wave equation are constantly stable, enabling to significantly enhance model efficiency while ensuring its performance. Extensive experiments show that GWNs achieve state-of-the-art and efficient performance on benchmark datasets, and exhibit outstanding performance in addressing challenging graph problems, such as over-smoothing and heterophily. Our code is available at https://github.com/YueAWu/Graph-Wave-Networks. Juwei Yue, Haikuo Li, Jiawei Sheng, Xinghua Zhang 0001, Chuan Zhou 0001, Tingwen Liu, Li Guo 0001 |
WWW | 8 |
| 2024 | From Fingerprint to Footprint: Characterizing the Dependencies in Encrypted DNS Infrastructures
Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Li Guo 0001 |
ESORICS (2) | 7 |
| 2024 | Anti-Packet-Loss Encrypted Traffic Classification via Masked Autoencoder
Li Guo 0001, Gaopeng Gou, Gang Xiong 0001, Yangyang Guan |
WASA (1) | 2 |
| 2023 | Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State TrackingabstractQingyue Wang, Liang Ding, Yanan Cao, Yibing Zhan, Zheng Lin, Shi Wang, Dacheng Tao, Li Guo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qingyue Wang, Liang Ding 0006, Yanan Cao 0001, Yibing Zhan, Zheng Lin 0001, Shi Wang 0002, Dacheng Tao, Li Guo 0001 |
ACL (1) | 8 |
| 2023 | Confident Slot Iterative Learning for Multi-Domain Dialogue State Tracking
Qingyue Wang, Yanan Cao 0001, Piji Li, Yanhe Fu, Zheng Lin 0001, Cong Cao 0001, Shi Wang 0002, Li Guo 0001 |
CogSci | 8 |
| 2023 | Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsabstractExtracting structured information from all manner of webpages is an important problem with the potential to automate many real-world applications. Recent work has shown the effectiveness of leveraging DOM trees and pre-trained language models to describe and encode webpages. However, they typically optimize the model to learn the semantic co-occurrence of elements and labels in the same webpage, thus their effectiveness depends on sufficient labeled data, which is labor-intensive. In this paper, we further observe structural co-occurrences in different webpages of the same website: the same position in the DOM tree usually plays the same semantic role, and the DOM nodes in this position also share similar surface forms. Motivated by this, we propose a novel method, Structor, to effectively incorporate the structural co-occurrences over DOM tree and surface form into pre-trained language models. Such structural co-occurrences help the model learn the task better under low-resource settings, and we study two challenging experimental scenarios: website-level low-resource setting and webpage-level low-resource setting, to evaluate our approach. Extensive experiments on the public SWDE dataset show that Structor significantly outperforms the state-of-the-art models in both settings, and even achieves three times the performance of the strong baseline model in the case of extreme lack of training data. Zhenyu Zhang 0006, Bowen Yu 0002, Tingwen Liu, Tianyun Liu, Li Guo 0001 |
WWW | 6 |
| 2022 | Slot Dependency Modeling for Zero-Shot Cross-Domain Dialogue State Tracking
Qingyue Wang, Yanan Cao 0001, Piji Li, Yanhe Fu, Zheng Lin 0001, Li Guo 0001 |
COLING | 6 |
| 2022 | Knowledge Graph Embedding by Double Limit Scoring LossabstractKnowledge graph embedding is an effective way to represent knowledge graph, which greatly enhance the performances on knowledge graph completion tasks, e.g., entity or relation prediction. For knowledge graph embedding models, designing a powerful loss framework is crucial to the discrimination between correct and incorrect triplets. Margin-based ranking loss is a commonly used negative sampling framework to make a suitable margin between the scores of positive and negative triples. However, this loss can not ensure ideal low scores for the positive triplets and high scores for the negative triplets, which is not beneficial for knowledge completion tasks. In this paper, we present a double limit scoring loss to separately set upper bound for correct triplets and lower bound for incorrect triplets, which provides more effective and flexible optimization for knowledge graph embedding. Upon the presented loss framework, we present several knowledge graph embedding models including TransE-SS, TransH-SS, TransD-SS, ProjE-SS and ComplEx-SS. The experimental results on link prediction and triplet classification show that our proposed models have the significant improvement compared to state-of-the-art baselines. Xiaofei Zhou 0002, Lingfeng Niu, Qiannan Zhu, Xingquan Zhu 0001, Ping Liu 0001, Jianlong Tan, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2021 | UMVD-FSL: Unseen Malware Variants Detection Using Few-Shot LearningabstractAs the tool for launching cyber attacks, the ever-increasing malware variants pose a significant threat to the interconnected network community. The detection methods based on conventional machine learning techniques require lots of samples for training. However, in real-world scenarios, such as in the early stage of novel attacks appearance, only a small number of malicious samples can be obtained. Applying data-intensive traditional methods in the above scenarios will cause serious overfitting problems. Therefore, there is a need for few-shot detection. In his paper, we propose UMVD-FSL, a framework based on few-shot learning to detect unseen malware variants with a small set of data. We start with network traffic data generated by malware variants and benign applications and then convert them to grayscale images. The prototype-based few-shot learning model takes the grayscale images as the input and utilizes meta-training to generalize the meta-learner for adapting new tasks. When a new sample appears, the model performs classification by computing distances to prototype representation of each class. We evaluate different methods through a series of comparative experiments. Our method has the best performance on all subtasks. The experimental results indicate that our method is universal and robust in detecting malware variants from the same network environment and different network environments. The above points prove that our method can accomplish the task of few-shot unseen malware variants detection. Candong Rong, Gaopeng Gou, Chengshang Hou, Zhen Li 0011, Gang Xiong 0001, Li Guo 0001 |
IJCNN | 6 |
| 2021 | Incorporating Specific Knowledge into End-to-End Task-oriented Dialogue SystemsabstractExternal knowledge is vital to many natural language processing tasks. However, current end-to-end dialogue systems often struggle to interface knowledge bases(KBs) with response smoothly and effectively. In this paper, we convert the raw knowledge into relation knowledge and integrated knowledge and then incorporate them into end-to-end task-oriented dialogue systems. The relation knowledge extracted from knowledge triples is combined with dialogue history, aiming to enhance semantic inputs and support better language understanding. Integrated knowledge involves entities and relations by graph attention, assisting the model in generating informative responses. The experimental results on three public dialogue datasets show that our model improves over the previous state-of-the-art models in sentence fluency and informativeness. Qingyue Wang, Yanan Cao 0001, Junyan Jiang, Yafang Wang, Lingling Tong, Li Guo 0001 |
IJCNN | 6 |
| 2021 | Direction Relation Transformer for Image CaptioningabstractImage captioning is a challenging task that combines computer vision and natural language processing for generating a textual description of the content within an image. Recently, Transformer-based encoder-decoder architectures have shown great success in image captioning, where multi-head attention mechanism is utilized to capture the contextual interactions between object regions. However, such methods regard region features as a bag of tokens without considering the directional relationships between them, making it hard to understand the relative position between objects in the image and generate correct captions effectively. In this paper, we propose a novel Direction Relation Transformer to improve the orientation perception between visual features by incorporating the relative direction embedding into multi-head attention, termed DRT. We first generate the relative direction matrix according to the positional information of the object regions, and then explore three forms of direction-aware multi-head attention to integrate the direction embedding into Transformer architecture. We conduct experiments on challenging Microsoft COCO image captioning benchmark. The quantitative and qualitative results demonstrate that, by integrating the relative directional relation, our proposed approach achieves significant improvements over all evaluation metrics compared with baseline model, e.g., DRT improves task-specific metric CIDEr score from 129.7% to 133.2% on the offline ''Karpathy'' test split. Zeliang Song, Xiaofei Zhou 0002, Linhua Dong, Jianlong Tan, Li Guo 0001 |
ACM Multimedia | 5 |
| 2021 | Knowledge Base Reasoning with Convolutional-Based Recurrent Neural NetworksabstractRecurrent neural network(RNN) has achieved remarkable performances in complex reasoning on knowledge bases, which usually takes as inputs vector embeddings of relations along a path between an entity pair. However, it is insufficient to extract local correlations of a path due to RNN is better at capturing global sequential information of a path. In this paper, we take full advantages of convolutional neural network that can effectively extract local features, and propose a convolutional-based RNN architecture denoted as C-RNN to perform reasoning. C-RNN first utilizes CNN to extract local high-level correlation features of a path, and then feeds the correlation features into recurrent neural network to model the path representation. Our C-RNN architecture is adaptable to obtain not only local features but also global sequential features of a path. Based on C-RNN architecture, we devise two models, the unidirectional C-RNN and bidirectional C-RNN. We empirically evaluate them on a large-scale FreeBase+ClueWeb prediction task. Experimental results show that C-RNN models achieve state-of-the-art predictive performance. Qiannan Zhu, Xiaofei Zhou 0002, Jianlong Tan, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Distilling Knowledge from Well-Informed Soft Labels for Neural Relation ExtractionabstractExtracting relations from plain text is an important task with wide application. Most existing methods formulate it as a supervised problem and utilize one-hot hard labels as the sole target in training, neglecting the rich semantic information among relations. In this paper, we aim to explore the supervision with soft labels in relation extraction, which makes it possible to integrate prior knowledge. Specifically, a bipartite graph is first devised to discover type constraints between entities and relations based on the entire corpus. Then, we combine such type constraints with neural networks to achieve a knowledgeable model. Furthermore, this model is regarded as teacher to generate well-informed soft labels and guide the optimization of a student network via knowledge distillation. Besides, a multi-aspect attention mechanism is introduced to help student mine latent information from text. In this way, the enhanced student inherits the dark knowledge (e.g., type constraints and relevance among relations) from teacher, and directly serves the testing scenarios without any extra constraints. We conduct extensive experiments on the TACRED and SemEval datasets, the experimental results justify the effectiveness of our approach. Zhenyu Zhang 0006, Xiaobo Shu, Bowen Yu 0002, Tingwen Liu, Jiapeng Zhao, Quangang Li, Li Guo 0001 |
AAAI | 7 |
| 2020 | A Knowledge-Aware Attentional Reasoning Network for RecommendationabstractKnowledge-graph-aware recommendation systems have increasingly attracted attention in both industry and academic recently. Many existing knowledge-aware recommendation methods have achieved better performance, which usually perform recommendation by reasoning on the paths between users and items in knowledge graphs. However, they ignore the users' personal clicked history sequences that can better reflect users' preferences within a period of time for recommendation. In this paper, we propose a knowledge-aware attentional reasoning network KARN that incorporates the users' clicked history sequences and path connectivity between users and items for recommendation. The proposed KARN not only develops an attention-based RNN to capture the user's history interests from the user's clicked history sequences, but also a hierarchical attentional neural network to reason on paths between users and items for inferring the potential user intents on items. Based on both user's history interest and potential intent, KARN can predict the clicking probability of the user with respective to a candidate item. We conduct experiment on Amazon review dataset, and the experimental results demonstrate the superiority and effectiveness of our proposed KARN model. Qiannan Zhu, Xiaofei Zhou 0002, Jia Wu 0001, Jianlong Tan, Li Guo 0001 |
AAAI | 5 |
| 2020 | BPA: The Optimal Placement of Interdependent VNFs in Many-Core System
Youbing Zhong, Zhou Zhou 0007, Xuan Liu 0006, Da Li 0002, Meijun Guo, Shuai Zhang 0007, Qingyun Liu 0001, Li Guo 0001 |
CollaborateCom (2) | 8 |
| 2020 | Document-level Relation Extraction with Dual-tier Heterogeneous GraphabstractDocument-level relation extraction (RE)poses new challenges over its sentence-level counterpart since it requires an adequate comprehension of the whole document and the multi-hop reasoning ability across multiple sentences to reach the final result.In this paper, we propose a novel graphbased model with Dual-tier Heterogeneous Graph (DHG) for document-level RE.In particular, DHG is composed of a structure modeling layer followed by a relation reasoning layer.The major advantage is that it is capable of not only capturing both the sequential and structural information of documents but also mixing them together to benefit for multi-hop reasoning and final decisionmaking.Furthermore, we employ Graph Neural Networks (GNNs) based message propagation strategy to accumulate information on DHG.Experimental results demonstrate that the proposed method achieves state-of-the-art performance on two widely used datasets, and further analyses suggest that all the modules in our model are indispensable for document-level RE. Zhenyu Zhang 0006, Bowen Yu 0002, Xiaobo Shu, Tingwen Liu, Hengzhu Tang, Li Guo 0001 |
COLING | 7 |
| 2020 | A Relation-Specific Attention Network for Joint Entity and Relation ExtractionabstractJoint extraction of entities and relations is an important task in natural language processing (NLP), which aims to capture all relational triplets from plain texts. This is a big challenge due to some of the triplets extracted from one sentence may have overlapping entities. Most existing methods perform entity recognition followed by relation detection between every possible entity pairs, which usually suffers from numerous redundant operations. In this paper, we propose a relation-specific attention network (RSAN) to handle the issue. Our RSAN utilizes relation-aware attention mechanism to construct specific sentence representations for each relation, and then performs sequence labeling to extract its corresponding head and tail entities. Experiments on two public datasets show that our model can effectively extract overlapping triplets and achieve state-of-the-art performance. Xiaofei Zhou 0002, Shirui Pan, Qiannan Zhu, Zeliang Song, Li Guo 0001 |
IJCAI | 6 |
| 2020 | MalFinder: An Ensemble Learning-based Framework For Malicious Traffic DetectionabstractMalicious events pose a significant threat to the current increasingly interconnected Internet community. Detection based on features of network traffic and machine learning algorithms is a common approach to identify malicious events. The performance of approaches is associated with the used features and algorithms. In this paper, we propose MalFinder, an ensemble learning-based framework for malicious traffic detection. Considering the trend of network traffic encryption and the complexity of decrypting traffic, we utilize statistical features and sequence features to describe network traffic. We extend the dimensions of these two types of features to enhance their capability for representing traffic data. Feature importance analysis and contrast experiments illustrate the effectiveness of our new features. Among our selected classifiers suitable for malicious traffic detection, boosting-based classifiers XGBoost and LightGBM can reduce bias, and bagging-based classifier Random Forest can reduce variance. Stacking, which is the integration method of the classification results used in our framework, can improve the generalization ability of the method. MalFinder can achieve 96.58% F-measure and 95.44% accuracy in the malicious traffic detection task on a real-world dataset, whose results are better than those of comparison methods. In terms of unseen malicious traffic discovery, MalFinder still provides good performance with 93.46% F-measure and 91.04% accuracy, which even surpasses the results in the task of known malicious traffic detection of other comparative methods. With consideration of the scarcity of public data sets used for malicious traffic detection, we have exposed our self-built dataset for more extensive researches. Candong Rong, Gaopeng Gou, Mingxin Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
ISCC | 6 |
| 2020 | SLGAT: Soft Labels Guided Graph Attention Networks
Zhenyu Zhang 0006, Tingwen Liu, Li Guo 0001 |
PAKDD (1) | 4 |
| 2020 | TransNet: Unseen Malware Variants Detection Using Deep Transfer Learning
Candong Rong, Gaopeng Gou, Mingxin Cui, Gang Xiong 0001, Zhen Li 0011, Li Guo 0001 |
SecureComm (2) | 6 |
| 2020 | Fine-Grained Semantics-Aware Heterogeneous Graph Neural Networks
Zhenyu Zhang 0006, Tingwen Liu, Li Guo 0001 |
WISE (1) | 6 |
| 2020 | A Compare-Aggregate Model with External Knowledge for Query-Focused Summarization
Jing Ya, Tingwen Liu, Li Guo 0001 |
WISE (2) | 3 |
| 2019 | DAN: Deep Attention Neural Network for News RecommendationabstractWith the rapid information explosion of news, making personalized news recommendation for users becomes an increasingly challenging problem. Many existing recommendation methods that regard the recommendation procedure as the static process, have achieved better recommendation performance. However, they usually fail with the dynamic diversity of news and user’s interests, or ignore the importance of sequential information of user’s clicking selection. In this paper, taking full advantages of convolution neural network (CNN), recurrent neural network (RNN) and attention mechanism, we propose a deep attention neural network DAN for news recommendation. Our DAN model presents to use attention-based parallel CNN for aggregating user’s interest features and attention-based RNN for capturing richer hidden sequential features of user’s clicks, and combines these features for new recommendation. We conduct experiment on real-world news data sets, and the experimental results demonstrate the superiority and effectiveness of our proposed DAN model. Qiannan Zhu, Xiaofei Zhou 0002, Zeliang Song, Jianlong Tan, Li Guo 0001 |
AAAI | 5 |
| 2019 | NTS: A Scalable Virtual Testbed Architecture with Dynamic Scheduling and Backpressure
Youbing Zhong, Zhou Zhou 0007, Da Li 0002, Wenliang He, Chao Zheng 0001, Qingyun Liu 0001, Li Guo 0001 |
CollaborateCom | 7 |
| 2019 | Chinese Social Media Entity Linking Based on Effective Context with Topic SemanticsabstractOn social media, entity linking is very important for natural language processing tasks, such as Sentiment Analysis, Question Answering (QA) and Machine Translation. Compared to English-oriented entity linking, Chinese entity linking has its special difficulties. Just like the entity linking for short text, Chinese microblogs have lots of noise and the mention lacks effective context information. In order to solve these problems, we present a new model for Chinese microblogs entity linking. Entity linking usually includes two steps: candidate entities generation and candidate entities ranking. First, based on the characteristics of Chinese, we put forward multi-method fusion strategies for candidate generation to improve the recall rate of candidate entities. Second, we propose a new neural network model called TAS (Topic attention Siamese) for candidate entities ranking. In TAS model, we add effective topic semantics on Siamese network to learn representations of context, mention and entity, and rank the mention-entity similarity. The representation of mention incorporates information from multiple sentences on the same topic, which can effectively solve the problem of the lack of contextual information. We also use Character-enhanced Word Embedding model (CWE) to pre-train both word embedding and characters embedding to work out noise and word segmentation impact. Experimental results demonstrate that our method significantly outperforms the state-of-the-art results for entity linking on Chinese social media. Chengfang Ma, Ying Sha, Jianlong Tan, Li Guo 0001, Huailiang Peng |
COMPSAC (1) | 4 |
| 2019 | Hunting for Invisible SmartCam: Characterizing and Detecting Smart Camera Based on Netflow AnalysisabstractNowadays, the rapid growth of cloud computing and IoT enabled services among multiple organizations brings both promising prospects and security & privacy challenges. IP cameras have become a top target for hackers because of their relatively high computing power and throughput. To understand the risks of these threats requires learning about IP cameras-where are they, how many are there? Active scanning is considered to be an effective way, like SHODAN. However, deployment of smart cameras in the network address translation (NAT) environments with dynamic locations is usually desired. To find these Invisible Cameras, CamHunter: (i) introduces three statements of smart cameras when they are online, (ii) concludes the most popular smart cameras in China have very similar communication patterns, (iii) proposes a model to detect smart cameras in a passive way constructed by nineteen feature sets, and (iv) raises alarms for IoT manufacturers. Our real-world experiments demonstrate the effectiveness of CamHunter in finding smart cameras even if they are behind NATs and using encrypted connections like SSL/TLS or private protocols. We argue that CamHunter represents an important view of IoT security and privacy, and it can guide the effort of designing and protecting smart cameras. Baiyang Li, Yujia Zhu, Qingyun Liu 0001, Zhou Zhou 0007, Li Guo 0001 |
ICC | 5 |
| 2019 | Deep Active Learning for Anchor User PredictionabstractPredicting pairs of anchor users plays an important role in the cross-network analysis. Due to the expensive costs of labeling anchor users for training prediction models, we consider in this paper the problem of minimizing the number of user pairs across multiple networks for labeling as to improve the accuracy of the prediction. To this end, we present a deep active learning model for anchor user prediction (DALAUP for short). However, active learning for anchor user sampling meets the challenges of non-i.i.d. user pair data caused by network structures and the correlation among anchor or non-anchor user pairs. To solve the challenges, DALAUP uses a couple of neural networks with shared-parameter to obtain the vector representations of user pairs, and ensembles three query strategies to select the most informative user pairs for labeling and model training. Experiments on real-world social network data demonstrate that DALAUP outperforms the state-of-the-art approaches. Anfeng Cheng, Chuan Zhou 0001, Hong Yang 0003, Jia Wu 0001, Lei Li 0002, Jianlong Tan, Li Guo 0001 |
IJCAI | 7 |
| 2019 | Neighborhood-Aware Attentional Representation for Multilingual Knowledge GraphsabstractMultilingual knowledge graphs constructed by entity alignment are the indispensable resources for numerous AI-related applications. Most existing entity alignment methods only use the triplet-based knowledge to find the aligned entities across multilingual knowledge graphs, they usually ignore the neighborhood subgraph knowledge of entities that implies more richer alignment information for aligning entities. In this paper, we incorporate neighborhood subgraph-level information of entities, and propose a neighborhood-aware attentional representation method NAEA for multilingual knowledge graphs. NAEA devises an attention mechanism to learn neighbor-level representation by aggregating neighbors' representations with a weighted combination. The attention mechanism enables entities not only capture different impacts of their neighbors on themselves, but also attend over their neighbors' feature representations with different importance. We evaluate our model on two real-world datasets DBP15K and DWY100K, and the experimental results show that the proposed model NAEA significantly and consistently outperforms state-of-the-art entity alignment models. Qiannan Zhu, Xiaofei Zhou 0002, Jia Wu 0001, Jianlong Tan, Li Guo 0001 |
IJCAI | 5 |
| 2019 | NeuralAS: Deep Word-Based Spoofed URLs Detection Against Strong Similar SamplesabstractSpoofed URLs are associated with various cyber crimes such as phishing and ransomware etc. Most existing detection approaches design a set of hand-crafted features and feed them to machine learning classifiers. However, designing such features is a time consuming and labor intensive process. This paper proposes an approach named NeuralAS (Neural Anti-Spoofing) by segmenting URLs into word sequences and detecting spoofed URLs with recurrent neural networks. As a result, NeuralAS can perform detection with high-abstract and poor-interpretable features learned automatically, and achieve accurate detection with contextual information in sequences. We also propose a novel method to construct indistinguishable data sets of strong similar samples, which can be used to evaluate the robustness of different approaches. Extensive experimental results show that NeuralAS works well on spoofed URLs detection, and has a significant effectiveness and robustness even on strong similar data sets. Jing Ya, Tingwen Liu, Jinqiao Shi, Li Guo 0001, Zhaojun Gu |
IJCNN | 5 |
| 2019 | Smooth Deep Network EmbeddingabstractNetwork embedding is an efficient method to learn low-dimensional representations of vertexes in networks since the network structure can be captured and preserved through this process. Unlike shallow models, deep neural network framework is able to capture the highly non-linear network structure. Therefore, it can achieve much better performance in comparison of traditional network embedding methods. However, few attention has been paid to the smoothness of such models, in contrast to numerous research works for image and text fields. Methods without smoothness are not robust enough, which means that slight changes on network may lead dramatic changes on the embedding results. Hence, how to find a smooth deep framework is still an open yet important problem. To this end, in this paper, we propose a Smooth Deep Network Embedding method, namely SmNE, which generates stable and reliable embedding results. Empirically, we conduct experiments on real-world networks. The results show that compared to the state-of-the-art methods, our proposed method can achieve significant gains in several applications. Mengyu Zheng, Chuan Zhou 0001, Jia Wu 0001, Li Guo 0001 |
IJCNN | 4 |
| 2018 | Knowledge Graph Embedding With Iterative Guidance From Soft RulesabstractEmbedding knowledge graphs (KGs) into continuous vector spaces is a focus of current research. Combining such an embedding model with logic rules has recently attracted increasing attention. Most previous attempts made a one-time injection of logic rules, ignoring the interactive nature between embedding learning and logical inference. And they focused only on hard rules, which always hold with no exception and usually require extensive manual effort to create or validate. In this paper, we propose Rule-Guided Embedding (RUGE), a novel paradigm of KG embedding with iterative guidance from soft rules. RUGE enables an embedding model to learn simultaneously from 1) labeled triples that have been directly observed in a given KG, 2) unlabeled triples whose labels are going to be predicted iteratively, and 3) soft rules with various confidence levels extracted automatically from the KG. In the learning process, RUGE iteratively queries rules to obtain soft labels for unlabeled triples, and integrates such newly labeled triples to update the embedding model. Through this iterative procedure, knowledge embodied in logic rules may be better transferred into the learned embeddings. We evaluate RUGE in link prediction on Freebase and YAGO. Experimental results show that: 1) with rule knowledge injected iteratively, RUGE achieves significant and consistent improvements over state-of-the-art baselines; and 2) despite their uncertainties, automatically extracted soft rules are highly beneficial to KG embedding, even those with moderate confidence levels. The code and data used for this paper can be obtained from https://github.com/iieir-km/RUGE. Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
AAAI | 5 |
| 2018 | Social Recommendation with an Essential Preference SpaceabstractSocial recommendation, which aims to exploit social information to improve the quality of a recommender system, has attracted an increasing amount of attention in recent years. A large portion of existing social recommendation models are based on the tractable assumption that users consider the same factors to make decisions in both recommender systems and social networks. However, this assumption is not in concert with real-world situations, since users usually show different preferences in different scenarios. In this paper, we investigate how to exploit the differences between user preference in recommender systems and that in social networks, with the aim to further improve the social recommendation. In particular, we assume that the user preferences in different scenarios are results of different linear combinations from a more underlying user preference space. Based on this assumption, we propose a novel social recommendation framework, called social recommendation with an essential preferences space (SREPS), which simultaneously models the structural information in the social network, the rating and the consumption information in the recommender system under the capture of essential preference space. Experimental results on four real-world datasets demonstrate the superiority of the proposed SREPS model compared with seven state-of-the-art social recommendation methods. Chun-Yi Liu 0003, Chuan Zhou 0001, Jia Wu 0001, Yue Hu 0002, Li Guo 0001 |
AAAI | 5 |
| 2018 | Improving Knowledge Graph Embedding Using Simple ConstraintsabstractEmbedding knowledge graphs (KGs) into continuous vector spaces is a focus of current research.Early works performed this task via simple models developed over KG triples.Recent attempts focused on either designing more complicated triple scoring models, or incorporating extra information beyond triples.This paper, by contrast, investigates the potential of using very simple constraints to improve KG embedding.We examine non-negativity constraints on entity representations and approximate entailment constraints on relation representations.The former help to learn compact and interpretable representations for entities.The latter further encode regularities of logical entailment between relations into their distributed representations.These constraints impose prior beliefs upon the structure of the embedding space, without negative impacts on efficiency or scalability.Evaluation on WordNet, Freebase, and DBpedia shows that our approach is simple yet surprisingly effective, significantly and consistently outperforming competitive baselines.The constraints imposed indeed improve model interpretability, leading to a substantially increased structuring of the embedding space.Code and data are available at https://github.com/i ieir-km/ComplEx-NNE_AER. Boyang Ding, Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
ACL (1) | 4 |
| 2018 | SASD: A Self-Adaptive Stateful Decompression ArchitectureabstractDue to the increasing threats in the current network environment, many researchers have shifted their interests to network content audit, which combines deep packet inspection and natural language processing. However, the performance of network content audit systems is becoming the bottle-neck because of the demand on processing fast growing compressed traffic. While compressed traffic is often split into multiple out-of-order packets for transmission, stateful decompression ensures that the compressed data are processed in a timely manner without waiting for all the compressed traffic to arrive before decompressing. In the meanwhile, hardware innovations lead to new type of devices being invented, which shows promise to fully handle the offloaded traffic for complex calculations at higher throughput than software-based solutions. We consider both software-based and hardware-based solutions for decompressing traffic from network content audit systems and study the workload. We notice that the performance is data-dependant: hardware-based decompression solutions perform better for longer compressed data than software method. On the contrary, software-based decompressing methods are more preferred for the short content in terms of the processing speed. So there is no one-size-fits-all solution. In this paper, we combine the advantages of hardware and software and propose a novel self-adaptive stateful decompression architecture to support fast decompression in accordance with the traffic status and system state. Experiments on real-world traffic show that our proposed architecture can achieve about three times of the data decompression efficiency, compared to the best pure software and hardware algorithm, which can significantly improve the detection efficiency of many network content audit systems. Zhou Zhou 0007, Qingyun Liu 0001, Yujia Zhu, Da Li 0002, Li Guo 0001 |
GLOBECOM | 6 |
| 2018 | Hierarchical Attention Networks for User Profile Inference in Social Media Systems
Zhezhou Kang, Yanan Cao 0001, Yanmin Shang, Yanbing Liu 0007, Li Guo 0001 |
ICANN (3) | 6 |
| 2018 | User Alignment via Structural Interaction and PropagationabstractUser alignment between different social networks is a fundamental issue for many applications, such as information diffusion and recommendation. In actuality, the observed anchor users are normally sparse due to the expensiveness of labeling data. Hence how to make the best use of these sparse anchor information is an important open issue. To this end, we proposed a User Alignment via Structural Interaction and Propagation (UASIP) model to capture the structural information interaction across two social networks, which exploits deep structural infor- mation to enhance the representations of users. UASIP learns vector representation by automatically keeping the consistency between this additional structural information and intrinsic structural information of the two social networks. Experiments on real-world social network datasets demonstrate the effectiveness of UASIP compared with several state-of-the-art methods. Anfeng Cheng, Chun-Yi Liu 0003, Chuan Zhou 0001, Jianlong Tan, Li Guo 0001 |
IJCNN | 5 |
| 2018 | iWalk: Interest-Aware Random Walk for Network EmbeddingabstractNetwork embedding plays a key role in network analysis, due to its ability to represent features of network structure in a low-dimensional Euclidean space, making it possible to directly utilize the of f-the-shelf mining techniques in a variety of analysis tasks. Although fruitful research papers on network embedding have sprung up in recent years, most of them neglect an important fact that nodes and edges in real-world networks are of diverse interests especially when the network contains little side information such as labels. To tackle this challenge, we propose a novel iWalk model to learn interest-aware network embedding in an unsupervised fashion. iWalk can automatically assign interest to nodes and edges based on network topology and construct custom paths navigated by assigned interest, then Skip-gram is used to learn network embedding from these paths. Sufficient experiments are conducted on different tasks and three typical datasets, the empirical results demonstrate that our model outperform the stat-of-art methods in most instances. Wen Zan, Chuan Zhou 0001, Hong Yang 0003, Yue Hu 0002, Li Guo 0001 |
IJCNN | 5 |
| 2018 | FraudNE: a Joint Embedding Approach for Fraud DetectionabstractDetecting fraudsters is a meaningful problem for both users and e-commerce platform. Existing graph-based approaches mainly adopt shallow models, which cannot capture the highly non-linear relationship between vertexes in a bipartite graph composed of users and items. To address this issue, in this paper we propose a joint deep structure embedding approach FraudNE for fraud detection that (a) can preserve the highly non-linear structural information of networks, (b) is robust to sparse networks, (c) embeds different types of vertexes jointly in the same latent space. It is worth mentioning that we can detect multiple fraudulent groups without the number of groups as a priori. Compared with baselines, our method achieved significant accuracy improvement. Mengyu Zheng, Chuan Zhou 0001, Jia Wu 0001, Shirui Pan, Jinqiao Shi, Li Guo 0001 |
IJCNN | 6 |
| 2018 | Fine-Grained Correlation Learning with Stacked Co-attention Networks for Cross-Modal Information Retrieval
Jing Yu 0007, Yanbing Liu 0007, Jianlong Tan, Li Guo 0001, Weifeng Zhang 0002 |
KSEM (1) | 5 |
| 2018 | A Sequence Transformation Model for Chinese Named Entity Recognition
Qingyue Wang, Yanjing Song, Yanan Cao 0001, Yanbing Liu 0007, Li Guo 0001 |
KSEM (1) | 6 |
| 2017 | Learning Knowledge Embeddings by Combining Limit-based Scoring LossabstractIn knowledge graph embedding models, the margin-based ranking loss as the common loss function is usually used to encourage discrimination between golden triplets and incorrect triplets, which has proved effective in many translation-based models for knowledge graph embedding. However, we find that the loss function cannot ensure the fact that the scoring of correct triplets must be low enough to fulfill the translation. In this paper, we present a limit-based scoring loss to provide lower scoring of a golden triplet, and then to extend two basic translation models TransE and TransH, separately to TransE-RS and TransH-RS by combining limit-based scoring loss with margin-based ranking loss. Both the presented models have low complexities of parameters benefiting for application on large scale graphs. In experiments, we evaluate our models on two typical tasks including triplet classification and link prediction, and also analyze the scoring distributions of positive and negative triplets by different models. Experimental results show that the introduced limit-based scoring loss is effective to improve the capacities of knowledge graph embedding. Xiaofei Zhou 0002, Qiannan Zhu, Ping Liu 0001, Li Guo 0001 |
CIKM | 4 |
| 2017 | Flexible Expert Finding on the Web via Semantic Hypergraph Learning and Affinity Propagation ModelabstractExpert finding (EF) task has received widespread attention as an important task of information retrieval.One key category of EF is expert finding on the web, which seeks to rank influential public figures from diverse webpage sources with respect to given query.Previous web expert finding approach relies on casting webpages to hypergraph structure and run heat diffusion process to find the top ranking person vertices according to their heat of popularity.Such approach suffer from two major drawbacks:First, previous web expert finding approach (CoDiffusion) suffers from unflexibility of selecting queries.This means that all corresponding queries must be stored as vertices in hypergraph index beforehand, otherwise CoDiffusion cannot run the expert finding process.Such defect make it ungeneric and incapable of handling the newly invented technical terms or phrases in real world scenarios.Second, the performance of previous approach is less satisfying.We incorporate semantic relatedness information with Hypergraph Learning Framework and Affinity Propagation ({HLFAP}) to handle the above drawbacks.In order to overcome the first disadvantage, we distribute initial heat to the related vertices according to their semantic similarity on given query. In order to solve the second disadvantage, we propose semantic labeled hypergraph learning framework and person influence affinity propagation model to make high quality candidates can receive more heat transition. Experimental results shows that our generic methodology achieves more satisfying results than the non-semantics state-of-the-art baseline method. Tingwen Liu, Jinqiao Shi, Qiuyan Wang, Li Guo 0001 |
ICTAI | 5 |
| 2017 | CPMF: A collective pairwise matrix factorization model for upcoming event recommendationabstractDue to the rapid growth of event-based social networks (EBSNs), event recommendation which helps users find their preferred events has become a popular topic. Different from movies or books in conventional recommendation problem, events usually have recommendation lifetimes and almost all the events to be recommended are upcoming, which brings a severe cold start problem. To achieve better event recommendation performance, we formulates multiple interactions among users, events, groups and locations into an unified framework and propose a collective pairwise matrix factorization (CPMF) model to estimate users' pairwise preferences on events, groups and locations. We further develop an efficient stochastic gradient descent algorithm for the model learning. We conduct experiments on real-world Meetup datasets and the experimental results demonstrate that our CPMF model can outperform the state-of-the-art methods. Chun-Yi Liu 0003, Chuan Zhou 0001, Jia Wu 0001, Hongtao Xie 0001, Yue Hu 0002, Li Guo 0001 |
IJCNN | 6 |
| 2017 | Improving Password Guessing Using Byte Pair Encoding
Dakui Wang, Xiaojun Chen 0004, Jinqiao Shi, Li Guo 0001 |
ISC | 6 |
| 2017 | Inferring User Profiles in Online Social Networks Based on Convolutional Neural Network
Yanan Cao 0001, Yanmin Shang, Yanbing Liu 0007, Jianlong Tan, Li Guo 0001 |
KSEM | 6 |
| 2017 | Boosting imbalanced data learning with Wiener process oversampling
Qian Li 0003, Gang Li 0009, Wenjia Niu, Yanan Cao 0001, Liang Chang 0003, Jianlong Tan, Li Guo 0001 |
Frontiers Comput. Sci. | 7 |
| 2017 | SSE: Semantically Smooth Embedding for Knowledge GraphsabstractThis paper considers the problem of embedding Knowledge Graphs (KGs) consisting of entities and relations into low-dimensional vector spaces. Most of the existing methods perform this task based solely on observed facts. The only requirement is that the learned embeddings should be compatible within each individual fact. In this paper, aiming at further discovering the intrinsic geometric structure of the embedding space, we proposeSemantically Smooth Embedding(SSE). The key idea of SSE is to take full advantage of additional semantic information and enforce the embedding space to be semantically smooth, i.e., entities belonging to the same semantic category will lie close to each other in the embedding space. Two manifold learning algorithms Laplacian Eigenmaps and Locally Linear Embedding are used to model the smoothness assumption. Both are formulated as geometrically based regularization terms to constrain the embedding task. Two lines of embedding strategies are tested, i.e., strategies based on latent distance models and strategies based on tensor factorization techniques. We empirically evaluate SSE on two benchmark tasks of link prediction and triple classification, and achieve significant and consistent improvements over state-of-the-art methods. The results demonstrate the superiority and generality of SSE. Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2017 | Knowledge Graph Embedding: A Survey of Approaches and ApplicationsabstractKnowledge graph (KG) embedding is to embed components of a KG including entities and relations into continuous vector spaces, so as to simplify the manipulation while preserving the inherent structure of the KG. It can benefit a variety of downstream tasks such as KG completion and relation extraction, and hence has quickly gained massive attention. In this article, we provide a systematic review of existing techniques, including not only the state-of-the-arts but also those with latest trends. Particularly, we make the review based on the type of information used in the embedding task. Techniques that conduct embedding using only facts observed in the KG are first introduced. We describe the overall framework, specific model design, typical training procedures, as well as pros and cons of such techniques. After that, we discuss techniques that further incorporate additional information besides facts. We focus specifically on the use of entity types, relation paths, textual descriptions, and logical rules. Finally, we briefly introduce how KG embedding can be applied to and benefit a wide variety of downstream tasks such as KG completion, relation extraction, question answering, and so forth. Quan Wang 0002, Zhendong Mao 0001, Bin Wang 0004, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | On the Minimum Differentially Resolving Set Problem for Diffusion Source Inference in NetworksabstractIn this paper we theoretically study the minimum Differentially Resolving Set (DRS) problem derived from the classical sensor placement optimization problem in network source locating. A DRS of a graph G = (V, E) is defined as a subset S ⊆ V where any two elements in V can be distinguished by their different differential characteristic sets defined on S. The minimum DRS problem aims to find a DRS S in the graph G with minimum total weight Σv∈S w(v). In this paper we establish a group of Integer Linear Programming (ILP) models as the solution. By the weighted set cover theory, we propose an approximation algorithm with the Θ(ln n) approximability for the minimum DRS problem on general graphs, where n is the graph size. Chuan Zhou 0001, Weixue Lu, Peng Zhang 0001, Jia Wu 0001, Yue Hu 0002, Li Guo 0001 |
AAAI | 6 |
| 2016 | Location-aware Friend Recommendation in Event-based Social Networks: A Bayesian Latent Factor ApproachabstractIn this paper we study the friend recommendation problem in event-based social networks (EBSNs). Effective friend recommendation is of benefit to EBSNs, since it can promote user interaction and accelerate information diffusion for promoted events. Different from usual friend recommendations, the aim of making friends in EBSNs is to better participate offline events and enhance user experience. Meanwhile friend recommendation in EBSNs encounters three types of data, i.e. geographical information, implicate user rating, and user behavior. These differences imply that existing friend recommendation approaches are not adequate any more for EBSNs. Under this background, in this paper we propose a Bayesian latent factor model, which can jointly formulate above three types of data, for friend recommendation with better event promotion and user experience. Results on real-world datasets show the efficacy of our approach. Zhi Qiao 0005, Chuan Zhou 0001, Yue Hu 0002, Li Guo 0001 |
CIKM | 5 |
| 2016 | Jointly Embedding Knowledge Graphs and Logical RulesabstractEmbedding knowledge graphs into continuous vector spaces has recently attracted increasing interest.Most existing methods perform the embedding task using only fact triples.Logical rules, although containing rich background information, have not been well studied in this task.This paper proposes a novel method of jointly embedding knowledge graphs and logical rules.The key idea is to represent and model triples and rules in a unified framework.Specifically, triples are represented as atomic formulae and modeled by the translation assumption, while rules represented as complex formulae and modeled by t-norm fuzzy logics.Embedding then amounts to minimizing a global loss over both atomic and complex formulae.In this manner, we learn embeddings compatible not only with triples but also with rules, which will certainly be more predictive for knowledge acquisition and inference.We evaluate our method with link prediction and triple classification tasks.Experimental results show that joint embedding brings significant and consistent improvements over stateof-the-art methods.Particularly, it enhances the prediction of new facts which cannot even be directly inferred by pure logical inference, demonstrating the capability of our method to learn more predictive embeddings. Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
EMNLP | 5 |
| 2016 | Riemannian optimization with subspace tracking for low-rank recoveryabstractLow-rank matrix recovery (MR) has been widely used in data analysis and dimensionality reduction. As a direct heuristic to MR, convex relaxation is usually degraded by the repeated calling of singular value decomposition (SVD), especially in large-scale applications. In this paper, we propose a novel Riemannian optimization method (ROAM) for MR problem by exploiting the Riemannian geometry of the searching space. In particular, ROAM utilizes an efficient subspace tracking schema that automatically detects the unknown rank to identify the preferable geometry space. Moreover, a gradient-based optimization algorithm is proposed to obtain the latent low-rank component, which avoids the expensive full dimension of SVD. More significantly, ROAM algorithm is proved to converge under mild assumptions, which also verifies the effectiveness of ROAM. Extensive empirical results demonstrate the improved accuracy and efficiency of ROAM over convex-relaxation approaches. Qian Li 0003, Wenjia Niu, Gang Li 0009, Jianlong Tan, Gang Xiong 0001, Li Guo 0001 |
IJCNN | 6 |
| 2016 | An Unsupervised Framework Towards Sci-Tech Compound Entity Recognition
Tingwen Liu, Li Guo 0001, Jiapeng Zhao, Jinqiao Shi |
KSEM | 3 |
| 2016 | Securing cyberspaceabstractSecuring cyberspaceCyberspace, the ubiquitous space that exists in relation to the Internet, is usually referred to as a dynamic broad domain ranging from Internet and its infrastructures to social networks.More research work in security has been extended from securing computers to securing Cyberspace, which includes the physical-level security, the network-level security, and the application-level security and addresses improvements in Cyberspace management.As a result, recent years have witnessed increasing research attention on securing Cyberspace, and many interesting methods have been proposed to locate suspicious IP, detect gossip content, prevent illegal information publication and distribution, manage social software and applications, and profile user behavior and opinion.This trend has provided the motivation to launch this special issue.Based on an open call in this area and invited best papers from The Fifth International Conference on Applications and Techniques for Information Security (ATIS 2014) and The first International Workshop on Curbing Cyber-Crimes (C 3 2014), five submissions have been accepted to best illustrate the main development and perspectives.The papers in this issue report a variety of methods used to tackle the security issues in cyberspace.They aim at improving security in applications ranging from RFID systems, location based service, discovery of software vulnerability to private medical records, and outsourcing in Multi-Cloud.The problems discussed in these papers are also related to disciplines including data mining, network security, digital forensics, and behavioral and psychological sciences.Here, we provide an integrative perspective of this special issue by summarizing each contribution contained therein.In [1], to address security and privacy issue in RFID systems, a new off-line reading orderindependent grouping-proof protocol is proposed to generate a proof that a group of tags have been scanned simultaneously in the range of a reader.The proposed protocol defines an ideal groupingproof functionality aiming at capturing the secure grouping-proof generation for a group of RFID tags in the UC framework.The new protocol maintains its security properties when composed concurrently with an unbounded number of instances of arbitrary protocol.In addition, the protocol conforms to the computational constraints of EPC Class-Gen-2 passive RFID tags.It is suitable for low-cost passive RFID tags, which are widely used in practical applications.In [2], the authors proposed an algorithm to address the problem of preserving privacy for individual users in location-aware applications.They define a novel distance measurement that combines the semantic and Euclidean distance to address the privacy-preserving issue.They conduct performance experiments on the proposed algorithm and distance metric, and results suggest that they can successfully retain the utility of the location services.In [3], to discover software vulnerability, an effective and efficient mechanism is proposed.The method also helps programmers to write secure code to avoid the existence of vulnerability at the early stage of software development.The proposed mechanism uses code clone verification to discover vulnerability in software programs and reduces the false positive of detection by combining the advantages of static and dynamic analysis.In addition, it also mitigates the path explosion problem in the testing process when verifying the existence of vulnerability.As a result, the proposed approach effectively improves the security of software systems, applications, and utilities in various areas of Cyberspace.In particular, it helps to create a reliable environment for the communications of all the social media participants.In [4], the authors analyze the security of Fair Remote Retrieval (FRR) model that is used to ensure the integrity of remote medical records.They show that FRR model fails to achieve its security goals, therefore present an improved protocol, called IFR 2, to fix the security minor faults Gang Li 0009, Wenjia Niu, Li Guo 0001, Lynn Margaret Batten, Yinlong Liu, Guoyong Cai |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Exploring probabilistic follow relationship to prevent collusive peer-to-peer piracy
Wenjia Niu, Endong Tong, Qian Li 0003, Gang Li 0009, Xuemin Wen, Jianlong Tan, Li Guo 0001 |
Knowl. Inf. Syst. | 7 |
| 2015 | Semantically Smooth Knowledge Graph EmbeddingabstractShu Guo, Quan Wang, Bin Wang, Lihong Wang, Li Guo. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
ACL (1) | 5 |
| 2015 | Sentiment Word Identification with Sentiment Contextual Factors
Jiguang Liang, Xiaofei Zhou 0002, Yue Hu 0002, Li Guo 0001, Shuo Bai |
APWeb | 4 |
| 2015 | A Self-learning Rule-Based Approach for Sci-tech Compound Phrase Entity Recognition
Tingwen Liu, Jinqiao Shi, Li Guo 0001 |
APWeb | 5 |
| 2015 | Lingo: Linearized Grassmannian Optimization for Nuclear Norm MinimizationabstractAs a popular heuristic to the matrix rank minimization problem, nuclear norm minimization attracts intensive research attentions. Matrix factorization based algorithms can reduce the expensive computation cost of SVD for nuclear norm minimization. However, most matrix factorization based algorithms fail to provide the theoretical guarantee for convergence caused by their non-unique factorizations. This paper proposes an efficient and accurate Linearized Grassmannian Optimization (Lingo) algorithm, which adopts matrix factorization and Grassmann manifold structure to alternatively minimize the subproblems. More specially, linearization strategy makes the auxiliary variables unnecessary and guarantees the close-form solution for low per-iteration complexity. Lingo then converts linearized objective function into a nuclear norm minimization over Grassmannian manifold, which could remedy the non-unique of solution for the low-rank matrix factorization. Extensive comparison experiments demonstrate the accuracy and efficiency of Lingo algorithm. The global convergence of Lingo is guaranteed with theoretical proof, which also verifies the effectiveness of Lingo. Qian Li 0003, Wenjia Niu, Gang Li 0009, Yanan Cao 0001, Jianlong Tan, Li Guo 0001 |
CIKM | 6 |
| 2015 | Context-Dependent Knowledge Graph EmbeddingabstractWe consider the problem of embedding knowledge graphs (KGs) into continuous vector spaces.Existing methods can only deal with explicit relationships within each triple, i.e., local connectivity patterns, but cannot handle implicit relationships across different triples, i.e., contextual connectivity patterns.This paper proposes context-dependent KG embedding, a twostage scheme that takes into account both types of connectivity patterns and obtains more accurate embeddings.We evaluate our approach on the tasks of link prediction and triple classification, and achieve significant and consistent improvements over state-of-the-art methods. Yuanfei Luo, Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
EMNLP | 4 |
| 2015 | Data-oriented multi-index hashingabstractMulti-index hashing (MIH) is the state-of-the-art method for indexing binary codes, as it divides long codes into substrings and builds multiple hash tables. However, MIH is based on the dataset codes uniform distribution assumption, and will lose efficiency in dealing with non-uniformly distributed codes. Besides, there are lots of results sharing the same Hamming distance to a query, which makes the distance measure ambiguous. In this paper, we propose a data-oriented multi-index hashing method. We first compute the covariance matrix of bits and learn adaptive projection vector for each binary substring. Instead of using substrings as direct indices into hash tables, we project them with corresponding projection vectors to generate new indices. With adaptive projection, the indices in each hash table are near uniformly distributed. Then with covariance matrix, we propose a ranking method for the binary codes. By assigning different bit-level weights to different bits, the returned binary codes are ranked at a finer-grained binary code level. Experiments conducted on reference large scale datasets show that compared to MIH the time performance of our method can be improved by 36.9%-87.4%, and the search accuracy can be improved by 22.2%. Qingyun Liu 0001, Hongtao Xie 0001, Li Guo 0001 |
ICME | 5 |
| 2015 | What is the next step of binary features?abstractVarious binary features have been recently proposed in literature, aiming at improving the computational efficiency and storage efficiency of image retrieval applications. However, the most common way of using binary features is voting strategy based on brute-force matching, since binary features are discrete data points distributed in Hamming space, so that models based on clustering such as BoW are unsuitable for them. Although indexing mechanism substantially decreases the time cost, the brute-force matching strategy becomes a bottleneck that restricts the performance of binary features. To address this issue, we propose a simple but effective method, namely COIP (Coding by Order-independent Projection), which projects binary features into a binary code of limited bits. As a result, each image is represented by one single binary code that can be indexed for computational and storage efficiency. We prove that the similarity between the COIP codes of two images with probability proportional to the ratio of their matched features. A comprehensive evaluation with several state-of-the-art binary features is performed on benchmark dataset. Experimental results reveal that for binary feature based image retrieval, our approach improves the storage/time efficiency by one/two orders of magnitude, while the retrieval performance remains almost unchanged. Zhendong Mao 0001, Lei Zhang 0119, Bin Wang 0004, Li Guo 0001 |
ICME | 4 |
| 2015 | Knowledge Base Completion Using Embeddings and Rules
Quan Wang 0002, Bin Wang 0004, Li Guo 0001 |
IJCAI | 3 |
| 2015 | Towards misdirected email detection based on multi-attributesabstractEmail has become widely used in recent years bringing with it new problems. Although this event doesn't happen often, misdirected emails can bring out great information leakage. It is not easy to detect these misdirected emails from legitimate ones since they may be only distinguishable in the sender's perspective. Existing methods discover misdirected emails from user agent or gateway but are not appropriate for varied application environment. This paper proposes a misdirected mail detection method based on multi-attributes which can be deployed on server side. Three type of attributes including email content fingerprinting, social relationship and meta information are considered in this method. Based on SVM classification algorithm, experiments show that it can detect misdirected emails with up to 91.6% accuracy. Yiguo Pu, Jinqiao Shi, Xiaojun Chen 0004, Li Guo 0001, Tingwen Liu |
ISCC | 4 |
| 2015 | Evolving Chinese Restaurant Processes for Modeling Evolutionary Traces in Temporal Data
Peng Wang 0028, Chuan Zhou 0001, Peng Zhang 0001, Weiwei Feng, Li Guo 0001, Binxing Fang |
PAKDD (2) | 5 |
| 2015 | Modelling semantics across multiple time series and its applications
Zhi Qiao 0005, Guangyan Huang, Jing He 0004, Peng Zhang 0001, Yanchun Zhang, Li Guo 0001 |
Knowl. Based Syst. | 6 |
| 2015 | E-Tree: An Efficient Indexing Structure for Ensemble Models on Data StreamsabstractEnsemble learning is a common tool for data stream classification, mainly because of its inherent advantages of handling large volumes of stream data and concept drifting. Previous studies, to date, have been primarily focused on building accurate ensemble models from stream data. However, a linear scan of a large number of base classifiers in the ensemble during prediction incurs significant costs in response time, preventing ensemble learning from being practical for many real-world time-critical data stream applications, such as Web traffic stream monitoring, spam detection, and intrusion detection. In these applications, data streams usually arrive at a speed of GB/second, and it is necessary to classify each stream record in a timely manner. To address this problem, we propose a novel Ensemble-tree (E-tree for short) indexing structure to organize all base classifiers in an ensemble for fast prediction. On one hand, E-trees treat ensembles as spatial databases and employ an R-tree like height-balanced structure to reduce the expected prediction time from linear to sub-linear complexity. On the other hand, E-trees can be automatically updated by continuously integrating new classifiers and discarding outdated ones, well adapting to new trends and patterns underneath data streams. Theoretical analysis and empirical studies on both synthetic and real-world data streams demonstrate the performance of our approach. Peng Zhang 0001, Chuan Zhou 0001, Peng Wang 0028, Byron J. Gao, Xingquan Zhu 0001, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2015 | On the Upper Bounds of Spread for Greedy Algorithms in Social Network Influence MaximizationabstractInfluence maximization, defined as finding a small subset of nodes that maximizes spread of influence in social networks, is NP-hard under both Independent Cascade (IC) and Linear Threshold (LT) models, where many greedy-based algorithms have been proposed with the best approximation guarantee. However, existing greedy-based algorithms are inefficient on large networks, as it demands heavy Monte-Carlo simulations of the spread functions for each node at the initial step [7]. In this paper, we establish new upper bounds to significantly reduce the number of Monte-Carlo simulations in greedy-based algorithms, especially at the initial step. We theoretically prove that the bound is tight and convergent when the summation of weights towards (or from) each node is less than 1. Based on the bound, we propose a new Upper Bound based Lazy Forward algorithm (UBLF in short) for discovering the top-k influential nodes in social networks. We test and compare UBLF with prior greedy algorithms, especially CELF [30]. Experimental results show that UBLF reduces more than 95 percent Monte-Carlo simulations of CELF and achieves about 2-10 times speedup when the seed set is small. Chuan Zhou 0001, Peng Zhang 0001, Wenyu Zang, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2014 | Combining Heterogenous Social and Geographical Information for Event RecommendationabstractWith the rapid growth of event-based social networks (EBSNs) like Meetup, the demand for event recommendation becomes increasingly urgent. In EBSNs, event recommendation plays a central role in recommending the most relevant events to users who are likely to participate in. Different from traditional recommendation problems, event recommendation encounters three new types of information, i.e., heterogenous online+offline social relationships, geographical features of events and implicit rating data from users. Yet combining the three types of data for offline event recommendation has not been considered. Therefore, we present a Bayesian latent factor model that can unify these data for event recommendation. Experimental results on real-world data sets show the performance of our method. Zhi Qiao 0005, Peng Zhang 0001, Yanan Cao 0001, Chuan Zhou 0001, Li Guo 0001, Binxing Fang |
AAAI | 5 |
| 2014 | Event Recommendation in Event-Based Social NetworksabstractWith the rapid growth of event-based social networks, the demand of event recommendation becomes increasingly important. Different from classic recommendation problems, event recommendation generally faces the problems of heterogenous online and offline social relationships among users and implicit feedback data. In this paper, we present a baysian probability model that can fully unleash the power of heterogenous social relations and efficiently tackle with implicit feedback characteristic for event recommendation. Experimental results on several real-world datasets demonstrate the utility of our method. Zhi Qiao 0005, Peng Zhang 0001, Chuan Zhou 0001, Yanan Cao 0001, Li Guo 0001, Yanchuan Zhang |
AAAI | 5 |
| 2014 | Delta-K 2-tree for Compact Representation of Web Graphs
Gang Xiong 0001, Yanbing Liu 0007, Ping Liu 0001, Li Guo 0001 |
APWeb | 6 |
| 2014 | POSTER: Mining Elephant Applications in Unknown Traffic by Service ClusteringabstractNetwork traffic classification is of great importance for fine-grained network management and network security. However, with the rapid development of new network applications in recent years, traffic that cannot be identified by classifiers accounts for an increasing ratio, which brings a great challenge for network operators. Most of the unknown traffic is usually generated by only a few or some certain kinds of applications. We call this kind of traffic as the elephant traffic. It is generally recognized that traffic sharing the same server IP and server port is generated by the same application. In this paper, we say that they are belonging to the same service. Therefore, we propose a novel method, in which service-based statistical features are used for cluster analysis, to classify these elephant traffic. Preliminary results on a real network traffic dataset show that our method is able to automatically identify similar unknown applications. We believe that classifying unknown traffic in service perspective is a promising direction. Gang Xiong 0001, Li Guo 0001, Zhen Li 0011, Yong Wang 0032 |
CCS | 4 |
| 2014 | CONR: A Novel Method for Sentiment Word IdentificationabstractSentiment word identification (SWI) is of high relevance to sentiment analysis technologies and applications. Currently most SWI methods heavily rely on sentiment seed words that have limited sentiment information. Even though there emerge non-seed approaches based on sentiment labels of documents, but in which the context information has not been fully considered. In this paper, based on matrix factorization with co-occurrence neighbor regularization which is derived from context, we propose a novel non-seed model called CONR for SWI. Instead of seed words, CONR exploits two important factors: sentiment matching and sentiment consistency for sentiment word identification. Experimental results on four publicly available datasets show that CONR can outperform the state of-the-art methods. Jiguang Liang, Xiaofei Zhou 0002, Yue Hu 0002, Li Guo 0001, Shuo Bai |
CIKM | 4 |
| 2014 | A Regularized Competition Model for Question Difficulty Estimation in Community Question Answering ServicesabstractEstimating questions ’ difficulty levels is an important task in community question answering (CQA) services. Previous stud-ies propose to solve this problem based on the question-user comparisons extract-ed from the question answering threads. However, they suffer from data sparseness problem as each question only gets a lim-ited number of comparisons. Moreover, they cannot handle newly posted question-s which get no comparisons. In this pa-per, we propose a novel question difficul-ty estimation approach called Regularized Competition Model (RCM), which natu-rally combines question-user comparisons and questions ’ textual descriptions into a unified framework. By incorporating tex-tual information, RCM can effectively deal with data sparseness problem. We further employ a K-Nearest Neighbor approach to estimate difficulty levels of newly post-ed questions, again by leveraging textu-al similarities. Experiments on two pub-licly available data sets show that for both well-resolved and newly-posted question-s, RCM performs the estimation task sig-nificantly better than existing methods, demonstrating the advantage of incorpo-rating textual information. More interest-ingly, we observe that RCMmight provide an automatic way to quantitatively mea-sure the knowledge levels of words. 1 Quan Wang 0002, Jing Liu 0022, Bin Wang 0004, Li Guo 0001 |
EMNLP | 4 |
| 2014 | A factor-searching-based multiple string matching algorithm for intrusion detectionabstractMultiple string matching plays a fundamental role in network intrusion detection systems. Automata-based multiple string matching algorithms like AC, SBDM and SBOM are widely used in practice, but the huge memory usage of automata prevents them from being applied to a large-scale pattern set. Meanwhile, poor cache locality of huge automata degrades the matching speed of algorithms. Here we propose a space-efficient multiple string matching algorithm BVM, which makes use of bit-vector and succinct hash table to replace the automata used in factor-searching-based algorithms. Space complexity of the proposed algorithm is O(rm2+ ΣpϵP|p|), that is more space-efficient than the classic automata-based algorithms. Experiments on datasets including Snort, ClamAV, URL blacklist and synthetic rules show that the proposed algorithm significantly reduces memory usage and still runs at a fast matching speed. Above all, BVM costs less than 0.75% of the memory usage of AC, and is capable of matching millions of patterns efficiently. Yanbing Liu 0007, Qingyun Liu 0001, Ping Liu 0001, Jianlong Tan, Li Guo 0001 |
ICC | 5 |
| 2014 | A probabilistic approach towards modeling email network with realistic featuresabstractEmail plays a very important role in our daily life. Much work have been put into practice on email network. Those studies mostly require real email network datasets and reliable models to analyze user information and understand the mechanisms of network evolution. However, much research work is constrained by the absence of real large-scale email datasets. Although email communication is ubiquitous, there are very few large-scale available email datasets satisfied different research purposes. Due to privacy policy and restricted permissions, it is arduous to collect a real large-scale email dataset in a short time. Various social network models are usually used to create synthetic email networks. However, these models focus on modeling several structural properties of network without considering user behaviour patterns. They are not appropriate to generate large-scale realistic synthetic email network datasets. Towards this end, we propose a probabilistic model by which we can construct large-scale synthetic email datasets with a small captured email log. What is more important is that the generated synthetic dataset matches real email network properties and individual communication patterns. Moreover, it has linear complexity, and can be paralleled easily. Experimental results on Enron dataset demonstrate the above benefits of our model. Quangang Li, Jinqiao Shi, Tingwen Liu, Li Guo 0001, Zhiguang Qin |
ICCCN | 4 |
| 2014 | Online Nonparametric Max-Margin Matrix Factorization for Collaborative PredictionabstractMax-margin matrix factorization (M3F) has been popularly applied to collaborative filtering for personalized recommendations. The nonparametric M3F model represents the latest progress of the M3F methods, which can auto-select the number of factors by using nonparametric techniques. However, existing non-parametric M3F methods assume a collection of user rating data can be fully obtained before training, and they are inapplicable for on-the-fly recommender systems where user rating data arrive continuously. In this paper, we present a new efficient online nonparametric 3F model for flexible recommendation. Specifically, we design an online nonparametric M3F model (OnM3F) based on the online Passive-Aggressive learning and solve the corresponding optimization problem by using the online stochastic gradient descent. Empirical studies on two large real-world data sets verify the effectiveness of the proposed method. Zhi Qiao 0005, Peng Zhang 0001, Wenjia Niu, Chuan Zhou 0001, Peng Wang 0028, Li Guo 0001 |
ICDM | 6 |
| 2014 | A Moving Target Framework to Improve Network Service AccessibilityabstractNowadays, the problem of Internet services accessibility has become a hot topic with more and more cyber attacks and censorship. This has prompted the rapid development of jamming-resistance infrastructure consisting of multiple dynamic access points such as proxies, anonymous communication nodes and covert communication nodes. However, the channel between user and access point has become an emerging attacking target for the adversary. Once the channel is detected and identified by the adversary, the channel will be interrupted. Though users can require new access points from the infrastructure and resume the communication to their destination, the service quality will be downgraded dramatically due to the time-consuming bootstrapping process. In this paper, a moving target framework is proposed to improve the network service accessibility, which combines both time-consuming bootstrapping and frequent and low-cost channel refreshing operations. The constant channel refreshing operations is imported to limit the adversary's ability of detecting and blocking communications, thus to make the lifespan of effective communication longer. With theoretical and simulating analysis, the optimized refreshing strategy is proposed, which can help improve service accessibility. Jinqiao Shi, Xiao Wang 0001, Binxing Fang, Li Guo 0001 |
NAS | 4 |
| 2014 | Forward Classification on Data Streams
Peng Wang 0028, Peng Zhang 0001, Yanan Cao 0001, Li Guo 0001, Binxing Fang |
PAKDD (1) | 4 |
| 2014 | Topic Block: Mining User Inner Interests for Text and Link Analysis in Social NetworksabstractText corpus and link network are interrelated data in social networks. Discovering the inner relationship between these two kinds of data can help better understand the evolution mechanism underneath social networks. Moreover, social networks exhibit unique characteristics such as sparse and noisy in both text and link data. Thus, it is imperative to combine both text and link data to complement and correct mining results. However, previous work did not explore a uniform generative model that can unveil their inner relationship probably because of the difficulty to harness the heterogenous data in social networks. To address this issue, in this paper we present a generative model Topic Block that clearly pinpoints the latent concept underlying the text corpus and link network, i.e., User inner interests. In our generative model, user inner interests guide the generation of the topic and community distributions underlying the text corpus and link data. We can infer the topic and community distributions based on the user inner interests through both content and topology information. Compared to existing popular models, our method experimentally outperforms on three real world social network data sets. Wenyu Zang, Chuan Zhou 0001, Xiao Wang 0001, Li Guo 0001 |
PDCAT | 4 |
| 2014 | Towards Improving Service Accessibility by Adaptive Resource Distribution Strategy
Jinqiao Shi, Xiao Wang 0001, Binxing Fang, Qingfeng Tan, Li Guo 0001 |
SecureComm (1) | 5 |
| 2014 | Automated Power Control for Virtualized Infrastructures
Yu Wen 0001, Weiping Wang 0005, Li Guo 0001, Dan Meng 0002 |
J. Comput. Sci. Technol. | 3 |
| 2014 | A block-aware hybrid data dissemination with hotspot elimination in wireless sensor network
Wenjia Niu, Gang Li 0009, Endong Tong, Quan Z. Sheng, Qian Li 0003, Yue Hu 0002, Athanasios V. Vasilakos, Li Guo 0001 |
J. Netw. Comput. Appl. | 8 |
| 2014 | Towards Fast and Optimal Grouping of Regular Expressions via DFA Size EstimationabstractRegular Expression (RegEx) matching, as a core operation in many network and security applications, is typically performed on Deterministic Finite Automata (DFA) to process packets at wire speed; however, DFA size is often exponential in the number of RegExes. RegEx grouping is the practical way to address DFA state explosion. Prior RegEx grouping algorithms are extremely slow and memory intensive. In this paper, we first propose DFAestimator, an algorithm that can quickly estimate DFA size for a given RegEx set without building the actual DFA. Second, we propose RegexGrouper, a RegEx grouping algorithm based on DFA size estimation. In terms of speed and memory consumption, our work is orders of magnitude more efficient than prior art because DFA size estimation is much faster and memory efficient than DFA construction. In terms of the resulting size sum of DFAs, our work is significantly more effective than prior art because we use a much finer grained quantification of the degree of interaction between two RegExes. For example, to divide the RegEx set of the L7-filter system into 7 groups, prior art uses 279.3 minutes and the resulting 7 DFAs have a total of 29047 states, whereas RegexGrouper uses 3.2 minutes and the resulting 7 DFAs have a total of 15578 states. Tingwen Liu, Alex X. Liu, Jinqiao Shi, Li Guo 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2014 | Contextual Query Expansion for Image RetrievalabstractIn this paper, we study the problem of image retrieval by introducing contextual query expansion to address the shortcomings of bag-of-words based frameworks: semantic gap of visual word quantization, and the efficiency and storage loss due to query expansion. Our method is built on common visual patterns (CVPs), which are the distinctive visual structures between two images and have rich contextual information. With CVPs, two contextual query expansions on visual word-level and image-level are explored, respectively. For visual word-level expansion, we find contextual synonymous visual words (CSVWs) and expand a word in the query image with its CSVWs to boost retrieval accuracy. CSVWs are the words that appear in the same CVPs and have same contextual meaning, i.e. similar spatial layout and geometric transformations. For image-level expansion, the database images that have the same CVPs are organized by linked list and the images that have the same CVPs as the query image, but not included in the results are automatically expanded. The main computation of these two expansions is carried out offline, and they can be integrated into the inverted file and efficiently applied to all images in the dataset. Experiments conducted on three reference datasets and a dataset of one million images demonstrate the effectiveness and efficiency of our method. Hongtao Xie 0001, Yongdong Zhang 0001, Jianlong Tan, Li Guo 0001, Jintao Li 0001 |
IEEE Trans. Multim. | 4 |
| 2013 | Design and Evaluation of Access Control Model Based on Classification of Users' Network Behaviors
Peipeng Liu, Jinqiao Shi, Li Guo 0001 |
APWeb | 5 |
| 2013 | Parallel auto-encoder for efficient outlier detectionabstractDetecting outliers from big data plays an important role in network security. Previous outlier detection algorithms are generally incapable of handling big data. In this paper we present an parallel outlier detection method for big data, based on a new parallel auto-encoder method. Specifically, we build a replicator model of the input data to obtain the representation of sample data. Then, the replicator model is used to measure the replicability of test data, where records having higher reconstruction errors are classified as outliers. Experimental results show the performance of the proposed parallel algorithm. Peng Zhang 0001, Yanan Cao 0001, Li Guo 0001 |
IEEE BigData | 4 |
| 2013 | Personalized influence maximization on social networksabstractIn this paper, we study a new problem on social network influence maximization. The problem is defined as, given a target user $w$, finding the top-k most influential nodes for the user. Different from existing influence maximization works which aim to find a small subset of nodes to maximize the spread of influence over the entire network (i.e., global optima), our problem aims to find a small subset of nodes which can maximize the influence spread to a given target user (i.e., local optima). The solution is critical for personalized services on social networks, where fully understanding of each specific user is essential. Although some global influence maximization models can be narrowed down as the solution, these methods often bias to the target node itself. To this end, in this paper we present a local influence maximization solution. We first provide a random function, with low variance guarantee, to randomly simulate the objective function of local influence maximization. Then, we present efficient algorithms with approximation guarantee. For online social network applications, we also present a scalable approximate algorithm by exploring the local cascade structure of the target user. We test the proposed algorithms on several real-world social networks. Experimental results validate the performance of the proposed algorithms. Peng Zhang 0001, Chuan Zhou 0001, Yanan Cao 0001, Li Guo 0001 |
CIKM | 5 |
| 2013 | An empirical analysis of family in the Tor networkabstractAs one of the most popular anonymous communication systems, Tor has become a research hotspot in this area. Recently, Tor nodes from Tor families (referred to as family nodes) have played an increasingly important role and caused significant influence on the Tor network. However, existing research about Tor mostly focuses on the entire Tor network without much consideration about the difference between family nodes and the others. In order to analyze family nodes' contribution to the entire Tor network as well as their influence, this paper distinguishes family nodes from the others, and gives an empirical analysis of family nodes based on the live Tor network data of 3 years. Results show that, family nodes compose a small but full functional subset of Tor nodes; and compared with the other Tor nodes, they can provide relatively stable and high-performance service to Tor users. Furthermore, family nodes naturally form a hot area in the Tor network, relaying increasingly high-density traffic through a small number of nodes. Compared with random node targets, selective attacks focusing on family nodes can cause serious availability downgrade of the Tor network with much lower cost. Xiao Wang 0001, Jinqiao Shi, Binxing Fang, Li Guo 0001 |
ICC | 4 |
| 2013 | UBLF: An Upper Bound Based Approach to Discover Influential Nodes in Social NetworksabstractInfluence maximization, defined as finding a small subset of nodes that maximizes spread of influence in social networks, is NP-hard under both Linear Threshold (LT) and Independent Cascade (IC) models, where a line of greedy/heuristic algorithms have been proposed. The simple greedy algorithm [14] achieves an approximation ratio of 1-1/e. The advanced CELF algorithm [16], by exploiting the sub modular property of the spread function, runs 700 times faster than the simple greedy algorithm on average. However, CELF is still inefficient [4], as the first iteration calls for N times of spread estimations (N is the number of nodes in networks), which is computationally expensive especially for large networks. To this end, in this paper we derive an upper bound function for the spread function. The bound can be used to reduce the number of Monte-Carlo simulation calls in greedy algorithms, especially in the first iteration of initialization. Based on the upper bound, we propose an efficient Upper Bound based Lazy Forward algorithm (UBLF in short), by incorporating the bound into the CELF algorithm. We test and compare our algorithm with prior algorithms on real-world data sets. Experimental results demonstrate that UBLF, compared with CELF, reduces more than 95% Monte-Carlo simulations and achieves at least 2-5 times speed-raising when the seed set is small. Chuan Zhou 0001, Peng Zhang 0001, Xingquan Zhu 0001, Li Guo 0001 |
ICDM | 5 |
| 2013 | Behavioral targeting with social regularizationabstractBehavioral targeting (BT) is a valuable tool for online advertising. In this paper, we study a new problem of incorporating social information into traditional behavior targeting models. Specifically, we present a social regularization based Poisson regression framework for behavior targeting. Based on the observation that social information can be diverse and competing, we furthermore present two specific social regularization terms: the average-based social regularization term and the individual-based social regularization term. To validate the effectiveness of the proposed models, we use the KDDCUP'12 behavior targeting data, issued by the Tecent company in China, as the test bed. The results demonstrate that the proposed models, by incorporating additional social network information, can achieve at least 5% improvement compared to the traditional Poisson regression based model from the CTR lift viewpoint, especially when the historical behavior data is sparse and insufficient. Yanmin Shang, Peng Zhang 0001, Yanan Cao 0001, Li Guo 0001 |
ISI | 4 |
| 2013 | Discovering Semantics from Multiple Correlated Time Series Stream
Zhi Qiao 0005, Guangyan Huang, Jing He 0004, Peng Zhang 0001, Li Guo 0001, Jie Cao 0001, Yanchun Zhang |
PAKDD (2) | 5 |
| 2013 | Research of Intrusion Detection System on AndroidabstractIn this paper, we proposed an intrusion detection system for detecting anomaly on Android smartphones. The intrusion detection system continuously monitors and collects the information of smartphone under normal conditions and attack state. It extracts various features obtained from the Android system, such as the network traffic of smartphones, battery consumption, CPU usage, the amount of running processes and so on. Then, it applies Bayes Classifying Algorithm to determine whether there is an invasion. In order to further analyze the Android system abnormalities and locate malicious software, along with system state monitoring the intrusion detection system monitors the process and network flow of the smartphone. Finally, experiments on the system which was designed in this paper have been carried out. Empirical results suggest that the proposed intrusion detection system is effective in detecting anomaly on Android smartphones. Fangfang Yuan, Lidong Zhai, Yanan Cao 0001, Li Guo 0001 |
SERVICES | 4 |
| 2013 | Robust common visual pattern discovery using graph matching
Hongtao Xie 0001, Yongdong Zhang 0001, Ke Gao 0012, Sheng Tang, Kefu Xu, Li Guo 0001, Jintao Li 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2012 | A Prefiltering Approach to Regular Expression Matching for Network Security Systems
Tingwen Liu, Alex X. Liu, Li Guo 0001, Binxing Fang |
ACNS | 4 |
| 2012 | Efficient Behavior Targeting Using SVM Ensemble IndexingabstractBehavior targeting (BT) is a promising tool for online advertising. The state-of-the-art BT methods, which are mainly based on regression models, have two limitations. First, learning regression models for behavior targeting is difficult since user clicks are typically several orders of magnitude fewer than views. Second, the user interests are not fixed, but often transient and influenced by media and pop culture. In this paper, we propose to formulate behavior targeting as a classification problem. Specifically, we propose to use an SVM ensemble for behavior prediction. The challenge of using ensemble SVM for BT stems from the computational complexity (it takes 53 minutes in our experiments to predict behavior for 32 million users, which is inadequate for online application). To this end, we propose a fast ensemble SVM prediction framework, which builds an indexing structure for SVM ensemble to achieve sub-linear prediction time complexity. Experimental results on real-world large scale behavior targeting data demonstrate that the proposed method is efficient and outperforms existing linear regression based BT models. Jun Li 0016, Peng Zhang 0001, Yanan Cao 0001, Ping Liu 0001, Li Guo 0001 |
ICDM | 5 |
| 2012 | A semantics aware approach to automated reverse engineering unknown protocolsabstractExtracting the protocol message format specifications of unknown applications from network traces is important for a variety of applications such as application protocol parsing, vulnerability discovery, and system integration. In this paper, we propose ProDecoder, a network trace based protocol message format inference system that exploits the semantics of protocol messages without the executable code of application protocols. ProDecoder is based on the key insight that the n-grams of protocol traces exhibit highly skewed frequency distribution that can be leveraged for accurate protocol message format inference. In ProDecoder, we first discover the latent relationship among n-grams by first grouping protocol messages with the same semantics and then inferring message formats by keyword based clustering and cluster sequence alignment. We implemented and evaluated ProDecoder to infer message format specifications of SMB (a binary protocol) and SMTP (a textual protocol). Our experimental results show that ProDecoder accurately parses and infers SMB protocol with 100% precision and recall. For SMTP, ProDecoder achieves approximately 95% precision and recall. Yipeng Wang 0001, Xiao-chun Yun, Zubair Shafiq, Alex X. Liu, Danfeng Yao, Yongzheng Zhang 0002, Li Guo 0001 |
ICNP | 9 |
| 2012 | On Accuracy of Early Traffic ClassificationabstractThe widely employment of traffic encryption, tunneling and other protection/obfuscation mechanisms in modern network applications, prompts the emergence of traffic behavior (i.e., packet direction pattern, size, and inter-arrival time) based classification approaches. Some proposals even demonstrate its potential for on-line early traffic classification - using the first 4-6 data packets at the beginning of a TCP connection to identify the corresponding application. Nevertheless, the related accuracy issues on early classification are still unclear when forged packets exist. The performance of such mechanism under malicious environment, where sophisticated forged data packets injection techniques are presented, had not been addressed. This work aims to touch the above issues, especially when forged packets are inserted before actual application transaction started. Our contributions are two-folded: (1) confirm the discrimination power of early classification as revealed by previous study; (2) explore it's accuracy vulnerability to forged packets the experiments on both simulated and real SSH tunnel traces show the accuracy declines when forged packets are injected. Our findings show that the intellective early classification methods still deserve further investigation before actual deployment. Buyun Qu, Li Guo 0001, Dan Meng 0002 |
NAS | 3 |
| 2012 | A Task-Based Model for the Lifespan of Peer-to-Peer Swarms
Ting He 0001, Alex X. Liu, Li Guo 0001, Binxing Fang |
Networking (2) | 5 |
| 2012 | EFA for Efficient Regular Expression Matching in NIDS (Poster Abstract)
Dengke Qiao, Tingwen Liu, Li Guo 0001 |
RAID | 4 |
| 2012 | Mining Multi-Label Data Streams Using Ensemble-Based Active LearningabstractData stream classification has drawn increasing attention from the data mining community in recent years, where a large number of stream classification models were proposed. However, most existing models were merely focused on mining from single-label data streams. Mining from multi-label data streams has not been fully addressed yet. On the other hand, although some recent work touched the multi-label stream mining problem, they never consider the expensive labeling cost issue, preventing them from real-world applications. To this end, we study, in this paper, a challenging problem that mining from multi-label data streams with limited labeling resource. Specifically, we propose an ensemble-based active learning framework to handle the large volume of stream data, expensive labeling cost and concept drifting problems on multi-label data streams. Experiments on both synthetic and real world data sets demonstrate the performance of the proposed method. Peng Wang 0028, Peng Zhang 0001, Li Guo 0001 |
SDM | 3 |
| 2012 | SDFA: Series DFA for Memory-Efficient Regular Expression Matching
Tingwen Liu, Li Guo 0001, Binxing Fang |
CIAA | 3 |
| 2012 | A novel logic-based automatic approach to constructing compliant security policies
Yibao Bao, Lihua Yin, Binxing Fang, Li Guo 0001 |
Sci. China Inf. Sci. | 4 |
| 2012 | A framework for application-driven classification of data streams
Peng Zhang 0001, Byron J. Gao, Ping Liu 0001, Yong Shi 0001, Li Guo 0001 |
Neurocomputing | 5 |
| 2011 | Inferring Protocol State Machine from Network Traces: A Probabilistic Approach
Yipeng Wang 0001, Danfeng Yao, Buyun Qu, Li Guo 0001 |
ACNS | 5 |
| 2011 | Performance evaluation of Xunlei peer-to-peer network: A measurement studyabstractXunlei is a P2P file sharing application that is popular in China. The performance of previous P2P applications is limited by the selfishness of peers and the widely use of firewall. Xunlei applies implicit uploading strategy and firewall bypassing technologies to conquer above drawbacks. To evaluate the effects of above solutions, we perform a series of measurements on Xunlei and BitTorrent networks. Our study provides three new findings by comparing the situation in the two networks. (1)Compared with the free control strategy in BitTorrent, the implicit uploading strategy improves not only the number of active peers but also the seed ratio. Benefit from the strategy, Xunlei has much better downloading performance in our measurement. (2)The strategy extends the seed service time that is related to swarm lifespan. According to our result, Xunlei swarm lives longer than BitTorrent one. (3)The connectivity influences the performance greatly. The result of our measurement shows that only 5% peers can be connected directly. Our theoretical analysis and measurement results show that better connectivity also leads to higher downloading speed and longer lifespan. Yipeng Wang 0001, Li Guo 0001, Binxing Fang |
CCNC | 4 |
| 2011 | Mining frequent patterns across multiple data streamsabstractMining frequent patterns from data streams has drawn increasing attention in recent years. However, previous mining algorithms were all focused on a single data stream. In many emerging applications, it is of critical importance to combine multiple data streams for analysis. For example, in real-time news topic analysis, it is necessary to combine multiple news report streams from dierent media sources to discover collaborative frequent patterns which are reported frequently in all media, and comparative frequent patterns which are reported more frequently in a media than others. To address this problem, we propose a novel frequent pattern mining algorithm Hybrid-Streaming, H-Stream for short. H-Stream builds a new Hybrid-Frequent tree to maintain historical frequent and potential frequent itemsets from all data streams, and incrementally updates these itemsets for efficient collaborative and comparative pattern mining. Theoretical and empirical studies demonstrate the utility of the proposed method. Peng Zhang 0001, Jianlong Tan, Li Guo 0001 |
CIKM | 4 |
| 2011 | Continuous data stream query in the cloudabstractCloud computing represents one of the most important research directions for modern computing systems. Existing research efforts on Cloud computing were all focused on designing advanced storage and query techniques for static data. None of them consider the problem that data in a Cloud may appear as continuous and rapid data streams. To address this problem, in this paper we propose a new LCN-Index framework to handle continuous data stream queries in the Cloud. LCN-Index uses the Map-Reduce computing paradigm to process all the queries. In the Mapping stage, it divides all the queries into a batch of predicate sets which are then deployed onto mapping nodes using interval predicate index. In the reducing stage, it merges results from the mapping nodes using multi attribute hash index. In so doing, a data stream can be efficiently evaluated by traversing through the LCN-Index framework. Experiments demonstrate the utility of the proposed method. Jun Li 0016, Peng Zhang 0001, Jianlong Tan, Ping Liu 0001, Li Guo 0001 |
CIKM | 5 |
| 2011 | Enabling Fast Lazy Learning for Data StreamsabstractLazy learning, such as k-nearest neighbor learning, has been widely applied to many applications. Known for well capturing data locality, lazy learning can be advantageous for highly dynamic and complex learning environments such as data streams. Yet its high memory consumption and low prediction efficiency have made it less favorable for stream oriented applications. Specifically, traditional lazy learning stores all the training data and the inductive process is deferred until a query appears, whereas in stream applications, data records flow continuously in large volumes and the prediction of class labels needs to be made in a timely manner. In this paper, we provide a systematic solution that overcomes the memory and efficiency limitations and enables fast lazy learning for concept drifting data streams. In particular, we propose a novel Lazy-tree (Ltree for short) indexing structure that dynamically maintains compact high-level summaries of historical stream records. L-trees are M-Tree [5] like, height-balanced, and can help achieve great memory consumption reduction and sub-linear time complexity for prediction. Moreover, L-trees continuously absorb new stream records and discard outdated ones, so they can naturally adapt to the dynamically changing concepts in data streams for accurate prediction. Extensive experiments on real-world and synthetic data streams demonstrate the performance of our approach. Peng Zhang 0001, Byron J. Gao, Xingquan Zhu 0001, Li Guo 0001 |
ICDM | 4 |
| 2011 | A Covert Communication Method Based on User-Generated Content SitesabstractWith the worldwide increasing of Internet censorship, censorship-resistance technology has attracted more and more attentions, some famous systems, such as Tor and JAP, have been deployed to provide public service for censorship-resistance. However, these systems all rely on dedicated infrastructure and entry points for service accessibility. The network infrastructure and entry points may become the target of censorship attack. In this paper, a UGC-based method is proposed (called user-generated content based covert communication, UGC3) for covert communication in a friends-to-friends (F2F) manner. It uses existing infrastructures (i.e., UGC sites ) to form a fully distributed overlay network. An efficient resource discovery algorithm is proposed to negotiate the rendezvous point. Analysis shows that this method is able to circumvent internet censorship with user repudiation and fault tolerance. Qingfeng Tan, Peipeng Liu, Jinqiao Shi, Xiao Wang 0001, Li Guo 0001 |
ICTAI | 5 |
| 2011 | An efficient regular expressions compression algorithm from a new perspectiveabstractDeep packet inspection plays a increasingly important role in network security devices and applications, which use more regular expressions to depict patterns. DFA engine is usually used as a classical representation for regular expressions to perform pattern matching, because it only need O(1) time to process one input character. However, DFAs of regular expression sets require large amount of memory, which limits the practical application of regular expressions in high-speed networks. Some compression algorithms have been proposed to address this issue in recent literatures. In this paper, we reconsider this problem from a new perspective, namely observing the characteristic of transition distribution inside each state, which is different from previous algorithms that observe transition characteristic among states. Furthermore, we introduce a new compression algorithm which can reduce 95% memory usage of DFA stably without significant impact on matching speed. Moreover, our work is orthogonal to previous compression algorithms, such as D2FA, δFA. Our experiment results show that applying our work to them will have several times memory reduction, and matching speed of up to dozens of times comparing with original δFA in software implementation. Tingwen Liu, Yifu Yang, Yanbing Liu 0007, Li Guo 0001 |
INFOCOM | 5 |
| 2011 | Revisiting the swarm evolution: A long term perspectiveabstractAlthough peer-to-peer system is scalable for content distribution, its lifespan is shorter than traditional systems. More and more researchers pay close attention to swarm lifespan and propose specialized methods to extend it. Bundling technique and SRE (Share Ratio Enforcement) mechanism in Private Tracker system are two successful enhancements which are proved by practical systems. However, they have side effects to system users. One of the reason is that the enhancements do not take effect on the key factors. So there is a question, what factors influence lifespan? A model which can answer this question is needed to help system designers. Facing the challenges, we measure a popular BitTorrent system with a macro view which is different from previous works. Based on our measurement, we model the long term evolution to find out factors that influence swarm lifespan. In order to unveil the trend of a swarm, we propose a half-life based metric. According to the metric, our model fits the real swarms more closely. By the help of our model, we analyze the strengths and shortcomings of bundling technique and SRE incentives. Li Guo 0001, Binxing Fang |
ISCC | 3 |
| 2011 | Accelerating DFA Construction by Hierarchical MergingabstractRegular expression matching is widely used in many network applications to analyze suspicious traffic against predefined signatures, and to discover anomalous events. Deterministic Finite Automaton (DFA), which recognizes a set of regular expressions, is the basic data structure to scan input traffic byte by byte. Though DFA meets the requirement of real-time processing of network traffic, constructing a combined DFA for a set of regular expression signatures is very time-consuming, especially when the signature set is large. To attack this problem, we propose new strategies to accelerate DFA construction. The basic idea of our method is to construct the combined DFA by hierarchical merging of the DFAs of each single regular expression. Our method runs in $O(|Q| |\Sigma|\ln n)$ time, which is substantially superior to the time complexity $O(|Q| |\Sigma|(\overset{n}{\underset{i=1}{\sum}}|Q_i|)^2)$of classical subset construction algorithm\cite{Aho1986}. Experiment on real signatures from open-source systems, such as L7-filter, BRO and SNORT, demonstrates that our method performs 45 times faster than the subset construction algorithm on average. Yanbing Liu 0007, Li Guo 0001, Muyi Guo, Ping Liu 0001 |
ISPA | 2 |
| 2011 | Enabling fast prediction for ensemble models on data streamsabstractEnsemble learning has become a common tool for data stream classification, being able to handle large volumes of stream data and concept drifting. Previous studies focus on building accurate prediction models from stream data. However, a linear scan of a large number of base classifiers in the ensemble during prediction incurs significant costs in response time, preventing ensemble learning from being practical for many real world time-critical data stream applications, such as Web traffic stream monitoring, spam detection, and intrusion detection. In these applications, data streams usually arrive at a speed of GB/second, and it is necessary to classify each stream record in a timely manner. To address this problem, we propose a novel Ensemble-tree (E-tree for short) indexing structure to organize all base classifiers in an ensemble for fast prediction. On one hand, E-trees treat ensembles as spatial databases and employ an R-tree like height-balanced structure to reduce the expected prediction time from linear to sub-linear complexity. On the other hand, E-trees can automatically update themselves by continuously integrating new classifiers and discarding outdated ones, well adapting to new trends and patterns underneath data streams. Experiments on both synthetic and real-world data streams demonstrate the performance of our approach. Peng Zhang 0001, Jun Li 0016, Peng Wang 0028, Byron J. Gao, Xingquan Zhu 0001, Li Guo 0001 |
KDD | 6 |
| 2011 | Using Entropy to Classify Traffic More DeeplyabstractThe network community always pays its attention to find better methods for traffic classification, which is crucial for Internet Service Providers (ISPs) to provide better QoS for users. Prior works on traffic classification mainly focus their attentions on dividing Internet traffic into different categories based on application layer protocols (such as HTTP, Bit Torrent etc.). Making traffic classification from another point of view, we divide Internet traffic into different content types. Our technology is an attempt to solve the classification problem of network traffic, which contains unknown and proprietary protocols (i.e., no publicly available protocol specification). In this paper, we design a classifier which can distinguish Internet traffic into different content types using machine learning techniques. Features of our classifier are entropy of consecutive bytes and frequencies of characters. Our method is capable of classifying real-world traces into different content types (including Text, Picture, Audio, Video, Compressed, Base 64-encoded image, Base 64-encoded text and Encrypted). The chief features of our classifier are small computing space (about 1K Bytes) and high classification accuracy (about 81%). Yipeng Wang 0001, Li Guo 0001 |
NAS | 3 |
| 2011 | Understanding Long-Term Evolution and Lifespan in Peer-to-Peer SystemsabstractAlthough peer-to-peer system is scalable for content distribution, its lifespan is shorter than traditional systems. More and more researchers began to pay close attention to lifespan and proposed specialized methods to extend it. However, what factors influence lifespan and how the effects are? To answer these questions, we analyze the long-term characteristics of swarms by our sampling measurement. The results show that swarms have different behaviors comparing with previous studies that focused on flash-crowd. Based on these new findings, we propose a long term model to depict the swarm evolution in large time scale. According to our model, the attenuation parameter of arrive rate and the average task length influence lifespan linearly while the initial arrival rate and peer availability have logarithmic influence. In order to validate our model, we compare them with real swarm and simulations. Our experiments show that the model captures the evolution closely and our results are valid in different situations. Li Guo 0001 |
NAS | 3 |
| 2011 | Biprominer: Automatic Mining of Binary Protocol FeaturesabstractApplication-level protocol specifications are helpful for network security management, including intrusion detection and intrusion prevention which rely on monitoring technologies such as deep packet inspection. Moreover, detailed knowledge of protocol specifications is also an effective way of detecting malicious code. However, current methods for obtaining unknown and proprietary protocol message formats (i.e., no publicly available protocol specification), especially binary protocols, highly rely on manual operations, such as reverse engineering which is time-consuming and laborious. In this paper, we propose Biprominer, a tool that can automatically extract binary protocol message formats of an application from its real-world network trace. In addition, we present a transition probability model for a better description of the protocol. The chief feature of Biprominer is that it does not need to have any priori knowledge of protocol formats, because Biprominer is based on the statistical nature of the protocol format. We evaluate the efficacy of Biprominer over three binary protocols, with an average precision more than 99% and a recall better than 96.7%. Yipeng Wang 0001, Xingjian Li 0002, Jiao Meng, Li Guo 0001 |
PDCAT | 6 |
| 2011 | XunleiProbe: A Sensitive and Accurate Probing on a Large-Scale P2SP SystemabstractXunlei [1] is a new P2P content distribution system which is popular in China. It composes traditional HTTP/FTP downloading and P2P content distribution features which attract many people including researchers. Xunlei's network is Bit Torrent-like and the measurement is more difficult than other P2P networks [2]. There are many constrains on Xunlei tracker, so we can not obtain information from tracker easily. Besides this, the measurement will encounter some challenges that skew the results. Most of previous works on Bit Torrent system are based on tracker logs. However, we can not obtain tracker logs of Xunlei. As far as we know, there is no proposal about precise and detailed measurement method that probes the network directly in this area. Face to the challenges, we analyze the constrains that appear in most Bit Torrent-like systems and propose a stratified random selection model to describe the behavior of tracker. Based on the model, we design a measurement tool called XunleiProbe. With the help of our solutions that increase the accuracy of our results, we measure a popular swarm for about 22 hours. The results show that the average peer coverage of our tool can reach about 93%. Li Guo 0001, Binxing Fang |
PDCAT | 3 |
| 2011 | Giant complete automaton for uncertain multiple string matching and its high speed construction algorithm
Yue Hu 0002, Qingshi Gao, Li Guo 0001, PeiFeng Wang |
Sci. China Inf. Sci. | 3 |
| 2011 | Robust ensemble learning for mining noisy data streams
Peng Zhang 0001, Xingquan Zhu 0001, Yong Shi 0001, Li Guo 0001, Xindong Wu 0001 |
Decis. Support Syst. | 4 |
| 2010 | SKIF: a data imputation framework for concept drifting data streamsabstractMissing data commonly occurs in many applications. While many data imputation methods exist to handle the missing data problem for large scale databases, when applied to concept drifting data streams, these methods face some common difficulties. First, due to large and continuous data volumes, we are unable to maintain all stream records to form a candidate pool and estimate missing values, as most existing methods commonly do. Second, even if we could maintain all complete stream records using a summary structure, the concept drifting problem would make some information obsolete, and thus deteriorate the imputation accuracy. Third, in data streams, it is necessary to develop a fast yet accurate algorithm to find the most similar data for imputation. Fourth, due to the dynamic and sophisticated data collection environments, the missing rate of most stream data may be much higher than that in generic static databases, so the imputation method should be able to accommodate high missing rate in the data. To tackle these challenges, we propose, in this paper, a Streaming k-Nearest-Neighbors Imputation Framework (SKIF) for concept drifting data streams. To handle concept drifting and large volume problems in data streams, SKIF first summarizes historical complete records in some micro-resources (which are high-level statistical data structures), and maintains these micro-resources in a candidate pool as benchmark data. After that, SKIF employs a novel hybrid-kNN imputation procedure, which uses a hybrid similarity search mechanism, to find the most similar micro-resources from the large scale candidate pool efficiently. Experimental results demonstrate the effectiveness of the proposed SKIF framework for data stream imputation tasks. Peng Zhang 0001, Xingquan Zhu 0001, Jianlong Tan, Li Guo 0001 |
CIKM | 4 |
| 2010 | Classifier and Cluster Ensembles for Mining Concept Drifting Data StreamsabstractEnsemble learning is a commonly used tool for building prediction models from data streams, due to its intrinsic merits of handling large volumes stream data. Despite of its extraordinary successes in stream data mining, existing ensemble models, in stream data environments, mainly fall into the ensemble classifiers category, without realizing that building classifiers requires labor intensive labeling process, and it is often the case that we may have a small number of labeled samples to train a few classifiers, but a large number of unlabeled samples are available to build clusters from data streams. Accordingly, in this paper, we propose a new ensemble model which combines both classifiers and clusters together for mining data streams. We argue that the main challenges of this new ensemble model include (1) clusters formulated from data streams only carry cluster IDs, with no genuine class label information, and (2) concept drifting underlying data streams makes it even harder to combine clusters and classifiers into one ensemble framework. To handle challenge (1), we present a label propagation method to infer each cluster's class label by making full use of both class label information from classifiers, and internal structure information from clusters. To handle challenge (2), we present a new weighting schema to weight all base models according to their consistencies with the up-to-date base model. As a result, all classifiers and clusters can be combined together, through a weighted average mechanism, for prediction. Experiments on real-world data streams demonstrate that our method outperforms simple classifier ensemble and cluster ensemble for stream data mining. Peng Zhang 0001, Xingquan Zhu 0001, Jianlong Tan, Li Guo 0001 |
ICDM | 4 |
| 2010 | Learning from Multiple Related Data Streams with Asynchronous Flowing SpeedsabstractRelated data streams refer to data streams that can be joined together by matching their join attributes. Existing research on learning from related data streams is based on an assumption that all streams arrive at a central processing unit in a synchronous way, such that in an arbitrary sliding window, all tuples of the streams can be perfectly joined together. This assumption, however, does not hold when related data streams are generated or transferred at different speeds, and thus may arrive in the central processing unit in an asynchronous manner. In this paper, we argue that for asynchronous data streams, there exist a small portion of perfectly joined examples (i.e., complete examples) and a large portion of partially joined examples (i.e., incomplete examples). Accordingly, we present a new Learning from Complete and Fixed Examples (LCFE) framework that can fix incomplete examples to boost the learning. Experiments on both synthetic and real-world data streams demonstrate that LCFE is able to achieve a higher prediction accuracy for learning from related data streams than other simple solutions can offer. Zhi Qiao 0005, Peng Zhang 0001, Jing He 0004, Jinghua Yan, Li Guo 0001 |
ICMLA | 5 |
| 2010 | Fast and Memory-Efficient Traffic Classification with Deep Packet Inspection in CMP ArchitectureabstractTraffic classification is important to many network applications, such as network monitoring. The classic way to identify flows, e.g., examining the port numbers in the packet headers, becomes ineffective. In this context, deep packet inspection technology, which does not only inspect the packet headers but also the packet payloads, plays a more important role in traffic classification. Meanwhile regular expressions are replacing strings to represent patterns because of their expressive power, simplicity and flexibility. However, regular expressions mathcing technique causes a high memory usage and processing cost, which result in low throughout. In this paper, we analyze the application-level protocol distribution of network traffic and conclude its characteristic. Furthermore, we design a fast and memory-efficient system of a two-layer architecture for traffic classification with the help of regular expressions in multi-core architecture, which is different from previous one-layer architecture. In order to reduce the memory usage of DFA, we use a compression algorithm called CSCA to perform regular expressions matching, which can reduce 95% memory usage of DFA. We also introduce some optimizations to accelerate the matching speed. We use real-world traffic and all L7-filter protocol patterns to make our experiments, and the results show that the system achieves at Gbps level throughout in 4-cores Servers. Tingwen Liu, Li Guo 0001 |
NAS | 3 |
| 2010 | Inferring Protocol State Machine from Real-World Trace
Yipeng Wang 0001, Li Guo 0001 |
RAID | 3 |
| 2010 | Compressing Regular Expressions' DFA Table by Matrix Decomposition
Yanbing Liu 0007, Li Guo 0001, Ping Liu 0001, Jianlong Tan |
CIAA | 2 |
| 2009 | Optimizing Network Anomaly Detection Scheme Using Instance Selection MechanismabstractNetwork anomaly detection is a classically difficult research topic in intrusion detection. However, existing research has been solely focused on the detection algorithm. An important issue that has not been well studied so far is the selection of normal training data for network anomaly detection algorithm, which is highly related to the detection performance and computational complexity. Based on our previous proposed TCM-KNN (Transductive Confidence Machines for K-Nearest Neighbors) anomaly detection method, which can detect anomalies with high detection rate and low false positive rate, we develop an instance selection mechanism for TCM-KNN based on EFCM (Extended Fuzzy C-Means) clustering algorithm in this paper, aiming at limiting the size of training dataset, thus reducing the computational cost of TCM-KNN and boosting its detection performance. We report the experimental results over real network traffic. The results demonstrate the instance selection method presented in this paper is effective for TCM-KNN and thus optimizing it as an effectively lightweight network anomaly detection scheme. Yang Li 0002, Tianbo Lu, Li Guo 0001, Zhihong Tian 0001 |
GLOBECOM | 3 |
| 2009 | Load Balancing for Flow-Based Parallel Processing Systems in CMP ArchitectureabstractLoad balancing is critical to the performance of parallel processing systems. It is more difficult for network systems such as NIDS and Web Servers, because they must preserve flow order. But traditional flow-based load balancing schemes of network parallel processing systems, such as LLF, cost much resource and introduce lots of communication overhead. With the rapid popularization of multi-core system, it is a good choice to apply NIDS in CMP architecture to achieve higher performance. Some companies have taken the first step. Their scheduling algorithm operates at a custom NIC based on FPGA technology. LLF algorithm can't be used in such environment because it needs much more memory than NIC owns. In this paper, we propose a scheduling scheme that re-maps the new arrival flows when the system is unbalanced. To make effective adjustments we design a new triggering policy based on waiting lengths and their difference. Compared with LLF, our algorithm costs about 5% memory to get the same performance. Tingwen Liu, Li Guo 0001 |
GLOBECOM | 4 |
| 2009 | Mining Data Streams with Labeled and Unlabeled Training ExamplesabstractIn this paper, we propose a framework to build prediction models from data streams which contain both labeled and unlabeled examples. We argue that due to the increasing data collection ability but limited resources for labeling, stream data collected at hand may only have a small number of labeled examples, whereas a large portion of data remain unlabeled but can be beneficial for learning. Unleashing the full potential of the unlabeled instances for stream data mining is, however, a significant challenge, consider that even fully labeled data streams may suffer from the concept drifting, and inappropriate uses of the unlabeled samples may only make the problem even worse. To build prediction models, we first categorize the stream data into four different categories, each of which corresponds to the situation where concept drifting may or may not exist in the labeled and unlabeled data. After that, we propose a relational k-means based transfer semi-supervised SVM learning framework (RK-TS3VM), which intends to leverage labeled and unlabeled samples to build prediction models. Experimental results and comparisons on both synthetic and real-world data streams demonstrate that the proposed framework is able to help build prediction models more accurate than other simple approaches can offer. Peng Zhang 0001, Xingquan Zhu 0001, Li Guo 0001 |
ICDM | 3 |
| 2009 | Simulation Analysis of Probabilistic Timing Covert ChannelsabstractIt is very important to analyze the bandwidth and transmission error rate in the study of probabilistic timing covert channels. For the purpose, a simulation system of probabilistic timing covert channels has been set up in the paper. The simulation results show that (1) the bandwidth and the transmission error rate of probabilistic timing covert channels are closely related to the hardware/software environment, probability factor, time factor and/or coding methods as well as scheduling times; (2) the approximate transmission error rate can be measured with the central limit theorem; (3) it is not accurate to estimate the amount of information leakage based on weak probabilistic bisimulation; and (4) in probabilistic timing covert channels, there exist some characteristics which are different from non-deterministic covert channels. Yunchuan Guo, Lihua Yin, Yuan Zhou 0008, Chao Li 0027, Li Guo 0001 |
NAS | 5 |
| 2009 | An Experimental Study on Instance Selection Schemes for Efficient Network Anomaly Detection
Yang Li 0002, Li Guo 0001, Binxing Fang, Xiangtao Liu |
RAID | 2 |
| 2009 | Towards lightweight and efficient DDOS attacks detection for web serverabstractIn this poster, based on our previous work in building a lightweight DDoS (Distributed Denial-of-Services) attacks detection mechanism for web server using TCM-KNN (Transductive Confidence Machines for K-Nearest Neighbors) and genetic algorithm based instance selection methods, we further propose a more efficient and effective instance selection method, named E-FCM (Extend Fuzzy C-Means). By using this method, we can obtain much cheaper training time for TCM-KNN while ensuring high detection performance. Therefore, the optimized mechanism is more suitable for lightweight DDoS attacks detection in real network environment. Yang Li 0002, Tianbo Lu, Li Guo 0001, Zhihong Tian 0001, Qin-Wu Nie |
WWW | 3 |
| 2008 | UBSF: A novel online URL-Based Spam FilterabstractSpam fighting is a classic puzzle in network security. In the past decades, many filtering spam solutions have been proposed. However, currently the conventional techniques still suffer from high false positives and false negatives, especially the former which usually cannot be accepted by the end users. This paper proposes a novel online URL-based spam filter (UBSF) on the basis of analyses over the conventional especially the URL-based anti-spam techniques. UBSF identifies spam by comparing the similarity of the extracted URLs (universal resource locator) from them with the URLs from our user-oriented standard URL Library (SUL). Experimental evaluations both from the contrast experiment and the prototype based on UBSF demonstrate it can significantly raise the filtering accuracy, effectively reduce false positives and can be applied to online process and heavy traffic environment by reducing the computational cost than the state-of-the-art techniques. Yang Li 0002, Binxing Fang, Li Guo 0001, Zhihong Tian 0001, Yongzheng Zhang 0002, Zhi-Gang Wu |
ISCC | 3 |
| 2008 | Incremental web page template detectionabstractMost template detection methods process web pages in batches that a newly crawled page can not be processed until enough pages have been collected. This results in large storage consumption and a huge delay of data refreshing. In this paper, we present an incremental framework to detect templates in which a page is processed as soon as it has been crawled. In this framework, we don't need to cache any web page. Experiments show that our framework consumes less than 7% storage than traditional methods. And also the speed of data refreshing is accelerated because of the incremental manner. Yu Wang 0009, Binxing Fang, Xueqi Cheng 0001, Li Guo 0001 |
WWW | 4 |
| 2008 | A lightweight web server anomaly detection method based on transductive scheme and genetic algorithms
Yang Li 0002, Li Guo 0001, Zhihong Tian 0001, Tianbo Lu |
Comput. Commun. | 2 |
| 2007 | Network anomaly detection based on TCM-KNN algorithmabstractIntrusion detection is a critical component of secure information systems. Network anomaly detection has been an active and difficult research topic in the field of Intrusion Detection for many years. However, it still has some problems unresolved. They include high false alarm rate, difficulties in obtaining exactly clean data for the modeling of normal patterns and the deterioration of detection rate because of some "noisy" data in the training set. In this paper, we propose a novel network anomaly detection method based on improved TCM-KNN (Transductive Confidence Machines for K-Nearest Neighbors) machine learning algorithm. A series of experimental results on the well-known KDD Cup 1999 dataset demonstrate it can effectively detect anomalies with high true positive rate, low false positive rate and high confidence than the state-of-the-art anomaly detection methods. In addition, even interfered by "noisy" data (unclean data), the proposed method is robust and effective. Moreover, it still retains good detection performance after employing feature selection aiming at avoiding the "curse of dimensionality". Yang Li 0002, Binxing Fang, Li Guo 0001, You Chen 0005 |
AsiaCCS | 3 |
| 2007 | Parallelizing Protocol Processing on SMT Processor Efficiently: A FSM Decomposition ApproachabstractWith the increase of network bandwidth, high performance protocol processing plays more and more important role in high speed network security. Recent studies show that current computer architecture advances and CPU performance improvements have limited impact on network protocol processing performance. Some studies find that in real SMT processor like Intel Xeon processor with hyper-threadings, the sharing resources (like cache) contention between threads can hurt the processing performance of network applications like servers or IDS. How to make protocol processing cope with the advances in computer architecture has been widely studied. In this paper, we put our focus on the processing performance of TCP automata phases, using execution based simulations to model the relationship between each phase performance and cache size, and then measuring the cache contention between threads. We find (1) the load/store units can be the bottleneck of protocol processing; and (2) in connection establishing phase of TCP processing, cache contention between threads is more aggressive than any other phase. We also suggest a FSM decomposition based parallel processing approach to use sharing cache of SMT processors effectively. Li Guo 0001, Binxing Fang, Xiaojun Chen 0004 |
IPCCC | 2 |
| 2007 | LASF: A Flow Scheduling Policy in Stateful Packet Inspection SystemsabstractCurrent increase in network bandwidth raised an aggressive challenge in network security, and stateful packet inspection based security systems is playing a more and more important role. Recent advances in scheduling theory show that it is possible to reduce the expected mean response time of a queuing system, simply by changing the order in which we schedule the requests according to the job size, which is so called size-based scheduling policy. In this paper, we start by an analysis of connection sojourn time distribution of network traffic. Based on this analysis, first we design a two level session table in order to avoid session table explosion. Then we propose a connection scheduling policy in stateful packet inspection systems called LASF (Least Attained Sojourn First). We show that our policy can improve mean response time and flow throughput especially when system is overloaded. Finally we assess the costs of LASF in terms of unfairness. Li Guo 0001, Binxing Fang |
ISCC | 3 |
| 2007 | A Novel Data Mining Method for Network Anomaly Detection Based on Transductive Scheme
Yang Li 0002, Binxing Fang, Li Guo 0001 |
ISNN (1) | 3 |
| 2007 | An active learning based TCM-KNN algorithm for supervised network intrusion detection
Yang Li 0002, Li Guo 0001 |
Comput. Secur. | 2 |
| 2006 | Survey and Taxonomy of Feature Selection Algorithms in Intrusion Detection System
You Chen 0005, Yang Li 0002, Xueqi Cheng 0001, Li Guo 0001 |
Inscrypt | 4 |
| 2006 | Traffic classification-based spam filterabstractWe propose an unsupervised spam filter called Bulk Mail Traffic Classification (BMTC) for filtering junk mails from the perspective of ISPs. Our insight is that spammers generally sent mass unsolicited emails with few alterations to a common message content, which can be found at an extensive traffic environment. In our approach, we classify email delivery traffic into different categories by the similarity of message contents. Then we can decide whether or not a particular email category is spam by the number of similar mails of this category and take measures to filter it. We also design a simulator, two sketches data structure, and a series of algorithms to support our method. We have applied BMTC to email traffic data captured at one of the largest commercial Internet service providers in China, and the experimental result indicates that a 70.4% reduction of emails can be achieved with our method. The results also show that BMTC is practical. We can implement it in a high-volume traffic environment handling over millions of mails every day with small memory consumption. Binxing Fang, Xueqi Cheng 0001, Li Guo 0001 |
ICC | 5 |
| 2006 | A traffic-classified technique for filtering spam from bulk delivery E-mailsabstractTremendous increases in spam traffic have made an imperative to develop effective techniques that can filter Internet spam traffic for using in network operation and security management. In this paper, we propose a more effective scheme called improved bulk mail traffic classification (IBMTC) for filtering spam from bulk delivery E-mails traffic. Based on our earlier designed scheme, we further address some issues related to spam attack and running performance. We also develop a series of more effective techniques to support our method. We have applied IBMTC to the E-mail traffic data captured at one of the largest commercial Internet service providers in China, and the experimental result indicates that the new scheme is more effective and can further improve the performance of filter Binxing Fang, Li Guo 0001, Xueqi Cheng 0001 |
IPCCC | 4 |
| 2006 | TTSF: A Novel Two-Tier Spam FilterabstractSpam prevention is a classic puzzle in the research area of network security. The conventional spam filtering techniques still result to high false positives and have weak on-line processing ability. This paper presents a novel two-tier spam filter (TTSF) on the basis of analyses over the conventional anti-spam techniques. TTSF on-line filters spam by using URLs and off-line filters spam using digest-based approach. Experimental evaluations from the constructed prototype based on TTSF demonstrate it can significantly raise the filtering accuracy, effectively reduce false positives and can be applied to on-line processing and heavy traffic environment by reducing the computational cost than the state-of-the-art techniques Yang Li 0002, Binxing Fang, Li Guo 0001 |
PDCAT | 3 |