Kai Lei

dblp:64/9060 · DBLP profile ↗
← Back
95ranked-venue papers
29as first author
20since 2021 · last 2026
0000-0001-9197-895XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 9 first-author · 4 since 2021Computer networks · 25 · 11 first-author · 9 since 2021Databases, data management, data science and information retrieval · 25 · 2 first-author · 2 since 2021Systems, architecture and hardware · 12 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
abstract
The stability of Internet services is persistently challenged by large volumetric TCP SYN floods, for which conventional defenses such as SYN Cookies preserve server state but still amplify bandwidth pressure. This paper presents EdgePoW, an ingress aware defense architecture that integrates non interactive Proof of Work with an SDN control plane for managed edge networks. The controller monitors per ingress SYN pressure and raises PoW difficulty when flooding is detected. If traffic mainly originates from a stable source region, enforcement is refined to the offending source prefix to reduce overhead on benign co located clients; otherwise, ingress wide enforcement is retained under randomized or spoofed sources. We further design a conservative Difficulty Discovery Protocol that reuses TCP retransmissions and commits difficulty updates only after a successful handshake. Experiments on a custom SDN testbed show restored application QoS under concentrated and spoofed floods, 11.7% higher benign client throughput than ingress only enforcement, and below 0.8% transient false escalations under 2% random loss.
Wenyang Jia, Xianneng Zou, Kai Lei
APNet4
2026 Octopus: An ABR-RAN Closed-Loop Approach Towards High QoE Multi-User 5G VR Gaming
Chengke Wang, Junchen Guo, Zidong Yang, Yinian Zhou, Tao Sun 0010, Kai Lei, Yunhuai Liu, Chenren Xu
SECON7
2025 Multi-Task Model Fusion via Adaptive Merging
abstract
Multi-task model fusion (MTMF) aims to integrate the capabilities of individual models into a unified model. Past approaches either require extensive training and fine-tuning or necessitate that models share the same pre-training and initialization. Recently, several fusion methods have been proposed that do not require extensive training or fine-tuning. These methods can merge multiple independently trained models with different task capabilities into a single multi-task model without increasing the number of parameters. In this work, we identify a common flaw in these fusion methods: they tend to focus on how well the modules of the individual models match before merging while neglecting the representation bias of the merged model. To address this problem, we propose a simple yet effective mitigation method called adaptive merging by representation alignment (AdMbRA). Specifically, we improve the method of weight matching by using representation bias as a constraint and optimize the merging process. Experiments demonstrate that our method can effectively mitigate the representation bias of the merged model, thus improving the performance of each task.
Ziwei Xiang, Kai Lei, Xu-Yao Zhang
ICASSP3
2025 Mamba4Net: Distilled Hybrid Mamba Large Language Models For Networking
abstract
Transformer-based large language models (LLMs) are increasingly being adopted in networking research to address domain-specific challenges. However, their quadratic time complexity and substantial model sizes often result in significant computational overhead and memory constraints, particularly in resource-constrained environments. Drawing inspiration from the efficiency and performance of the Deepseek-R1 model within the knowledge distillation paradigm, this paper introduces Mamba4Net, a novel cross-architecture distillation framework. Mamba4Net transfers networking-specific knowledge from transformer-based LLMs to student models built on the Mamba architecture, which features linear time complexity. This design substantially enhances computational efficiency compared to the quadratic complexity of transformer-based models, while the reduced model size further minimizes computational demands, improving overall performance and resource utilization. To evaluate its effectiveness, Mamba4Net was tested across three diverse networking tasks: viewport prediction, adaptive bitrate streaming, and cluster job scheduling. Compared to existing methods that do not leverage LLMs, Mamba4Net demonstrates superior task performance. Furthermore, relative to direct applications of transformer-based LLMs, it achieves significant efficiency gains, including a throughput 3.96 times higher and a storage footprint of only 5.48% of that required by previous LLM-based approaches. These results highlight Mamba4Net’s potential to enable the cost-effective application of LLM-derived knowledge in networking contexts. The source code is openly available to support further research and development.
Linhan Xia, Mingzhan Yang, Ziwei Yan, Yakun Ren, Kai Lei
ICNP7
2025 PolyBERT: Fine-Tuned Poly Encoder BERT-Based Model for Word Sense Disambiguation
Linhan Xia, Mingzhan Yang, Guohui Yuan, Shengnan Tao, Yujing Qiu, Kai Lei
KSEM (4)7
2025 LLM-Enhanced Heterogeneous Graph Embedding Model for Multi-Task DNS Security
Wenyang Jia, Ziwei Yan, Tanren Liu, Kai Lei
NPC (1)5
2025 BlockSDN-VC: A SDN-Based Virtual Coordinate-Enhanced Transaction Broadcast Framework for High-Performance Blockchains
Wenyang Jia, Ziwei Yan, Guohui Yuan, Tanren Liu, Yakun Ren, Kai Lei
NPC (1)7
2025 A plug-and-play data-driven approach for anti-money laundering in bitcoin
Yuzhi Liang, Weijing Wu, Ruiju Liang, Kai Lei, Guo Zhong, Qingqing Gan, Jinsheng Huang
Expert Syst. Appl.5
2024 Cross-Domain Few-Shot Learning with Equiangular Embedding and Dynamic Adversarial Augmentation
Ziwei Xiang, Kai Lei, Xu-Yao Zhang
ICONIP (3)3
2024 Blockchain-based cooperative game bilateral matching architecture for shared storage
Guanjie Lin, Mingyuan Zeng, Zhiguang Shan, Kaishun Wu, Kai Lei
Future Gener. Comput. Syst.6
2023 On the Profitability of Selfish Mining Attack Under the Checkpoint Mechanism
abstract
Though designed with security in mind, blockchains are vulnerable to various kinds of attacks, especially when the network computational power is low. Selfish mining is one of the most rudimentary and notorious attacks, which maliciously renders blocks found by honest miners orphaned by strategically withholding and revealing the found blocks. In this paper, we analyze the profitability of selfish mining under the checkpoint mechanism—a mechanism that has been adopted as a finality gadget by many blockchains like Ethereum and Bitcoin Cash. We develop a rigorous analysis method and conduct quantitative evaluations in various scenarios to explore the mechanism's suppression effect on selfish mining. The results illustrate that the checkpoint mechanism can restrict the profit of selfish mining and increase the threshold of computational power that makes selfish mining profitable, suggesting that it is a practical defense mechanism against selfish mining.
Yu Zhou 0047, Shang Gao 0006, Weiwei Qiu, Kai Lei, Bin Xiao 0001
GLOBECOM4
2023 Measuring the Consistency Between Data and Control Plane in SDN
abstract
Software Defined Networking (SDN) simplifies network control and management by decoupling the control plane from the data plane. However, the actual packet behaviors, conforming to the rules in the data plane flow tables, may violate the original policies in the controller due to the inconsistency between the data plane and control plane. To address this problem, we propose 2MVeri, a framework for measuring the consistency between the Data and Control plane, defined as the consistency between the control plane policies and data plane rules. 2MVeri uses a modules, a Bloom filter and a two-dimensional vector as a tag which is inserted in the packet header and is updated in each switch that the packet traverses. By exploiting path information compressed in the tag, 2MVeri can verify the consistency between the data and control plane. Moreover, when verification fails, 2MVeri is able to localize the faulty switch. Experimental results show that in the k = 4 fat tree topology, the verification accuracy of 2MVeri is as high as 100%. In addition, when the actual path is inconsistent with the expected path, 2MVeri can locate the wrong switch with an accuracy of 99.8%.
Kai Lei, Guanjie Lin, Meimei Zhang, Xiaojun Jing
IEEE/ACM Trans. Netw.1
2022 Slicer: Verifiable, Secure and Fair Search over Encrypted Numerical Data Using Blockchain
abstract
Verifiable Searchable Symmetric Encryption (SSE) enables reliable search over encrypted, privacy-preserving data on untrusted clouds. Most existing SSE designs only focus on keyword-file search. However, a more difficult but useful search, range search over encrypted numerical values remains unsolved. Moreover, the fairness of search in the mutual distrusted scenario without public verification, where data users may maliciously deny the results after the local result verification, is not well addressed yet. In this paper, we take the first step to study the public verification problem atop the blockchain for encrypted numerical search. We design a novel verifiable SSE scheme named Slicer based on a Succinct Order-Revealing Encryption (SORE) scheme to achieve range search on numerical data. Our search results are verifiable, updated and privacy-preserving by SSE and maintaining the forward security. We illustrate the security and practicality of our design through rigorous analysis and extensive evaluations respectively.
Haotian Wu 0001, Rui Song 0010, Kai Lei, Bin Xiao 0001
ICDCS3
2022 Meta-Regulation: Adaptive Adjustment to Block Size and Creation Interval for Blockchain Systems
abstract
Once deployed, a decentralized blockchain system ensures that it will operate faithfully so that no one can interfere with or manipulate its predefined regulations, such as block size and block creation interval investigated in this paper. However, fixed regulations prevent that system from adapting to the change of the environment, such as increasing the underlying network capacity, and result in sub-optimal performance. For example, Bitcoin remains at 7 TPS (transactions per second), even operating over the current Internet. In this paper, we propose a new paradigm for defining the behavior of a consensus system, named as Meta-Regulation, which allows autonomous evolution of the system behavior. A meta-regulation adjusts the actual behavior of a consensus system in response to the changing capacity of the underlying infrastructure and the community of participants. We demonstrate the effectiveness of the proposed meta-regulation by achieving significantly improved throughput and latency for Bitcoin, adapted to the current capacity of the Internet. Our experimental results show that Meta-Regulation can achieve at least$7\times $performance improvement over Bitcoin network deployed in 2009, resulting in 49.7 TPS or 68% reduction confirmation latency by fully utilizing the bandwidth and the computing power of average network nodes.
Mingpei Cao, Hao Wang 0002, Tailing Yuan, Kun Xu 0003, Kai Lei, Jiaping Wang
IEEE J. Sel. Areas Commun.5
2022 GBRM: a graph embedding and blockchain-based resource management framework for 5G MEC
Kai Lei, Junjie Fang, Peiwu Chen, Liangjie Zhang, Jing Xiao 0006
J. Supercomput.1
2021 Towards a Translation-Based Method for Dynamic Heterogeneous Network Embedding
abstract
Network embedding, which aims to map the discrete network topology to a continuous low-dimensional representation space with the major topological properties preserved, has emerged as an essential technique to support various network inference tasks. However, incorporating both the evolutionary nature and the network's heterogeneity remains a challenge for existing network embedding methods. In this study, we propose a novel Translation-Based Dynamic Heterogeneous Network Embedding (TransDHE) approach to consider both the aspects simultaneously. For a dynamic heterogeneous network with a sequence of snapshots and multiple types of nodes and edges, we introduce a translation-based embedding module to capture the heterogeneous characteristics (e.g., type information) of each single snapshot. An orthogonal alignment module and RNN-based aggregation module are then applied to explore the evolutionary patterns among multiple successive snapshots for the final representation learning. Extensive experiments on a set of real-world networks demonstrate that TransDHE can derive the more informative embedding result for the network dynamic and heterogeneity over state-of-the-art network embedding baselines.
Kai Lei, Yuzhi Liang, Jing Xiao 0006, Peiwu Chen
ICC1
2021 Dual-channel hybrid community detection in attributed networks
Meng Qin 0002, Kai Lei
Inf. Sci.2
2021 HOPASS: A two-layer control framework for bandwidth and delay guarantee in datacenters
Kai Lei, Bo Bai 0001, Fan Zhang 0016, Gong Zhang 0001, Jingjie Jiang
J. Netw. Comput. Appl.1
2021 Comparative analysis of probabilistic forwarding strategies in ICN for edge computing
Meimei Zhang, Liangjie Zhang, Xiquan Yu, Kai Lei
Peer-to-Peer Netw. Appl.6
2021 Multitask Learning and Reinforcement Learning for Personalized Dialog Generation: An Empirical Study
abstract
Open-domain dialog generation, which is a crucial component of artificial intelligence, is an essential and challenging problem. In this article, we present a personalized dialog system, which leverages the advantages of multitask learning and reinforcement learning for personalized dialogue generation (MRPDG). Specifically, MRPDG consists of two subtasks: 1) an author profiling module that recognizes user characteristics from the input sentence (auxiliary task) and 2) a personalized dialog generation system that generates informative, grammatical, and coherent responses with reinforcement learning algorithms (primary task). Three kinds of rewards are proposed to generate high-quality conversations. We investigate the effectiveness of three widely used reinforcement learning methods [i.e., Q-learning, policy gradient, and actor-critic (AC) algorithm] in a personalized dialog generation system and demonstrate that the AC algorithm achieves the best results on the underlying framework. Comprehensive experiments are conducted to evaluate the performance of the proposed model on two real-life data sets. Experimental results illustrate that MRPDG is able to produce high-quality personalized dialogs for users with different characteristics. Quantitatively, the proposed model can achieve better performance than the compared methods across different evaluation metrics, such as the human evaluation, BiLingual Evaluation Understudy (BLEU), and perplexity.
Min Yang 0007, Wenting Tu, Qiang Qu 0001, Ying Shen 0001, Kai Lei
IEEE Trans. Neural Networks Learn. Syst.6
2020 Attentive User-Engaged Adversarial Neural Network for Community Question Answering
abstract
We study the community question answering (CQA) problem that emerges with the advent of numerous community forums in the recent past. The task of finding appropriate answers to questions from informative but noisy crowdsourced answers is important yet challenging in practice. We present an Attentive User-engaged Adversarial Neural Network (AUANN), which interactively learns the context information of questions and answers, and enhances user engagement with the CQA task. A novel attentive mechanism is incorporated to model the semantic internal and external relations among questions, answers and user contexts. To handle the noise issue caused by introducing user context, we design a two-step denoise mechanism, including a coarse-grained selection process by similarity measurement, and a fine-grained selection process by applying an adversarial training module. We evaluate the proposed method on large-scale real-world datasets SemEval-2016 and SemEval-2017. Experimental results verify the benefits of incorporating user information, and show that our proposed model significantly outperforms the state-of-the-art methods.
Yuexiang Xie, Ying Shen 0001, Yaliang Li, Min Yang 0007, Kai Lei
AAAI5
2020 Relabel the Noise: Joint Extraction of Entities and Relations via Cooperative Multiagents
abstract
Distant supervision based methods for entity and relation extraction have received increasing popularity due to the fact that these methods require light human annotation efforts.In this paper, we consider the problem of shifted label distribution, which is caused by the inconsistency between the noisy-labeled training set subject to external knowledge graph and the human-annotated test set, and exacerbated by the pipelined entity-then-relation extraction manner with noise propagation.We propose a joint extraction approach to address this problem by re-labeling noisy instances with a group of cooperative multiagents.To handle noisy instances in a fine-grained manner, each agent in the cooperative group evaluates the instance by calculating a continuous confidence score from its own perspective; To leverage the correlations between these two extraction tasks, a confidence consensus module is designed to gather the wisdom of all agents and re-distribute the noisy training set with confidence-scored labels.Further, the confidences are used to adjust the training losses of extractors.Experimental results on two realworld datasets verify the benefits of re-labeling noisy instance, and show that the proposed model significantly outperforms the state-ofthe-art entity and relation extraction methods.
Daoyuan Chen, Yaliang Li, Kai Lei, Ying Shen 0001
ACL3
2020 IGAN-IDS: An imbalanced generative adversarial network towards intrusion detection system in ad-hoc networks
Shuokang Huang, Kai Lei
Ad Hoc Networks2
2020 Time-driven feature-aware jointly deep reinforcement learning for financial signal representation and algorithmic trading
Kai Lei, Min Yang 0007, Ying Shen 0001
Expert Syst. Appl.1
2020 Blockchain-Based Cache Poisoning Security Protection and Privacy-Aware Access Control in NDN Vehicular Edge Computing Networks
Kai Lei, Junjie Fang, Junjun Lou, Maoyu Du, Jiyue Huang, Kuai Xu
J. Grid Comput.1
2020 Tag recommendation by text classification with attention-based capsule network
Kai Lei, Qiuai Fu, Min Yang 0007, Yuzhi Liang
Neurocomputing1
2020 Reachability preserving compression for dynamic graph
Yuzhi Liang, Kai Lei, Min Yang 0007, Ziyu Lyu
Inf. Sci.4
2020 Generative adversarial fusion network for class imbalance credit scoring
Kai Lei, Yuexiang Xie, Shangru Zhong, Jingchao Dai, Min Yang 0007, Ying Shen 0001
Neural Comput. Appl.1
2020 Path-based reasoning with constrained type attention for knowledge graph completion
Kai Lei, Yuexiang Xie, Desi Wen, Daoyuan Chen, Min Yang 0007, Ying Shen 0001
Neural Comput. Appl.1
2020 Cross-domain aspect/sentiment-aware abstractive review summarization by combining topic modeling and deep reinforcement learning
Min Yang 0007, Qiang Qu 0001, Ying Shen 0001, Kai Lei, Jia Zhu 0003
Neural Comput. Appl.4
2020 Groupchain: Towards a Scalable Public Blockchain in Fog Computing of IoT Services Computing
abstract
Powered by a number of smart devices distributed throughout the whole network, the Internet of Things (IoT) is supposed to provide services computing for massive data from devices. Fog computing, an extension of cloud-based IoT-oriented solutions, has emerged with requirements for distribution and decentralization. In this respect, the conjunction with Blockchain provides a natural solution for decentralization, as well as potentially helps fog computing overcome some deficiencies such as security and privacy then consequently expand the application scope of IoT. However, one of the key challenges of blockchain's integration with fog computing is scalability. To end this, this work proposes Groupchain, a novel scalable public blockchain of a two-chain structure suitable for fog computing of IoT services computing. Groupchain employs the leader group to collectively commit blocks for highertransaction efficiency and introduces bonus and deposit into the incentive mechanism to supervise behaviors of members in the leader group. Our security analysis shows that Groupchain retains the security of Bitcoin-like blockchain and enhances defense against attacks such as double-spend and selfish mining. We implement a prototype of Groupchain and conducted experiments. The experimental results demonstrate that Groupchain achieves optimization on transaction throughput and confirmation latency which are argued in Bitcoin.
Kai Lei, Maoyu Du, Jiyue Huang
IEEE Trans. Serv. Comput.1
2019 Multi-Task Learning with Multi-View Attention for Answer Selection and Knowledge Base Question Answering
abstract
Answer selection and knowledge base question answering (KBQA) are two important tasks of question answering (QA) systems. Existing methods solve these two tasks separately, which requires large number of repetitive work and neglects the rich correlation information between tasks. In this paper, we tackle answer selection and KBQA tasks simultaneously via multi-task learning (MTL), motivated by the following motivations. First, both answer selection and KBQA can be regarded as a ranking problem, with one at text-level while the other at knowledge-level. Second, these two tasks can benefit each other: answer selection can incorporate the external knowledge from knowledge base (KB), while KBQA can be improved by learning contextual information from answer selection. To fulfill the goal of jointly learning these two tasks, we propose a novel multi-task learning scheme that utilizes multi-view attention learned from various perspectives to enable these tasks to interact with each other as well as learn more comprehensive sentence representations. The experiments conducted on several real-world datasets demonstrate the effectiveness of the proposed method, and the performance of answer selection and KBQA is improved. Also, the multi-view attention scheme is proved to be effective in assembling attentive information from different representational perspectives.
Yang Deng 0002, Yuexiang Xie, Yaliang Li, Min Yang 0007, Nan Du 0001, Wei Fan 0001, Kai Lei, Ying Shen 0001
AAAI7
2019 MedTruth: A Semi-supervised Approach to Discovering Knowledge Condition Information from Multi-Source Medical Data
abstract
Knowledge Graph (KG) contains entities and the relations between entities. Due to its representation ability, KG has been successfully applied to support many medical/healthcare tasks. However, in the medical domain, knowledge holds under certain conditions. Such conditions for medical knowledge are crucial for decision-making in various medical applications, which is missing in existing medical KGs. In this paper, we aim to discovery medical knowledge conditions from texts to enrich KGs. Electronic Medical Records (EMRs) are systematized collection of clinical data and contain detailed information about patients, thus EMRs can be a good resource to discover medical knowledge conditions. Unfortunately, the amount of available EMRs is limited due to reasons such as regularization. Meanwhile, a large amount of medical question answering (QA) data is available, which can greatly help the studied task. However, the quality of medical QA data is quite diverse, which may degrade the quality of the discovered medical knowledge conditions. In the light of these challenges, we propose a new truth discovery method, MedTruth, for medical knowledge condition discovery, which incorporates prior source quality information into the source reliability estimation procedure, and also utilizes the knowledge triple information for trustworthy information computation. We conduct series of experiments on real-world medical datasets to demonstrate that the proposed method can discover meaningful and accurate conditions for medical knowledge by leveraging both EMR and QA data. Further, the proposed method is tested on synthetic datasets to validate its effectiveness under various scenarios.
Yang Deng 0002, Yaliang Li, Ying Shen 0001, Nan Du 0001, Wei Fan 0001, Min Yang 0007, Kai Lei
CIKM7
2019 Self-Adaptive Scaling for Learnable Residual Structure
abstract
Residual has been widely applied to build deep neural networks with enhanced feature propagation and improved accuracy.In the literature, multiple variants of residual structure are proposed.However, most of them are manually designed for particular tasks and datasets and the combination of existing residual structures has not been well studied.In this work, we propose the Self-Adaptive Scaling (SAS) approach that automatically learns the design of residual structure from data.The proposed approach makes the best of various residual structures, resulting in a general architecture covering several existing ones.In this manner, we construct a learnable residual structure which can be easily integrated into a wide range of residual-based models.We evaluate our approach on various tasks concerning different modalities, including machine translation (IWSLT-2015 EN-VI and WMT-2014 EN-DE, EN-FR), image classification (CIFAR-10 and CIFAR-100), and image captioning (MSCOCO).Empirical results show that the proposed approach consistently improves the residual-based models and exhibits desirable generalization ability.In particular, by incorporating the proposed approach to the Transformer model, we establish new state-of-thearts on the IWSLT-2015 EN-VI low-resource machine translation dataset.
Yuanxin Liu, Kai Lei
CoNLL4
2019 Sparse Gradient Compression for Distributed SGD
Haobo Sun, Yingxia Shao, Jiawei Jiang 0001, Bin Cui 0001, Kai Lei
DASFAA (2)5
2019 HOMMO: A Hierarchical Flow Management Framework for Multi-Objective Data Center Networks
abstract
In data center networks, flows with different objectives coexist and compete for limited resources. From the application-level perspective, it is hard to satisfy the demands of different flows without effective resource planning. To address the bandwidth allocation problem under multi-objective scenarios, we study a multi-objective network utility maximization (NUM) problem and propose our practical bandwidth allocation framework HOMMO, which consists of an upper layer algorithm (ULA) and a lower layer algorithm (LLA). This hierarchical design helps to strike a balance between accuracy and efficiency. Implemented in network switches, ULA is an online learning-based scheme that allocates bandwidth for aggregated flows with different performances objectives. Taking the outputs from ULA as capacity constraints, LLA acts as a fast scheduling method at packet-level. We extend NUMFabric by implementing the design of isolate queues and corresponding dequeue strategy in switches and make it a feasible solution of the LLA. Therefore, HOMMO achieves satisfactory isolation across flows with different objectives, which is equivalent to provide a network slicing solution. To evaluate the proposed framework, we implement it in ns-3 and verify the performance under various scenarios. The simulation results show that HOMMO not only quickly converges to a near-optimal solution of the multi-objective NUM problem but also guarantees a Pareto-optimal solution. Moreover, it outperforms the well-known transport protocol (i.e., DCTCP) with 2x increase on average bandwidth utilization and 1.96x improvement in global network utility at the aggregated flow level.
Kai Lei, Fan Zhang 0016, Hengky Susanto, Bo Bai 0001, Gong Zhang 0001
GLOBECOM1
2019 Detecting Malicious Domains with Behavioral Modeling and Graph Embedding
abstract
The last decade has witnessed the explosive growth of malicious Internet domains which serve as the fundamental infrastructure for establishing advanced persistent threat command and control communication channels or hosting phishing Web sites. Given the big data nature of Internet traffic data and the ability of algorithmically generating domains and acquiring and registering the domains in a near-automated fashion, detecting malicious domains in real-time is a daunting task for security analysts and network operators. In this paper, we introduce bipartite graphs to capture the interactions between end hosts and domains, identify associated IP addresses of domains, and characterize time-series patterns of DNS queries for domains, and explore one-mode projections of these bipartite graphs for modeling the behavioral, IP-structural, and temporal similarities between domains. We employ graph embedding technique to automatically learn dynamic and discriminative feature representations for over 10,000 labeled domains, and develop an SVM-based classification algorithm for predicting malicious or benign domains. Our model makes the progress towards adapting to the changing and evolving strategies of malicious domains. The experimental results have shown that our proposed algorithm achieves an area under the curve (AUC) of 0.94 based on k-fold cross-validation. To the best of our knowledge, this is the first effort to apply the combination of behavioral modeling and graph embedding for effectively and accurately detecting malicious domains.
Kai Lei, Qiuai Fu, Jiake Ni, Min Yang 0007, Kuai Xu
ICDCS1
2019 Exploring and Distilling Cross-Modal Information for Image Captioning
abstract
Recently, attention-based encoder-decoder models have been used extensively in image captioning. Yet there is still great difficulty for the current methods to achieve deep image understanding. In this work, we argue that such understanding requires visual attention to correlated image regions and semantic attention to coherent attributes of interest. To perform effective attention, we explore image captioning from a cross-modal perspective and propose the Global-and-Local Information Exploring-and-Distilling approach that explores and distills the source information in vision and language. It globally provides the aspect vector, a spatial and relational representation of images based on caption contexts, through the extraction of salient region groupings and attribute collocations, and locally extracts the fine-grained regions and attributes in reference to the aspect vector for word selection. Our fully-attentive model achieves a CIDEr score of 129.3 in offline COCO evaluation with remarkable efficiency in terms of accuracy, speed, and parameter budget.
Xuancheng Ren, Yuanxin Liu, Kai Lei, Xu Sun 0001
IJCAI4
2019 Multi-Task Learning with Capsule Networks
abstract
Multi-task learning is a machine learning approach learning multiple tasks jointly while exploiting commonalities and differences across tasks. A shared representation is learned by multi-task learning, and what is learned for each task can help other tasks be learned better. Most of existing multi-task learning methods adopt deep neural network as the classifier of each task. However, a deep neural network can exploit its strong curve-fitting capability to achieve high accuracy in training data even when the learned representation is not good enough. This is contradictory to the purpose of multi-task learning. In this paper, we propose a framework named multi-task capsule (MT-Capsule) which improves multi-task learning with capsule network. Capsule network is a new architecture which can intelligently model part-whole relationships to constitute viewpoint invariant knowledge and automatically extend the learned knowledge to different new scenarios. The experimental results on large real-world datasets show MT-Capsule can significantly outperform the state-of-the-art methods.
Kai Lei, Qiuai Fu, Yuzhi Liang
IJCNN1
2019 GCN-GAN: A Non-linear Temporal Link Prediction Model for Weighted Dynamic Networks
abstract
In this paper, we generally formulate the dynamics prediction problem of various network systems (e.g., the prediction of mobility, traffic and topology) as the temporal link prediction task. Different from conventional techniques of temporal link prediction that ignore the potential non-linear characteristics and the informative link weights in the dynamic network, we introduce a novel non-linear model GCN-GAN to tackle the challenging temporal link prediction task of weighted dynamic networks. The proposed model leverages the benefits of the graph convolutional network (GCN), long short-term memory (LSTM) as well as the generative adversarial network (GAN). Thus, the dynamics, topology structure and evolutionary patterns of weighted dynamic networks can be fully exploited to improve the temporal link prediction performance. Concretely, we first utilize GCN to explore the local topological characteristics of each single snapshot and then employ LSTM to characterize the evolving features of the dynamic networks. Moreover, GAN is used to enhance the ability of the model to generate the next weighted network snapshot, which can effectively tackle the sparsity and the wide-value-range problem of edge weights in real-life dynamic networks. To verify the model's effectiveness, we conduct extensive experiments on four datasets of different network systems and application scenarios. The experimental results demonstrate that our model achieves impressive results compared to the state-of-the-art competitors.
Kai Lei, Meng Qin 0002, Bo Bai 0001, Gong Zhang 0001, Min Yang 0007
INFOCOM1
2019 Multitask Learning for Cross-Domain Image Captioning
abstract
Recent artificial intelligence research has witnessed great interest in automatically generating text descriptions of images, which are known as the image captioning task. Remarkable success has been achieved on domains where a large number of paired data in multimedia are available. Nevertheless, annotating sufficient data is labor-intensive and time-consuming, establishing significant barriers for adapting the image captioning systems to new domains. In this study, we introduc a novel Multitask Learning Algorithm for cross-Domain Image Captioning (MLADIC). MLADIC is a multitask system that simultaneously optimizes two coupled objectives via a dual learning mechanism: image captioning and text-to-image synthesis, with the hope that by leveraging the correlation of the two dual tasks, we are able to enhance the image captioning performance in the target domain. Concretely, the image captioning task is trained with an encoder-decoder model (i.e., CNN-LSTM) to generate textual descriptions of the input images. The image synthesis task employs the conditional generative adversarial network (C-GAN) to synthesize plausible images based on text descriptions. In C-GAN, a generative model $G$ synthesizes plausible images given text descriptions, and a discriminative model $D$ tries to distinguish the images in training data from the generated images by $G$. The adversarial process can eventually guide $G$ to generate plausible and high-quality images. To bridge the gap between different domains, a two-step strategy is adopted in order to transfer knowledge from the source domains to the target domains. First, we pre-train the model to learn the alignment between the neural representations of images and that of text data with the sufficient labeled source domain data. Second, we fine-tune the learned model by leveraging the limited image-text pairs and unpaired data in the target domain. We conduct extensive experiments to evaluate the performance of MLADIC by using the MSCOCO as the source domain data, and using Flickr30k and Oxford-102 as the target domain data. The results demonstrate that MLADIC achieves substantially better performance than the strong competitors for the cross-domain image captioning task.
Min Yang 0007, Wei Zhao 0033, Yabing Feng, Zhou Zhao 0001, Xiaojun Chen 0006, Kai Lei
IEEE Trans. Multim.7
2019 MARES: multitask learning algorithm for Web-scale real-time event summarization
Min Yang 0007, Wenting Tu, Qiang Qu 0001, Kai Lei, Xiaojun Chen 0006, Jia Zhu 0003, Ying Shen 0001
World Wide Web4
2018 Drug2Vec: Knowledge-aware Feature-driven Method for Drug Representation Learning
Ying Shen 0001, Kaiqi Yuan, Yaliang Li, Buzhou Tang, Min Yang 0007, Nan Du 0001, Kai Lei
BIBM7
2018 Knowledge as A Bridge: Improving Cross-domain Answer Selection with External Knowledge
abstract
Answer selection is an important but challenging task. Significant progresses have been made in domains where a large amount of labeled training data is available. However, obtaining rich annotated data is a time-consuming and expensive process, creating a substantial barrier for applying answer selection models to a new domain which has limited labeled data. In this paper, we propose Knowledge-aware Attentive Network (KAN), a transfer learning framework for cross-domain answer selection, which uses the knowledge base as a bridge to enable knowledge transfer from the source domain to the target domains. Specifically, we design a knowledge module to integrate the knowledge-based representational learning into answer selection models. The learned knowledge-based representations are shared by source and target domains, which not only leverages large amounts of cross-domain data, but also benefits from a regularization effect that leads to more general representations to help tasks in new domains. To verify the effectiveness of our model, we use SQuAD-T dataset as the source domain and three other datasets (i.e., Yahoo QA, TREC QA and InsuranceQA) as the target domains. The experimental results demonstrate that KAN has remarkable applicability and generality, and consistently outperforms the strong competitors by a noticeable margin for cross-domain answer selection.
Yang Deng 0002, Ying Shen 0001, Min Yang 0007, Yaliang Li, Nan Du 0001, Wei Fan 0001, Kai Lei
COLING7
2018 Cooperative Denoising for Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction greatly reduces human efforts in extracting relational facts from unstructured texts. However, it suffers from noisy labeling problem, which can degrade its performance. Meanwhile, the useful information expressed in knowledge graph is still underutilized in the state-of-the-art methods for distantly supervised relation extraction. In the light of these challenges, we propose CORD, a novelCOopeRativeDenoising framework, which consists two base networks leveraging text corpus and knowledge graph respectively, and a cooperative module involving their mutual learning by the adaptive bi-directional knowledge distillation and dynamic ensemble with noisy-varying instances. Experimental results on a real-world dataset demonstrate that the proposed method reduces the noisy labels and achieves substantial improvement over the state-of-the-art methods.
Kai Lei, Daoyuan Chen, Yaliang Li, Nan Du 0001, Min Yang 0007, Wei Fan 0001, Ying Shen 0001
COLING1
2018 Distributed Information-Agnostic Flow Scheduling in Data Centers Based on Wait-Time
abstract
Existing flow scheduling schemes in Data Center Network (DCN) are designed mainly to minimize the flow complete time (FCT) of short flows and do not consider optimizing the FCT of latency-sensitive long flows (e.g. VR video streaming, interactive artificial intelligence question&answer stream). Besides, among these traffic scheduling schemes, the information-aware schemes (e.g. L2DCT, D2TCP) are hard to deploy in practice since they assume prior knowledge of flow information (e.g, flow size); and the information-agnostic scheme (i.e. PIAS), which is based on the premise that flow size is not known a priori, requires a central server, causing a poor scalability in large network scales. Given the limitations of existing solutions, in this paper, we propose a distributed information-agnostic flow scheduling scheme (DIAS), which minimizes the FCT of both short flows and latency-sensitive long flows. In DIAS, packets are forwarded complying with their priorities, which are determined based on packets wait-time that is defined as staying time in end hosts' send buffers, and the longer a packet stays in a send buffer, the lower its priority. Meanwhile, instead of utilizing a central server to collect traffic load information, each switch feeds traffic load information which is used to adjust the thresholds of determining packets priority back to end hosts via ACK packets. The experimental results in ns-3 simulator show that DIAS reduces FCT by up to 54.7% and 50.1 % over DCTCP and L2DCT, respectively. Besides, DIAS ensures a smaller FCT of latency-sensitive long flows and performs better than PIAS.
Kai Lei, Bo Jin 0002, Yi Wang 0004
GLOBECOM1
2018 Measuring the Control-Data Plane Consistency in Software Defined Networking
abstract
Software Defined Networking (SDN) simplifies network management by separating the control plane from the data plane in performance networks. However, the actual packet behaviors, conforming to the rules in the data plane flow tables, may violate the original policies in the controller due to the inconsistency between the data plane and control plane. To address this problem, we propose 2MVeri, a new framework for measuring the control-data plane consistency, defined as the consistency between the control plane policies and data plane rules. 2MVeri uses a Bloom filter and a two-dimensional vector as a tag which is inserted in the packet header and is updated in each switch that the packet traverses. By exploiting path information compressed in the tag, 2MVeri can verify the control-data plane consistency. Moreover, when verification fails, 2MVeri can localize the faulty switch. Evaluations conducted on a datacenter network with the fat tree topology k=4 demonstrate that 2MVeri achieves 100% accuracy in consistency verification and high fault localization performance.
Kai Lei, Weichao Li 0001, Yi Wang 0004
ICC1
2018 Fast Parallel Path Concatenation for Graph Extraction
abstract
In this paper, we study the problem of extracting a homogeneous graph from a heterogeneous graph. The key challenges of the extraction problem are how to efficiently enumerate paths matched by the provided line pattern and aggregate values for each pair of vertices from the matched paths. To address above two challenges, we propose a parallel graph extraction framework (PGE), where we use vertex-centric model to enumerate paths and compute aggregate functions in parallel. The framework compiles the line pattern into a path concatenation plan and generates the final weighted edges in a divide-and-conquer manner. The new solution outperforms the state-of-the-art ones through the comprehensive experiments.
Yingxia Shao, Kai Lei, Lei Chen 0002, Zi Huang, Bin Cui 0001, Zhongyi Liu 0001, Yunhai Tong, Jin Xu 0002
ICDE2
2018 Semantic Similarity Measures to Disambiguate Terms in Medical Text
Kai Lei, Jiyue Huang, Shangchun Si, Ying Shen 0001
ICONIP (7)1
2018 OptCaching: A Stackelberg Game and Belief Propagation Based Caching Scheme for Joint Utility Optimization in Fog Computing
abstract
Fog Computing which extends the cloud computing paradigm to the edge of the network provides great opportunities for applications with stringent latency requirement. How to allocate the limited caching resources of Fog Nodes (FNs)influences the performance of the fog computing system. In contrast to previous works on caching resource allocation with users' utility as the only consideration, we propose OptCaching which jointly optimize the utility of all network participants including Content Provider (CP), Internet Service Provider (ISP)and users. With caching incentive introduced, utility functions of these three roles are defined. Our joint utility optimization caching scheme is conducted in two stages combining global and local decision making. Firstly, interaction between CP and ISP is modeled as a non-cooperative hierarchy Stackelberg game to make decision on incentive caching prices and global caching amount aiming at optimizing the utility of all network participants. Secondly, for the purpose of further optimizing the utility of users, a belief propagation based cache placement algorithm which utilizes global caching amount constraint and local information is conducted by FNs to reduce users' average download delay. Mathematical analysis and simulation results show that the utility of CP, ISP and users are jointly optimized at Stackelberg equilibrium. The utility of users is further optimized by belief propagation based cache placement algorithm with users' average download delay reduced by 33.7% compared with global popularity based caching strategy.
Kai Lei, Haijun Zhang 0001, Gong Zhang 0001, Bo Bai 0001
ICPADS1
2018 Reputation-Based Byzantine Fault-Tolerance for Consortium Blockchain
abstract
The Practical Byzantine Fault Tolerance algorithm (PBFT)has been highly applied in consortium blockchain systems, however, this kind of consensus algorithm can hardly identify and remove faulty nodes in time, and also vulnerable to many attacks against the primary node of PBFT. The equality of consortium members' discourse rights is inapplicable to some real scenarios where dominating members are likely to have a larger discourse rights in the voting process. To address these problems, this paper presents Reputation-based Byzantine Fault Tolerance (RBFT)algorithm that incorporates a reputation model to evaluate the operations of each node in the consensus process. The faulty nodes will get lower discourse rights in the voting process if any malicious behavior is detected, with their reputation decreased. Furthermore, this paper presents an innovative reputation-based primary change scheme. The node with higher reputation obtains greater opportunities to be a primary to generate new valid blocks, which reduces the security risk of the primary. The experimental results demonstrate that RBFT gains better performance and ensures system security and reliability. Compared with PBFT, it increases the average throughput by 15% and reduces delay by 10%, and the faulty node rate of the system can continue to decrease over time.
Kai Lei, Limei Xu, Zhuyun Qi
ICPADS1
2018 NDN Producer Mobility Management Based on Echo State Network: A Lightweight Machine Learning Approach
abstract
NDN is one of promising underlying network architectures that supporting 5G because of its characteristics such as decoupling senders and receivers, hop by hop transmission, in-network caching, etc. However, it still faces challenges in producer mobility management like the triangle routing problem (non-optimal routing path) and global centralization of the home agent, causing a poor scalability in large network scales and long handover delay. In this paper, we propose the ESN - PBA, a NDN producer mobility management scheme using the ESN in prediction to realize a lightweight machine learning based seamless handover algorithm. Better than the existing fixed and post-adjustment management schemes of producer mobility management, the ESN-PBA can perceive nodes movements heuristically and pre-configure the adjustment in advance to reduce overall processing overhead. In addition, with fine-grained home router status feedbacks and NDN content data oriented philosophy, the training process of normal machine learning method can be mutually enhanced. The experimental results in ndnSIM show that, in the case of successful prediction, the effect of seamless handover can be achieved straightly on the fly. In order to improve the hit rate of cache, we take advantage of NDN's multipath forwarding support, the utilization of ESN prediction of multiple candidates and synchronous forwarding. Compared with PIT-based approach and DNS-based approach, the handover delay of ESN-PBA reduces by 66.7% and 75% respectively. Besides, its handover overhead reduces by 38.4 %, compared with DNS-based approach.
Xuewei Piao, Haijun Zhang 0001, Kai Lei
ICPADS4
2018 MedSim: A Novel Semantic Similarity Measure in Bio-medical Knowledge Graphs
Kai Lei, Kaiqi Yuan, Qiang Zhang 0015, Ying Shen 0001
KSEM (1)1
2018 IP/NDN: A multi-level translation and migration mechanism
abstract
Different from TCP/IP architecture, Named Data Networking (NDN) adopts a data-centric and in-network caching approach to achieve improved network efficiency and reduced traffic redundancy. As NDN begins to incrementally deploy on the real-world, there will be a hybrid network of TCP/IP and NDN during the transitional period. However, redeveloping TCP/IP- based applications for NDN is a daunting and time-consuming task. Thus, a more sensible and economical approach is to design an effective strategies for leveraging NDN to deliver IP datagrams and porting TCP/IP-based applications on NDN without changing its original code. Towards this end, this paper introduces three different migration methods at Internet layer, TCP layer and application layer, respectively to translate TCP/IP-based packets into NDN packets and run TCP/IP-based applications on NDN. By implementing these methods on a real NDN test-bed, we demonstrate the multi-level IP/NDN translation and migration mechanism is feasible. The analysis of testing results shows that each of the proposed methods is valid with varying benefits and costs: translating at TCP layer with more time overhead can have a less packets overhead than Internet layer. And when translated at application layer, the network load is reduced by half because of the in-network cache.
Shangru Zhong, Kai Lei
NOMS3
2018 Investigating Deep Reinforcement Learning Techniques in Personalized Dialogue Generation
abstract
In this paper, we propose a personalized dialogue generation system, which combines reinforcement learning techniques with an attention-based hierarchical recurrent encoderdecoder model. Firstly, we incorporate user-specific information into the decoder to capture user's background information and speaking style. Secondly, we employ reinforcement learning techniques to maximize future reward in dialogue, which enables our system to generate topic-coherent, informative and grammatical responses. Moreover, we propose three types of rewards to characterize good conversations. Finally, we compare the performance of the following reinforcement learning methods in dialogue generation: policy gradient, Q-learning, and actor-critic algorithms. We conduct experiments to verify the effectiveness of the proposed model on two dialogue datasets. Experimental results demonstrate that our model can generate better personalized dialogues for different users. Quantitatively, our method achieves better performance than the state-of-the-art dialogue systems in terms of BLEU score, perplexity, and human evaluation.
Min Yang 0007, Qiang Qu 0001, Kai Lei, Jia Zhu 0003, Zhou Zhao 0001, Xiaojun Chen 0006, Joshua Zhexue Huang
SDM3
2018 Ontology Evaluation with Path-based Text-aware Entropy Computation
abstract
With the rising importance of knowledge exchange, ontologies have become a key technology in the development of shared knowledge models for semantic-driven applications, such as knowledge interchange and semantic integration. Significant progress has been made in the use of entropy to measure the predictability and redundancy of knowledge bases, particularly ontologies. However, the current entropy applications used to evaluate ontologies consider only single-point connectivity rather than path connectivity, assign equal weights to each entity and path, and assume that vertices are static. To address these deficiencies, the present study proposes a Path-based Text-aware Entropy Computation method, PTEC, by considering the path information between different vertices and the textual information within the path to calculate the connectivity path of the whole network and the different weights between various nodes. Information obtained from structure-based embedding and text-based embedding is multiplied by the connectivity matrix of the entropy computation. An experimental evaluation of three real-world ontologies is performed based on ontology statistical information (data quantity), entropy evaluation (data quality), and a case study (ontology structure and text visualization). These aspects mutually demonstrate the reliability of our method. Experimental results demonstrate that PTEC can effectively evaluate ontologies, particularly those in the medical field.
Ying Shen 0001, Daoyuan Chen, Min Yang 0007, Yaliang Li, Nan Du 0001, Kai Lei
SIGIR6
2018 Knowledge-aware Attentive Neural Network for Ranking Question Answer Pairs
abstract
Ranking question answer pairs has attracted increasing attention recently due to its broad applications such as information retrieval and question answering (QA). Significant progresses have been made by deep neural networks. However, background information and hidden relations beyond the context, which play crucial roles in human text comprehension, have received little attention in recent deep neural networks that achieve the state of the art in ranking QA pairs. In the paper, we propose KABLSTM, a Knowledge-aware Attentive Bidirectional Long Short-Term Memory, which leverages external knowledge from knowledge graphs (KG) to enrich the representational learning of QA sentences. Specifically, we develop a context-knowledge interactive learning architecture, in which a context-guided attentive convolutional neural network (CNN) is designed to integrate knowledge embeddings into sentence representations. Besides, a knowledge-aware attention mechanism is presented to attend interrelations between each segments of QA pairs. KABLSTM is evaluated on two widely-used benchmark QA datasets: WikiQA and TREC QA. Experiment results demonstrate that KABLSTM has robust superiority over competitors and sets state-of-the-art.
Ying Shen 0001, Yang Deng 0002, Min Yang 0007, Yaliang Li, Nan Du 0001, Wei Fan 0001, Kai Lei
SIGIR7
2018 An ontology-driven clinical decision support system (IDDAP) for infectious disease diagnosis and antibiotic prescription
Ying Shen 0001, Kaiqi Yuan, Daoyuan Chen, Joël Colloc, Min Yang 0007, Yaliang Li, Kai Lei
Artif. Intell. Medicine7
2018 An event summarizing algorithm based on the timeline relevance model in Sina Weibo
Kai Lei, Lizhu Zhang, Ying Liu 0021, Ying Shen 0001, Chenwei Liu, WeiTao Weng
Sci. China Inf. Sci.1
2018 Feature-enhanced attention network for target-dependent sentiment classification
Min Yang 0007, Qiang Qu 0001, Xiaojun Chen 0006, Chaoxue Guo, Ying Shen 0001, Kai Lei
Neurocomputing6
2018 CBN: Constructing a clinical Bayesian network based on data from the electronic medical record
Ying Shen 0001, Lizhu Zhang, Min Yang 0007, Buzhou Tang, Yaliang Li, Kai Lei
J. Biomed. Informatics7
2018 Adaptive community detection incorporating topology and content in social networks✰
abstract
In social network analysis , community detection is a basic step to understand the structure and function of networks. Some conventional community detection methods may have limited performance because they merely focus on the networks’ topological structure . Besides topology, content information is another significant aspect of social networks. Although some state-of-the-art methods started to combine these two aspects of information for the sake of the improvement of community partitioning, they often assume that topology and content carry similar information. In fact, for some examples of social networks, the hidden characteristics of content may unexpectedly mismatch with topology. To better cope with such situations, we introduce a novel community detection method under the framework of non-negative matrix factorization (NMF). Our proposed method integrates topology as well as content of networks and has an adaptive parameter (with two variations) to effectively control the contribution of content with respect to the identified mismatch degree. Based on the disjoint community partition result, we also introduce an additional overlapping community discovery algorithm, so that our new method can meet the application requirements of both disjoint and overlapping community detection. The case study using real social networks shows that our new method can simultaneously obtain the community structures and their corresponding semantic description , which is helpful to understand the semantics of communities. Related performance evaluations on both artificial and real networks further indicate that our method outperforms some state-of-the-art methods while exhibiting more robust behavior when the mismatch between topology and content is observed.
Meng Qin 0002, Di Jin 0001, Kai Lei, Bogdan Gabrys, Katarzyna Musial
Knowl. Based Syst.3
2018 An NDN IoT Content Distribution Model With Network Coding Enhanced Forwarding Strategy for 5G
abstract
The challenging requirements of fifth-generation (5G) Internet-of-Things (IoT) applications have motivated a desired need for feasible network architecture, while Named Data Networking (NDN) is a suitable candidate to support the high density IoT applications. To effectively distribute increasingly large volumes of data in large-scale IoT applications, this paper applies network coding techniques into NDN to improve IoT network throughput and efficiency of content delivery for 5G. A probability-based multipath forwarding strategy is designed for network coding to make full use of its potential. To quantify performance benefits of applying network coding in 5G NDN, this paper integrates network coding into a NDN streaming media system implemented in the ndnSIM simulator. The experimental results clearly and fairly demonstrate that considering network coding in 5G NDN can significantly improve the performance, reliability, and QoS. Besides, this is a general solution as it is applicable for most cache approaches. More importantly, our approach has promising potentials in delivering growing IoT applications including high-quality streaming video services.
Kai Lei, Shangru Zhong, Fangxing Zhu, Kuai Xu, Haijun Zhang 0001
IEEE Trans. Ind. Informatics1
2017 Statistical Optimal Hash-Based Longest Prefix Match
abstract
Longest Prefix Match (LPM) is a basic and important function for current network devices. Hash-based approaches appear to be excellent candidate solutions for LPM with the capability of fast lookup speed and low latency. The number of hash table probes, i.e. the search path of a hash-based LPM algorithm, directly determines the lookup performance. In this paper, we propose Ω-LPM to improve the lookup performance by optimizing the search path of the hash-based LPM. Ω-LPM first reconstructs the forwarding table to support random search [19], then it applies a dynamic programming algorithm to find the shortest search path based on the statistics of the matching probabilities. Ω-LPM concretely reduces the number of hash table probes via searching most of the packets in optimal search paths. Even in the worst case, the upper bound of the average search path of Ω-LPM is 1 + log2(N), here N is the length of the longest prefix in the routing table. The case studies of the name lookup in Named Data Networking and the IP lookup in current Internet demonstrate that Ω-LPM can shorten 61.04% and 86.88% search paths compared with the basic hash-based methods of name lookup [22] and IP lookup [12], respectively, furthermore Ω-LPM reduces 32.3% probes of the name lookup and 73.55% probes of the IP lookup compared with the optimal linear search. The experimental results conducted on extensional name tables and IP tables also show that Ω-LPM has both low memory overhead and excellent scalability.
Yi Wang 0004, Zhuyun Qi, Huichen Dai, Hao Wu 0023, Kai Lei, Bin Liu 0001
ANCS5
2017 A trust networks recommender algorithm based on Latent Factor Model
abstract
Recommender system in social networks has been a solution to the problem of information overload, but easily suffers the challenge of data sparsity and cold start. Meanwhile, since the concept of trust has gained wide concerns, many online social networks begin to introduce trust relationship between users, which motivates us to incorporate trust information to address above problems and improve the quality of recommendation in social networks. Latent Factor Model (LFM) such as SVD++ has been proven effectiveness in social recommendation task. Therefore, this paper proposes a trust networks recommender method based on LFM for rating prediction of recommending items. Our method exploits both explicit and implicit trust information to describe users' relationship and model their social behaviors. Furthermore, we combine LFM with neighborhood model to overcome its inability of interpretation. The experimental results on Epinions can demonstrate that trust information helps alleviate the problem of data sparsity and cold start. Besides, our combination algorithm achieves higher accuracy for prediction and outperforms other trust-based algorithms, including the state-of-the-art method.
Shangru Zhong, Weiyang Zhang, Qiang Zhang 0015, Kai Lei
ICC4
2017 Spam comments detection with self-extensible dictionary and text-based features
abstract
The new social media have become popular for information spreading, allowing online users to publish latest events and personal opinions. However, massive spam comments seriously decrease users' reading experience. To detect spam comments in Chinese social media, we employ semantic analysis to build the self-extensible dictionary which updates and extends itself with new cyber words automatically. The Semantic analysis brings extra semantic features which helps in text classification. Based on the statistical analysis of microblogging comments, we select four text-based features, which basically represent characteristics of Chinese spam comments. We use spam dictionary and text-based features to construct classifiers for detecting spam comments. Finally, we achieve an average detection accuracy of 93.6%, which is preferable to existing spam comments detection methods. Experimental results demonstrate that our method can effectively detect spam comments in Chinese microblogging field.
Qiang Zhang 0015, Chenwei Liu, Shangru Zhong, Kai Lei
ISCC4
2017 Anchor-Chain: A Seamless Producer Mobility Support Scheme in NDN
abstract
Named Data Networking (NDN) has natural advantages in the consumer mobility support for its design characteristics, while the producer mobility support was left unspecified. Some producer mobility support schemes have been proposed in recent years, which better solved the huge overhead caused by the route aggregation in normal NDN, yet critical issues like long handover latency and high packet loss rate still remain. In this paper, we improved and optimized a quite popular method anchor-based solution by setting multi-level anchor nodes strategically based on the network topology hierarchy, which we called Anchor-Chain, and the Interest packet forwarded to the producer would pass through it. After the producer moved, the newly generated Anchor-Chain would reuse the previous forwarding path as far as possible to guide more Interests to the new location of producer and reduce the Interest packets drop rate. Besides, we introduce a mobility preprocess mechanism that takes the advantage of connectionless and multi-path forwarding in NDN to establish the future Interest forwarding path in advance, which assists the producer performs "seamless" handover. By numerical analysis, our approach can reduce the handover latency and improve the response ratio compared with the main current solutions.
Yizhe Zheng, Xuewei Piao, Kai Lei
VTC Spring3
2017 Fast Parallel Path Concatenation for Graph Extraction
abstract
Heterogeneous graph is a popular data model to represent the real-world relations with abundant semantics. To analyze heterogeneous graphs, an important step is extracting homogeneous graphs from the heterogeneous graphs, called homogeneous graph extraction. In an extracted homogeneous graph, the relation is defined by a line pattern on the heterogeneous graph and the new attribute values of the relation are calculated by user-defined aggregate functions. The key challenges of the extraction problem are how to efficiently enumerate paths matched by the line pattern and aggregate values for each pair of vertices from the matched paths. To address above two challenges, we propose a parallel graph extraction framework, where we use vertex-centric model to enumerate paths and compute aggregate functions in parallel. The framework compiles the line pattern into a path concatenation plan, which determines the order of concatenating paths and generates the final paths in a divide-and-conquer manner. We introduce a cost model to estimate the cost of a plan and discuss three plan selection strategies, among which the best plan can enumerate paths in O(log)(l) iterations, where l is the length of a pattern. Furthermore, to improve the performance of evaluating aggregate functions, we classify the aggregate functions into three categories, i.e., distributive aggregation, algebraic aggregation, and holistic aggregation. Since the distributive and algebraic aggregations can be computed from the partial paths, we speed up the aggregation by computing partial aggregate values during the path enumeration.
Yingxia Shao, Kai Lei, Lei Chen 0002, Zi Huang, Bin Cui 0001, Zhongyi Liu 0001, Yunhai Tong, Jin Xu 0002
IEEE Trans. Knowl. Data Eng.2
2017 GVoS: A General System for Near-Duplicate Video-Related Applications on Storm
abstract
The exponential increase of online videos greatly enriches the life of users but also brings huge numbers of near-duplicate videos (NDVs) that seriously challenge the video websites. The video websites entail NDV-related applications such as detection of copyright violation, video monitoring, video re-ranking, and video recommendation. Since these applications adopt different features and different processing procedures due to diverse scenarios, constructing separate and special-purpose systems for them incurs considerable costs on design, implementation, and maintenance. In this article, we propose a general NDV system on Storm (GVoS)—a popular distributed real-time stream processing platform—to simultaneously support a wide variety of video applications. The generality of GVoS is achieved in two aspects. First, we extract the reusable components from various applications. Second, we conduct the communication between components via a mechanism called Stream Shared Message (SSM) that contains the video-related data. Furthermore, we present an algorithm to reduce the size of SSM in order to avoid the data explosion and decrease the network latency. The experimental results demonstrate that GVoS can achieve performance almost the same as the customized systems. Meanwhile, GVoS accomplishes remarkably higher systematic versatility and efficiently facilitates the development of various NDV-related applications.
Jiawei Jiang 0001, Yunhai Tong, Hua Lu 0001, Bin Cui 0001, Kai Lei, Lele Yu
ACM Trans. Inf. Syst.5
2016 Detecting spam comments posted in micro-blogs using the self-extensible spam dictionary
abstract
The high popularity of Weibo has greatly enriched people's lives, allowing online users to share their feelings through posting comments. However, more and more spam comments are also being posted in users' blogs on this social media. In this paper, in order to effectively detect spam comments in Chinese micro-blogs, we introduce semantic analysis to construct a Self-Extensible Spam Dictionary which automatically expands itself when new words emerge on the micro-blogs frequently. The use of semantic analysis can provide us with additional features which are beneficial to detecting spam comments. A Proportion-Weight Filter (PWF) model is also proposed to detect two kinds of spam comments (AD and vulgar comments), by filtering the spam-weight and the spam-proportion of the Weibo comments based on our Self-Extensible Spam Dictionary criteria. Our experimental results demonstrate that when detecting a combination of both AD and vulgar spam comments, we can achieve an average detection accuracy of 87.9%. Particularly for AD spam comments detection, we can achieve an average accuracy of 96.2%, which is preferable compared to when using machine learning methods. The statistical analysis of the results verifies that our proposed methods can identify the spam comments effectively and to relatively high degrees of accuracy.
Chenwei Liu, Jiawei Wang 0002, Kai Lei
ICC3
2016 A Peer-to-Peer File Sharing System over Named Data Networking
abstract
Named Data Networking (NDN), a promising Future Internet Architecture design, requires new experimental applications to demonstrate its performance and feasibility. Through designing, implementing, and evaluating NDNMaze, i.e., an NDN version of a widely deployed peer-to-peer file sharing application called IPMaze, we find that NDNMaze has a simpler system architecture with improved performance and flexibility. The innovative messaging mechanism and distributed hash tables (DHT) in NDNMaze simplified the implementations of key system components such as user management, nearest neighbor discovery, file discovery and distributions. To systematically evaluate the performance, we simulate both versions with NS-3 simulator, and collect a broad range of performance metrics including hop count, data request latency, data request efficiency, and network transmission efficiency. Our experimental results show that NDNMaze achieves better performance than IPMaze due to the NDN's advantages in content-centric data distribution and sharing. Our work shedslight for distributed application design in NDN.
Xuewei Piao, Yunbo Xun, Kai Lei
ICPADS5
2016 Online Learning for Accurate Real-Time Map Matching
Biwei Liang, Tengjiao Wang 0003, Shun Li 0001, Wei Chen 0021, Hongyan Li 0002, Kai Lei
PAKDD (2)6
2016 Dboost: A Fast Algorithm for DBSCAN-based Clustering on High Dimensional Data
Xiaorong Wang, Wei Chen 0021, Tengjiao Wang 0003, Kai Lei
PAKDD (2)6
2016 A Novel Trust Model for Activity Social Network Based on PeerTrust
abstract
Recently, kinds of Social Network Services (SNS) have gained enough popularity among Internet users. Activity Social Network (ASN) service, as a new kind of SNS with interest as the core, is dominated by the activities and it connects users much closer through social activities. However, cheating behaviors appear frequently in SNS especially ASN because of their anonymity, which makes a large lack of trust in ASN and becomes a stumbling block that hinders the development of ASN. The researches on trust mechanism have become a key issue in recent years, but most studies focused on E-commerce and traditional social networks. The existing models are not completely suitable for social activities. Motivated by the idea of PeerTrust to compute trust values, we propose ActivityTrust model based on PeerTrust according to its unequal interaction characteristics to ensure the security and reliability of the activity social platform. Meanwhile, we build a simulative ASN platform on NetLogo, and make contrast experiments on it. We verify the effectiveness and adaptability of trust model with regards to activity success rate and trust evaluation rate.
Limei Xu, Kai Lei
PDCAT3
2016 Internet Traffic Analysis in a Large University Town: A Graphical and Clustering Approach
WeiTao Weng, Kai Lei, Kuai Xu, Xiaoyou Liu, Tao Sun 0010
WAIM (1)2
2015 ASEM: Mining Aspects and Sentiment of Events from Microblog
abstract
Microblogs contain the most up-to-date and abundant opinion information on current events. Aspect-based opinion mining is a good way to get a comprehensive summarization of events. The most popular aspect based opinion mining models are used in the field of product and service. However, existing models are not suitable for event mining. In this paper we propose a novel probabilistic generative model (ASEM) to simultaneously discover aspects and the specified opinions. ASEM incorporate a sequence labeling model(CRF) into a generative topic model. Additionally, we adopt a set of features for separating aspects and sentiments. Moreover, we novelly present a continuously learning model. It can utilize the knowledge of one event to learn another, and get a better performance. We use five real world events to do experiment. The experimental results show that ASEM extracts aspects and sentiments well, and ASEM outperforms other state-of-art models and the intuitive two-step method.
Ruhui Wang, Weijing Huang, Wei Chen 0021, Tengjiao Wang 0003, Kai Lei
CIKM5
2015 MDPF: An NDN Probabilistic Forwarding Strategy Based on Maximizing Deviation Method
abstract
Forwarding strategy is the key feature of Named Data Networking (NDN) to realize dynamic, adaptive and intelligent forwarding, but work in this area is still at a very preliminary stage. In this paper, selecting which forwarding interface among multiple alternatives in NDN is defined as a multiple attribute decision making (MADM) problem and a maximizing deviation based probabilistic forwarding (MDPF) strategy is proposed to select forwarding interface on probability. Since multiple network metrics such as interface status, pending Interest numbers are considered together, each alternative interface's availability is obtained more accurately. Thus, better content delivery efficiency can be achieved. In addition, MDPF provides good extensibility, as any appropriate metric can be added to enhance or customize it. We implement the proposal in ndnSIM and compare it with BestRoute and PI-based strategies under various topologies and scenarios. Experimental results show that MDPF strategy is more responsive and sensitive to network changes, and can realize higher throughput, lower drop rate as well as better load balance.
Kai Lei, Jiawei Wang 0002
GLOBECOM1
2015 An entropy-based probabilistic forwarding strategy in Named Data Networking
abstract
The forwarding strategy is the key to the resiliency and efficiency of Named Data Networking (NDN), which is a new and fundamental research area. For forwarding strategy, dynamically selecting an optimal interface from multiple alternative interfaces to forward an Interest packet is indeed a multiple attribute decision making (MADM) problem. In this paper, an entropy-based probabilistic forwarding (EPF) strategy is proposed to make a stochastic interface selection based on the combination of interfaces' dynamic availabilities and static routing information, which achieves better load balance in comparison with deterministic interface selection. By objectively assigning weights to attributes and considering multiple real-time network condition metrics, EPF can obtain the availabilities of interfaces more accurately and comprehensively. Since additional network metrics can be easily added and integrated into interfaces' assessment model, EPF provides good extensibility. In addition, we innovatively define two parameters (γ, δ) which can be used to trade off the effect factors between static routing information and dynamic running status of interfaces to customize EPF strategy for different network and application scenarios. Experiments show that EPF can realize preferable load balance and achieve higher throughput compared to the representative BestRoute forwarding strategy.
Kai Lei, Jiawei Wang 0002
ICC1
2015 Extracting unknown words from Sina Weibo via data clustering
abstract
Sina Weibo, a Twitter-like microblogging site attracting over 240 million monthly active users to tweet, retweet, and comment, has rapidly become one of the most popular social media sites in China. As many users create new and innovative words on their tweets and comments, it is necessary to extract these emerging words, which do not exist in today's Chinese vocabulary or dictionary. Towards this end, this paper proposes a novel method based on data clustering of Weibo users and tweets for extracting unknown words from Weibo tweets and comments. Specifically, relying on the similarity of the users who post the tweets, we apply a hierarchical clustering to divide Weibo data into distinct groups, e.g., sports, news stories, movies, before extraction. Comparing with the method of unclustered Weibo data, our experimental results have successfully demonstrated the benefits of the proposed data clustering scheme for improving the recall and accuracy of extracting unknown Chinese words from tweets and comments.
Kai Lei, Weiyang Zhang, Kuai Xu
ICC1
2015 Profiling the followers of the most influential and verified users on Sina Weibo
abstract
The new social media such as Twitter and Sina Weibo has become an increasingly popular channel for spreading influence, challenging traditional media such as TVs and newspapers. The most influential and verified users, also called big-V accounts on Sina Weibo often attract million of followers and fans, creating massive “celebrity-centric” social networks on the social media, which play a key role in disseminating breaking news, latest events, and controversial opinions on social issues. Given the importance of these accounts, it is very crucial to understand social networks and user influence of these accounts and profile their followers' behaviors. Towards this end, this paper monitors a selected group of influential users on Sina Weibo and collects their tweet streams as well as retweeting and commenting activities on these tweets from their followers. Our analysis on tweet data streams from Sina Weibo reveals when and what the followers comment on the tweets of these influential users, and discovers different temporal patterns and word diversity in the comments. Based on the insight gained from follower characteristics, we further develop simple and intuitive algorithms for classifying the followers into spammers and normal fans. Our experimental results demonstrate that the proposed algorithms are able to achieve an average accuracy of 95.20% in detecting spammers from the followers who have commented on the tweets of these influential accounts.
Kai Lei, Kuai Xu
ICC2
2015 Network Coding for Effective NDN Content Delivery: Models, Experiments, and Applications
abstract
How to effectively distribute and share increasingly large volumes of data in large-scale network applications is a key challenge for Internet infrastructure. Although NDN, a promising new future internet architecture which takes data oriented transfer approaches, aims to better solve such needs than IP, it still faces problems like data redundancy transmission and inefficient in-network cache utilization. This paper combines network coding techniques to NDN to improve network throughput and efficiency. The merit of our design is that it is able to avoid duplicate and unproductive data delivery while transferring disjoint data segments along multiple paths and with no excess modification to NDN fundamentals. To quantify performance benefits of applying network coding in NDN, we integrate network coding into an NDN streaming media system implemented in the ndn SIM simulator. Basing on BRITE generated network topologies in our simulation, the experimental results clearly and fairly demonstrate that considering network coding in NDN can significantly improve the performance, reliability and QoS. More importantly, our approach is capable of and well fit for delivering growing Big Data applications including high-performance and high-density video streaming services.
Kai Lei, Fangxing Zhu, Kuai Xu
ICPP1
2015 Overlapping Community Detection in Directed Heterogeneous Social Network
Changhe Qiu, Wei Chen 0021, Tengjiao Wang 0003, Kai Lei
WAIM4
2015 An Adaptive Skew Handling Join Algorithm for Large-scale Data Analysis
Tengjiao Wang 0003, Shun Li 0001, Hongyan Li 0002, Kai Lei
WAIM6
2015 Emerging medical informatics with case-based reasoning for aiding clinical decision in multi-agent system
Ying Shen 0001, Joël Colloc, Armelle Jacquet-Andrieu, Kai Lei
J. Biomed. Informatics4
2014 An Adaptive Skew Insensitive Join Algorithm for Large Scale Data Analytics
Wenjing Liao, Tengjiao Wang 0003, Hongyan Li 0002, Dongqing Yang, Kai Lei
APWeb6
2014 Topic-Based Sentiment Analysis Incorporating User Interactions
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei
APWeb6
2014 An encryption and probability based access control model for named data networking
abstract
The new named data networking (NDN) has shifted the Internet from today's IP-based packet-delivery model to the name-based data retrieval model. The architecture shift from IP addresses to named data results in effective content delivery via in-networking cache and direct object retrieval. However, this shift has also created challenges and obstacles for securing data objects and providing appropriate access control on named data due to broad data replications and the loss of network perimeters. This paper designs, implements, and evaluates an encryption and probability based access control model for NDN with video streaming service as a case study. In particularly, we explore a combination of public-key cryptography and symmetric ciphers to encrypt video data for preventing unauthorized access. In addition, we build a bloom-filter probabilistic data structure for pre-filtering Interests from consumers without desired credentials. Our experimental results have demonstrated the capabilities of the proposed model for providing access control while incurring low system and performance overhead on producers and consumers.
Kai Lei, Kuai Xu
IPCCC2
2014 Hot topic analysis and content mining in social media
abstract
Sina Weibo has become an increasingly critical social media in China for sharing latest news, marketing new products, and discussing controversial issues. The rising importance of Sina Weibo on the society makes it very important to understand “what”, “when”, “who” on hot topics that are being continuously tweeted and searched by millions of active users. In this paper, we develop a systematic approach to characterize temporal distribution of hot topics searched by Sina Weibo users over a four-month time-span and to uncover correlated hot topics that are not only tweeted by the same users, but also appear in the similar set of tweet messages. We analyze real-time Sina Weibo tweet data streams and study volume correlations and temporal gaps between user searches and tweeting activities on hot topics. In addition, we examine the correlations between hot topic searches on social media and on search engines to understand hot topics and user behaviors across different platforms. Given the challenges of analyzing massive amount of tweet data, we explore Hadoop MapReduce framework to effectively process millions of tweets from the collected data-sets, and quantify the performance benefits of MapReduce on analyzing tweet streams. To the best of our knowledge, this paper is the first effort to characterize temporal search patterns of hot topics on Sina Weibo and to study their correlations with tweeting data streams as well as search engine statistics.
WeiTao Weng, Kai Lei, Kuai Xu
IPCCC4
2014 Sarcasm Detection in Social Media Based on Imbalanced Classification
Wei Chen 0021, Gaoyan Ou, Tengjiao Wang 0003, Dongqing Yang, Kai Lei
WAIM6
2014 Characterizing Tweeting Behaviors of Sina Weibo Users via Public Data Streaming
Kai Lei, Kuai Xu
WAIM3
2013 Minimum storage BASIC codes: A system perspective
abstract
The explosion of big data stored in distributed file systems calls for more efficient storage paradigms. While replication is widely used to ensure data availability, erasure codes provide a much better tradeoff between storage and availability. Reed-Solomon (RS) codes are the standard design choice, however, their high repair cost is often considered an unavoidable price to pay for high storage efficiency and high reliability. BASIC codes can achieve the optimal tradeoff between storage capacity and repair bandwidth with much less complexity of regenerating codes, which is first proposed in [1]. This paper integrate one construction of the minimum storage BASIC (MS-BASIC) codes [2] into a Hadoop HDFS cluster testbed with up to 22 storage nodes. We demonstrate that MS-BASIC codes conform to the theoretical findings and achieve recovery bandwidth saving compared to the conventional recovery approach based on RS codes.
Xianxia Huang, Hui Li 0022, Tai Zhou, Hanxu Hou, Kai Lei
IEEE BigData8
2013 Understanding Sina Weibo online social network: A community approach
abstract
Sina Weibo, one of the most popular online social networks in China, has recently become a critical medium for Internet users to disseminate and discuss breaking news, social events and other information. Although online social networks and social media have received significant attention from the research community, few studies have focused on Sina Weibo due to the lack of data collection. Given the sheer size of Sina Weibo online social network and vast amount of tweets, retweets and comments, this paper introduces a novel community approach for understanding Sina Weibo online social network. Specifically, we collect all Weibo users registered with Shenzhen as primary geographic location, and build a Shenzhen Weibo community graph based on their following or follower relationships. Our experimental results describe interesting graphical characteristics such as clustering coefficients of this community graph, and reveal the impact of user popularity on tweet influence. Through modeling interactions of Shenzhen Weibo users and their tweeted messages with bipartite graphs and one-mode projections, we analyze the similarity of retweeting and commenting activities among these users, and discuss the implications of the findings on understanding different types of user accounts and the motivations of their following and retweeting behaviors. To the best of our knowledge, this study is the first effort to introduce a community approach for understanding the community characteristics of Sina Weibo and characterizing the similarity of retweeting behaviors and following relationships.
Kai Lei, Kuai Xu
GLOBECOM1
2013 Massively parallel learning of Bayesian networks with MapReduce for factor relationship analysis
abstract
Bayesian Network (BN) is one of the most popular models in data mining technologies. Most of the algorithms of BN structure learning are developed for the centralized datasets, where all the data are gathered into a single computer node. They are often too costly or impractical for learning BN structures from large scale data. Through a simple interface with two functions, map and reduce, MapReduce facilitates parallel implementation of many real-world tasks such as data processing for search engines and machine learning. In this paper, we present a parallel algorithm for BN structure leaning from large-scale dateset by using a MapReduce cluster. We discuss the benefits of using MapReduce for BN structure learning, and demonstrate the performance of this approach by applying it to a real world financial factor relationships learning task from the domain of financial analysis.
Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu
IJCNN4
2013 Size-Constrained Clustering Using an Initial Points Selection Method
Kai Lei, Sibo Wang 0003
KSEM1
2013 Aspect-Specific Polarity-Aware Summarization of Online Reviews
Gaoyan Ou, Wei Chen 0021, Tengjiao Wang 0003, Dongqing Yang, Kai Lei, Yueqin Liu
WAIM6