EDBT 2026 Demo / reviewers in the wild / expert
Jianlong Tan
dblp:19/336
· DBLP profile ↗
66ranked-venue papers
1as first author
16since 2021 · last 2023
0009-0001-4558-560XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 6 since 2021Databases, data management, data science and information retrieval · 20 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 since 2021Computer networks · 7 · 1 since 2021Security and privacy · 3Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Research on the construction of event corpus with document-level causal relations for social securityabstractEvent corpora are imperative to train event extraction models. Currently, most existing event corpora suffer from being available only in English, and their construction is limited by high annotation costs. This paper aims to construct a corpus that concerns social security causality events in Chinese and proposes a faster and less expensive construction method. The contributions are as follows: (i) An event corpus SSECau for the social security field in Chinese is constructed. They are from 2,235 web texts and microblogs, with event causality annotated at the document level. (ii) A corpus construction method with manual annotation and machine pre-tagging is proposed to improve accuracy and speed. (iii) A pre-tagging method based on BiLSTM-CRF (bidirectional long short-term memory and conditional random field) is deployed to extract events automatically. The experimental results show the best consistency between automatic pre-tagging and manual annotation can reach up to 82 %; while the dynamic tagging process improves both the labeling speed and accuracy. The SSECau corpus can aid the development and evaluation of event extraction models for the social security field; annotated cause-effect relationships at the document level can potentially enhance the training of complex extraction models; the proposed dynamic process with pre-tagging can serve as a reference for future corpus construction. Ga Xiang, Yangsen Zhang, Jianlong Tan, Zihan Ran, En Shi |
Inf. Process. Manag. | 3 |
| 2023 | Multi-View Tensor Graph Neural Networks Through Reinforced AggregationabstractGraph Neural Networks (GNNs) have yielded fruitful results in learning multi-view graph data. However, it is challenging for existing GNNs to capture the potential correlation information (PCI) among the graph structure features of multiple views. It is also challenging to adaptively identify valuable neighbors for node feature fusion in different views. To this end, we propose a novelReinforcedTensorGraphNeuralNetwork (RTGNN) framework to more effectively perform multi-view graph representation learning through reinforcing inter- and intra-graph aggregation. Specifically, RTGNN first uses tensor decomposition to extract the graph structure features (GSFs) of each view in the common feature space. These GSFs contain the PCI of multiple views and alleviate fusion conflicts that may be caused by differences between view feature spaces in cross-view feature fusion. Since fusing the features of all neighbor nodes may harm the features of the center node, we filter the irrelevant neighbors to improve the performance of intra-graph aggregation in each view. Concretely, a reinforcement learning (RL)-guided scheme is developed to automatically calculate the optimal filtering threshold for each view, avoiding tedious manual updates and infeasible back propagation updates. Experimental results and analysis on five datasets show that RTGNN surpasses the best multi-view graph representation baselines and achieves the maximum 14.26% performance improvement in terms of F1. The code link ishttps://github.com/RingBDStack/RTGNN. Xusheng Zhao, Qiong Dai, Jia Wu 0001, Hao Peng 0001, Mingsheng Liu, Jianlong Tan, Senzhang Wang, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Dynamic Nonlinear Mixup with Distance-based Sample SelectionabstractData augmentation with mixup has shown to be effective on the NLP tasks. Although its great success, the mixup still has shortcomings. First, vanilla mixup randomly selects one sample to generate the mixup sample for a given sample. It remains unclear how to best choose the input samples for the mixup. Second, linear interpolation limits the space of synthetic data and its regularization effect. In this paper, we propose the dynamic nonlinear mixup with distance-based sample selection, which not only generates multiple sample pairs based on the distance between each sample but also enlarges the space of synthetic samples. Specifically, we compute the distance between each input data by cosine similarity and select multiple samples for a given sample. Then we use the dynamic nonlinear mixup to fuse sample pairs. It does not use a linear, scalar mixing strategy, but a nonlinear interpolation strategy, where the mixing strategy is adaptively updated for the input and label pairs. Experiments on the multiple public datasets demonstrate that dynamic nonlinear mixup outperforms state-of-the-art methods. Shaokang Zhang, Lei Jiang 0003, Jianlong Tan |
COLING | 3 |
| 2022 | Heterogeneous Graph Attention Network for Malicious Domain Detection
Zhiping Li, Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Fang Fang 0009, Jianlong Tan |
ICANN (2) | 6 |
| 2022 | A Bert Based Joint Learning Model with Feature Gated Mechanism for Spoken Language UnderstandingabstractIntent detection (ID) and slot filling (SF) are two major tasks for spoken language understanding (SLU). Recent joint learning approaches consider the relationship between intent detection and slot filling, which leverage the shared knowledge across two tasks to benefit each other. However, most existing methods do not make full use of the BERT model and gate mechanisms to improve the semantic correlation between slot filling and intent detection tasks. In this paper, we propose a joint learning model based on BERT, which introduce dual encoder structure and utilizes semantic information by performing feature gate mechanisms in predicting intents and slots. Experimental results demonstrate that our proposed method provides very competitive results on CAIS and DDoST datasets. Lei Jiang 0003, Shaokang Zhang, Jianlong Tan |
ICASSP | 5 |
| 2022 | DOM2R-Graph: A Web Attribute Extraction Architecture with Relation-Aware Heterogeneous Graph Transformer
Jiali Feng, Cong Cao 0001, Fangfang Yuan, Zhiping Li, Yanbing Liu 0007, Jianlong Tan |
ICONIP (1) | 7 |
| 2022 | Malicious Domain Detection with Heterogeneous Graph Propagation Network
Fangfang Yuan, Yanbing Liu 0007, Cong Cao 0001, Jianlong Tan |
WASA (1) | 6 |
| 2022 | Cross-domain knowledge distillation for text classification
Shaokang Zhang, Lei Jiang 0003, Jianlong Tan |
Neurocomputing | 3 |
| 2022 | Knowledge Graph Embedding by Double Limit Scoring LossabstractKnowledge graph embedding is an effective way to represent knowledge graph, which greatly enhance the performances on knowledge graph completion tasks, e.g., entity or relation prediction. For knowledge graph embedding models, designing a powerful loss framework is crucial to the discrimination between correct and incorrect triplets. Margin-based ranking loss is a commonly used negative sampling framework to make a suitable margin between the scores of positive and negative triples. However, this loss can not ensure ideal low scores for the positive triplets and high scores for the negative triplets, which is not beneficial for knowledge completion tasks. In this paper, we present a double limit scoring loss to separately set upper bound for correct triplets and lower bound for incorrect triplets, which provides more effective and flexible optimization for knowledge graph embedding. Upon the presented loss framework, we present several knowledge graph embedding models including TransE-SS, TransH-SS, TransD-SS, ProjE-SS and ComplEx-SS. The experimental results on link prediction and triplet classification show that our proposed models have the significant improvement compared to state-of-the-art baselines. Xiaofei Zhou 0002, Lingfeng Niu, Qiannan Zhu, Xingquan Zhu 0001, Ping Liu 0001, Jianlong Tan, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2021 | Image Captioning with Context-Aware Auxiliary GuidanceabstractImage captioning is a challenging computer vision task, which aims to generate a natural language description of an image. Most recent researches follow the encoder-decoder framework which depends heavily on the previous generated words for the current prediction. Such methods can not effectively take advantage of the future predicted information to learn complete semantics. In this paper, we propose Context-Aware Auxiliary Guidance (CAAG) mechanism that can guide the captioning model to perceive global contexts. Upon the captioning model, CAAG performs semantic attention that selectively concentrates on useful information of the global predictions to reproduce the current generation. To validate the adaptability of the method, we apply CAAG to three popular captioners and our proposal achieves competitive performance on the challenging Microsoft COCO image captioning benchmark, e.g. 132.2 CIDEr-D score on Karpathy split and 130.7 CIDEr-D (c40) score on official online evaluation server. Zeliang Song, Zhendong Mao 0001, Jianlong Tan |
AAAI | 4 |
| 2021 | Multiphish: Multi-Modal Features Fusion Networks for Phishing DetectionabstractPhishing is an increasingly serious cybercrime. Phishers create phishing websites by mimicking legitimate websites to confuse users and steal their personal information. The proliferation of phishing websites and more advanced camouflage techniques are problems faced by most existing methods. In this paper, we propose a features fusion networks (MultiPhish) which is the first study on fusing multi-modal features with neural networks for the phishing detection task. In this end-to-end network, the domain and favicon of the website are represented via deep neural networks, and the representation of the website identity is obtained through multi-modal features fusion. In addition, the variation autoencoder (VAE) is introduced to optimize the representation. In the phishing detection module, we incorporate URL features to improve situations where phishing websites cannot be detected only by estimating whether the website identity is disguised. Based on the latest collected dataset, we have carried out extensive experiments and proved that our model is superior to the relevant methods. In addition, MultiPhish is a completely language-independent strategy, so it can perform phishing detection regardless of the text language. Luchen Liu, Jianlong Tan |
ICASSP | 4 |
| 2021 | Malicious Domain Detection on Imbalanced Data with Deep Reinforcement Learning
Fangfang Yuan, Teng Tian, Yanmin Shang, Yuhai Lu, Yanbing Liu 0007, Jianlong Tan |
ICONIP (4) | 6 |
| 2021 | End-to-end Boundary Exploration for Weakly-supervised Semantic SegmentationabstractIt is full of challenges for weakly supervised semantic segmentation (WSSS) acquiring the pixel-level object location with only image-level annotations. Especially, the single-stage methods learn image- and pixel-level labels simultaneously to avoid complicated multi-stage computations and sophisticated training procedures. In this paper, we argue that using a single model to accomplish image- and pixel-level classification will fall into the balance of multi-target and consequently weakens the recognition capability. Because the image-level task tends to learn position-independent features, but the pixel-level task tends to be position-sensitive. Hence, we propose an effective encoder-decoder framework to explore object boundaries and solve the above dilemma. The encoder and decoder learn position-independent and position-sensitive features independently during the end-to-end training. In addition, a global soft pooling is suggested to suppress background pixels' activation for the encoder training and further improve the class activation map (CAM) performance. The edge annotations for the decoder training are synthesized by the high confidence CAMs, which do not requires extra supervision. The extensive experiments on the Pascal VOC12 dataset demonstrate that our method achieves state-of-the-art compared to the end-to-end approaches. It gets 63.6% and 65.7% mIoU scores on val and test sets respectively. Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Yue Hu 0002, Jianlong Tan |
ACM Multimedia | 6 |
| 2021 | Direction Relation Transformer for Image CaptioningabstractImage captioning is a challenging task that combines computer vision and natural language processing for generating a textual description of the content within an image. Recently, Transformer-based encoder-decoder architectures have shown great success in image captioning, where multi-head attention mechanism is utilized to capture the contextual interactions between object regions. However, such methods regard region features as a bag of tokens without considering the directional relationships between them, making it hard to understand the relative position between objects in the image and generate correct captions effectively. In this paper, we propose a novel Direction Relation Transformer to improve the orientation perception between visual features by incorporating the relative direction embedding into multi-head attention, termed DRT. We first generate the relative direction matrix according to the positional information of the object regions, and then explore three forms of direction-aware multi-head attention to integrate the direction embedding into Transformer architecture. We conduct experiments on challenging Microsoft COCO image captioning benchmark. The quantitative and qualitative results demonstrate that, by integrating the relative directional relation, our proposed approach achieves significant improvements over all evaluation metrics compared with baseline model, e.g., DRT improves task-specific metric CIDEr score from 129.7% to 133.2% on the offline ''Karpathy'' test split. Zeliang Song, Xiaofei Zhou 0002, Linhua Dong, Jianlong Tan, Li Guo 0001 |
ACM Multimedia | 4 |
| 2021 | Discriminative Representation Learning for Cross-Domain Sentiment Classification
Shaokang Zhang, Lei Jiang 0003, Huailiang Peng, Qiong Dai, Jianlong Tan |
PAKDD (2) | 5 |
| 2021 | Knowledge Base Reasoning with Convolutional-Based Recurrent Neural NetworksabstractRecurrent neural network(RNN) has achieved remarkable performances in complex reasoning on knowledge bases, which usually takes as inputs vector embeddings of relations along a path between an entity pair. However, it is insufficient to extract local correlations of a path due to RNN is better at capturing global sequential information of a path. In this paper, we take full advantages of convolutional neural network that can effectively extract local features, and propose a convolutional-based RNN architecture denoted as C-RNN to perform reasoning. C-RNN first utilizes CNN to extract local high-level correlation features of a path, and then feeds the correlation features into recurrent neural network to model the path representation. Our C-RNN architecture is adaptable to obtain not only local features but also global sequential features of a path. Based on C-RNN architecture, we devise two models, the unidirectional C-RNN and bidirectional C-RNN. We empirically evaluate them on a large-scale FreeBase+ClueWeb prediction task. Experimental results show that C-RNN models achieve state-of-the-art predictive performance. Qiannan Zhu, Xiaofei Zhou 0002, Jianlong Tan, Li Guo 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Type-Aware Anchor Link Prediction across Heterogeneous Networks Based on Graph Attention NetworkabstractAnchor Link Prediction (ALP) across heterogeneous networks plays a pivotal role in inter-network applications. The difficulty of anchor link prediction in heterogeneous networks lies in how to consider the factors affecting nodes alignment comprehensively. In recent years, predicting anchor links based on network embedding has become the main trend. For heterogeneous networks, previous anchor link prediction methods first integrate various types of nodes associated with a user node to obtain a fusion embedding vector from global perspective, and then predict anchor links based on the similarity between fusion vectors corresponding with different user nodes. However, the fusion vector ignores effects of the local type information on user nodes alignment. To address the challenge, we propose a novel type-aware anchor link prediction across heterogeneous networks (TALP), which models the effect of type information and fusion information on user nodes alignment from local and global perspective simultaneously. TALP can solve the network embedding and type-aware alignment under a unified optimization framework based on a two-layer graph attention architecture. Through extensive experiments on real heterogeneous network datasets, we demonstrate that TALP significantly outperforms the state-of-the-art methods. Yanmin Shang, Yanan Cao 0001, Yangxi Li, Jianlong Tan, Yanbing Liu 0007 |
AAAI | 5 |
| 2020 | A Knowledge-Aware Attentional Reasoning Network for RecommendationabstractKnowledge-graph-aware recommendation systems have increasingly attracted attention in both industry and academic recently. Many existing knowledge-aware recommendation methods have achieved better performance, which usually perform recommendation by reasoning on the paths between users and items in knowledge graphs. However, they ignore the users' personal clicked history sequences that can better reflect users' preferences within a period of time for recommendation. In this paper, we propose a knowledge-aware attentional reasoning network KARN that incorporates the users' clicked history sequences and path connectivity between users and items for recommendation. The proposed KARN not only develops an attention-based RNN to capture the user's history interests from the user's clicked history sequences, but also a hierarchical attentional neural network to reason on paths between users and items for inferring the potential user intents on items. Based on both user's history interest and potential intent, KARN can predict the clicking probability of the user with respective to a candidate item. We conduct experiment on Amazon review dataset, and the experimental results demonstrate the superiority and effectiveness of our proposed KARN model. Qiannan Zhu, Xiaofei Zhou 0002, Jia Wu 0001, Jianlong Tan, Li Guo 0001 |
AAAI | 4 |
| 2020 | DistilSum: : Distilling the Knowledge for Extractive SummarizationabstractA popular choice for extractive summarization is to conceptualize it as sentence-level classification, supervised by binary labels. While the common metric ROUGE prefers to measure the text similarity, instead of the performance of classifier. For example, BERTSUMEXT, the best extractive classifier so far, only achieves a precision of 32.9% at the top 3 extracted sentences ([email protected]) on CNN/DM dataset. It is obvious that current approaches cannot model the complex relationship of sentences exactly with 0/1 targets. In this paper, we introduce DistilSum, which contains teacher mechanism and student model. Teacher mechanism produces high entropy soft targets at a high temperature. Our student model is trained with the same temperature to match these informative soft targets and tested with temperature of 1 to distill for ground-truth labels. Compared with large version of BERTSUMEXT, our experimental result on CNN/DM achieves a substantial improvement of 0.99 ROUGE-L score (text similarity) and 3.95 [email protected] score (performance of classifier). Our source code will be available on Github. Ruipeng Jia, Yanan Cao 0001, Haichao Shi, Fang Fang 0009, Yanbing Liu 0007, Jianlong Tan |
CIKM | 6 |
| 2020 | Data Augmentation for Insider Threat Detection with GANabstractIn insider threat detection domain, the datasets are highly imbalanced, where the number of user's normal behavior is higher than that of insider's anomalous behavior. A direct approach to handle the class imbalance problem is using data augmentation on the minority class. Existing data augmentation methods mainly produce synthetic samples according with the linear operation based on samples of the minority class. Hence, these methods just focus on local information which leads to the unitarily of the synthetic samples, resulting in overfitting. To enrich the diversity of the synthetic samples, we propose a deep adversarial insider threat detection (DAITD) framework using the Generative Adversarial Networks (GAN) to approximate the true anomalous behavior distribution. Specifically, we first obtain anomalous user behavior representations from the anomalous behavior data (minority class), and then use the generator of the GAN to model the actual anomalous behavior distribution, use the discriminator of the GAN to distinguish whether the synthetic sample from the generator is real or not. In this way, our method is able to generate high quality synthetic samples that are close to the anomalous user behavior. Experimental results show that the DAITD framework outperforms other comparative inside threat detection algorithms. Fangfang Yuan, Yanmin Shang, Yanbing Liu 0007, Yanan Cao 0001, Jianlong Tan |
ICTAI | 5 |
| 2020 | Automatic Cerebral Artery System Labeling Using Registration and Key Points Tracking
Mengjun Shen, Jianyong Wei, Jitao Fan, Jianlong Tan, Zhenchang Wang, Zhenghan Yang, Penggang Qiao, Fangzhou Liao |
KSEM (1) | 4 |
| 2020 | Category-Level Adversarial Network for Cross-Domain Sentiment Classification
Shaokang Zhang, Huailiang Peng, Yanan Cao 0001, Lei Jiang 0003, Qiong Dai, Jianlong Tan |
KSEM (2) | 6 |
| 2020 | An Interactive Two-Pass Decoding Network for Joint Intent Detection and Slot Filling
Huailiang Peng, Mengjun Shen, Lei Jiang 0003, Qiong Dai, Jianlong Tan |
NLPCC (2) | 5 |
| 2020 | Cross-modal knowledge reasoning for knowledge-based visual question answering
Jing Yu 0007, Yujing Wang 0002, Weifeng Zhang 0002, Yue Hu 0002, Jianlong Tan |
Pattern Recognit. | 6 |
| 2020 | Reasoning on the Relation: Enhancing Visual Representation for Visual Question Answering and Cross-Modal RetrievalabstractCross-modal analysis has become a promising direction for artificial intelligence. Visual representation is crucial for various cross-modal analysis tasks that require visual content understanding. Visual features which contain semantical information can disentangle the underlying correlation between different modalities, thus benefiting the downstream tasks. In this paper, we propose a Visual Reasoning and Attention Network (VRANet) as a plug-and-play module to capture rich visual semantics and help to enhance the visual representation for improving cross-modal analysis. Our proposed VRANet is built based on the bilinear visual attention module which identifies the critical objects. We propose a novel Visual Relational Reasoning (VRR) module to reason about pair-wise and inner-group visual relationships among objects guided by the textual information. The two modules enhance the visual features at both relation level and object level. We demonstrate the effectiveness of the proposed VRANet by applying it to both Visual Question Answering (VQA) and Cross-Modal Information Retrieval (CMIR) tasks. Extensive experiments conducted on VQA 2.0, CLEVR, CMPlaces, and MS-COCO datasets indicate superior performance comparing with state-of-the-art work. Jing Yu 0007, Weifeng Zhang 0002, Zengchang Qin, Yue Hu 0002, Jianlong Tan, Qi Wu 0001 |
IEEE Trans. Multim. | 6 |
| 2019 | DSINE: Deep Structural Influence Learning via Network EmbeddingabstractStructural representations of user social influence are critical for a variety of applications such as viral marketing and recommendation products. However, existing studies only focus on capturing and preserving the structure of relations, and ignore the diversity of influence relations patterns among users. To this end, we propose a deep structural influence learning model to learn social influence structure via mining rich features of each user, and fuse information from the aligned selfnetwork component for preserving global and local structure of the influence relations among users. Experiments on two real-world datasets demonstrate that the proposed model outperforms the state-of-the-art algorithms for learning rich representations in multi-label classification task. Jianjun Wu 0004, Ying Sha, Bo Jiang 0013, Jianlong Tan |
AAAI | 4 |
| 2019 | DAN: Deep Attention Neural Network for News RecommendationabstractWith the rapid information explosion of news, making personalized news recommendation for users becomes an increasingly challenging problem. Many existing recommendation methods that regard the recommendation procedure as the static process, have achieved better recommendation performance. However, they usually fail with the dynamic diversity of news and user’s interests, or ignore the importance of sequential information of user’s clicking selection. In this paper, taking full advantages of convolution neural network (CNN), recurrent neural network (RNN) and attention mechanism, we propose a deep attention neural network DAN for news recommendation. Our DAN model presents to use attention-based parallel CNN for aggregating user’s interest features and attention-based RNN for capturing richer hidden sequential features of user’s clicks, and combines these features for new recommendation. We conduct experiment on real-world news data sets, and the experimental results demonstrate the superiority and effectiveness of our proposed DAN model. Qiannan Zhu, Xiaofei Zhou 0002, Zeliang Song, Jianlong Tan, Li Guo 0001 |
AAAI | 4 |
| 2019 | Chinese Social Media Entity Linking Based on Effective Context with Topic SemanticsabstractOn social media, entity linking is very important for natural language processing tasks, such as Sentiment Analysis, Question Answering (QA) and Machine Translation. Compared to English-oriented entity linking, Chinese entity linking has its special difficulties. Just like the entity linking for short text, Chinese microblogs have lots of noise and the mention lacks effective context information. In order to solve these problems, we present a new model for Chinese microblogs entity linking. Entity linking usually includes two steps: candidate entities generation and candidate entities ranking. First, based on the characteristics of Chinese, we put forward multi-method fusion strategies for candidate generation to improve the recall rate of candidate entities. Second, we propose a new neural network model called TAS (Topic attention Siamese) for candidate entities ranking. In TAS model, we add effective topic semantics on Siamese network to learn representations of context, mention and entity, and rank the mention-entity similarity. The representation of mention incorporates information from multiple sentences on the same topic, which can effectively solve the problem of the lack of contextual information. We also use Character-enhanced Word Embedding model (CWE) to pre-train both word embedding and characters embedding to work out noise and word segmentation impact. Experimental results demonstrate that our method significantly outperforms the state-of-the-art results for entity linking on Chinese social media. Chengfang Ma, Ying Sha, Jianlong Tan, Li Guo 0001, Huailiang Peng |
COMPSAC (1) | 3 |
| 2019 | Deep Active Learning for Anchor User PredictionabstractPredicting pairs of anchor users plays an important role in the cross-network analysis. Due to the expensive costs of labeling anchor users for training prediction models, we consider in this paper the problem of minimizing the number of user pairs across multiple networks for labeling as to improve the accuracy of the prediction. To this end, we present a deep active learning model for anchor user prediction (DALAUP for short). However, active learning for anchor user sampling meets the challenges of non-i.i.d. user pair data caused by network structures and the correlation among anchor or non-anchor user pairs. To solve the challenges, DALAUP uses a couple of neural networks with shared-parameter to obtain the vector representations of user pairs, and ensembles three query strategies to select the most informative user pairs for labeling and model training. Experiments on real-world social network data demonstrate that DALAUP outperforms the state-of-the-art approaches. Anfeng Cheng, Chuan Zhou 0001, Hong Yang 0003, Jia Wu 0001, Lei Li 0002, Jianlong Tan, Li Guo 0001 |
IJCAI | 6 |
| 2019 | Learning to Draw Text in Natural Images with Conditional Adversarial NetworksabstractIn this work, we propose an entirely learning-based method to automatically synthesize text sequence in natural images leveraging conditional adversarial networks. As vanilla GANs are clumsy to capture structural text patterns, directly employing GANs for text image synthesis typically results in illegible images. Therefore, we design a two-stage architecture to generate repeated characters in images. Firstly, a character generator attempts to synthesize local character appearance independently, so that the legible characters in sequence can be obtained. To achieve style consistency of characters, we propose a novel style loss based on variance-minimization. Secondly, we design a pixel-manipulation word generator constrained by self-regularization, which learns to convert local characters to plausible word image. Experiments on SVHN dataset and ICDAR, IIIT5K datasets demonstrate our method is able to synthesize visually appealing text images. Besides, we also show the high-quality images synthesized by our method can be used to boost the performance of a scene text recognition algorithm. Shancheng Fang, Hongtao Xie 0001, Jianlong Tan, Yongdong Zhang 0001 |
IJCAI | 4 |
| 2019 | Neighborhood-Aware Attentional Representation for Multilingual Knowledge GraphsabstractMultilingual knowledge graphs constructed by entity alignment are the indispensable resources for numerous AI-related applications. Most existing entity alignment methods only use the triplet-based knowledge to find the aligned entities across multilingual knowledge graphs, they usually ignore the neighborhood subgraph knowledge of entities that implies more richer alignment information for aligning entities. In this paper, we incorporate neighborhood subgraph-level information of entities, and propose a neighborhood-aware attentional representation method NAEA for multilingual knowledge graphs. NAEA devises an attention mechanism to learn neighbor-level representation by aggregating neighbors' representations with a weighted combination. The attention mechanism enables entities not only capture different impacts of their neighbors on themselves, but also attend over their neighbors' feature representations with different importance. We evaluate our model on two real-world datasets DBP15K and DWY100K, and the experimental results show that the proposed model NAEA significantly and consistently outperforms state-of-the-art entity alignment models. Qiannan Zhu, Xiaofei Zhou 0002, Jia Wu 0001, Jianlong Tan, Li Guo 0001 |
IJCAI | 4 |
| 2019 | UAFA: Unsupervised Attribute-Friendship Attention Framework for User Representation
Yanmin Shang, Yaman Cao, Yanbing Liu 0007, Jianlong Tan |
KSEM (1) | 5 |
| 2018 | Reinforcement Learning for Joint Extraction of Entities and Relations
Wenpeng Liu, Yanan Cao 0001, Yanbing Liu 0007, Yue Hu 0002, Jianlong Tan |
ICANN (2) | 5 |
| 2018 | Attention-Based RNN Model for Joint Extraction of Intent and Word Slot Based on a Tagging Strategy
Zheng Fang 0002, Yanan Cao 0001, Yanbing Liu 0007, Xiaojun Chen 0004, Jianlong Tan |
ICANN (3) | 6 |
| 2018 | A Data-Deduplication-Based Matching Mechanism for URL FilteringabstractURL filtering plays an important role in various network security applications. URL filtering usually requires high matching performance, but the performance of the classical multiple string matching algorithms have been difficult to be significantly improved. In this article, we found that the online URLs to be filtered contain a large number of duplicate URLs. According to this observation, we propose a novel deduplication-based matching mechanism (DBM) for URL filtering. The DBM caches information of the duplicate URLs in a hash table to avoid duplicate URLs being repeatedly scanned by URL filtering system. The DBM can be used in conjunction with any multiple string matching algorithms. Experimental results show that when a multiple string matching algorithm used in conjunction with the DBM, the matching speed of the URL filtering system can be increased by 9\%-68\%. So DBM can significantly accelerate the speed of URL filtering system. Besides increasing speed of URL filtering system, DBM is a mechanism independent of the specific matching algorithm and can be easily used in other field. Yuhai Lu, Yanbing Liu 0007, Jianlong Tan |
ICC | 4 |
| 2018 | Sequence Generative Adversarial Network for Long Text SummarizationabstractIn this paper, we propose a new adversarial training framework for text summarization task. Although sequence-to-sequence models have achieved state-of-the-art performance in abstractive summarization, the training strategy (MLE) suffers from exposure bias in the inference stage. This discrepancy between training and inference makes generated summaries less coherent and accuracy, which is more prominent in summarizing long articles. To address this issue, we model abstractive summarization using Generative Adversarial Network (GAN), aiming to minimize the gap between generated summaries and the ground-truth ones. This framework consists of two models: a generator that generates summaries, a discriminator that evaluates generated summaries. Reinforcement learning (RL) strategy is used to guarantee the co-training of generator and discriminator. Besides, motivated by the nature of summarization task, we design a novel Triple-RNNs discriminator, and extend the off-the-shelf generator by appending encoder and decoder with attention mechanism. Experimental results showed that our model significantly outperforms the state-of-the-art models, especially on long text corpus. Yanan Cao 0001, Ruipeng Jia, Yanbing Liu 0007, Jianlong Tan |
ICTAI | 5 |
| 2018 | User Alignment via Structural Interaction and PropagationabstractUser alignment between different social networks is a fundamental issue for many applications, such as information diffusion and recommendation. In actuality, the observed anchor users are normally sparse due to the expensiveness of labeling data. Hence how to make the best use of these sparse anchor information is an important open issue. To this end, we proposed a User Alignment via Structural Interaction and Propagation (UASIP) model to capture the structural information interaction across two social networks, which exploits deep structural infor- mation to enhance the representations of users. UASIP learns vector representation by automatically keeping the consistency between this additional structural information and intrinsic structural information of the two social networks. Experiments on real-world social network datasets demonstrate the effectiveness of UASIP compared with several state-of-the-art methods. Anfeng Cheng, Chun-Yi Liu 0003, Chuan Zhou 0001, Jianlong Tan, Li Guo 0001 |
IJCNN | 4 |
| 2018 | Fine-Grained Correlation Learning with Stacked Co-attention Networks for Cross-Modal Information Retrieval
Jing Yu 0007, Yanbing Liu 0007, Jianlong Tan, Li Guo 0001, Weifeng Zhang 0002 |
KSEM (1) | 4 |
| 2018 | Group Outlying Aspects Mining
Shaoni Wang, Haiyang Xia 0001, Gang Li 0009, Jianlong Tan |
KSEM (1) | 4 |
| 2018 | WDMTI: Wireless Device Manufacturer and Type Identification Using Hierarchical Dirichlet ProcessabstractWireless devices have been widely adopted across all domains. With the convenience brought by wireless communication technology, increasing number of conventional (wired) devices are evolving to become wireless. However, significant security issues arise with the popularity of wireless devices. To start an attack, the adversary usually performs a network reconnaissance to discover exposed devices, identify device manufacturers and types, and then scan for vulnerabilities. From the defense side, network administrators are expected to identify the potential vulnerabilities/risks and enforce Network Access Control (or Network Admission Control, NAC) on all the connecting devices. To do this, it is essential to accurately identify the make/model/type of each device that attempts to connect to the network, e.g., MacBooks, Samsung smart phones (Android), Amazon kindles, DLink surveillance cameras, TP-Link smart plugs, etc. In this paper, we present a novel approach, namely WDMTI, for the identification of wireless device manufacturer and type. We tackle the challenge from two aspects: the features and the classification model. First, we claim that it is critical to discover the device manufacturer and type as soon as the device requests to join the WLAN, and it is unrealistic to make other assumptions on the status of the device, e.g., assuming that the device is booting up or initializing a new connection to corresponding servers/clouds. We primarily depend on the features extracted from the network connection phase, while features from device booting are considered "bonus". In particular, we propose to utilize features from the raw HDCP packets, which is shown to be sufficient for device manufacturer and type recognition with high accuracy. Meanwhile, in the WDMTI system, we employ the Hierarchical Dirichlet Process (HDP), which is a nonparametric Bayesian model for grouped data. HDP allows new groups to be introduced with new data being added, i.e. previously unknown devices connect to the network and the extracted features receive new labels. The WDMTI mechanism is dynamically retrained on-line, instead of requiring a time-consuming off-line retraining process. Our experiments show that WDMTI identifies known types of devices with average accuracy of 0.89, and new types of devices with average accuracy of 0.96, both of which is higher than the state-of-art approaches. In summary, we present a wireless device manufacturer and type identification (WDMTI) system that is both scalable and accurate, and capable of adapting to unknown types of devices on-the-fly. Lingjing Yu, Zhaoyu Zhou, Yujia Zhu, Qingyun Liu 0001, Jianlong Tan |
MASS | 6 |
| 2018 | Attention and Language Ensemble for Scene Text Recognition with Convolutional Sequence ModelingabstractRecent dominant approaches for scene text recognition are mainly based on convolutional neural network (CNN) and recurrent neural network (RNN), where the CNN processes images and the RNN generates character sequences. Different from these methods, we propose an attention-based architecture1 which is completely based on CNNs. The distinctive characteristics of our method include: (1) the method follows encoder-decoder architecture, in which the encoder is a two-dimensional residual CNN and the decoder is a deep one-dimensional CNN. (2) An attention module that captures visual cues, and a language module that models linguistic rules are designed equally in the decoder. Therefore the attention and language can be viewed as an ensemble to boost predictions jointly. (3) Instead of using a single loss from language aspect, multiple losses from attention and language are accumulated for training the networks in an end-to-end way. We conduct experiments on standard datasets for scene text recognition, including Street View Text, IIIT5K and ICDAR datasets. The experimental results show our CNN-based method has achieved state-of-the-art performance on several benchmark datasets, even without the use of RNN. Shancheng Fang, Hongtao Xie 0001, Zhengjun Zha, Nannan Sun, Jianlong Tan, Yongdong Zhang 0001 |
ACM Multimedia | 5 |
| 2018 | My Friend Leaks My Privacy: Modeling and Analyzing Privacy in Social NetworksabstractWith the dramatically increasing participation in online social networks (OSNs), huge amount of private information becomes available on such sites. It is critical to preserve users' privacy without preventing them from socialization and sharing. Unfortunately, existing solutions fall short meeting such requirements. We argue that the key component of OSN privacy protection is protecting (sensitive) content -- privacy as having the ability to control information dissemination. We follow the concepts of private information boundaries and restricted access and limited control to introduce a social circle model. We articulate the formal constructs of this model and the desired properties for privacy protection in the model. We show that the social circle model is efficient yet practical, which provides certain level of privacy protection capabilities to users, while still facilitates socialization. We then utilize this model to analyze the most popular social network platforms on the Internet (Facebook, Google+, WeChat, etc), and demonstrate the potential privacy vulnerabilities in some social networks. Finally, we discuss the implications of the analysis, and possible future directions. Lingjing Yu, Sri Mounica Motipalli, Dongwon Lee 0001, Peng Liu 0005, Qingyun Liu 0001, Jianlong Tan, Bo Luo |
SACMAT | 7 |
| 2017 | Inferring Social Network User's Interest Based on Convolutional Neural Network
Yanan Cao 0001, Shi Wang 0002, Cong Cao 0001, Yanbing Liu 0007, Jianlong Tan |
ICONIP (5) | 6 |
| 2017 | Acceleration of RSA processes based on hybrid ARM-FPGA clusterabstractCooperation of software and hardware with hybrid architectures, such as Xilinx Zynq SoC combining ARM CPU and FPGA fabric, is a high-performance and low-power platform for accelerating RSA Algorithm. This paper adopts the none-subtraction Montgomery algorithm and the Chinese Remainder Theorem (CRT) to implement high-speed RSA processors, and deploys a 48-node cluster infrastructure based on Zynq SoC to achieve extremely high scalability and throughput of RSA computing. In this design, we use the ARM to implement node-to-node communication with the Message Passing Interface (MPI) while use the FPGA to handle complex calculation. Finally, the experimental results show that the overall performance is linear with the number of nodes. And the cluster achieves 6×~9× speedup against a multi-core desktop (Intel i7-3770) and comparable performance to a many-core server (288-core). In addition, we gain up to 2.5× energy efficiency compared to these two traditional platforms. Lei Jiang 0003, Qiong Dai, Jianlong Tan |
ISCC | 5 |
| 2017 | Inferring User Profiles in Online Social Networks Based on Convolutional Neural Network
Yanan Cao 0001, Yanmin Shang, Yanbing Liu 0007, Jianlong Tan, Li Guo 0001 |
KSEM | 5 |
| 2017 | Identification of Influential Users Based on Topic-Behavior Influence Tree in Social Networks
Jianjun Wu 0004, Ying Sha, Rui Li 0022, Qi Liang 0002, Bo Jiang 0013, Jianlong Tan, Bin Wang 0004 |
NLPCC | 6 |
| 2017 | Boosting imbalanced data learning with Wiener process oversampling
Qian Li 0003, Gang Li 0009, Wenjia Niu, Yanan Cao 0001, Liang Chang 0003, Jianlong Tan, Li Guo 0001 |
Frontiers Comput. Sci. | 6 |
| 2016 | PiDFA: A practical multi-stride regular expression matching engine based On FPGAabstractDPI technology has been widely deployed in networking intrusion detection system (NIDS) to detect attacks or viruses. State-of-the-art NIDS uses deterministic finite automata (DFA) algorithms to perform regular expression matching for its stable matching speed. However, traditional DFA algorithm's throughput is limited by the input character's width (usually one character per time). Although the multi-stride method (process multiple characters per time) can increase the throughput, it leads the DFA transition table to an exponentially increased memory consumption. In this paper, we propose a novel multi-stride regular expression matching engine called PiDFA based on Field-Programmable Gate Array (FPGA). It applies two methods to solve traditional multi-stride algorithms' memory explosion problem: DFA Transition Merging method and top-k state extraction method. Experiment results show that PiDFA achieves more than 30-fold better performance than original DFA algorithm. Whats more, PiDFA is orthogonal to existing transition table compression algorithms. Implemented with PiDFA algorithm, ClusterFA's matching speed is increased by 6-50 times while maintaining ClusterFA's low memory consumption. Lei Jiang 0003, Qiu Tang, Qiong Dai, Jianlong Tan |
ICC | 5 |
| 2016 | Riemannian optimization with subspace tracking for low-rank recoveryabstractLow-rank matrix recovery (MR) has been widely used in data analysis and dimensionality reduction. As a direct heuristic to MR, convex relaxation is usually degraded by the repeated calling of singular value decomposition (SVD), especially in large-scale applications. In this paper, we propose a novel Riemannian optimization method (ROAM) for MR problem by exploiting the Riemannian geometry of the searching space. In particular, ROAM utilizes an efficient subspace tracking schema that automatically detects the unknown rank to identify the preferable geometry space. Moreover, a gradient-based optimization algorithm is proposed to obtain the latent low-rank component, which avoids the expensive full dimension of SVD. More significantly, ROAM algorithm is proved to converge under mild assumptions, which also verifies the effectiveness of ROAM. Extensive empirical results demonstrate the improved accuracy and efficiency of ROAM over convex-relaxation approaches. Qian Li 0003, Wenjia Niu, Gang Li 0009, Jianlong Tan, Gang Xiong 0001, Li Guo 0001 |
IJCNN | 4 |
| 2016 | Exploring probabilistic follow relationship to prevent collusive peer-to-peer piracy
Wenjia Niu, Endong Tong, Qian Li 0003, Gang Li 0009, Xuemin Wen, Jianlong Tan, Li Guo 0001 |
Knowl. Inf. Syst. | 6 |
| 2015 | Lingo: Linearized Grassmannian Optimization for Nuclear Norm MinimizationabstractAs a popular heuristic to the matrix rank minimization problem, nuclear norm minimization attracts intensive research attentions. Matrix factorization based algorithms can reduce the expensive computation cost of SVD for nuclear norm minimization. However, most matrix factorization based algorithms fail to provide the theoretical guarantee for convergence caused by their non-unique factorizations. This paper proposes an efficient and accurate Linearized Grassmannian Optimization (Lingo) algorithm, which adopts matrix factorization and Grassmann manifold structure to alternatively minimize the subproblems. More specially, linearization strategy makes the auxiliary variables unnecessary and guarantees the close-form solution for low per-iteration complexity. Lingo then converts linearized objective function into a nuclear norm minimization over Grassmannian manifold, which could remedy the non-unique of solution for the low-rank matrix factorization. Extensive comparison experiments demonstrate the accuracy and efficiency of Lingo algorithm. The global convergence of Lingo is guaranteed with theoretical proof, which also verifies the effectiveness of Lingo. Qian Li 0003, Wenjia Niu, Gang Li 0009, Yanan Cao 0001, Jianlong Tan, Li Guo 0001 |
CIKM | 5 |
| 2015 | An efficient sparse matrix format for accelerating regular expression matching on field-programmable gate arraysabstractRegular expression matching is widely used in many programming languages and applications. A regular expression is transformed into a deterministic finite automata DFA for processing. However, the DFA requires large memory resources because of the state blowup problem. Many algorithms have been proposed to compress the DFA storage and generally store the compressed DFA in sparse matrix format. For field-programmable gate array FPGA-based implementations, operations on sparse matrix consume multiple clock cycles, thus reducing the flexibility and performance of applications. To accelerate the regular expression matching, we present a compact sparse matrix format for storing the compressed DFA transition table on the FPGA. Taking advantage of the special properties of sparse matrices generated by DFAs, we can accomplish one access within a single clock cycle. Furthermore, we develop a regular expression matching engine on a Xilinx Xilinx Inc. Location: 2100 Logic Dr, San Jose, CA 95124-3400, USA Virtex-6 FPGA chip using this sparse matrix format. Compared with previous solutions, this regular expression matching engine has more flexibility while keeping high compression ratio. The results show that this regular expression matching engine saves 94% of memory space compared with the original DFA structure while keeping a fast matching speed. By running multiple engines in parallel, our design achieves a throughput up to 29Gbps. Copyright ©2013 John Wiley & Sons, Ltd. Lei Jiang 0003, Jianlong Tan, Qiu Tang |
Secur. Commun. Networks | 2 |
| 2014 | A factor-searching-based multiple string matching algorithm for intrusion detectionabstractMultiple string matching plays a fundamental role in network intrusion detection systems. Automata-based multiple string matching algorithms like AC, SBDM and SBOM are widely used in practice, but the huge memory usage of automata prevents them from being applied to a large-scale pattern set. Meanwhile, poor cache locality of huge automata degrades the matching speed of algorithms. Here we propose a space-efficient multiple string matching algorithm BVM, which makes use of bit-vector and succinct hash table to replace the automata used in factor-searching-based algorithms. Space complexity of the proposed algorithm is O(rm2+ ΣpϵP|p|), that is more space-efficient than the classic automata-based algorithms. Experiments on datasets including Snort, ClamAV, URL blacklist and synthetic rules show that the proposed algorithm significantly reduces memory usage and still runs at a fast matching speed. Above all, BVM costs less than 0.75% of the memory usage of AC, and is capable of matching millions of patterns efficiently. Yanbing Liu 0007, Qingyun Liu 0001, Ping Liu 0001, Jianlong Tan, Li Guo 0001 |
ICC | 4 |
| 2014 | A fast regular expression matching engine for NIDS applying prediction schemeabstractRegular expression matching is considered important as it lies at the heart of many networking applications using deep packet inspection (DPI) techniques. For example, modern networking intrusion detection systems (NIDSs) typically accomplish regular expression matching using deterministic finite automata (DFA) algorithm. However, DFA suffers from the high memory consumption for the state blowup problem. Many algorithms have been proposed to compress the DFA memory storage space, meanwhile, they usually pay the price of low matching speed and high memory bandwidth. In this paper, we first propose an effective DFA compression algorithm by exploiting the similarity between DFA states. Then, we apply a next-state prediction strategy and present a fast DFA matching engine. Carefully designing the DFA matching circuit, we keep the prediction success rate by more than 99,5%, thus get a comparable matching speed with original DFA algorithm. On the side of memory consumption, experimental results show that with typical NIDS rule sets, our algorithm compressed the original DFA by more than 99%. Mapping this algorithm on Xilinx Virtex-7 FPGA chip, we get a throughput of more than 200Gbps. Lei Jiang 0003, Qiong Dai, Qiu Tang, Jianlong Tan, Binxing Fang |
ISCC | 4 |
| 2014 | Contextual Query Expansion for Image RetrievalabstractIn this paper, we study the problem of image retrieval by introducing contextual query expansion to address the shortcomings of bag-of-words based frameworks: semantic gap of visual word quantization, and the efficiency and storage loss due to query expansion. Our method is built on common visual patterns (CVPs), which are the distinctive visual structures between two images and have rich contextual information. With CVPs, two contextual query expansions on visual word-level and image-level are explored, respectively. For visual word-level expansion, we find contextual synonymous visual words (CSVWs) and expand a word in the query image with its CSVWs to boost retrieval accuracy. CSVWs are the words that appear in the same CVPs and have same contextual meaning, i.e. similar spatial layout and geometric transformations. For image-level expansion, the database images that have the same CVPs are organized by linked list and the images that have the same CVPs as the query image, but not included in the results are automatically expanded. The main computation of these two expansions is carried out offline, and they can be integrated into the inverted file and efficiently applied to all images in the dataset. Experiments conducted on three reference datasets and a dataset of one million images demonstrate the effectiveness and efficiency of our method. Hongtao Xie 0001, Yongdong Zhang 0001, Jianlong Tan, Li Guo 0001, Jintao Li 0001 |
IEEE Trans. Multim. | 3 |
| 2012 | ClusterFA: a memory-efficient DFA structure for network intrusion detectionabstractNetwork intrusion detection systems (NIDS) plays an increasing important role in the field of network security. Current NIDS, such as Bro and Snort, mainly use signatures to represent and detect networking attacks. Traditionally the signatures are depicted by exact string patterns. However, new worms and viruses emerge endlessly in recent years. As a result, the scale of signatures increases sharply. Compared with exact strings, regular expressions have more powerful expressiveness, and are replacing exact strings gradually in state-of-the-art NIDS. Lei Jiang 0003, Jianlong Tan, Yanbing Liu 0007 |
AsiaCCS | 2 |
| 2011 | Mining frequent patterns across multiple data streamsabstractMining frequent patterns from data streams has drawn increasing attention in recent years. However, previous mining algorithms were all focused on a single data stream. In many emerging applications, it is of critical importance to combine multiple data streams for analysis. For example, in real-time news topic analysis, it is necessary to combine multiple news report streams from dierent media sources to discover collaborative frequent patterns which are reported frequently in all media, and comparative frequent patterns which are reported more frequently in a media than others. To address this problem, we propose a novel frequent pattern mining algorithm Hybrid-Streaming, H-Stream for short. H-Stream builds a new Hybrid-Frequent tree to maintain historical frequent and potential frequent itemsets from all data streams, and incrementally updates these itemsets for efficient collaborative and comparative pattern mining. Theoretical and empirical studies demonstrate the utility of the proposed method. Peng Zhang 0001, Jianlong Tan, Li Guo 0001 |
CIKM | 3 |
| 2011 | Continuous data stream query in the cloudabstractCloud computing represents one of the most important research directions for modern computing systems. Existing research efforts on Cloud computing were all focused on designing advanced storage and query techniques for static data. None of them consider the problem that data in a Cloud may appear as continuous and rapid data streams. To address this problem, in this paper we propose a new LCN-Index framework to handle continuous data stream queries in the Cloud. LCN-Index uses the Map-Reduce computing paradigm to process all the queries. In the Mapping stage, it divides all the queries into a batch of predicate sets which are then deployed onto mapping nodes using interval predicate index. In the reducing stage, it merges results from the mapping nodes using multi attribute hash index. In so doing, a data stream can be efficiently evaluated by traversing through the LCN-Index framework. Experiments demonstrate the utility of the proposed method. Jun Li 0016, Peng Zhang 0001, Jianlong Tan, Ping Liu 0001, Li Guo 0001 |
CIKM | 3 |
| 2011 | Adaptive Shared-Filter Ordering for Efficient Multimedia Stream MonitoringabstractMultimedia stream monitoring refers to removing unwanted and malicious records from multimedia streams. In this application, a large number of filtering queries are registered on time-critical multimedia streams. Each filtering query contains multiple meta filters and a meta filter is shared among multiple filtering queries. The filtering queries and meta filters form a bipartite graph, and the objective is to minimize the overall evaluation time of the queries in the bipartite graph. In order to achieve this goal, some heuristic algorithms were proposed to order the shared meta filters in the graph to reduce the overall evaluation cost. While these methods can achieve near-optimal solutions in ideal stream environments that have stationary probability distributions, in this paper we propose an Adaptive Shared-filter Ordering Model (ASOM) for efficient filtering in dynamic data stream environments. To capture new trends and patterns along dynamic data streams, ASOM uses a time-based exponential smoothing forecasting method to adaptively order the shared meta filters for fast estimation. Experiments demonstrate that ASOM outperforms existing heuristic ordering methods in dynamic stream environments. Jun Li 0016, Peng Wang 0028, Peng Zhang 0001, Jianlong Tan |
ICTAI | 4 |
| 2010 | SKIF: a data imputation framework for concept drifting data streamsabstractMissing data commonly occurs in many applications. While many data imputation methods exist to handle the missing data problem for large scale databases, when applied to concept drifting data streams, these methods face some common difficulties. First, due to large and continuous data volumes, we are unable to maintain all stream records to form a candidate pool and estimate missing values, as most existing methods commonly do. Second, even if we could maintain all complete stream records using a summary structure, the concept drifting problem would make some information obsolete, and thus deteriorate the imputation accuracy. Third, in data streams, it is necessary to develop a fast yet accurate algorithm to find the most similar data for imputation. Fourth, due to the dynamic and sophisticated data collection environments, the missing rate of most stream data may be much higher than that in generic static databases, so the imputation method should be able to accommodate high missing rate in the data. To tackle these challenges, we propose, in this paper, a Streaming k-Nearest-Neighbors Imputation Framework (SKIF) for concept drifting data streams. To handle concept drifting and large volume problems in data streams, SKIF first summarizes historical complete records in some micro-resources (which are high-level statistical data structures), and maintains these micro-resources in a candidate pool as benchmark data. After that, SKIF employs a novel hybrid-kNN imputation procedure, which uses a hybrid similarity search mechanism, to find the most similar micro-resources from the large scale candidate pool efficiently. Experimental results demonstrate the effectiveness of the proposed SKIF framework for data stream imputation tasks. Peng Zhang 0001, Xingquan Zhu 0001, Jianlong Tan, Li Guo 0001 |
CIKM | 3 |
| 2010 | Classifier and Cluster Ensembles for Mining Concept Drifting Data StreamsabstractEnsemble learning is a commonly used tool for building prediction models from data streams, due to its intrinsic merits of handling large volumes stream data. Despite of its extraordinary successes in stream data mining, existing ensemble models, in stream data environments, mainly fall into the ensemble classifiers category, without realizing that building classifiers requires labor intensive labeling process, and it is often the case that we may have a small number of labeled samples to train a few classifiers, but a large number of unlabeled samples are available to build clusters from data streams. Accordingly, in this paper, we propose a new ensemble model which combines both classifiers and clusters together for mining data streams. We argue that the main challenges of this new ensemble model include (1) clusters formulated from data streams only carry cluster IDs, with no genuine class label information, and (2) concept drifting underlying data streams makes it even harder to combine clusters and classifiers into one ensemble framework. To handle challenge (1), we present a label propagation method to infer each cluster's class label by making full use of both class label information from classifiers, and internal structure information from clusters. To handle challenge (2), we present a new weighting schema to weight all base models according to their consistencies with the up-to-date base model. As a result, all classifiers and clusters can be combined together, through a weighted average mechanism, for prediction. Experiments on real-world data streams demonstrate that our method outperforms simple classifier ensemble and cluster ensemble for stream data mining. Peng Zhang 0001, Xingquan Zhu 0001, Jianlong Tan, Li Guo 0001 |
ICDM | 3 |
| 2010 | Compressing Regular Expressions' DFA Table by Matrix Decomposition
Yanbing Liu 0007, Li Guo 0001, Ping Liu 0001, Jianlong Tan |
CIAA | 4 |
| 2009 | Measuring the Influence of Active Measurement on Unstructured Peer-to-Peer NetworkabstractAlthough intensive researches have been performed regarding P2P network measurement, it is still unknown to what extent the measurement system influences the final measurement results. As an initial study, we investigated the influence of a measurement system on degree distribution of a P2P network. Theoretical analysis and simulation results suggest an interesting phase-transition phenomena when the size of the measurement system increases. A P2P network mixed with a small active measurement system remains as a scale-free network; however, the mixture P2P network will not remain as scale-free when the size of the measurement system exceeds a threshold. We also observed that the number of measuring peers usually has stronger influence than the number of connections among these peers. Briefly speaking, a small but dense active measurement is better than a large but loose measurement system. Ying Sha, Jianlong Tan |
ICPADS | 3 |
| 2009 | A Table Compression Method for Extended Aho-Corasick Automaton
Yanbing Liu 0007, Yifu Yang, Ping Liu 0001, Jianlong Tan |
CIAA | 4 |
| 2008 | Accelerating Multiple String Matching by Using Cache-Efficient StrategyabstractString matching plays a fundamental role in many network security applications such as NIDS, virus detection and information filtering. In this paper, we proposed cache-efficient methods to accelerate classical multiple string matching algorithms. We observed that most classical algorithms perform poorly as pattern set grows due to their high memory requirement and the poor cache behavior. Based on this observation, we proposed efficient methods employing cache-efficient strategies, i.e., to accelerate string matching by minimizing memory usage and maximizing cache locality. Experimental results on random datasets demonstrated that our new methods are substantially faster than classical methods. Jianlong Tan, Yanbing Liu 0007, Ping Liu 0001 |
WAIM | 1 |
| 2005 | A Partition-Based Efficient Algorithm for Large Scale Multiple-Strings Matching
Ping Liu 0001, Yanbing Liu 0007, Jianlong Tan |
SPIRE | 3 |