Zhuoyi Wang

dblp:194/7513 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RECS: LSTM-Based Cognitive Status Estimation for Human-Robot Interaction
abstract
Effective communication is critical to the success of many types of human-robot interaction. A key capability for enabling effective communication is accurate modeling of the cognitive status that entities hold (e.g., modeling what objects interlocutors are currently thinking about, or are generally aware of). However, existing models of cognitive status estimation are not well suited for situated and embodied interactions, as they do not account for nonverbal cues, which are a key way in which humans moderate cognitive status. To address this gap, we make three primary contributions. First, we introduce the BOWTIE corpus of dialogues from a multi-modal open-world referential task, annotated with cognitive status, gesture type, linguistic roles, and grammatical roles of entities across utterances. Second, we introduce RECS, the first LSTM-based model of cognitive status, which we train on the BOWTIE corpus. Third, we present empirical evidence for the success of RECS. These contributions stand to accelerate the future use of cognitively informed algorithms for robot language understanding and generation.
Mark Higger, Zhuoyi Wang, Polina Rygina, Lara Ferreira Bezerra, Logan Daigler, Zane Aloia, Sheena Wu, Ishani Pandey, Amanda Chen, Tom Williams 0001
HRI2
2026 An Automated Vulnerability Detection Framework for Smart Contracts
abstract
With the increase of the adoption of blockchain technology in providing decentralized solutions to various problems, smart contracts have become more popular to the point that billions of US Dollars are currently exchanged every day through such technology. Meanwhile, various vulnerabilities in smart contracts have been exploited by attackers to steal cryptocurrencies worth millions of dollars. The automatic detection of smart contract vulnerabilities therefore is an essential research problem. Existing solutions to this problem particularly rely on human experts to define features or different rules to detect vulnerabilities. However, this often causes many vulnerabilities to be ignored, and they are inefficient in detecting new vulnerabilities. In this study, to overcome such challenges, we propose a framework to automatically detect vulnerabilities in smart contracts on the blockchain. More specifically, first, we utilize novel feature vector generation techniques from bytecode of smart contract as source code is rarely publicly available. These feature vectors are then analyzed using our innovative metric learning-based Deep Neural Networks (DNNs) to produce detection results. The framework’s predictions are further refined through a voting mechanism to achieve consensus. We conduct comprehensive experiments on large-scale benchmarks, and the quantitative results demonstrate the effectiveness and efficiency of our approach.
Feng Mi, Chen Zhao 0010, Zhuoyi Wang, Sadaf Md. Halim, Xiaodi Li 0002, Zhouxiang Wu, Latifur Khan, Bhavani Thuraisingham
Distributed Ledger Technol. Res. Pract.3
2024 TabLog: Test-Time Adaptation for Tabular Data Using Logic Rules
abstract
We consider the problem of test-time adaptation of predictive models trained on tabular data. Effective solution of this problem requires adaptation of predictive models trained on the source domain to a target domain, using only unlabeled target domain data, without access to source domain data. Existing test-time adaptation methods for tabular data have difficulty coping with the heterogeneous features and their complex dependencies inherent in tabular data. To overcome these limitations, we consider test-time adaptation in the setting wherein the logical structure of the rules is assumed to remain invariant despite distribution shift between source and target domains whereas the numerical parameters associated with the rules and the weights assigned to them can vary to accommodate distribution shift. TabLog discretizes numerical features, models dependencies between heterogeneous features, introduces a novel contrastive loss for coping with distribution shift, and presents an end-to-end framework for efficient training and test-time adaptation by taking advantage of a logical neural network representation of a rule ensemble. We present results of experiments using several benchmark data sets that demonstrate TabLog is competitive with or improves upon the state-of-the-art methods for test-time adaptation of predictive models trained on tabular data. Our code is available at https://github.com/WeijieyingRen/TabLog.
Weijieying Ren, Xiaoting Li 0001, Huiyuan Chen, Vineeth Rakesh, Zhuoyi Wang, Mahashweta Das, Vasant G. Honavar
ICML5
2022 Latent Coreset Sampling based Data-Free Continual Learning
abstract
Catastrophic forgetting poses a major challenge in continual learning where the old knowledge is forgotten when the model is updated on new tasks. Existing solutions tend to solve this challenge through generative models or exemplar-replay strategies. However, such methods may not alleviate the issue that the low-quality samples are generated or selected for the replay, which would directly reduce the effectiveness of the model, especially in the class imbalance, noise, or redundancy scenarios. Accordingly, how to select a suitable coreset during continual learning becomes significant in such setting. In this work, we propose a novel approach that leverages continual coreset sampling (CCS) to address these challenges. We aim to select the most representative subsets during each iteration. When the model is trained on new tasks, it closely approximates/matches the gradient of both the previous and current tasks with respect to the model parameters. This way, adaptation of the model to new datasets could be more efficient. Furthermore, different from the old data storage for maintaining the old knowledge, our approach choose to preserving them in the latent space. We augment the previous classes in the embedding space as the pseudo sample vectors from the old encoder output, strengthened by the joint training with selected new data. It could avoid data privacy invasions in a real-world application when we update the model. Our experiments validate the effectiveness of our proposed approach over various CV/NLP datasets under against current baselines, and we also indicate the obvious improvement of model adaptation and forgetting reduction in a data-free manner.
Zhuoyi Wang, Dingcheng Li, Ping Li 0001
CIKM1
2021 Single View Point Cloud Generation via Unified 3D Prototype
abstract
As 3D point clouds become the representation of choice for multiple vision and graphics applications, such as autonomous driving, robotics, etc., the generation of them by deep neural networks has attracted increasing attention in the research community. Despite the recent success of deep learning models in classification and segmentation, synthesizing point clouds remains challenging, especially from a single image. State-of-the-art (SOTA) approaches can generate a point cloud from a hidden vector, however, they treat 2D and 3D features equally and disregard the rich shape information within the 3D data. In this paper, we address this problem by integrating image features with 3D prototype features. Specifically, we propose to learn a set of 3D prototype features from a real point cloud dataset and dynamically adjust them through the training. These prototypes are then integrated with incoming image features to guide the point cloud generation process. Experimental results show that our proposed method outperforms SOTA methods on single image based 3D reconstruction tasks.
Yu Lin 0002, Yigong Wang, Yifan Li 0003, Zhuoyi Wang, Yang Gao 0027, Latifur Khan
AAAI4
2021 Contextual Rephrase Detection for Reducing Friction in Dialogue Systems
abstract
For voice assistants like Alexa, Google Assistant and Siri, correctly interpreting users' intentions is of utmost importance.However, users sometimes experience friction with these assistants, caused by errors from different system components or user errors such as slips of the tongue.Users tend to rephrase their query until they get a satisfactory response.Rephrase detection is used to identify the rephrases and has long been treated as a task with pairwise input, which does not fully utilize the contextual information (e.g.users' implicit feedback).To this end, we propose a contextual rephrase detection model ContReph to automatically identify rephrases from multiturn dialogues.We showcase how to leverage the dialogue context and user-agent interaction signals, including user's implicit feedback and the time gap between different turns, which can help significantly outperform the pairwise rephrase detection models.
Zhuoyi Wang, Saurabh Gupta 0008, Dingcheng Li, Alexander Hanbo Li, Chenlei Guo
EMNLP (1)1
2021 An Episodic Learning based Geolocation Detection Framework for Imbalanced Data
abstract
A social media user's geographical location is vital to many applications like local search and event detection. The scarcity of publicly available location information motivates researchers to predict user geolocation based on information such as tweet text and social interaction data. In this paper, we investigate and improve on the task of predicting a Twitter user's city-level location based on the content of the user's historical tweets. In order to train a reliable location classifier, previous studies on this topic have typically assumed that there are sufficient amount of users living in each cities. However, they simply ignore the fact that different demographic groups may participate in social media platforms, which results in a highly imbalanced data distribution. Being aware of this population imbalance issue, we propose an episodic learning based framework to extract a single representative for each class (location), so that classifiers can later be trained on a balanced class distribution. To examine the effectiveness of our method, we design experiments which involve two kinds of baselines, the state-of-the-art geolocation detection methods and the well-known approaches handling imbalanced data in classification. The results of experiments on the data collected from Twitter demonstrated the superiority of our method when compared with baselines.
Hemeng Tao, Yang Gao 0027, Zhuoyi Wang, Latifur Khan, Bhavani Thuraisingham
IJCNN3
2021 Generating Point Cloud from Single Image in The Few Shot Scenario
abstract
Reconstructing point clouds from images would extremely benefit many practical CV applications, such as robotics, automated vehicles, and Augmented Reality. Fueled by the advances of deep neural network, many deep learning frameworks are proposed to address this problem recently. However, these frameworks generally rely on a large amount of labeled training data (e.g., image and point cloud pairs). Although we usually have numerous 2D images, corresponding 3D shapes are insufficient in practice. In addition, most available 3D data covers only a limited amount of classes, which further restricts the models' generalization ability to novel classes. To mitigate these issues, we propose a novel few-shot single-view point cloud generation framework by considering both class-specific and class-agnostic 3D shape priors. Specifically, we abstract each class by a prototype vector that embeds class-specific shape priors. Class-agnostic shape priors are modeled by a set of learnable shape primitives that encode universal 3D shape information shared across classes. Later, we combine the input image with class-specific prototypes and class-agnostic shape primitives to guide the point cloud generation process. Experiments on the popular ModelNet and ShapeNet datasets demonstrate that our method outperforms state-of-the-art methods in the few-shot setting.
Yu Lin 0002, Jinghui Guo, Yang Gao 0027, Yifan Li 0003, Zhuoyi Wang, Latifur Khan
ACM Multimedia5
2021 CIFDM: Continual and Interactive Feature Distillation for Multi-Label Stream Learning
abstract
Multi-label learning algorithms have attracted more and more attention as of recent. This is mainly because real-world data is generally associated with multiple and non-exclusive labels, which could correspond to different objects, scenes, actions, and attributes. In this paper, we consider the following challenging multi-label stream scenario: the new labels emerge continuously in the changing environments, and are assigned to the previous data. In this setting, data mining solutions must be able to learn the new concepts and avoid catastrophic forgetting simultaneously. We propose a novel continual and interactive feature distillation-based learning framework (CIFDM), to effectively classify instances with novel labels. We utilize the knowledge from the previous tasks to learn new knowledge to solve the current task. Then, the system compresses historical and novel knowledge and preserves it while waiting for new emerging tasks. CIFDM consists of three components: 1) a knowledge bank that stores the existing feature-level compressed knowledge, and predicts the observed labels so far; 2) a pioneer module that aims to learn and predict new emerged labels based on knowledge bank.; 3) an interactive knowledge compression function which is used to compress and transfer the new knowledge to the bank, and then apply the current compressed knowledge to initialize the label embedding of the pioneer for the next task.
Yigong Wang, Zhuoyi Wang, Yu Lin 0002, Latifur Khan, Dingcheng Li
SIGIR2
2021 Attention-Based Spatial Guidance for Image-to-Image Translation
abstract
The aim of image-to-image translation algorithms is to tackle the challenges of learning a proper mapping function across different domains. Generative Adversarial Networks (GANs) have shown superior ability to handle this problem in both supervised and unsupervised ways. However, one critical problem of GAN in practice is that the discriminator is typically much stronger than the generator, which could lead to failures such as mode collapse, diminished gradient, etc. To address these shortcomings, we propose a novel framework, which incorporates a powerful spatial attention mechanism to guide the generator. Specifically, our designed discriminator estimates the probability of realness of a given image, and provides an attention map regarding this prediction. The generated attention map contains the informative regions to distinguish the real and fake images, from the perspective of the discriminator. Such information is particularly valuable for the translation because the generator is encouraged to focus on those areas and produce more realistic images. We conduct extensive experiments and evaluations, and show that our proposed method is both qualitatively and quantitatively better than other state-of-the-art image translation frameworks.
Yu Lin 0002, Yigong Wang, Yifan Li 0003, Yang Gao 0027, Zhuoyi Wang, Latifur Khan
WACV5
2021 CLEAR: Contrastive-Prototype Learning with Drift Estimation for Resource Constrained Stream Mining
abstract
Non-stationary data stream mining aims to classify large scale online instances that emerge continuously. The most apparent challenge compared with the offline learning manner is the issue of consecutive emergence of new categories, when tackling non-static categorical distribution. Non-stationary stream settings often appear in real-world applications, e.g., online classification in E-commerce systems that involves the incoming productions, or the summary of news topics on social networks (Twitter). Ideally, a learning model should be able to learn novel concepts from labeled data (in new tasks) and reduce the abrupt degradation of model performance on the old concept (also named catastrophic forgetting problem). In this work, we focus on improving the performance of the stream mining approach under the constrained resources, where both the memory resource of old data and labeled new instances are limited/scarce. We propose a simple yet efficient resource-constrained framework CLEAR to facilitate previous challenges during the one-pass stream mining. Specifically, CLEAR focuses on creating and calibrating the class representation (the prototype) in the embedding space. We first apply the contrastive-prototype learning on large amount of unlabeled data, and generate the discriminative prototype for each class in the embedding space. Next, for updating on new tasks/categories, we propose a drift estimation strategy to calibrate/compensate for the drift of each class representation, which could reduce the knowledge forgetting without storing any previous data. We perform experiments on public datasets (e.g., CUB200, CIFAR100) under stream setting, our approach is consistently and clearly better than many state-of-the-art methods, along with both the memory and annotation restriction.
Zhuoyi Wang, Yuqiao Chen, Chen Zhao 0010, Yu Lin 0002, Xujiang Zhao, Hemeng Tao, Yigong Wang, Latifur Khan
WWW1
2020 Few Sample Learning without Data Storage for Lifelong Stream Mining (Student Abstract)
abstract
Continuously mining complexity data stream has recently been attracting an increasing amount of attention, due to the rapid growth of real-world vision/signal applications such as self-driving cars and online social media messages. In this paper, we aim to address two significant problems in the lifelong/incremental stream mining scenario: first, how to make the learning algorithms generalize to the unseen classes only from a few labeled samples; second, is it possible to avoid storing instances from previously seen classes to solve the catastrophic forgetting problem? We introduce a novelty stream mining framework to classify the infinite stream of data with different categories that occurred during different times. We apply a few-sample learning strategy to make the model recognize the novel class with limited samples; at the same time, we implement an incremental generative model to maintain old knowledge when learning new coming categories, and also avoid the violation of data privacy and memory restrictions simultaneously. We evaluate our approach in the continual class-incremental setup on the classification tasks and ensure the sufficient model capacity to accommodate for learning the new incoming categories.
Zhuoyi Wang, Yigong Wang, Yu Lin 0002, Hemeng Tao, Latifur Khan
AAAI1
2020 A Primal-Dual Subgradient Approach for Fair Meta Learning
abstract
The problem of learning to generalize on unseen classes during the training step, also known as few-shot classification, has attracted considerable attention. Initialization based methods, such as the gradient-based model agnostic meta-learning (MAML) [1], tackle the few-shot learning problem by “learning to fine-tune”. The goal of these approaches is to learn proper model initialization, so that the classifiers for new classes can be learned from a few labeled examples with a small number of gradient update steps. Few shot meta-learning is well-known with its fast-adapted capability and accuracy generalization onto unseen tasks [2]. Learning fairly with unbiased outcomes is another significant hallmark of human intelligence, which is rarely touched in few-shot meta-learning. In this work, we propose a Primal-Dual Fair Meta-learning framework, namely PDFM, which learns to train fair machine learning models using only a few examples based on data from related tasks. The key idea is to learn a good initialization of a fair model's primal and dual parameters so that it can adapt to a new fair learning task via a few gradient update steps. Instead of manually tuning the dual parameters as hyperparameters via a grid search, PDFM optimizes the initialization of the primal and dual parameters jointly for fair meta-learning via a subgradient primal-dual approach. We further instantiate an example of bias controlling using decision boundary covariance (DBC) [3] as the fairness constraint for each task, and demonstrate the versatility of our proposed approach by applying it to classification on a variety of three realworld datasets. Our experiments show substantial improvements over the best prior work for this setting. Our code and datasets are available at https://github.com/charliezhaoyinpeng/PDFM.git.
Chen Zhao 0010, Feng Chen 0001, Zhuoyi Wang, Latifur Khan
ICDM3
2020 Adaptive Multi-Region Network For Medical Image Analysis
abstract
Automated diagnosis of significant abnormalities (or lesions) from radiology images has been well exploited in Deep Learning (DL) because of the ability to model sophisticated features. However, a deep neural network should be trained on a huge amount of data to infer the parameter values. Unfortunately, for the problems in lesion diagnosis, there is only a limited amount of data annotated in a manner that is suitable to learn powerful deep models. Moreover, the lesion in the radiology image is often vague and hard to identify without expert knowledge. In this paper, we focus on previous challenges in the automated diagnosis and propose the approach named Adaptive Multi-region Network (AdapNet). The key idea is that we adaptively encode the similarity of lesions in different context regions through margin-max learning strategy, which incorporates the metrics learned on those regions to enhance the effectiveness of the model. Our experiments show that the proposed method can effectively obtain superior performance compared to the existing methods, on the DeepLesion data sets.
Hemeng Tao, Zhuoyi Wang, Yang Gao 0027, Yigong Wang, Latifur Khan
ICIP2
2020 Few-Sample and Adversarial Representation Learning for Continual Stream Mining
abstract
Deep Neural Networks (DNNs) have primarily been demonstrated to be useful for closed-world classification problems where the number of categories is fixed. However, DNNs notoriously fail when tasked with label prediction in a non-stationary data stream scenario, which has the continuous emergence of the unknown or novel class (categories not in the training set). For example, new topics continually emerge in social media or e-commerce. To solve this challenge, a DNN should not only be able to detect the novel class effectively but also incrementally learn new concepts from limited samples over time. Literature that addresses both problems simultaneously is limited. In this paper, we focus on improving the generalization of the model on the novel classes, and making the model continually learn from only a few samples from the novel categories. Different from existing approaches that rely on abundant labeled instances to re-train/update the model, we propose a new approach based on Few Sample and Adversarial Representation Learning (FSAR). The key novelty is that we introduce the adversarial confusion term into both the representation learning and few-sample learning process, which reduces the over-confidence of the model on the seen classes, further enhance the generalization of the model to detect and learn new categories with only a few samples. We train the FSAR operated in two stages: first, FSAR learns an intra-class compacted and inter-class separated feature embedding to detect the novel classes; next, we collect a few labeled samples belong to the new categories, utilize episode-training to exploit the intrinsic features for few-sample learning. We evaluated FSAR on different datasets, using extensive experimental results from various simulated stream benchmarks to show that FSAR effectively outperforms current state-of-the-art approaches.
Zhuoyi Wang, Yigong Wang, Yu Lin 0002, Evan Delord, Latifur Khan
WWW1
2019 Regression Prediction For Geolocation Aware Through Relative Density Ratio Estimation
abstract
Typically, a traditional regression model over a data set is trained on instances by assuming a stationary distribution (the training data and test data has the same distribution). The model learned from the training set is later used to predict value of response-variable in future test instances. In this paper, we study the regression problem in a novelty problem setting, referred as a transfer learning involves training set and test set with different data distribution. Here, the training set has a biased data distribution with respect to the test set. The regression problem with the transfer learning approach aims to predict the dependent variable of test data, through utilizing the training data with the biased distribution. Availability of sufficient training data is of prime importance. Unfortunately, such problem settings are sparingly available in many real-word applications, such as you only have the flight delay information of DFW (DallasWorth) airport and prefer to predict the flight delay information of JFK (John F. Kennedy) airport. In our solution, we utilize relative density ratio to evaluate the difference between the data from different locations. We revise the relative density ratio estimation based on different location types, and propose an effective approach for predicting arrival delay time for commercial flights by geolocation aware relative kernel density estimation. Extensive experimental results on commercial flight datasets with different locations show that our approach effectively boosts the performance of other state-of-the art.
Jinghui Guo, Zhuoyi Wang, Yang Gao 0027, Latifur Khan
IEEE BigData3
2019 Co-Representation Learning Framework For the Open-Set Data Classification
abstract
Deep Neural Network (DNN) has been largely demonstrated to be effective for real-world classification problems. However, such model requires a huge amount of training samples to get more accurate result. When limited samples allowed for the training step, the model may perform weak generalization ability on the test set, especially when the novel/unseen class may occur during the test period (we call it open-set classification). This severely limits its further utility in many real-world large scale applications, such as the open-set image and text classification scenarios. In this paper, we focus on addressing this key challenge by developing a DNN based co-representation learning approach RLCN. It utilizes limited samples for training a model then applies it to classify normal instances and detect the emergence of novel class over time. The key novelty is that we design a weighted pairwise-constraint loss (WPC) function to learn an enhanced generalization and robust feature embedding, where the intra-class (same class) compactness and inter-class (different class) separation are achieved. Moreover, we apply the temperature scaling scheme on the softmax function to replace traditional softmax output in our open-world classifier to achieve the classification and novel class detection simultaneously. Our extensive empirical evaluation on benchmark datasets demonstrate the effectiveness of our framework compared to other competing techniques.
Zhuoyi Wang, Yu Lin 0002, Yigong Wang, Md Shihabul Islam, Latifur Khan
IEEE BigData1
2019 Robust High Dimensional Stream Classification with Novel Class Detection
abstract
A primary challenge in label prediction over a data stream is the emergence of instances belonging to unknown or novel class over time. Traditionally, studies addressing this problem aim to detect such instances using cluster-based mechanisms. They typically assume that instances from the same class are closer to each other than those belonging to different classes in observed feature space. Unfortunately, this may not hold true in higher-dimensional feature space such as images. In recent years, Convolutional neural network (CNN) have emerged as a leading system to be employed in many real-world application. Yet, based on the assumption of closed world dataset with a fixed number of categories, CNN lacks robustness for novel class detection, so it is unclear on how such models can be used to deal with novel class instances along a high-dimensional image stream. In this paper, we focus on addressing this challenge by proposing an effective learning framework called CNN-based Prototype Ensemble (CPE) for novel class detection and correction. Our framework includes a prototype ensemble loss (PE) to improve the intra-class compactness and expand inter-class separateness in the output feature representation, thereby enabling the robustness of novel class detection. Moreover, we provide an incremental learning strategy which maintains a constant amount of exemplars to update the network, making it more practical for real-world application. We empirically demonstrate the effectiveness of our framework by comparing its performance over multiple realworld image benchmark data streams with existing state-of-theart data stream detection techniques. The implementation of CPE is on: https://github.com/Vitvicky/Convolutional-Net-PrototypeEnsemble
Zhuoyi Wang, Zelun Kong, Swarup Chandra, Hemeng Tao, Latifur Khan
ICDE1
2019 COMC: A Framework for Online Cross-domain Multistream Classification
abstract
With the tremendous increase of the online data, training a single classifier may suffer because of the large variety of data domains. One solution could be to learn separate classifiers for each domain. However, this would arise a huge cost to gather annotated training data for a large number of domains and ignore similarity shared across domains. Hence, it leads to our problem setting: can labeled data from a related source domain help predict the unlabeled data in the target domain? In this paper, we consider two independent simultaneous data streams, which are referred to as the source and target streams. The target stream continuously generates data instances from one domain where the label is unknown, while the source stream continuously generates labeled data instances from another domain. Most likely, the two data streams would have different but related feature spaces and different data distributions. Moreover, these streams may have asynchronous concept drifts between them. Our problem setting, which is called Cross-domain Multistream Classification, is to predict the class labels of data instances in the target stream using a classifier trained on the labeled source stream. In this paper, we propose an efficient solution for cross-domain multistream classification by integrating change detection into online data stream adaptation. The class labels of data instances in the target stream are predicted using the sufficient amount of label information in the related source stream. And the concept drifts along the two independent streams are continuously being addressed at the same time. Experimental results on real-world data sets indicate significantly improved performance over baseline methods.
Hemeng Tao, Zhuoyi Wang, Yifan Li 0003, Mahmoud Zamani, Latifur Khan
IJCNN2
2019 Metric Learning based Framework for Streaming Classification with Concept Evolution
abstract
A primary challenge in label prediction over a stream of continuously occurring data instances is the emergence of instances belonging to unknown or novel classes. It is imperative to detect such novel-class instances quickly along the stream for a superior prediction performance. Existing techniques that perform novel class detection typically employ a clustering-based mechanism by observing that instances belonging to the same class (intra-class) are closer to each other (cohesion) than inter-class samples (separation). While this is generally true in low dimensional feature spaces, we observe that such a property is not intrinsic among instances in complex real-world high-dimensional feature space such as images and text. In this paper, we focus on addressing this key challenge that negatively affects prediction performance of a data stream classifier. Concretely, we develop a metric learning mechanism that transforms high-dimensional features into a latent feature space to make above property holds true. Unlike existing metric learning method which only focus on classification task, our approach address the novel class detection and stream classification simultaneously. We showcase a framework along the stream to achieve larger prediction performance compared to existing state-of-the-art detection techniques while using the least amount of labeled data during detection. Extensive experimental results on simulated and real-world stream demonstrate the effectiveness of our approach.
Zhuoyi Wang, Hemeng Tao, Zelun Kong, Swarup Chandra, Latifur Khan
IJCNN1
2017 FUSION: An Online Method for Multistream Classification
abstract
Traditional data stream classification assumes that data is generated from a single non-stationary process. On the contrary, multistream classification problem involves two independent non-stationary data generating processes. One of them is the source stream that continuously generates labeled data. The other one is the target stream that generates unlabeled test data from the same domain. The distribution represented by the source stream data is biased compared to that of the target stream. Moreover, these streams may have asynchronous concept drifts between them. The multistream classification problem is to predict the class labels of target stream instances by utilizing labeled data from the source stream. This kind of scenario is often observed in real-world applications due to scarcity of labeled data. The only existing approach for multistream classification uses separate drift detection on the streams for addressing the asynchronous concept drift problem. If a concept drift is detected in any of the streams, it uses an expensive batch technique for data shift adaptation. These add significant execution overhead, and limit its usability. In this paper, we propose an efficient solution for multistream classification by fusing drift detection into online data shift adaptation. We study the theoretical convergence rate and computational complexity of the proposed approach. Moreover, empirical results on benchmark data sets indicate significantly improved performance over the baseline methods.
Ahsanul Haque, Zhuoyi Wang, Swarup Chandra, Latifur Khan, Kevin W. Hamlen
CIKM2
2016 Sampling-based distributed Kernel mean matching using spark
abstract
Limited access to supervised information may forge scenarios in real-world data mining applications, where training and test data are interconnected by a covariate shift, i.e., having equal class conditional distribution with unequal covariate distribution. Traditional data mining techniques assume that both training and test data represent an identical distribution, therefore suffer in presence of a covariate shift. Kernel Mean Matching (KMM) is a well known approach that addresses covariate shift by weighing training instances appropriately. However, it has time complexity cubic in the size of training data, which is computationally impractical for large or streaming datasets due to limited scalability. In this paper, we present a sampling-based algorithm to address the limited scalability problem of KMM. Moreover, we show that the approach is highly parallelizable, and therefore propose a distributed algorithm for estimating training instance weights efficiently using Spark. Experiment results on benchmark datasets show that the proposed approach achieves competitive estimation accuracy within much lower execution time compared to the KMM algorithm. Moreover, it indicates that larger size of training data results into a higher accuracy with minimal effect on execution time of the proposed approach.
Ahsanul Haque, Zhuoyi Wang, Swarup Chandra, Yupeng Gao, Latifur Khan, Charu C. Aggarwal
IEEE BigData2